Magento TTFB: Varnish, Redis and the Backend Cost

lead

Magento TTFB is the number that decides whether any of your other performance work is measurable. Largest Contentful Paint cannot be faster than the moment the first byte arrives, so a store sitting at 1.8 seconds of server time has an LCP floor of 1.8 seconds no matter how good its images are — and the usual consequence is worse than slow pages. A team optimises images, sees LCP move by 0.1s, and concludes images do not matter on this store. They were measuring the floor.

This is one of four deep dives under Magento 2 Core Web Vitals: From Fails to Passes, which carries the metrics, the level ordering and the cost-benefit framing. Start there if you have not picked an order to work in — doing these out of sequence produces measurements you will misread.

Order of operations, and the reason for it: production mode, then full-page cache, then Redis, then the database, then application PHP. Each step changes what the next measurement means. Profiling PHP on a store running in developer mode measures the developer mode, not the code.


baseline: where the time before the first byte actually goes

👥 For: DevOps, Backend leads

⏱ TL;DR: Production mode (30 min) + Varnish (2 hr) + Database indexes (2–4 hr) = 50–60% TTFB improvement. Highest ROI quick wins.

A. Who owns this work

Ownership and effort for the whole programme live in the pillar article rather than being restated on each deep dive. Short version: DevOps owns caching and infrastructure, backend owns queries and code, and the two meet at profiling — backend reads it, DevOps deploys what it implies.

B. Profiling: what is safe to run where

You need a call graph, not a guess. The choice is mostly about where you may run it.

ToolLocalProductionCost
XHProf✅ First choice, ships with DDEV⚠️ Last resort, see belowFree
XDebug profiler✅ Fine❌ NeverFree
Blackfire✅ Fine✅ Built for itCommercial
Sentry profiling—✅ Sampled by designCommercial

Locally, reach for XHProf — DDEV ships it: ddev xhprof on, make the request, read the flame graph (next section). Production is where the two free options stop being interchangeable.

XDebug’s profiler must never run in production. It writes a cachegrind file per request and can multiply response times several-fold.

XHProf in production: the trigger condition is the danger, not the profiler. Misconfigure the sample rate or the enable flag and every request gets profiled, which will take the server down. Last resort only: gate it with a sampling percentage plus a header or cookie check, and verify on staging that an ordinary request is not profiled.

If you already run Sentry for exceptions, its profiling is the lower-risk route to continuous production data: it samples by design, the rate is a config value rather than something you implement, and it needs no extra agent. It also puts the profile beside the error trace, where an incident investigation wants it.

What you read the profile for matters more, because three shapes account for most Magento TTFB problems:

  • A huge call count — getName() invoked ten thousand times on a category page is a collection loaded without the attributes it needs, resolved one entity at a time.
  • Cache load() tens of thousands of times — cheap individually, not free in aggregate; usually a block asking per item instead of per page.
  • query() counts in the thousands — the classic N+1, and the fix is batching in a repository rather than tuning the query.

Self time separates "this is slow" from "this calls something slow": sort by self to find the work, by cumulative to find who asked.

C. DDEV built-in profiling (zero setup)

For DDEV users: Profiling comes out of the box

If you’re using DDEV (.ddev/ present), XDebug and Blackfire profiling are pre-installed and ready to use. No manual setup needed.

XDebug profiling via DDEV (already enabled):

# Inside DDEV container, enable profiling for a request
ddev exec curl -H "X-Debug-Profile: 1" http://localhost/product/test-product > /dev/null

# Find the profile
ddev exec ls -lh /tmp/xdebug/cachegrind.out.* | tail -1
ddev exec ls -lh /var/tmp/xdebug/cachegrind.out.* | tail -1

# Copy profile to host for analysis
ddev exec cat /tmp/xdebug/cachegrind.out.LATEST > /tmp/profile.cachegrind

# Visualize locally
qcachegrind /tmp/profile.cachegrind  # On macOS: install via `brew install qcachegrind`

Blackfire via DDEV (pre-configured):

DDEV includes Blackfire agent. To profile:

# 1. Authenticate Blackfire (one-time)
ddev blackfire auth --client-id=YOUR_ID --client-token=YOUR_TOKEN

# 2. Profile a page
ddev blackfire curl http://localhost/product/test-product

# 3. Output: Flamegraph in terminal (or JSON for programmatic analysis)
ddev blackfire curl http://localhost/product/test-product --json > profile.json

# 4. View interactive web interface (if Blackfire account linked)
# Results appear in https://blackfire.io dashboard

Finding profiles in DDEV:

# XDebug profiles stored in container
ddev exec ls -lh /tmp/xdebug/

# Blackfire profiles stored in Blackfire cloud (or local cache)
ddev blackfire list  # Show recent profiles

Quick profiling without CLI (via browser):

If you prefer GUI-driven profiling:

1. XDebug Web UI (included in DDEV):

  • Enable in .ddev/config.yaml: xdebug_mode: debug,develop,profile
  • Access debugger at: http://localhost:9003 (depends on DDEV setup)

2. Blackfire Browser Extension (if you have Blackfire account):

  • Install Blackfire extension in Chrome/Firefox
  • Visit your DDEV site: http://localhost
  • Click Blackfire icon → "Profile" → View results instantly

Real-world DDEV profiling workflow (fast path):

# 1. Navigate to your DDEV-based Magento project
cd ~/development/brocode/wp-brocode

# 2. Start DDEV (if not running)
ddev start

# 3. Profile a slow page
ddev blackfire curl http://localhost/product/test-product

# 4. See output like:
# Performance profile:
# ├─ Total: 1200ms
# ├─ DB queries: 600ms (50% of time)
# ├─ Cache: 200ms (17%)
# └─ Rendering: 400ms (33%)

# 5. Identify bottleneck (DB is slowest) → go fix queries/add indexes

Why this matters:

  • No need to manually install Blackfire agent, XDebug, or QCacheGrind
  • Profile production-like behavior in your local environment
  • Catch regressions before merging (add to CI: ddev blackfire curl --json > baseline.json)

D. Database optimization: keys and indexes

Why database performance matters:

  • Unindexed queries = full table scan = 100–1000ms per query
  • Missing indexes on commonly filtered columns = 50–80% of TTFB
  • Product pages often do 10–50 queries per load
  • 10 queries × 100ms each = 1 second TTFB hit

Step 1: Identify slow queries with EXPLAIN

# Log slow queries (MySQL)
# my.cnf / my.ini
[mysqld]
slow_query_log = 1
slow_query_log_file = /var/log/mysql/slow.log
long_query_time = 0.5  # Log queries taking >500ms

# Monitor slow queries
tail -f /var/log/mysql/slow.log

Step 2: Analyze query performance with EXPLAIN

-- Check if query uses indexes efficiently
EXPLAIN SELECT * FROM catalog_product_entity WHERE status = 1 AND visibility > 0;

-- Output explanation:
-- | id | select_type | table                    | type  | key  | rows   |
-- |----|-------------|--------------------------|-------|------|--------|
-- | 1  | SIMPLE      | catalog_product_entity   | ALL   | NULL | 500000 | ← NO INDEX (bad!)
-- |    | (filtered)  | (using WHERE clause)     |       |      | 5000   |

-- ✗ BAD: type=ALL (full table scan), key=NULL (no index used)
-- ✓ GOOD: type=ref/eq_ref, key=status_idx (using index)

Step 3: Add missing indexes

-- Create index on commonly filtered columns
ALTER TABLE catalog_product_entity ADD INDEX idx_status_visibility (status, visibility);

-- Re-run EXPLAIN to verify
EXPLAIN SELECT * FROM catalog_product_entity WHERE status = 1 AND visibility > 0;

-- Output should now show:
-- | type  | key                  |
-- |-------|----------------------|
-- | ref   | idx_status_visibility| ← INDEX USED (good!)

Common Magento tables needing indexes:

TableColumn(s) to IndexWhyQuery Frequency
catalog_product_entitystatus, visibilityProduct listing, visibility filteringHigh
catalog_product_entityskuProduct lookup by SKUVery high
sales_ordercustomer_id, created_atOrder history, admin gridHigh
sales_order_itemorder_idOrder details lookupHigh
catalog_category_productcategory_id, positionCategory product listingVery high
reviewentity_pk_value, statusProduct reviews filterMedium
customer_entityemailCustomer lookupHigh
quotecustomer_id, created_atCart recovery, admin gridMedium

Real impact of missing indexes (product page load):

Before indexing:
├─ Load product (unindexed SKU lookup): 150ms (full table scan)
├─ Load reviews (unindexed entity_pk_value): 200ms
├─ Load categories (unindexed category_id): 100ms
└─ Total: 450ms + other queries = 1.5s TTFB

After indexing:
├─ Load product (indexed SKU lookup): 5ms
├─ Load reviews (indexed entity_pk_value): 10ms
├─ Load categories (indexed category_id): 3ms
└─ Total: 18ms + other queries = 700ms TTFB (53% faster!)

Step 4: Verify index usage in Magento

// In Magento observer or custom module
$connection = $this->resourceConnection->getConnection();

// Check existing indexes
$indexInfo = $connection->getIndexList('catalog_product_entity');
// Returns array of indexes on this table

// Verify index is being used
$explain = $connection->query("EXPLAIN SELECT * FROM catalog_product_entity WHERE sku = 'TEST-SKU'");
// Check 'key' field in result - should show your index name

Step 5: Monitor index health (ongoing)

# Check for unused indexes (waste space, slow inserts)
SELECT OBJECT_SCHEMA, OBJECT_NAME, COUNT(*) as times_used
FROM performance_schema.table_io_waits_summary_by_index_usage
WHERE OBJECT_SCHEMA != 'mysql'
GROUP BY OBJECT_SCHEMA, OBJECT_NAME
ORDER BY times_used ASC;

# Rebuild fragmented indexes
OPTIMIZE TABLE catalog_product_entity;

Implementation owner:

  • Backend or Database Admin: Identify missing indexes via EXPLAIN
  • DevOps: Deploy indexes (schedule for low-traffic time)
  • Testing: Verify no regressions after index addition

Checklist:

  • [ ] Enable slow query logging (0.5s threshold)
  • [ ] Review slow query log weekly
  • [ ] Run EXPLAIN on each slow query
  • [ ] Add indexes for frequently used WHERE/JOIN columns
  • [ ] Verify queries now use indexes (type ≠ ALL)
  • [ ] Test product/category/order pages after indexing
  • [ ] Monitor query time before/after (document improvement)

E. Quick wins, in order of return

1. Production mode (MAGE_MODE=production)

  • Disables dev logging, static file generation on-the-fly
  • ~30–50% TTFB improvement alone
  • Measure before: time curl http://localhost/ (TTFB in first line)
  • Enable: export MAGE_MODE=production in .bashrc or env.php
  • Measure after: TTFB should drop 30–50%

2. Full-page cache (Varnish) — THE MOST IMPACTFUL OPTIMIZATION

Why Varnish matters (it’s the difference maker):

  • Cache entire HTML for anonymous visitors
  • Anonymous = no personalization = cacheable (90% of traffic)
  • Logged-in users bypass cache (10% of traffic, okay to be slower)
  • Hit rate: 80–95% for typical stores
  • 80% cache hit = 80% of users get <100ms TTFB
  • 20% cache miss = 20% of users get 2s TTFB (acceptable)
  • TTFB on cache hit: <100ms (vs 2s on origin)

Business impact:

  • Cache miss (uncached): TTFB 2000ms, LCP 2.5s (passes barely)
  • Cache hit (cached): TTFB 50ms, LCP 0.8s (excellent)
  • Avg LCP across traffic: (80% × 0.8s) + (20% × 2.5s) = 1.5s overall ✓

Varnish detailed config:

Verify Varnish is working:

Measure impact:

3. Redis for sessions

  • Replace file-based sessions (slow disk I/O)
  • ~200ms TTFB improvement per request

Setup:

Verify Redis is working:

4. PHP 8.3 or 8.4

  • Older PHP = slower opcode execution
  • 8.4 is ~15% faster than 8.1
  • 8.2→8.4 migration: ~2 weeks (mostly testing)

5. Database query optimization

  • Identify N+1 queries in observer/plugin chains
  • Add indexes on frequently filtered columns
  • Profile with slow_query_log

Enable slow query log:

Example slow query detected:

F. Before and after, concretely

Starting state (typical Luma store):

Production mode: OFF
Varnish: disabled
Redis: file sessions
PHP: 8.1
Indexes: default only

Profiler output (Blackfire):
├─ Bootstrap (100ms)
├─ DB queries (1200ms) ← BOTTLENECK
├─ Cache lookups (150ms)
├─ Observers (300ms)
└─ Rendering (250ms)
Total TTFB: 2000ms ❌

After optimizations (same store):

Production mode: ON
Varnish: enabled
Redis: configured
PHP: 8.4
Indexes: added 15 critical ones

Profiler output (Blackfire):
├─ Bootstrap (20ms) [80% faster]
├─ DB queries (200ms) [85% faster, indexes + caching]
├─ Cache lookups (30ms) [80% faster, Redis]
├─ Observers (80ms) [73% faster, unused disabled]
└─ Rendering (100ms) [60% faster, PHP 8.4]
Total TTFB: 430ms ✅ [78% improvement]

Impact on LCP:

  • Before: TTFB 2000ms + render 500ms = LCP 2500ms (scrapes a pass — 2.5s is the boundary, so any regression fails)
  • After: TTFB 430ms + render 500ms = LCP 930ms (passes!)

G. Multi-storeview: Luma and Hyvä side by side

Realistic approach for existing stores with multiple storeviews:

Many teams use a phased migration strategy:

  • Existing storeviews: Stay on Luma (optimize aggressively)
  • New storeviews: Launch on Hyvä (greenfield advantage, no migration risk)

Rationale:

  • Luma optimization buys 6–12 months (sections III–XII can get you to 2.1s LCP)
  • Hyvä migration for existing store is risky (checkout, extensions, team learning)
  • New storefront? Hyvä is lower-risk (no legacy code to migrate)
  • Team learns Tailwind + Alpine on lower-stakes new store
  • Can migrate existing stores to Hyvä later (with Hyvä experience proven)

Example timeline:

Month 1–3: Optimize existing Luma store (apply sections III–XII)
           → TTFB 2.0s → 0.4s, LCP 3.5s → 1.8s (passes CWV)

Month 4–6: Build new storefront on Hyvä
           → Parallel launch for new product line
           → Team learns Tailwind, Hyvä patterns
           → Performance baseline: LCP 1.5s (outperforms Luma version)

Month 7+:  Evaluate migration of existing store
           → If business is stable on Luma, stay until next major release
           → If new Hyvä storeview is winning metric war, plan full migration
           → Learned lessons inform migration strategy

Advantage: You don’t put all risk in one bet. Luma optimization is proven; Hyvä migration is de-risked.

working_example: backend PHP — observers, APIs and the response-time tail

Why PHP Backend Optimization Matters

PHP backend performance directly affects TTFB (Time to First Byte) — the metric that determines LCP on non-cached pages.

The impact chain:

Slow PHP code → High TTFB → High LCP → Core Web Vitals fails

Example: Product page without cache
├─ Browser requests: /product/test-product
├─ Server processes (PHP): 2000ms (slow code)
├─ TTFB: 2000ms (browser waits 2 seconds before first byte)
├─ LCP: 2.5–3.5s (text renders after TTFB completes)
└─ Result: Fails Core Web Vitals ✗

VS optimized PHP:
├─ Server processes (PHP): 300ms (optimized code)
├─ TTFB: 300ms (browser gets first byte fast)
├─ LCP: 1.5–2.0s (text renders earlier)
└─ Result: Passes Core Web Vitals ✓

Real-world scenarios where this matters:

Page TypeCache Hit RateImpact of Slow PHP
Product page60% cached, 40% uncached40% of users hit slow PHP, see high TTFB
Search results20% cached, 80% uncached80% of users hit slow PHP code
Category filtering10% cached, 90% uncached90% of users hit slow PHP (filters are dynamic)
Checkout0% cached (personalized)100% of users hit PHP code (no cache)

Optimization ROI:

  • PHP processing time 2000ms → 300ms = 1.7s TTFB gain
  • LCP improvement: 3.5s → 2.0s (1.5s faster page)
  • Result: ~50% of users see passing LCP (instead of failing)

The Problem: PHP Loops Process Data Slowly

Observer chains and API responses often involve large data processing. Userland PHP loops are slow because:

  • Every iteration calls a function (overhead)
  • Each operation is interpreted, not compiled
  • Memory allocation per operation

Real impact on INP (observer processing 10k products):

Userland loop (foreach):
├─ Load 10k products from DB: 50ms
├─ Loop + validate each product: 120ms (12ms per 1000)
├─ Filter unwanted products: 80ms
├─ Extract prices/attributes: 90ms
└─ Total: 340ms (INP FAILS!)

Built-in array functions (compiled C):
├─ Load 10k products from DB: 50ms
├─ array_filter() + validation: 20ms
├─ array_column() extract prices: 15ms
├─ array_map() transform data: 25ms
└─ Total: 110ms (INP PASSES!) ✓

Solution 1: Use Built-in Array Functions (Compiled)

Instead of:

// BAD: Userland loop (interpreted PHP)
$filtered = [];
foreach ($products as $product) {
  if ($product->getStatus() == 1 && $product->getVisibility() > 0) {
    $filtered[] = $product;
  }
}
// ~80ms for 10k products

Use:

// GOOD: array_filter (compiled C function)
$filtered = array_filter($products, function($product) {
  return $product->getStatus() == 1 && $product->getVisibility() > 0;
});
// ~15ms for 10k products (5x faster!)

Common patterns (foreach → built-in):

// ❌ LOOP: Extract column values
$prices = [];
foreach ($products as $product) {
  $prices[] = $product->getPrice();
}

// ✅ BUILT-IN: array_column (10x faster)
$prices = array_column($products, 'price');  // ~5ms vs 50ms for 10k


// ❌ LOOP: Transform data
$formatted = [];
foreach ($products as $product) {
  $formatted[] = [
    'id' => $product->getId(),
    'name' => strtoupper($product->getName()),
    'price' => $product->getPrice() * 1.2
  ];
}

// ✅ BUILT-IN: array_map (10x faster)
$formatted = array_map(function($product) {
  return [
    'id' => $product->getId(),
    'name' => strtoupper($product->getName()),
    'price' => $product->getPrice() * 1.2
  ];
}, $products);  // ~20ms vs 200ms for 10k


// ❌ LOOP: Find specific item
$found = null;
foreach ($products as $product) {
  if ($product->getId() == 123) {
    $found = $product;
    break;
  }
}

// ✅ BUILT-IN: in_array or array_search (C-level)
$found = array_search(123, array_column($products, 'id'), true);  // <1ms

Real Magento observer example:

// BAD: Observer processing 10k products (blocks INP 340ms)
class InventoryObserver implements ObserverInterface
{
  public function execute(Observer $observer)
  {
    $products = $this->getAllProducts();  // 10k products
    
    // Loop through each product (120ms)
    $filtered = [];
    foreach ($products as $product) {
      if ($product->getStatus() && $this->hasStock($product)) {
        $filtered[] = $product;
      }
    }
    
    // Extract prices (80ms)
    $prices = [];
    foreach ($filtered as $product) {
      $prices[] = $product->getPrice();
    }
    
    // Update cache (140ms total)
    $this->cache->save(json_encode($prices), 'inventory_prices');
  }
}

// GOOD: Same observer, 5x faster (70ms)
class InventoryObserver implements ObserverInterface
{
  public function execute(Observer $observer)
  {
    $products = $this->getAllProducts();  // 10k products
    
    // Filter using array_filter (20ms, not 120ms)
    $filtered = array_filter($products, function($p) {
      return $p->getStatus() && $this->hasStock($p);
    });
    
    // Extract prices using array_column (15ms, not 80ms)
    $prices = array_column($filtered, 'price');
    
    // Update cache (25ms)
    $this->cache->save(json_encode($prices), 'inventory_prices');
  }
}

Impact on INP: Observer 340ms → 70ms (5x faster) ✓ PASSES


Streaming and generator patterns

Why Streaming Matters for TTFB

Streaming/generator patterns reduce memory usage AND improve TTFB for data-heavy operations (imports, API responses, bulk exports).

The impact on TTFB:

Scenario: Product import during flash sale (10,000 products)

Non-streaming (load all):
├─ Load 10k products into memory: 2000ms
├─ Parse JSON: 1000ms
├─ Process in PHP: 5000ms
├─ Total TTFB: 8000ms (8 seconds!)
├─ Memory: 2GB (crashes server if limit = 512MB)
└─ Result: Timeout / out of memory ✗

Streaming (one at a time):
├─ Load first product: 5ms
├─ Process and save: 50ms (loop for each)
├─ Total TTFB for first response: 55ms
├─ Memory: constant 5MB
├─ Total import time: 500 seconds (but streaming, not blocking)
└─ Result: API responds immediately, processes async ✓

Real impact on API response times:

API EndpointWithout StreamingWith StreamingTTFB Savings
POST /api/import (10k items)8000ms (timeout)50ms (immediate)7950ms
GET /api/export (50k items)5000ms (memory error)20ms (start streaming)4980ms
GET /api/catalog (pagination)2000ms (load whole dataset)200ms (load first page)1800ms
Bulk inventory update3000ms (all in memory)100ms (queue async)2900ms

The Problem: Loading Entire Datasets into Memory

When processing large CSV imports or API responses:

// BAD: Load entire file into memory
$data = json_decode(file_get_contents('/imports/products.json'));  // 500MB file = 2GB RAM!
foreach ($data as $row) {
  // Process each row (but entire file already in memory)
}
// Memory spike: 2GB (kills server, times out)

Solution: Generator + Streaming (constant memory)

// GOOD: Generator yields one row at a time
function streamProducts(string $filePath) {
  $file = fopen($filePath, 'r');
  
  while (!feof($file)) {
    $line = fgets($file);
    if ($line) {
      yield json_decode($line, true);  // Yield one row, not entire file
    }
  }
  
  fclose($file);
}

// Use generator (constant 5MB RAM, not 2GB!)
foreach (streamProducts('/imports/products.json') as $row) {
  $this->importProduct($row);  // Process row by row
  // Memory: always ~5MB (one row buffered)
}

Real impact (500MB file with 100k products):

Array approach (load all):
├─ Load file: 2000ms
├─ Parse JSON: 1000ms
├─ Memory usage: 2GB (exceeds PHP limit)
└─ Result: timeout / out of memory error ✗

Generator approach (streaming):
├─ Open file: 5ms
├─ Parse row-by-row: 100ms (while processing)
├─ Memory usage: 5MB (constant)
└─ Result: completes in 15 seconds ✓

Use phpleague/csv Library (Recommended)

Instead of DIY streaming, use the community-standard library:

composer require league/csv

Example: Import CSV without memory explosion

use League\Csv\Reader;

// BAD: Load entire CSV into memory
$csv = array_map('str_getcsv', file('/imports/products.csv'));
foreach ($csv as $row) {
  // 50MB file = 500MB+ memory
}

// GOOD: Stream CSV row by row (constant memory)
$reader = Reader::createFromPath('/imports/products.csv', 'r');
$reader->setHeaderOffset(0);  // First row is header

// Memory: constant ~2MB (one record buffered)
foreach ($reader->getRecords() as $record) {
  $this->importProduct([
    'sku' => $record['sku'],
    'name' => $record['product_name'],
    'price' => (float)$record['price']
  ]);
}

phpleague/csv Features:

use League\Csv\Reader;
use League\Csv\Writer;

// 1. Read with filtering (stream-safe)
$reader = Reader::createFromPath('/imports/products.csv', 'r');
$reader->setHeaderOffset(0);

// Filter on-the-fly (no memory bloat)
$filtered = $reader->filter(function($record) {
  return $record['status'] === 'active';  // Only yield active products
});

foreach ($filtered as $record) {
  $this->importProduct($record);
}

// 2. Write CSV with streaming (not buffered)
$writer = Writer::createFromPath('/exports/products.csv', 'w');
$writer->insertOne(['SKU', 'Name', 'Price']);  // Header

// Stream write (constant memory)
foreach ($this->getProductsIterator() as $product) {
  $writer->insertOne([
    $product->getSku(),
    $product->getName(),
    $product->getPrice()
  ]);
  // Memory: constant, one product at a time
}

// 3. Chunked processing (balance between memory & performance)
$reader = Reader::createFromPath('/imports/products.csv', 'r');

// Process in chunks of 100 (faster than one-by-one)
$chunk = [];
foreach ($reader->getRecords() as $record) {
  $chunk[] = $record;
  
  if (count($chunk) === 100) {
    $this->importBatch($chunk);  // Bulk insert
    $chunk = [];  // Reset
  }
}

Magento Observer Example (Streaming Products for Sync)

// BAD: Sync all 50k products into 3rd-party system (blocks observer 5 seconds)
class ExternalSyncObserver implements ObserverInterface
{
  public function execute(Observer $observer)
  {
    // Load all products (memory spike)
    $products = $this->productRepository->getList(...);  // 50k products
    
    // Send to external API (5 second request)
    $response = $this->apiClient->sync($products);
    
    // Observer blocks entire request
    return $this;
  }
}

// GOOD: Queue products for async sync (observer completes instantly)
class ExternalSyncObserver implements ObserverInterface
{
  public function __construct(
    private PublisherInterface $publisher
  ) {}
  
  public function execute(Observer $observer)
  {
    // Queue each product one-by-one (async)
    $this->streamProductsToQueue();
    
    // Observer returns immediately (async handler processes)
    return $this;
  }
  
  private function streamProductsToQueue(): void
  {
    // Generator yields one product at a time
    foreach ($this->getProductsStream() as $product) {
      // Queue for async processing
      $this->publisher->publish('external.sync.product', json_encode([
        'product_id' => $product->getId(),
        'sku' => $product->getSku()
      ]));
      
      // Memory: constant, never loads all 50k at once
    }
  }
}

// Message handler (processes async, doesn't block INP)
class ExternalSyncHandler
{
  public function execute(string $message): void
  {
    $data = json_decode($message, true);
    $product = $this->productRepository->getById($data['product_id']);
    
    // Sync single product (takes as long as needed)
    $this->apiClient->syncProduct($product);
  }
}

Impact:

  • Observer time: 5 seconds → 50ms (100x faster!)
  • Memory: 500MB (all products) → 5MB (one at a time)
  • INP: blocks entire checkout → passes instantly ✓

Scaling async workers

The Problem: Single Worker is Sequential

One async worker processes messages one-at-a-time:

Queue: [msg1, msg2, msg3, msg4, msg5]

Single worker:
├─ Process msg1: 2 seconds
├─ Process msg2: 2 seconds
├─ Process msg3: 2 seconds
├─ Process msg4: 2 seconds
├─ Process msg5: 2 seconds
└─ Total time: 10 seconds (messages pile up during traffic spikes)

Solution: Run Multiple Workers in Parallel

Queue: [msg1, msg2, msg3, msg4, msg5]

5 workers in parallel:
├─ Worker 1: Process msg1 (2 seconds)
├─ Worker 2: Process msg2 (2 seconds)  ← Parallel
├─ Worker 3: Process msg3 (2 seconds)  ← Parallel
├─ Worker 4: Process msg4 (2 seconds)  ← Parallel
├─ Worker 5: Process msg5 (2 seconds)  ← Parallel
└─ Total time: 2 seconds (5x faster!)

Real impact on queue backlog:

Scenario: 1000 product sync messages queued during flash sale

Single worker (1 process):
├─ Processes 1 message/second
├─ Backlog: 1000 messages
├─ Time to clear: 1000 seconds (16+ minutes!)
└─ Users wait: Inventory doesn't sync until sale ends ✗

5 workers (5 processes):
├─ Process 5 messages/second
├─ Backlog: 1000 messages
├─ Time to clear: 200 seconds (3 minutes)
└─ Users see: Inventory updates within sale window ✓

10 workers:
├─ Process 10 messages/second
├─ Time to clear: 100 seconds (1.7 minutes)
└─ Users see: Real-time inventory ✓✓

Implementation: Run Multiple Consumers

Step 1: Start multiple worker processes

# Start 5 workers for image conversion queue
for i in {1..5}; do
  php bin/magento queue:consumers:start brocode_image_convert \
    --max-messages=100 &
done

# Verify workers running
ps aux | grep queue:consumers:start
# Output:
# worker-1: queue:consumers:start brocode_image_convert --max-messages=100
# worker-2: queue:consumers:start brocode_image_convert --max-messages=100
# worker-3: queue:consumers:start brocode_image_convert --max-messages=100
# worker-4: queue:consumers:start brocode_image_convert --max-messages=100
# worker-5: queue:consumers:start brocode_image_convert --max-messages=100

Step 2: Supervisor Config (Keep Workers Running)

# /etc/supervisor/conf.d/magento-workers.conf
[group:magento_workers]
programs=magento_worker_1,magento_worker_2,magento_worker_3,magento_worker_4,magento_worker_5

[program:magento_worker_1]
process_name=%(program_name)s_%(process_num)02d
command=php /var/www/magento/bin/magento queue:consumers:start brocode_image_convert --max-messages=100
autostart=true
autorestart=true
numprocs=1
redirect_stderr=true
stdout_logfile=/var/log/magento/worker_1.log
user=www-data

[program:magento_worker_2]
process_name=%(program_name)s_%(process_num)02d
command=php /var/www/magento/bin/magento queue:consumers:start brocode_image_convert --max-messages=100
autostart=true
autorestart=true
numprocs=1
redirect_stderr=true
stdout_logfile=/var/log/magento/worker_2.log
user=www-data

# ... repeat for workers 3-5

Step 3: Monitor Worker Health

# Check if workers are processing
ps aux | grep "queue:consumers" | grep -v grep | wc -l
# Output: 5 (all 5 workers running)

# Monitor queue depth (messages waiting)
redis-cli
> LLEN magento2_default:queue:brocode_image_convert
(integer) 245  # 245 messages in queue

# Calculate processing time
# 245 messages ÷ 5 workers ÷ 1 message/sec per worker = ~50 seconds

Step 4: DDEV Users (Run Workers in Container)

# Run multiple workers in DDEV
ddev exec -d php bin/magento queue:consumers:start brocode_image_convert --max-messages=100 &
ddev exec -d php bin/magento queue:consumers:start brocode_image_convert --max-messages=100 &
ddev exec -d php bin/magento queue:consumers:start brocode_image_convert --max-messages=100 &

# Or use a helper script
cat > bin/start-workers.sh << 'EOF'
#!/bin/bash
WORKER_COUNT=${1:-5}
QUEUE_NAME=${2:-brocode_image_convert}

for i in $(seq 1 $WORKER_COUNT); do
  ddev exec -d php bin/magento queue:consumers:start $QUEUE_NAME --max-messages=100 &
  echo "Started worker $i"
done

echo "$WORKER_COUNT workers started for queue: $QUEUE_NAME"
EOF

chmod +x bin/start-workers.sh

# Usage
./bin/start-workers.sh 5 brocode_image_convert

How to Choose Worker Count

Formula: workers = (queue_depth / 60 seconds) / messages_per_second_per_worker

Example: 1000 messages, 2 seconds per message
├─ Messages per second per worker: 0.5 (1 ÷ 2)
├─ Workers needed: (1000 / 60) / 0.5 = ~33 workers
│  (To clear backlog within 1 minute)
│
├─ More realistic (clear within 5 minutes):
│  Workers needed: (1000 / 300) / 0.5 = ~7 workers
│
└─ Rule of thumb:
   - Light load (few messages): 2–3 workers
   - Normal load (100+ messages/minute): 5–10 workers
   - Heavy load (1000+ messages/minute): 10–20 workers
   - Peak load (5000+ messages/minute): 20–50 workers

Scaling Example: Brocode Image Optimizer

Scenario: 50k products, each needs AVIF/WebP conversion (2 seconds per product)

Single worker:
├─ Time to convert all: 50000 × 2 = 100,000 seconds (28 hours!)
└─ Images won't be ready until next day ✗

5 workers:
├─ Time to convert all: 100,000 ÷ 5 = 20,000 seconds (5.5 hours)
└─ Images ready for next morning ✓

10 workers:
├─ Time to convert all: 100,000 ÷ 10 = 10,000 seconds (2.8 hours)
└─ Images ready in evening ✓✓

20 workers:
├─ Time to convert all: 100,000 ÷ 20 = 5,000 seconds (1.4 hours)
└─ Images ready within an hour ✓✓✓

Resource Cost vs Benefit

Each worker uses:
├─ CPU: 1 core (50–100% utilization)
├─ Memory: 50–100MB (PHP process)
├─ Network: Minimal (queue + DB only)
└─ Total: 5 workers = ~250–500MB memory, 5 CPU cores

Cost trade-off:
├─ 1 worker: Cheap, slow (queue builds up)
├─ 5 workers: Moderate cost, good performance
├─ 10 workers: Higher cost, excellent performance
├─ 20+ workers: Very high cost, overkill for most stores
└─ Recommendation: Start with 5, scale up if queue builds >1000 messages

Monitor & Auto-Scale

// Health check: Alert if queue depth > threshold
class QueueHealthObserver implements ObserverInterface
{
  public function __construct(
    private Redis $redis,
    private AlertService $alertService
  ) {}

  public function execute(Observer $observer): void
  {
    // Check queue depth every minute (via cron)
    $queueDepth = $this->redis->llen('magento2_default:queue:brocode_image_convert');
    
    if ($queueDepth > 1000) {
      // Alert: Need more workers
      $this->alertService->notify(
        "Queue depth critical: $queueDepth messages. Start more workers."
      );
    }
    
    // Auto-scale recommendation
    $workersNeeded = ceil($queueDepth / 200);  // 200 messages per worker acceptable
    if ($workersNeeded > 5) {
      $this->alertService->notify(
        "Recommend $workersNeeded workers (currently 5)"
      );
    }
  }
}

Scaling the worker tier

For multi-server setups (high-traffic stores), workers distribute across servers.

Architecture Overview:

                          ┌─────────────────┐
                          │  Web Browsers   │
                          └────────┬────────┘
                                   │
                      ┌────────────┴────────────┐
                      │                         │
                  ┌───▼────┐              ┌────▼───┐
                  │ App 1  │              │ App 2  │
                  │(web)   │              │(web)   │
                  └───┬────┘              └────┬───┘
                      │ Publish              │ Publish
                      └────────────┬─────────┘
                                   │
                      ┌────────────▼────────────┐
                      │  Shared Message Queue   │
                      │  (RabbitMQ / Redis)     │
                      └────────────┬────────────┘
                                   │
                    ┌──────────────┼──────────────┐
                    │              │              │
                ┌───▼────┐     ┌──▼────┐     ┌──▼────┐
                │Worker 1 │     │Worker 2│     │Worker 3│
                │(Server 1)     │(Server 2)     │(Server 3)
                └────────┘     └────────┘     └────────┘
                
Key: All workers pull from same queue (no duplication)

Benefits of cluster worker distribution:

  • Messages processed in parallel across multiple servers
  • If Server 1 fails, Workers 2–3 still process messages
  • Load spreads automatically (no manual queue assignment)
  • Scales linearly with server count

Example: 100 messages in queue, 3 servers with 2 workers each

Queue: [msg1, msg2, msg3, ... msg100]

Single-server (5 workers):
├─ Processes 5 messages in parallel
├─ Time to clear: 100 ÷ 5 = 20 seconds

Cluster (3 servers × 2 workers = 6 workers):
├─ Processes 6 messages in parallel
├─ Time to clear: 100 ÷ 6 = 16.7 seconds (17% faster!)

Cluster (3 servers × 5 workers = 15 workers):
├─ Processes 15 messages in parallel
├─ Time to clear: 100 ÷ 15 = 6.7 seconds (3x faster!)

Setup: Shared Message Queue (RabbitMQ)

All servers connect to same RabbitMQ broker:

// env.php (same on all servers)
'queue' => [
  'amqp' => [
    'host' => 'rabbitmq.internal',    // Shared RabbitMQ server
    'port' => 5672,
    'user' => 'magento',
    'password' => 'secure_password',
    'virtualhost' => '/'
  ]
]

Start workers on each server:

# Server 1
for i in {1..5}; do
  php bin/magento queue:consumers:start brocode_image_convert &
done

# Server 2
for i in {1..5}; do
  php bin/magento queue:consumers:start brocode_image_convert &
done

# Server 3
for i in {1..5}; do
  php bin/magento queue:consumers:start brocode_image_convert &
done

# Result: 15 workers pulling from same RabbitMQ queue

Verify cluster distribution:

# On RabbitMQ server, check consumers
rabbitmqctl list_consumers -p /

# Output:
# queue                           consumerTag             ackRequired  prefetchCount
# brocode_image_convert           server1-worker-1        true         10
# brocode_image_convert           server1-worker-2        true         10
# brocode_image_convert           server2-worker-1        true         10
# brocode_image_convert           server2-worker-2        true         10
# brocode_image_convert           server3-worker-1        true         10
# brocode_image_convert           server3-worker-2        true         10
# Total: 6 consumers on same queue (messages distributed!)

Monitor queue depth across cluster:

# Check RabbitMQ queue depth (central metric)
rabbitmqctl list_queues name messages consumers

# Output:
# brocode_image_convert    2145    6
# ├─ Messages: 2145 (backlog)
# ├─ Consumers: 6 (workers processing)
# └─ Processing rate: depends on message complexity

# Calculate time to clear
# 2145 messages ÷ 6 workers ÷ 1 msg/sec = ~360 seconds (6 minutes)

# Alert if backlog grows
if [ "$queue_depth" -gt 5000 ]; then
  send_alert "Queue backlog critical: $queue_depth messages"
  # Action: Scale up workers or add servers
fi

Failover and health checks

Problem: Server dies, its workers stop

Scenario: Server 2 crashes
├─ Workers 2.1–2.5 stop processing
├─ Remaining workers: 1.1–1.5 + 3.1–3.5 = 10 workers
├─ Processing rate drops: 15 workers → 10 workers (33% slower)
├─ Queue backlog grows

Solution: Auto-restart workers via Supervisor

Setup Supervisor on each server:

# /etc/supervisor/conf.d/magento-cluster-workers.conf
[group:magento_workers]
programs=mage_worker_1,mage_worker_2,mage_worker_3,mage_worker_4,mage_worker_5

[program:mage_worker_1]
process_name=%(program_name)s_%(process_num)02d
command=php /var/www/magento/bin/magento queue:consumers:start brocode_image_convert --max-messages=100
autostart=true
autorestart=true        # Restart if worker crashes
numprocs=1
priority=999
redirect_stderr=true
stdout_logfile=/var/log/supervisor/mage_worker_1.log
startsecs=10
stopasgroup=true
user=www-data

Automatic failover flow:

1. Worker crashes on Server 2
2. Supervisor detects exit (within 1 second)
3. Supervisor restarts worker automatically
4. Worker reconnects to RabbitMQ queue
5. Continues processing messages (no messages lost)
6. Queue backlog + restart logged to monitoring

Monitor worker health (cross-cluster):

// Health check observer (runs on all servers)
class ClusterHealthObserver
{
  public function __construct(
    private Connection $connection,  // RabbitMQ connection
    private Alert $alertService
  ) {}

  public function execute(): void
  {
    // Check queue depth
    $queueDepth = $this->getQueueDepth('brocode_image_convert');
    $consumerCount = $this->getConsumerCount('brocode_image_convert');
    
    // Backlog growing?
    if ($queueDepth > 5000) {
      $this->alertService->notify(
        "Queue backlog: $queueDepth messages, $consumerCount workers. " .
        "Scale up workers on cluster."
      );
    }
    
    // Processing rate slow?
    $processingTime = $queueDepth / max($consumerCount, 1);  // seconds to clear
    if ($processingTime > 600) {  // >10 minutes to clear
      $this->alertService->notify(
        "Slow processing: $processingTime seconds to clear queue. " .
        "Add workers or optimize message handlers."
      );
    }
  }
}

Scaling strategy

Determine if you need more workers:

Queue metrics:
├─ Messages per second: 50 (during peak)
├─ Processing time per message: 2 seconds
├─ Processing capacity: 1 worker = 0.5 msg/sec
├─ Workers needed: 50 ÷ 0.5 = 100 workers!
│
├─ Current setup: 3 servers × 5 workers = 15 workers
├─ Deficit: 100 - 15 = 85 workers needed
└─ Action: Scale to 3 servers × 30 workers = 90 workers (≈100 needed)

Scaling options (best to worst):

Option 1: Add more servers to cluster (BEST)
├─ Add 2 more servers (5 servers total)
├─ Each runs 20 workers = 100 workers total
├─ Cost: $$$ (new hardware)
├─ Benefit: 10x capacity, future-proof
└─ Deployment time: 1–2 hours

Option 2: Increase workers per server (MIDDLE)
├─ Scale to 10 workers per server
├─ 3 servers × 10 = 30 workers (but still short of 100)
├─ Cost: $ (more CPU/memory on existing)
├─ Risk: Might overload server CPU
└─ Deployment time: 15 minutes

Option 3: Optimize message handlers (BEST ROI)
├─ Reduce processing time: 2 seconds → 1 second per message
├─ 15 workers × 2 msg/sec = 30 msg/sec (matches 50/sec peak at 50% util)
├─ Cost: $ (engineer time for optimization)
├─ Benefit: Same capacity, lower resource cost
└─ Deployment time: 2–4 weeks

tradeoff: cache everything, debug nothing

Every layer added here buys response time and costs you the ability to reason about the system. That is not an argument against caching — it is an argument for knowing which trade you are making at each step.

Varnish in front of Magento means a change to a template may not appear for a customer until something invalidates it, and "it works on my machine" acquires a second meaning: your machine has no Varnish. Redis for sessions means a Redis outage logs out every customer at once, converting a cache failure into a checkout failure. Aggressive block caching means a stale price is now possible, and the class of bug whose workaround is "flush the cache" starts appearing in your tracker.

The rule that keeps this honest: every cache layer needs a named invalidation story before it is enabled, not after the first stale-content bug. If nobody can say what invalidates a given cache and how long the worst-case staleness window is, that layer is not configured — it is merely switched on.

Two things follow. Keep a documented way to bypass the whole stack for debugging, so an engineer can see uncached truth in one command. And accept the genuine ceiling: past a certain point the remaining response time is real application work, and the answer is profiling rather than another cache. That is where the backend section below picks up.


verification: has TTFB actually moved?

Do this before and after each layer, not once at the end. The point of measuring per-layer is that it tells you when to stop.

# Server-side view, cold and warm, 10 samples
for i in $(seq 1 10); do
  curl -s -o /dev/null -w '%{time_starttransfer}\n' https://your-store.example/
done | sort -n | awk '{a[NR]=$1} END {print "p50:", a[int(NR/2)], " p90:", a[int(NR*0.9)]}'

# Is the full-page cache actually serving this URL?
curl -sI https://your-store.example/ | grep -i 'x-magento-cache-debug\|age\|x-cache'

A HIT on a page you expected to be cacheable is the single most valuable line of output in this whole article. A MISS on a category page means every other optimisation below is being measured against a moving baseline.

The infrastructure checklist

  • [ ] Production mode enabled (MAGE_MODE=production in env.php)
  • [ ] Varnish running: curl -I | grep X-Cache shows HIT
  • [ ] Redis for sessions: redis-cli DBSIZE > 0
  • [ ] PHP 8.3+ (check: php -v)
  • [ ] CDN for static assets (CSS, JS, images) – configure in System > Config > Web > Unsecure
  • [ ] Database on SSD (check: lsblk shows ssd, not rotational disk)
  • [ ] Slow query log enabled (MySQL, check >100ms queries weekly)
  • [ ] Blackfire or XDebug profiling enabled for benchmarking
  • [ ] Heatmap generation script deployed (for ongoing TTFB monitoring)


Related reading

Sources & References

Disclaimer

Timings are from stores of moderate size and are the shape of the curve, not a benchmark for yours. Every cache configuration here changes failure behaviour as well as performance — read the tradeoff section before enabling any of it in production.

Related Topics

← Previous
Next →

Written in collaboration with AI (Claude, by Anthropic). Ideas, verification, and accountability are mine; research and drafting are AI-assisted. Full disclosure → · Found an error? Tell me.