Magento 2 Incident Recovery: 6 Scenarios and a Runbook

lead

The gap between a one-hour incident and a one-day incident is decided in the first five minutes, and almost never by skill. It is decided by whether someone has to think about what to check, or merely read it. That is the whole argument for a runbook: under pressure, at 03:00, with the phone going, recall is the first thing to go and a checklist is the last thing to fail. Magento 2 incident recovery is not a special discipline — it is ordinary incident response applied to a stack with a few specific failure shapes that recur: a crashed InnoDB table, a deploy that half-landed, a payment gateway that went dark, an OpenSearch node that OOM’d, a bot flood, and the one nobody rehearses, a breach. Six scenarios, each with a diagnosis flowchart you can read in under two minutes and recovery steps with real timings attached.

TL;DR — The biggest MTTR lever is not the recovery procedure — it is the monitoring you set up while nothing was wrong. A store with structured logs, a trace ID per request and errors flowing into Sentry recovers in roughly 45 minutes on these six scenarios; the same store without them spends the first two hours establishing what broke. Treat recovery as a runbook someone reads under pressure, not something they reason out: the gap between a one-hour and a one-day incident is decided in the first five minutes.


baseline: the monitoring you need before anything breaks

Reality check: Most shop owners skip runbooks. They respond when incidents happen, then debug afterward. This section is for real shops: live monitoring (Sentry), log storage (Loki), readable dashboards (Grafana), and post-mortem debugging tools.

A. Live Event Monitoring (Sentry)

Why Sentry matters: Know about errors BEFORE customers report them.

Install Sentry:
1. Sign up at sentry.io (free tier covers small shops)
2. Add to Magento: composer require sentry/sentry-laravel
3. Configure: app/etc/env.php

# app/etc/env.php
'sentry' => [
  'dsn' => 'https://key@sentry.io/project-id',
  'environment' => 'production',
  'traces_sample_rate' => 0.1,  // Sample 10% of transactions
],

Real value:

  • Alert: "5 errors in 1 minute" → Slack notification (before phone rings)
  • Error context: Full stack trace + user session info (not "Something went wrong")
  • Performance: Track 95th percentile page load time (regression detection)
  • Release tracking: Correlate errors to deploys ("error started after v2.26.3")

Example Sentry alert in Slack:

🚨 5 errors in the last 1 minute
TypeError: undefined function getPrice()
  File: app/code/Custom/Module/Model/Product.php:145
  Release: v2.26.3
  Affected users: 12
  
→ Check Sentry

B. Log Storage & Querying (Loki)

Why Loki matters: You can’t debug an incident if logs are scattered across 10 servers with 2 days retention.

# Install Loki (on a centralized logging server)
curl https://raw.githubusercontent.com/grafana/loki/main/tools/install.sh | bash

# Configure Magento to send logs to Loki
# Via Filebeat (log shipper):
# /etc/filebeat/filebeat.yml
filebeat.inputs:
- type: log
  enabled: true
  paths:
    - /var/www/html/var/log/system.log
    - /var/www/html/var/log/exception.log
    - /var/log/php-fpm/error.log
    - /var/log/mysql/error.log

output.loki:
  hosts: ["http://loki-server:3100"]
  labels:
    job: magento
    host: production

Real value:

  • Search logs by tag: {job="magento"} | json | level="ERROR" (all errors across all servers)
  • Time correlation: {job="magento"} | json | status="500" between 14:30-14:45 (when incident happened)
  • Retention: Keep 30 days of logs (not 2 days)
  • Performance: Query 100GB of logs in <1 second (Loki is fast)

Example incident debugging with Loki:

Incident: "Site crashed at 14:32"

Search in Grafana (connected to Loki):
{job="magento", host="production"} 
| json 
| level="ERROR" 
| timestamp > "2026-08-21T14:30:00Z" 
| timestamp < "2026-08-21T14:35:00Z"

Result:
14:32:15 ERROR: InnoDB: Corruption detected in table `catalog_product_entity`
14:32:16 ERROR: MySQL has gone away (connection lost)
14:32:17 ERROR: Magento bootstrap failed
14:32:18 ERROR: (100 more...)

→ Root cause found: database corruption at 14:32:15

C. Logging Strategy (Before Incidents)

Most shops have broken logging:

❌ BAD:
- All logs → single file (system.log grows to 5GB)
- No log rotation (disk fills up, crashes happen)
- No log levels (error + info + debug all mixed)
- Plaintext passwords in logs (security risk)

✓ GOOD:
- Separate files per component (MySQL, PHP, Nginx)
- Automatic rotation (daily, keep 30 days)
- Log levels: ERROR (critical), WARN (watch), INFO (normal), DEBUG (dev only)
- Structured logs (JSON format, searchable by Loki)

Configure Magento logging:

// app/etc/env.php
'log' => [
  'monolog' => [
    'handlers' => [
      'system' => [
        'class' => 'Monolog\Handler\RotatingFileHandler',
        'level' => 'ERROR',  // Only errors
        'filename' => 'var/log/system.log',
        'maxFiles' => 30,  // Keep 30 days
      ],
      'exception' => [
        'class' => 'Monolog\Handler\RotatingFileHandler',
        'level' => 'ERROR',
        'filename' => 'var/log/exception.log',
        'maxFiles' => 30,
      ],
      'deprecation' => [
        'class' => 'Monolog\Handler\RotatingFileHandler',
        'level' => 'WARNING',  // Deprecations
        'filename' => 'var/log/deprecation.log',
        'maxFiles' => 7,
      ],
    ],
  ],
],

Log rotation (Nginx + MySQL + PHP):

# /etc/logrotate.d/magento
/var/www/html/var/log/*.log {
  daily
  rotate 30
  compress
  delaycompress
  notifempty
  create 0640 www-data www-data
  sharedscripts
  postrotate
    systemctl reload php-fpm > /dev/null 2>&1 || true
  endscript
}

D. Grafana Dashboards (Real-Time View)

Why: Dashboard replaces 10 SSH windows. See everything at a glance.

Key metrics to track during incidents:

1. Load Average (what it means):

Load average: 5.2, 3.8, 2.1
↑ 3 numbers: 1-min, 5-min, 15-min average

Interpretation:
├─ If CPU count = 4:
│  └─ Load 5.2 = 130% of CPU capacity (overloaded!)
├─ If CPU count = 8:
│  └─ Load 5.2 = 65% of CPU capacity (okay)
└─ Rising load (5.2 → 8.1 → 12.4) = getting worse

2. htop Reading (During Active Incident):

# Command to run during incident:
htop -o %CPU -d 1

# What you see:
PID  USER  PR   NI    VIRT    RES  %CPU  %MEM  TIME+    COMMAND
1234 www   20   0    2.5G   950M  95%   30%   1:23.45  php-fpm
5678 www   20   0    2.3G   900M  92%   28%   1:22.10  php-fpm
9012 mysql 20   0    4.0G   2.5G  45%   80%   0:45.23  mysqld

# Key columns during incident:
├─ %CPU: High (90%+) = single process eating CPU
├─ %MEM: High (80%+) = memory leak or large query result
├─ RES: Resident memory (actual RAM used)
└─ TIME+: Cumulative CPU time (which process ran longest)

# What to do:
├─ If PHP 95%+ → Restart PHP-FPM (but diagnose why first)
├─ If MySQL 80% MEM → Kill long-running query
└─ If load rising → Something started hammering the system

3. Grafana Panel Examples:

Panel 1: Response Time (p95, p99)
├─ Target: histogram_quantile(0.95, http_request_duration)
├─ Alert: If p95 > 500ms → Slack notification
└─ Normal: 50–200ms, Warning: 200–500ms, Critical: >500ms

Panel 2: Error Rate (5-min rolling)
├─ Target: rate(http_requests_total{status="5xx"}[5m])
├─ Alert: If > 1% → Page on-call engineer
└─ Normal: <0.1%, Warning: 0.1–1%, Critical: >1%

Panel 3: Active Database Connections
├─ Target: mysql_global_status_threads_connected
├─ Alert: If > 80% of max_connections → Investigate
└─ Normal: <50, Warning: 50–80%, Critical: >80%

Panel 4: Disk Usage
├─ Target: node_filesystem_avail_bytes
├─ Alert: If < 10% free → Clean up logs/old files
└─ Critical: <5% (MySQL auto-stops, crashes)

E. Post-Mortem Debugging Flow

After incident (even 2 weeks later), you can still debug:

1. Timeline reconstruction (from Loki logs):

# In Grafana Loki:
{job="magento"} | json | timestamp > "2026-08-20T14:30Z" | timestamp < "2026-08-20T15:00Z"
# Sorted timeline of EVERYTHING that happened

2. Correlate with Sentry events:

Loki shows: 14:32:15 MySQL error
Sentry shows: 12 users got 500 error at 14:32:16
Grafana shows: Load spiked from 2.0 to 8.5 at 14:32:14
→ All three point to same moment: root cause is MySQL, not app

3. Blame the right system:

WITHOUT logs: "Everything was slow" (vague, unmeasurable)
WITH logs + Sentry + Grafana:
  "MySQL crashed at 14:32:15 (from Loki),
   causing 12 users to get 500 errors (from Sentry),
   and load to spike to 8.5 (from Grafana).
   Prevention: add MySQL monitoring alert for connection errors."

F. Asking the tooling instead of grepping it (Sentry MCP + Loki MCP)

Sentry and Loki answer different questions, and the reason to run both is that an incident needs both answers within the same minute. Sentry knows what and who: 47 errors, 23 users affected, six checkouts lost, error rate 5/hr before the release and 47/15min after. Loki knows how: the exact ordered sequence of one request, observer to plugin to failure to fallback, with the timing at each hop. Sentry narrows the blast radius; Loki explains one instance of it.

Wiring both behind MCP servers turns the correlation step — historically the slowest part of a post-mortem — into a question you type. The pattern that pays for itself:

ToolDataTypical latencyAnswers
Sentry MCPEvents, releases, user impactsub-second"Which errors spiked after v2.26.3, and who did they hit?"
Loki MCPRaw structured log lines1–5 s"What exact sequence led to this one failure?"
Both, correlated on trace_idFull pictureseconds"Why did checkout fail for user 12345 — and did the other 46 recover?"

The correlation only works if every log line carries a request-scoped trace ID. That mechanism — a single ID stamped on every Monolog line, surfaced in nginx access logs and response headers — is a build in its own right, documented in Magento 2 request tracing and shipped as the log tracing module. Deploy it before you need it. Retrofitting trace IDs during an incident is not a thing you can do.

A worked correlation, the way it actually runs:

Q: "Why did checkout fail for user 12345?"

via Sentry MCP:
  → event for user 12345, trace_id = 5f8a2c3e4b1d6a9f
  → transaction: POST /checkout, HTTP 500, 234 ms

via Loki MCP, same trace_id:
  → ProductObserver → PricePlugin → undefined method → fallback
  → fallback_triggered = true, recovery_action = cached_price

Answer: user 12345 hit the ProductObserver error and recovered on
cached pricing. 47 users hit it; 17 recovered, 6 lost the checkout.
Fix the method signature and 30 errors/day disappear.

Scheduled queries close the loop: a 15-minute Loki sweep for new error signatures catches the class of incident that has no single loud failure — the slow bleed where 2% of checkouts have been failing since Tuesday’s deploy and nobody’s alert threshold noticed.

G. What centralised logging costs, and the PII problem

Two things get skipped in every "just ship logs to Loki" write-up, and both of them will find you.

Cost. Loki bills on ingest volume and retention, not on query. A mid-size Magento store at debug verbosity produces far more than teams expect once every request carries structured context. Rough shape, self-hosted on object storage: 100 GB/month of compressed log data at 30-day retention is a modest bill; the same volume at 90-day retention on a managed offering is several times that, and debug-level logging left on after an investigation is the usual reason a bill triples overnight. Decide retention per stream rather than globally — 90 days for order and payment paths where a chargeback dispute can arrive late, 14 days for the noisy request-level debug stream, and turn debug verbosity off with a config flag rather than a deploy.

PII. Magento logs are full of personal data by default: customer emails in exception context, addresses in quote dumps, full payloads in webhook failures. Shipping those to a central store means the store is now in GDPR scope — it needs the same access controls, region pinning and deletion path as the database. Concretely: mask email, address and payment fields at the Monolog processor level so raw values never leave the app; keep the log store in the same jurisdiction as the primary database; restrict read access to the people who carry the pager rather than to everyone with a Grafana login; and make sure a deletion request reaches the logs, not just customer_entity. A trace ID is the right thing to correlate on precisely because it is not personal data — resolve it to a customer only inside the application, at the moment someone needs to.

working_example: Scenario 1 — database corruption, the hardest recovery

Frequency: 22% of critical incidents

Typical cause: Crashed MySQL update, power failure mid-write, or buggy migration

Detection: "Error establishing database connection" or "database table damaged"

Recovery time: 30 min–4 hours (depending on backup age)

A. Diagnosis Flowchart (First 5 Minutes)

Is the store completely down?
├─ YES → Check MySQL error log (next step)
└─ NO → Check which tables are corrupted (run query below)

Check MySQL error log:
tail -50 /var/log/mysql/error.log

Look for:
├─ "Table is marked as crashed"
├─ "MySQL has gone away"
├─ "Disk full"
└─ "InnoDB: Corruption detected"

Is backup recent (<1 hour old)?
├─ YES → Proceed to restore (Sections B–C)
└─ NO → Try repair first (Section D)

Is this a dev/staging incident?
├─ YES → Accept data loss, restore from oldest backup
└─ NO → Notify leadership immediately (escalate)

B. Immediate Response (Minutes 1–5)

# 1. Stop traffic to prevent further corruption
sudo systemctl stop nginx
sudo systemctl stop php-fpm

# 2. Take snapshot of current state (evidence)
mysqldump --all-databases --single-transaction \
  > /backups/mysql-corrupted-$(date +%s).sql

# 3. Check disk space (common culprit)
df -h

# 4. Check MySQL status
sudo systemctl status mysql
mysql -u root -p -e "SHOW ENGINE INNODB STATUS;"

C. Restore from Backup (Minutes 5–30)

# 1. Locate latest backup
ls -lht /backups/mysql-*.sql | head -5

# 2. Stop MySQL
sudo systemctl stop mysql

# 3. Backup current corrupted database (keep for analysis)
mv /var/lib/mysql /var/lib/mysql.corrupted.$(date +%s)

# 4. Restore backup
mkdir -p /var/lib/mysql
chown mysql:mysql /var/lib/mysql
chmod 750 /var/lib/mysql

# 5. Restore data
mysql -u root -p < /backups/mysql-2026-08-20-23-59-59.sql

# 6. Verify restore
mysql -u root -p -e "SELECT COUNT(*) FROM catalog_product_entity;"

# 7. Start MySQL
sudo systemctl start mysql

# 8. Run consistency check
mysqlcheck --all-databases --repair --optimize

D. Repair Corrupted Tables (If Backup Old)

# Last resort: Try to repair without losing data
mysql -u root -p

> REPAIR TABLE catalog_product_entity;
> REPAIR TABLE sales_order;
> CHECK TABLE catalog_product_entity;

# If REPAIR fails:
> ALTER TABLE catalog_product_entity ENGINE=InnoDB;

E. Post-Recovery Verification

# 1. Check data integrity
SELECT COUNT(*) FROM catalog_product_entity;
SELECT COUNT(*) FROM sales_order;
SELECT COUNT(*) FROM customer_entity;

# 2. Test store functionality
curl -I http://localhost/index.php
bin/magento cache:flush

# 3. Verify backups exist and are recent
ls -lh /backups/mysql-*.sql | head -3

# 4. Check MySQL logs for errors
grep -i error /var/log/mysql/error.log | tail -20

F. Root Cause Analysis

Questions to answer:

1. When did corruption start?

  • Check MySQL error log timestamps
  • Compare with backup timestamps

2. Why wasn’t it detected earlier?

  • Was monitoring running? (check Grafana/Prometheus)
  • Were integrity checks scheduled? (mysqlcheck cron job)

3. Was this a one-time fluke or recurring?

  • Check if hardware (disk, RAM) is failing
  • Check if MySQL version has known bugs

4. Prevention for next time:

  • [ ] Implement daily integrity checks (cron job)
  • [ ] Automate backup to separate storage (S3)
  • [ ] Set up MySQL monitoring alerts
  • [ ] Test backup restore quarterly

Scenario 2: Failed Deployment (Code Won’t Load)

Frequency: 18% of critical incidents

Typical cause: Syntax error, broke existing functionality, or missing dependency

Detection: HTTP 500, "Parse error" in logs, or specific feature broken

Recovery time: 5–30 minutes

A. Diagnosis Flowchart (First 2 Minutes)

Is the whole site down?
├─ YES → Check PHP error log (production must have logging on!)
└─ NO → Which page is broken? (narrows scope)

Check PHP error log:
tail -30 /var/log/php-fpm/error.log
tail -30 /var/log/apache2/error.log

Look for:
├─ "Parse error"
├─ "Call to undefined function"
├─ "Class not found"
├─ "Out of memory"
└─ "Maximum execution time exceeded"

Is the error in custom code?
├─ YES → Revert last commit (Section B)
└─ NO → Check Magento logs (var/log/system.log)

Did deployment complete or fail mid-way?
├─ Completed → Revert (code issue)
└─ Failed → Re-run setup (missing files)

B. Quick Rollback (Minutes 1–5)

# 1. Identify last known good commit
git log --oneline | head -10

# 2. Revert to previous version
git revert --no-edit HEAD
git push origin main

# 3. Clear caches (Magento often caches old code paths)
bin/magento cache:flush
rm -rf var/view_preprocessed/*
rm -rf var/generation/*

# 4. Restart PHP-FPM
sudo systemctl restart php-fpm

# 5. Verify (access live site)
curl -I https://production.brocode.at/

C. Full Rollback (If Quick Revert Fails)

# Check what version was running before
git tag -l | grep deployment | tail -5

# Checkout previous stable commit
git checkout v2.26.3  # or specific commit hash

# Rebuild and redeploy
bin/magento setup:di:compile
bin/magento setup:static-content:deploy
bin/magento cache:flush

# Restart
sudo systemctl restart php-fpm

D. Root Cause Analysis

What went wrong?

# 1. Check deployment logs
tail -100 /var/log/deployment.log

# 2. Review code diff
git diff v2.26.2..v2.26.3 -- app/code/Custom/

# 3. Run syntax check on changed files
php -l app/code/Custom/Module/Model/Something.php

# 4. Check for missing dependencies
composer show | grep "name-of-new-package"

Prevention:

  • [ ] Test deployments on staging first (obvious, but missed often)
  • [ ] Run PHP syntax check in CI before merge
  • [ ] Run Magento’s static analysis (PHPStan, Psalm)
  • [ ] Automated smoke tests (page loads, checkout works)
  • [ ] Automated rollback if health checks fail

Scenario 3: Payment Processor Down

Frequency: 15% of incidents

Typical cause: Third-party API outage, rate limits hit, or credentials expired

Detection: "Payment failed" errors, no transactions processing

Recovery time: 5–60 minutes (depends on provider)

A. Diagnosis (2 Minutes)

# 1. Check payment processor status page
# Example: Stripe status (https://status.stripe.com)
# Example: PayPal status (https://www.paypalstatus.com)
curl -s https://status.stripe.com/api/v2/status.json | jq .

# 2. Test payment connectivity
curl -v -H "Authorization: Bearer sk_test_xxx" \
  https://api.stripe.com/v1/charges

# 3. Check Magento logs
tail -50 var/log/system.log | grep -i stripe

# 4. Check rate limits
# Stripe error "429 Too Many Requests" → hit rate limit
# PayPal error "connection timeout" → network issue

B. Immediate Response

If payment processor is down:

# 1. Check processor status page
# 2. Post message to status page (check for ETA)
# 3. Enable maintenance mode (optional, based on strategy)
# 4. Notify customers (see Communication Templates below)

If credentials expired:

# 1. Rotate credentials (generate new API key)
# 2. Update in System > Configuration > Payment > Stripe (or PayPal)
# 3. Clear cache
bin/magento cache:clean
# 4. Test payment (use test card)

If rate limit hit:

# 1. Check current request volume
# 2. Implement request queuing (defer non-critical calls)
# 3. Contact provider to request limit increase

C. Fallback Strategy (If Down >30 Min)

// Enable offline payment method temporarily
System > Configuration > Sales > Payment Methods
├─ Enable "Bank Transfer"
├─ Enable "Check/Money Order"
└─ Add note: "Stripe temporarily unavailable. We'll process payment when online."

// Then implement email follow-up workflow

D. Root Cause Analysis

  • [ ] Was this a provider outage? (check their status page history)
  • [ ] Did we hit rate limits? (check API logs)
  • [ ] Were credentials rotated recently? (check deployment logs)
  • [ ] Is monitoring in place? (Grafana/Prometheus for API health)

Scenario 4: search engine down (OpenSearch/Elasticsearch)

On Magento 2.4.8 and later this is always OpenSearch — Adobe removed the Elasticsearch engine options entirely, so a store that still points at an Elasticsearch cluster has a permanently broken catalog search rather than an incident. On 2.4.7 and earlier both engines are possible and the recovery steps below apply to either. Tuning either one is covered in Magento Elasticsearch tuning with real benchmarks.

Frequency: 12% of incidents

Typical cause: Memory exhausted, disk full, or JVM crash

Detection: Search queries time out, "no servers available" error

Recovery time: 5–30 minutes

A. Quick Diagnosis (1 Minute)

# 1. Check Elasticsearch status
curl -s http://localhost:9200/_cluster/health | jq .

# Expected response:
# {
#   "status": "green",
#   "number_of_nodes": 3,
#   "active_shards": 15
# }

# 2. If error, check ES logs
tail -50 /var/log/elasticsearch/elasticsearch.log

# 3. Check disk space
df -h

# 4. Check memory
free -h

B. Recovery

If ES process crashed:

sudo systemctl restart elasticsearch
# Wait 30 sec for startup
sleep 30
curl -s http://localhost:9200/_cluster/health

If disk full:

# Delete old indices
curl -X DELETE "localhost:9200/magento_product_2023-*"

# Or clean up logs/data
rm /var/lib/elasticsearch/nodes/0/indices/old_indices/*
sudo systemctl restart elasticsearch

If JVM out of memory:

# Increase heap size in elasticsearch.yml
-Xms2g  # min heap
-Xmx2g  # max heap

# Then restart
sudo systemctl restart elasticsearch

C. Reindex (If Data Lost)

# This can take 30 min–2 hours depending on catalog size
bin/magento indexer:reindex catalogsearch_fulltext

# Or full reindex
bin/magento indexer:reindex

Scenario 5: DDoS Attack (Site Unreachable)

Frequency: 8% of incidents

Typical cause: Volumetric attack (millions of requests), application layer attack

Detection: Sudden traffic spike, legitimate requests timing out, high CPU/network usage

Recovery time: 30 min–several hours

A. Quick Detection (1 Minute)

# Check traffic logs
tail -100 /var/log/nginx/access.log | awk '{print $1}' | sort | uniq -c | sort -rn | head -20

# If you see one IP making 1000+ requests/sec → likely DDoS

# Check server load
top -b -n1 | head -20
uptime

# Check network traffic
nethogs -n

B. Mitigation (First 10 Minutes)

Option 1: Block at firewall (fastest)

# Block suspicious IPs
iptables -I INPUT -s 203.0.113.5 -j DROP
iptables -I INPUT -s 198.51.100.0/24 -j DROP

# Save rules
iptables-save > /etc/iptables/rules.v4

Option 2: Enable CloudFlare (or similar)

1. Point DNS to CloudFlare
2. Enable "Under Attack" mode (aggressive challenge)
3. Wait 5 min for DNS propagation

Option 3: Rate-limit at Nginx

limit_req_zone $binary_remote_addr zone=general:10m rate=10r/s;
limit_req_zone $binary_remote_addr zone=checkout:10m rate=2r/s;

server {
  location / {
    limit_req zone=general burst=20 nodelay;
  }
  
  location /checkout/ {
    limit_req zone=checkout burst=5 nodelay;
  }
}

C. Long-Term Defense

  • [ ] Implement DDoS protection (CloudFlare, AWS Shield, Akamai)
  • [ ] Set up rate-limiting in nginx/Apache
  • [ ] Monitor for traffic spikes (Grafana alert)
  • [ ] Have load balancer ready (can auto-scale)

Scenario 6: "We’ve Been Hacked" (Security Breach Detection & Response)

Frequency: 5-8% of critical incidents

Typical cause: Compromised admin account, leaked API key, vulnerable extension, SQL injection

Detection: Unusual activity (unauthorized orders, data exports, file modifications)

Recovery time: 2-8 hours (depends on scope of compromise)

A. Breach Detection (First 15 Minutes)

Automated detection via Sansec eComscan:

eComscan scans the store’s files and database for malware (web shells, backdoors, payment skimmers), known-vulnerable extensions and unauthorized admin accounts, and alerts the team, for example in Slack. What it gives you in the first fifteen minutes:

Typical findings that start this scenario:
├─ Malware in a file: a PHP shell under pub/media/
├─ A skimmer injected into a CMS block or config value in the database
├─ An admin account nobody on the team created
└─ An installed extension version with a known vulnerability

What it does not give you is request-level forensics — which IP did what, when, with which payload. That comes from your own web server and application logs (section C). The broader detection-versus-prevention setup is Layer 9 of Magento 2 security hardening.

B. Immediate Containment (Minutes 15-30)

Step 1: Isolate the store

# BEFORE: Store still accessible (attacker might still be inside)
# AFTER: Store locked down (attacker locked out)

# Option A: Kill web access (nuclear, but safest)
sudo systemctl stop nginx
sudo systemctl stop php-fpm
# Result: 503 Service Unavailable for all users

# Option B: Block known attacker IP (surgical, if identified)
iptables -I INPUT -s 203.0.113.99 -j DROP

# Option C: Maintenance mode (customers see message)
bin/magento maintenance:enable
# Result: "We're performing maintenance" (better UX)

Step 2: Revoke compromised credentials

Lock everyone out except the person running the incident, end every live admin session, and invalidate every API token. Run it against the production database after the snapshot, with your own user_id in place of 1. Verified on Magento 2.4.9:

-- 1. Deactivate every admin account except the incident lead's (no new sign-ins)
UPDATE admin_user SET is_active = 0 WHERE user_id <> 1;

-- 2. End live admin sessions: status 3 = logged out manually.
--    Magento destroys the session on its next request.
UPDATE admin_user_session SET status = 3 WHERE status = 1;

-- 3. Invalidate every admin API token (JWT) issued until now.
--    user_type_id 2 = admin; use 3 for customer tokens.
INSERT INTO jwt_auth_revoked (user_type_id, user_id, revoke_before)
SELECT 2, user_id, UNIX_TIMESTAMP() FROM admin_user
ON DUPLICATE KEY UPDATE revoke_before = VALUES(revoke_before);

-- 4. Revoke integration (OAuth) access tokens
UPDATE oauth_token SET revoked = 1 WHERE type = 'access';

Do not overwrite password hashes by hand. Magento stores a salted, versioned hash, not an MD5, so a hand-written value matches no password anyone knows and only adds noise to the forensic record. Deactivating the account does the job, and passwords get reset properly in step D.3.

Step 3: Disconnect payment processor

In Adminhtml, go to Stores → Configuration → Sales → Payment Methods, disable each method whose credentials could have been read or replaced, and save. If the attacker swapped in their own merchant keys, every order placed now pays them. Being down for 30 minutes costs far less than that.

C. Forensic Investigation (Minutes 30-120)

Trace the attacker in your own logs. No security product reconstructs the attack for you; the evidence is in the web server access log and, if you run it, Loki. Start from what detection found — an admin account, a file, an IP — and walk outwards:

# Who called the admin user-save action, and from where
grep -E 'POST /admin_[^ ]*/admin/user/save' /var/log/nginx/access.log

# Everything that address did, in order
grep '^203.0.113.99 ' /var/log/nginx/access.log | less

# When a suspicious file appeared (mtime can be forged; compare with deploy records)
stat pub/media/catalog/shell.php

Snapshot the logs before rotation or cleanup overwrites them.

Investigation Checklist:

# 1. Find compromise date (when did attack start?)
# From Loki or the web server access log:
{job="magento"} | json | admin_action="user_save" | timestamp > "2026-08-20"
→ Ghost admin "backup_user" created at 2026-08-21 14:32:15

# 2. Check what was stolen
# From the database:
SELECT COUNT(*) FROM sales_order WHERE updated_at > "2026-08-21 14:32:00";
→ 300 orders accessed in 45 minutes

# 3. Determine breach type
├─ Compromised admin account? (check login logs)
├─ Leaked API key? (check integration configs)
├─ Vulnerable extension? (check the eComscan vulnerability findings)
└─ SQL injection? (check web server logs for suspicious URLs)

# 4. Notify payment processor
# "Possible compromise: 300 customer records accessed.
#  Please re-issue cards for orders placed after 2026-08-21 14:32.
#  We've disconnected our account. Will re-enable after security audit."

D. Remediation (Hours 2-4)

1. Clean the installation

# Sansec's Advanced and Enterprise plans include analysis and cleanup.
# Either way, review suspicious files manually:

# Find newly added admin accounts
mysql> SELECT * FROM admin_user WHERE created_at > "2026-08-21 14:32:00";
→ "backup_user" account (DELETE IT)

# Find modified core files (likely backdoor)
find app/code -name "*.php" -newermt "2026-08-21 14:00:00" -type f | head -20

# Remove suspected shells
rm -i app/code/Custom/Module/Shell.php
rm -i media/catalog/Shell.php

2. Patch the vulnerability

# From the eComscan findings: known vulnerability in extension Custom_DataExport
# Option A: Update extension
composer update custom/data-export

# Option B: Remove extension (if no longer needed)
bin/magento module:disable Custom_DataExport
rm -rf app/code/Custom/DataExport

# Option C: Patch vulnerability
# Apply the vendor's patch or a manual code fix

3. Reset credentials

  • Admin users: re-enable accounts one at a time, only after confirming each one is a real person who still needs access. Each user sets a new password through the admin "Forgot your password?" flow, or an administrator sets it under System → Permissions → All Users. Delete any account nobody can vouch for.
  • Integrations: under System → Extensions → Integrations, reauthorize each integration to issue new tokens, and rotate the credentials on the other side.
  • app/etc/env.php secrets: change the database password and every third-party key stored there, and rotate payment provider keys in the provider’s dashboard.
  • The encryption key: rotate it with bin/magento encryption:key:change, never by editing env.php. Editing the key by hand leaves every encrypted config value unreadable. The traps are in the key rotation article.
bin/magento encryption:key:change
bin/magento cache:flush

E. Verification & Re-enablement (Hours 4-6)

1. Put a compensating control in place, if the root cause is a known exploit

If the entry point was a published Magento or extension vulnerability and the patch is not deployed yet, Sansec Shield (a Composer module, sansec/magento2-module-shield, in Sansec’s Advanced plan and up) blocks known exploit patterns for it inside Magento until you patch. It is not rate limiting, brute-force protection or a general firewall, and it does not replace the patch in step D.2.

2. Re-enable services

# Only after a clean eComscan run and the manual review above

sudo systemctl start php-fpm
sudo systemctl start nginx
bin/magento maintenance:disable

# Monitor access logs
tail -f /var/log/nginx/access.log

# Keep eComscan in monitoring mode: a new finding alerts the team

3. Notify customers

Subject: Security Incident Update

We detected and contained a security issue affecting your account.
Our investigation shows:
├─ Your order data was accessed (email, shipping address)
├─ No credit card numbers were stolen (cards not stored)
├─ Store is now secured and monitored 24/7
└─ You don't need to take action

We've:
├─ Patched the vulnerability
├─ Reviewed all accounts for unauthorized access
├─ Implemented continuous malware and file monitoring
└─ Disabled unauthorized admin accounts

F. Post-Breach Hardening (Long-term)

1. Close the window between disclosure and patch

The breach started with an attack path somebody knew about before you did. Two things shorten that window next time: a patch SLA with a named owner (the hardening article’s patching layer), and — if your patch cycle is slower than a few days — Sansec Shield, which blocks known Magento exploit patterns inside the application until the fix is deployed. PolyShell (CVE-2026-48356) is the reference case: attacks from 19 March 2026, backports for older release lines on 14 July.

2. Enable 2FA (from Layer 1)

3. Deploy Log Tracer (from Part 2)

With Log Tracer + eComscan:
├─ Every admin action traced (who did what, when)
├─ Every API call logged (integration access)
├─ Unexpected file changes flagged (eComscan)
└─ Breach timeline reconstructible in minutes (not days)

4. Add monitoring alerts

Grafana alert: If admin_user count changes
├─ Threshold: INSERT/DELETE from admin_user table
├─ Action: Immediate Slack alert + lock admin panel
└─ Prevents ghost admin creation

Grafana alert: If database exports spike
├─ Threshold: SELECT COUNT(*) from sales_order > 10k in 1 hour
├─ Action: Kill connection + alert team
└─ Prevents massive data theft

Breach detection and response tooling

ToolWhat it doesWhat it does not doSansec plan
eComscanScans files and database for malware, vulnerable extensions and unauthorized admin accounts; alertsRequest-level forensics; blockingSecure (entry) and up
ShieldComposer module inside Magento; blocks known Magento exploit patterns with rules Sansec updates as attacks appearRate limiting, brute-force or bot protection; replacing patchesAdvanced and up

Sansec sells its plans by annual store-revenue tier; check sansec.io/pricing for current figures. PCI DSS asks for malware protection and change detection, not for a particular vendor — a scanner report is evidence in an assessment, not a requirement on its own.

When to use each:

ScenarioeComscanShieldWhy
Suspected breach✓✓eComscan finds what is on the server; Shield covers known exploit paths while you clean up
PCI DSS assessment✓optionalEvidence for the malware and change-detection requirements
Stores taking card data on-site✓✓Detect plus prevent
Patch cycle slower than a few days✓✓Shield exists for the gap between disclosure and deploy
Budget tight✓—Detection first

Communication Templates

Template 1: Incident Declared (To Internal Team)

SUBJECT: 🚨 Production Incident: [Service] Down

TO: Engineering, DevOps, Leadership

Incident started: 2026-08-21 14:32 UTC
Severity: CRITICAL
Current status: INVESTIGATING

What's happening:
- [Brief description: e.g., "Stripe API unreachable, payments failing"]
- Impact: ~€500/min revenue loss
- Estimated users affected: ~50% (those checking out)

Actions taken:
- [ ] Incident commander assigned (John Doe)
- [ ] Customers notified via email/Slack
- [ ] Fallback payment method enabled
- [ ] Payment team contacted (Stripe support ticket #123)

Next update: 14:45 UTC

Template 2: Status Update (To Customers)

We're experiencing issues with our payment processor. 

Status: We're working with our payment provider to restore service.
Timeline: We expect to be back online within 30 minutes.
What you can do: Try again in 10 minutes, or contact support.

We'll post updates every 15 minutes. Follow us @brocode_at for live updates.

Thank you for your patience.

Template 3: Post-Incident Summary (To Team + Stakeholders)

INCIDENT SUMMARY: Elasticsearch Crash (2026-08-20, 14:32–14:58 UTC)

Duration: 26 minutes
Severity: HIGH
Root cause: Disk full (90%), JVM garbage collection paused

Impact:
- Search was unavailable for 26 min
- ~200 customers complained (support ticket surge)
- ~€5k estimated revenue loss (lost conversions)

Timeline:
14:32 - Search stops working
14:35 - Team notified
14:40 - Root cause identified (disk full)
14:45 - Disk cleaned, ES restarted
14:50 - Search back online
14:58 - Full reindex complete

Prevention for next time:
- [x] Set up disk full alerting (Grafana)
- [x] Increase ES heap memory (2GB → 4GB)
- [x] Implement automatic cleanup of old indices
- [x] Add monitoring for JVM garbage collection pauses

Post-Incident Playbook (24 Hours After)

1. Post-Mortem (Team Meeting, 30–60 Min)

Questions to answer:

  • What was the root cause?
  • What early warning signs were missed?
  • What did we do well? (what slowed us down?)
  • How do we prevent this in the future?

Output: 1-page post-mortem document (timeline, lessons learned, action items)

2. Action Items (Owner + Due Date)

☐ Set up monitoring alert (e.g., "Elasticsearch CPU > 80%" → Slack)
  Owner: DevOps, Due: 48 hours

☐ Automate disk cleanup (old logs/indices)
  Owner: Backend, Due: 1 week

☐ Test backup restore procedure (quarterly)
  Owner: DevOps, Due: 1 month

☐ Update runbook with this scenario
  Owner: Tech Lead, Due: 48 hours

☐ Incident response training for new team members
  Owner: HR/Tech Lead, Due: ongoing

3. Metrics to Track

Mean Time To Detection (MTTD): 3 minutes
Mean Time To Recovery (MTTR): 26 minutes
Customer notification delay: 5 minutes

Target improvements:
- MTTD: 1 minute (better alerting)
- MTTR: <10 minutes (better runbooks)
- Notification: <2 minutes (automated)

Preventive Postmortems (Regular Checkups)

Reality: Most shops only do postmortems AFTER incidents. Better shops do them regularly, even without incidents.

Quarterly Incident Simulation (Tabletop Exercises)

Why: You find gaps before the real incident happens.

Example quarterly scenario (30 min simulation):

SCENARIO: Database crashed at 2 PM on a Tuesday (peak traffic)

Setup:
- Gather: DevOps lead, Backend lead, On-call engineer, Manager
- Timeline: 30 min (compressed from real 45 min)
- No actual code changes, just communication

Simulation timeline:
T+0   Min "Database crashed" → Who do we notify first?
T+5   Min "Error logs show InnoDB corruption" → What's the diagnosis?
T+10  Min "Is backup recent?" → Check backup process (did it run last night?)
T+15  Min "Start restore" → Any blockers?
T+25  Min "Verify data" → How do we confirm success?
T+30  Min "Notify customers" → What do we say?

Debrief:
├─ What went smoothly?
├─ What surprised us?
├─ What were we missing? (process, tools, knowledge)
└─ Action items for next quarter

Track over time:

  • Q1 2026: Discovered missing backup verification step → Fixed
  • Q2 2026: Realized on-call engineer didn’t know Loki → Trained
  • Q3 2026: Found we didn’t have customer notification template → Created
  • Q4 2026: Simulation ran smooth (lessons from Q1–Q3 are working)

Scheduled Monthly Reviews (Without Incident)

What if month has no incidents?

Still do reviews (learn from near-misses + prevent incidents):

MONTHLY REVIEW CHECKLIST:

Monitoring health:
├─ [ ] Sentry error rate normal? (target: <0.5%)
├─ [ ] Any alerts missed? (false negatives)
├─ [ ] Alerting fatigue? (too many false positives)
└─ [ ] New error types appearing? (research them)

Backup health:
├─ [ ] Last backup successful? (check logs)
├─ [ ] Backup size reasonable? (growing unexpectedly?)
├─ [ ] Can we restore in <30 min? (test monthly)
└─ [ ] Off-site backup exists? (not just local disk)

Logging health:
├─ [ ] Disk usage for logs? (growing too fast?)
├─ [ ] Log retention policy working? (old files deleted?)
├─ [ ] Can we find errors from 2 weeks ago in Loki? (test query)
└─ [ ] Any sensitive data in logs? (passwords, tokens)

Team readiness:
├─ [ ] On-call engineer familiar with runbooks? (quiz them)
├─ [ ] Do we have latest contact list? (phone numbers current?)
├─ [ ] New team members trained? (onboarding complete)
└─ [ ] Runbooks up-to-date? (changed since last incident?)

Output: If all checks pass, file report. If any fail, create action item.


tradeoff: what a runbook costs against what it saves

Track these across all incidents to identify patterns:

MetricCurrentTargetWhy It Matters
MTTD (Mean Time To Detect)8 min2 minHow fast we know something’s wrong
MTTR (Mean Time To Recover)45 min15 minHow fast we fix it
Time to notify customers10 min2 minHow transparent we are
First response time5 min2 minHow fast team reacts
Root cause identified2 hours30 minHow thorough we are

Track over time (monthly): Are MTTD/MTTR improving? If not, why?


verification: the incident response toolkit

Before Any Incident

  • [ ] Runbooks written + team trained (each scenario)
  • [ ] Backup restore tested quarterly
  • [ ] Monitoring + alerting in place (CPU, disk, memory, API latency)
  • [ ] On-call rotation defined (who responds first?)
  • [ ] Communication plan (customer notification flow)
  • [ ] Escalation path (when to wake leadership?)

During Incident (First 5 Min)

  • [ ] Declare incident ("CRITICAL" / "HIGH" / "MEDIUM")
  • [ ] Assign incident commander (one person, clear authority)
  • [ ] Assess impact (customers affected? revenue lost? data at risk?)
  • [ ] Notify internal team (Slack, email)
  • [ ] Start incident clock (timestamp everything)

During Recovery (5–60 Min)

  • [ ] Follow runbook for scenario (don’t improvise)
  • [ ] Document every action (paste commands into ticket)
  • [ ] Post status updates (every 15 min to team + customers)
  • [ ] Test recovery (verify service is working before claiming "resolved")

After Recovery (1–24 Hours)

  • [ ] Communicate all-clear (to customers, to team)
  • [ ] Preserve evidence (logs, database snapshots, error messages)
  • [ ] Schedule post-mortem (within 48 hours while fresh)
  • [ ] Draft post-incident summary (root cause, lessons learned)
  • [ ] Assign action items (with owners + deadlines)

Expected Timeline

1. Week 1: Create runbooks, test with team (dry run)

2. Week 2: Implement monitoring (Sentry + Loki + Grafana)

3. Week 3: Implement logging strategy (rotation, levels, structured)

4. Week 4: Incident response training for team

5. Ongoing: Simulate incidents quarterly, review monthly

Cost of not having this: €1k–50k per incident. RoI = second incident prevented.

Cost to implement: €200–500/month (Sentry + Loki + Grafana) + 40 hours setup = pays for itself on first incident prevented.



Related reading

Sources & References

Disclaimer

Timings in this article come from a small sample of real incidents on stores of moderate size, not from a controlled study — treat them as the shape of the curve, not as a benchmark. Recovery commands assume you have tested your restore path at least once outside an incident; the first restore should never be the one that matters. Where a step is destructive, it is marked, and the snapshot-first step before it is not optional.

← Previous

Written in collaboration with AI (Claude, by Anthropic). Ideas, verification, and accountability are mine; research and drafting are AI-assisted. Full disclosure → · Found an error? Tell me.