During incidents, preserve evidence, contain harm, change one thing, and verify both the SLO and the database mechanism.

Incident Runbooks for Replication Lag, Disk Pressure, Slow Queries, Elections, and Shard Imbalance

Build evidence-first incident runbooks for replication lag, disk pressure, slow queries, elections, and shard imbalance with safe containment and verification.

Advanced120–220 minutesIncident-runbook labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Use a common incident structure for replication lag, disk pressure, slow queries, elections, and shard imbalance.

02

Build evidence timelines before changing configuration and classify immediate containment separately from root-cause repair.

03

Practice safe failure injection without host firewall, system-clock, or destructive production changes.

04

Define stop/rollback conditions and verify recovery against both database and application SLO evidence.

05

Connect incident learning back to capacity, alerts, benchmarks, backup/restore, and upgrade controls.

Reproducible lab baseline

This chapter pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where client behavior or load generation matters. Lesson 5 uses one disposable MongoDB instance for slow-query/disk-capacity observation and deterministic simulators for election, replication-lag, and shard-imbalance timelines. Chapters 14–17 contain the full replica/sharded mechanisms. Host exposure is loopback-only on 127.0.0.1:27208. Authentication and TLS are disabled only for disposable local labs; production security remains the Chapter 22 prerequisite. Default read/write concern and primary read preference are used unless a step states otherwise. FCV is observed and never changed in mandatory labs. Atlas, Enterprise Advanced, Search, Vector Search, and KMS are optional unless explicitly labeled. No host firewall rules, real clock skew, forced replica-set reconfiguration, artificial production disk filling, or live shard corruption is used. Failure exercises are simulations or bounded disposable-query changes. Runtime performance, failover, capacity, and upgrade labs were not executed in the generation environment, so metric values, latency distributions, queue depths, oplog windows, replication lag, incident times, and upgrade durations must be measured locally rather than copied as invented output.

1. A runbook is an evidence protocol

During an incident, the worst first move is often “change something that usually helps.” A useful runbook starts by preserving an incident timeline: detection time, user impact, release/configuration events, topology state, key metrics, active operations, recent elections/migrations, disk/cache evidence, and application error distribution. Then it distinguishes containment (reduce immediate harm) from root-cause correction.

Phase Question Output
Detect Which SLO/user invariant is failing? Impact and start time
Scope Which tenants/nodes/shards/query shapes? Affected surface
Preserve What evidence will disappear if we restart/change state? Timeline/metrics/logs/profile samples
Contain What reversible action reduces harm? Traffic/rate/query/maintenance change
Diagnose Which mechanism explains the correlated evidence? Testable hypothesis
Repair What smallest change addresses the mechanism? One controlled intervention
Verify Did user SLO + server evidence recover? Post-change acceptance
Learn What capacity/alert/test/runbook gap allowed this? Preventive action

2. Slow-query runbook: shape before server-wide tuning

start the disposable incident node
docker rm -f atlasmart-ch26-l5 2>/dev/null || truedocker volume rm atlasmart-ch26-l5-db 2>/dev/null || truedocker run -d --name atlasmart-ch26-l5 -p 127.0.0.1:27208:27017 -v atlasmart-ch26-l5-db:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27208/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; done
create, diagnose, and repair one inefficient shape
const d=db.getSiblingDB("atlasmart");d.products_ch26_l5.drop();for (let b=0;b<30;b++) { const x=[]; for (let i=0;i<1000;i++){const n=b*1000+i;x.push({sku:`SKU-${n}`,tenantId:`tenant-${n%20}`,category:`cat-${n%40}`,priceCents:1000+n%100000,description:`portable atlasmart product ${n} `+"x".repeat(n%100)});} d.products_ch26_l5.insertMany(x); }let before=d.products_ch26_l5.find({tenantId:"tenant-7",category:"cat-7"}).sort({priceCents:1}).explain("executionStats");printjson({before:{nReturned:before.executionStats.nReturned,totalDocsExamined:before.executionStats.totalDocsExamined,totalKeysExamined:before.executionStats.totalKeysExamined,plan:before.queryPlanner.winningPlan}});d.products_ch26_l5.createIndex({tenantId:1,category:1,priceCents:1},{name:"idx_tenant_category_price"});let after=d.products_ch26_l5.find({tenantId:"tenant-7",category:"cat-7"}).sort({priceCents:1}).explain("executionStats");printjson({after:{nReturned:after.executionStats.nReturned,totalDocsExamined:after.executionStats.totalDocsExamined,totalKeysExamined:after.executionStats.totalKeysExamined,plan:after.queryPlanner.winningPlan}});

The repair is evidence-driven: the query contract stayed constant and one index changed. Do not generalize this exact index to another workload without checking field cardinality, sort/range order, write cost, and competing query shapes.

3. Replication lag and elections: keep availability and durability evidence separate

simulate a replication incident timeline and runbook decision
events=[ {"t":0,"lag_s":1,"oplog_window_h":24,"primary":True,"p99_ms":35}, {"t":5,"lag_s":18,"oplog_window_h":9,"primary":True,"p99_ms":62}, {"t":10,"lag_s":75,"oplog_window_h":3,"primary":True,"p99_ms":105}, {"t":15,"lag_s":140,"oplog_window_h":1.2,"primary":True,"p99_ms":170},]for e in events:    risk=e["lag_s"]>60 and e["oplog_window_h"]<4    print({**e,"secondary_resync_risk":risk})print("containment candidates: reduce nonessential write load; protect oplog window; investigate secondary disk/cache/network; do not force reconfig")

For elections, record which member became primary, election frequency, terms, heartbeat/network evidence, client error/retry behavior, and whether acknowledged writes meet the requested concern. Repeated elections are a symptom; increasing election timeout or forcing configuration without diagnosing connectivity/resource causes can trade one failure mode for another.

4. Disk pressure: project time-to-exhaustion before emergency cleanup

calculate disk exhaustion and intervention window
from datetime import timedeltafree_gib=85growth_gib_per_hour=3.5maintenance_lead_hours=12hours_to_full=free_gib/growth_gib_per_hourprint({"hours_to_full":round(hours_to_full,1),"maintenance_lead_hours":maintenance_lead_hours,"margin_hours":round(hours_to_full-maintenance_lead_hours,1)})

Containment may include pausing nonessential ingestion, moving exports off the database filesystem, reducing temporary maintenance work, or increasing storage through the supported platform mechanism. Never “fix disk pressure” by deleting unknown files inside dbPath. Preserve journal/data-file integrity and validate backup/recovery status before destructive cleanup.

5. Shard imbalance: distinguish data ownership from hot workload

A cluster can be byte-balanced and still hot if one shard receives a disproportionate query/write rate, and it can be range-imbalanced during an expected migration/zone policy. Collect sh.status(), range ownership, balancer state, per-shard operations/CPU/disk, query targeting, shard-key distribution, and recent migration/resharding events. Chapters 16–17 provide the safe local mechanisms.

Symptom Possible mechanism Evidence before action
One shard high CPU, equal bytes hot shard-key values or targeted workload skew per-shard ops + key/query distribution
Unequal bytes, balancer active migration convergence range metadata + balancer/migration state
Scatter/gather latency queries missing shard-key targeting mongos explain + query shapes
Migration churn poor key/zone design or insufficient headroom zone/range metadata + move history + disk/network
Jumbo/large ranges indivisible/skewed key interval range statistics + shard-key cardinality/frequency

6. Incident acceptance criteria and rollback

Every runbook needs a stop condition. For example: if a query/index change does not reduce documents examined and p99 under the same workload, roll it back; if a secondary cannot make positive catch-up progress before the oplog window closes, escalate toward resync/capacity protection; if an upgrade member fails to return healthy, stop the rollout. Verification must include the user-facing SLO and the mechanism-level metric that motivated the change.

small incident record schema
incident={ "impact":"checkout p99 above SLO", "start":"2026-09-03T10:00:00Z", "scope":["tenant-hot-1"], "evidence":["client-p99","serverStatus-delta","currentOp","explain","logs"], "hypothesis":"new query shape performs collection scan", "containment":"route/report endpoint disabled", "change":"add reviewed compound index", "rollback":"drop/hide new index if write cost or plan regression appears", "verify":["p99 recovered","docsExamined/returned normalized","write latency unchanged"]}print(incident)

Check your understanding

  1. Why preserve evidence before restarting a process?
  2. Why are replication lag and oplog window read together?
  3. Why can a byte-balanced sharded cluster still be overloaded on one shard?
  4. What should a runbook change first?
  5. What verifies recovery?
Review the answers

1. Restarting resets or removes some counters, active-operation state, cache state, and transient failure evidence.

2. Lag shows current delay while the oplog window shows how long the secondary can remain behind before history needed for catch-up disappears.

3. Workload/query-key skew can concentrate operations independently of stored byte distribution.

4. The smallest reversible action supported by the evidence, after immediate user harm is contained.

5. Both the user/application SLO and the mechanism-level evidence that was abnormal should return to an accepted range.

7. Production judgment

The operational loop for MongoDB is now complete: instrument the SLO, size recovery headroom, benchmark the real workload, gate upgrades, and encode incidents as evidence-driven runbooks. Avoid emergency folklore such as random parameter changes, forced replica reconfiguration, host firewall hacks, clock skew, or deleting storage files. Chapter 27 moves from operator evidence to application-driver architecture: topology discovery, pools, server selection, timeouts, retries, idempotency, observability, and a final production capstone.

Authoritative references

Operational fields, thresholds, upgrade paths, FCV behavior, Atlas metrics, and driver compatibility evolve. Re-check the current documentation for the exact server patch, deployment topology, driver, Atlas tier, and target upgrade/downgrade path before changing production systems.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.