During incidents, preserve evidence, contain harm, change one thing, and verify both the SLO and the database mechanism.
Incident Runbooks for Replication Lag, Disk Pressure, Slow Queries, Elections, and Shard Imbalance
Build evidence-first incident runbooks for replication lag, disk pressure, slow queries, elections, and shard imbalance with safe containment and verification.
Learning objectives
Use a common incident structure for replication lag, disk pressure, slow queries, elections, and shard imbalance.
Build evidence timelines before changing configuration and classify immediate containment separately from root-cause repair.
Practice safe failure injection without host firewall, system-clock, or destructive production changes.
Define stop/rollback conditions and verify recovery against both database and application SLO evidence.
Connect incident learning back to capacity, alerts, benchmarks, backup/restore, and upgrade controls.
This chapter pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and
PyMongo 4.17.0 where client behavior or load
generation matters. Lesson 5 uses one disposable MongoDB
instance for slow-query/disk-capacity observation and
deterministic simulators for election, replication-lag, and
shard-imbalance timelines. Chapters 14–17 contain the full
replica/sharded mechanisms. Host exposure is loopback-only on
127.0.0.1:27208. Authentication and TLS are
disabled only for disposable local labs; production security
remains the Chapter 22 prerequisite. Default read/write concern
and primary read preference are used unless a step states
otherwise.
FCV is observed and never changed in mandatory labs.
Atlas, Enterprise Advanced, Search, Vector Search, and KMS are
optional unless explicitly labeled. No host firewall rules, real
clock skew, forced replica-set reconfiguration, artificial
production disk filling, or live shard corruption is used.
Failure exercises are simulations or bounded disposable-query
changes. Runtime performance, failover, capacity, and upgrade
labs were not executed in the generation environment, so metric
values, latency distributions, queue depths, oplog windows,
replication lag, incident times, and upgrade durations must be
measured locally rather than copied as invented output.
1. A runbook is an evidence protocol
During an incident, the worst first move is often “change something that usually helps.” A useful runbook starts by preserving an incident timeline: detection time, user impact, release/configuration events, topology state, key metrics, active operations, recent elections/migrations, disk/cache evidence, and application error distribution. Then it distinguishes containment (reduce immediate harm) from root-cause correction.
| Phase | Question | Output |
|---|---|---|
| Detect | Which SLO/user invariant is failing? | Impact and start time |
| Scope | Which tenants/nodes/shards/query shapes? | Affected surface |
| Preserve | What evidence will disappear if we restart/change state? | Timeline/metrics/logs/profile samples |
| Contain | What reversible action reduces harm? | Traffic/rate/query/maintenance change |
| Diagnose | Which mechanism explains the correlated evidence? | Testable hypothesis |
| Repair | What smallest change addresses the mechanism? | One controlled intervention |
| Verify | Did user SLO + server evidence recover? | Post-change acceptance |
| Learn | What capacity/alert/test/runbook gap allowed this? | Preventive action |
2. Slow-query runbook: shape before server-wide tuning
docker rm -f atlasmart-ch26-l5 2>/dev/null || truedocker volume rm atlasmart-ch26-l5-db 2>/dev/null || truedocker run -d --name atlasmart-ch26-l5 -p 127.0.0.1:27208:27017 -v atlasmart-ch26-l5-db:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27208/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; done
const d=db.getSiblingDB("atlasmart");d.products_ch26_l5.drop();for (let b=0;b<30;b++) { const x=[]; for (let i=0;i<1000;i++){const n=b*1000+i;x.push({sku:`SKU-${n}`,tenantId:`tenant-${n%20}`,category:`cat-${n%40}`,priceCents:1000+n%100000,description:`portable atlasmart product ${n} `+"x".repeat(n%100)});} d.products_ch26_l5.insertMany(x); }let before=d.products_ch26_l5.find({tenantId:"tenant-7",category:"cat-7"}).sort({priceCents:1}).explain("executionStats");printjson({before:{nReturned:before.executionStats.nReturned,totalDocsExamined:before.executionStats.totalDocsExamined,totalKeysExamined:before.executionStats.totalKeysExamined,plan:before.queryPlanner.winningPlan}});d.products_ch26_l5.createIndex({tenantId:1,category:1,priceCents:1},{name:"idx_tenant_category_price"});let after=d.products_ch26_l5.find({tenantId:"tenant-7",category:"cat-7"}).sort({priceCents:1}).explain("executionStats");printjson({after:{nReturned:after.executionStats.nReturned,totalDocsExamined:after.executionStats.totalDocsExamined,totalKeysExamined:after.executionStats.totalKeysExamined,plan:after.queryPlanner.winningPlan}});
The repair is evidence-driven: the query contract stayed constant and one index changed. Do not generalize this exact index to another workload without checking field cardinality, sort/range order, write cost, and competing query shapes.
3. Replication lag and elections: keep availability and durability evidence separate
events=[ {"t":0,"lag_s":1,"oplog_window_h":24,"primary":True,"p99_ms":35}, {"t":5,"lag_s":18,"oplog_window_h":9,"primary":True,"p99_ms":62}, {"t":10,"lag_s":75,"oplog_window_h":3,"primary":True,"p99_ms":105}, {"t":15,"lag_s":140,"oplog_window_h":1.2,"primary":True,"p99_ms":170},]for e in events: risk=e["lag_s"]>60 and e["oplog_window_h"]<4 print({**e,"secondary_resync_risk":risk})print("containment candidates: reduce nonessential write load; protect oplog window; investigate secondary disk/cache/network; do not force reconfig")
For elections, record which member became primary, election frequency, terms, heartbeat/network evidence, client error/retry behavior, and whether acknowledged writes meet the requested concern. Repeated elections are a symptom; increasing election timeout or forcing configuration without diagnosing connectivity/resource causes can trade one failure mode for another.
4. Disk pressure: project time-to-exhaustion before emergency cleanup
from datetime import timedeltafree_gib=85growth_gib_per_hour=3.5maintenance_lead_hours=12hours_to_full=free_gib/growth_gib_per_hourprint({"hours_to_full":round(hours_to_full,1),"maintenance_lead_hours":maintenance_lead_hours,"margin_hours":round(hours_to_full-maintenance_lead_hours,1)})
Containment may include pausing nonessential ingestion, moving
exports off the database filesystem, reducing temporary
maintenance work, or increasing storage through the supported
platform mechanism. Never “fix disk pressure” by deleting
unknown files inside dbPath. Preserve
journal/data-file integrity and validate backup/recovery status
before destructive cleanup.
5. Shard imbalance: distinguish data ownership from hot workload
A cluster can be byte-balanced and still hot if one shard
receives a disproportionate query/write rate, and it can be
range-imbalanced during an expected migration/zone policy.
Collect sh.status(), range ownership, balancer
state, per-shard operations/CPU/disk, query targeting, shard-key
distribution, and recent migration/resharding events. Chapters
16–17 provide the safe local mechanisms.
| Symptom | Possible mechanism | Evidence before action |
|---|---|---|
| One shard high CPU, equal bytes | hot shard-key values or targeted workload skew | per-shard ops + key/query distribution |
| Unequal bytes, balancer active | migration convergence | range metadata + balancer/migration state |
| Scatter/gather latency | queries missing shard-key targeting | mongos explain + query shapes |
| Migration churn | poor key/zone design or insufficient headroom | zone/range metadata + move history + disk/network |
| Jumbo/large ranges | indivisible/skewed key interval | range statistics + shard-key cardinality/frequency |
6. Incident acceptance criteria and rollback
Every runbook needs a stop condition. For example: if a query/index change does not reduce documents examined and p99 under the same workload, roll it back; if a secondary cannot make positive catch-up progress before the oplog window closes, escalate toward resync/capacity protection; if an upgrade member fails to return healthy, stop the rollout. Verification must include the user-facing SLO and the mechanism-level metric that motivated the change.
incident={ "impact":"checkout p99 above SLO", "start":"2026-09-03T10:00:00Z", "scope":["tenant-hot-1"], "evidence":["client-p99","serverStatus-delta","currentOp","explain","logs"], "hypothesis":"new query shape performs collection scan", "containment":"route/report endpoint disabled", "change":"add reviewed compound index", "rollback":"drop/hide new index if write cost or plan regression appears", "verify":["p99 recovered","docsExamined/returned normalized","write latency unchanged"]}print(incident)
Check your understanding
- Why preserve evidence before restarting a process?
- Why are replication lag and oplog window read together?
- Why can a byte-balanced sharded cluster still be overloaded on one shard?
- What should a runbook change first?
- What verifies recovery?
Review the answers
1. Restarting resets or removes some counters, active-operation state, cache state, and transient failure evidence.
2. Lag shows current delay while the oplog window shows how long the secondary can remain behind before history needed for catch-up disappears.
3. Workload/query-key skew can concentrate operations independently of stored byte distribution.
4. The smallest reversible action supported by the evidence, after immediate user harm is contained.
5. Both the user/application SLO and the mechanism-level evidence that was abnormal should return to an accepted range.
7. Production judgment
The operational loop for MongoDB is now complete: instrument the SLO, size recovery headroom, benchmark the real workload, gate upgrades, and encode incidents as evidence-driven runbooks. Avoid emergency folklore such as random parameter changes, forced replica reconfiguration, host firewall hacks, clock skew, or deleting storage files. Chapter 27 moves from operator evidence to application-driver architecture: topology discovery, pools, server selection, timeouts, retries, idempotency, observability, and a final production capstone.
Authoritative references
Operational fields, thresholds, upgrade paths, FCV behavior, Atlas metrics, and driver compatibility evolve. Re-check the current documentation for the exact server patch, deployment topology, driver, Atlas tier, and target upgrade/downgrade path before changing production systems.
- serverStatus command
- db.stats() / dbStats
- $collStats aggregation stage
- Database Profiler
- db.setProfilingLevel()
- $currentOp aggregation stage
- MongoDB Log Messages
- Explain Results
- Replication
- Replica Set Oplog
- Check Replica Set Replication Lag
- WiredTiger Storage Engine
- Atlas Monitoring and Alerts
- Atlas Monitoring and Alert Guidance
- Atlas Alert Basics
- Atlas Metrics
- PyMongo Release Notes
- PyMongo Upgrade Guidance
- MongoDB 8.3 Release Notes
- Upgrade 8.2 to 8.3
- Upgrade 8.2 Replica Set to 8.3
- Upgrade 8.2 Sharded Cluster to 8.3
- MongoDB 8.3 Compatibility Changes
- Downgrade 8.3 to 8.2
- MongoDB Versioning
- Backup Methods
- mongosh Release Notes