Treat high availability as a testable application property by killing the primary and verifying server/client recovery.
Run Failure Drills: Kill the Primary, Observe Elections, and Verify Application Recovery
Execute a reversible primary-loss drill and verify both election and replica-set-aware application recovery.
Learning objectives
Run a hard-primary-loss drill on an isolated three-member replica set.
Record election/no-primary window, term changes, and post-failover majority write evidence.
Use a replica-set-aware PyMongo connection instead of pinning one host.
Distinguish retryable database operations from application-level idempotency.
Define production failover acceptance criteria for correctness, latency, observability, and recovery.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo
4.17.0 where the driver is used. The mandatory
topology is a disposable three-member replica set on one Docker
host, with loopback-published diagnostic ports 27097–27099.
Members communicate over a lesson-specific Docker bridge network
by container DNS names. Authentication and TLS are disabled only
for this isolated local learning topology. Feature Compatibility
Version (FCV) is observed and never changed. Unless explicitly
overridden, writes use the deployment default and reads target
the primary; examples that need stronger durability state
w:"majority" explicitly. Run only one Chapter 14
topology at a time and budget roughly 3–4 GB of free RAM plus
disk headroom. Atlas, Search, Vector Search, KMS, and Enterprise
Advanced are not mandatory. Member A starts with priority 2 to
make the kill target easier to control. The optional PyMongo
probe runs on the same Docker network so it can resolve
advertised member hostnames. Product commands were not executed
in this generation environment because Docker, mongod, mongosh,
and PyMongo are unavailable here; expected state is
documentation-derived and measured values must be recorded on
the learner’s machine.
1. AtlasMart drill: remove the primary, not a random node
A meaningful HA test asks whether surviving voting secondaries elect a new primary, clients discover it, eligible database operations retry within policy, majority-committed state survives, and the old primary rejoins cleanly. The no-primary interval is part of the result, not an embarrassment to hide.
docker rm -f atlasmart-mongo-ch14-l5-a atlasmart-mongo-ch14-l5-b atlasmart-mongo-ch14-l5-c 2>/dev/null || truedocker network rm atlasmart-ch14-l5-net 2>/dev/null || truedocker volume rm atlasmart-mongo-ch14-l5-a-data atlasmart-mongo-ch14-l5-b-data atlasmart-mongo-ch14-l5-c-data 2>/dev/null || truedocker network create atlasmart-ch14-l5-netdocker run -d --name atlasmart-mongo-ch14-l5-a --network atlasmart-ch14-l5-net -p 127.0.0.1:27097:27017 -v atlasmart-mongo-ch14-l5-a-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l5 --bind_ip_alldocker run -d --name atlasmart-mongo-ch14-l5-b --network atlasmart-ch14-l5-net -p 127.0.0.1:27098:27017 -v atlasmart-mongo-ch14-l5-b-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l5 --bind_ip_alldocker run -d --name atlasmart-mongo-ch14-l5-c --network atlasmart-ch14-l5-net -p 127.0.0.1:27099:27017 -v atlasmart-mongo-ch14-l5-c-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l5 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27097/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27097/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-rs14-l5",members:[ {_id:0,host:"atlasmart-mongo-ch14-l5-a:27017",priority:2}, {_id:1,host:"atlasmart-mongo-ch14-l5-b:27017",priority:1}, {_id:2,host:"atlasmart-mongo-ch14-l5-c:27017",priority:1}]})'until mongosh "mongodb://127.0.0.1:27097/admin?directConnection=true" --quiet --eval 'quit(db.hello().isWritablePrimary?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27097/admin?directConnection=true" --quiet --eval 'printjson({server:db.version(),hello:db.hello(),fcv:db.runCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion})'
2. Seed majority state and prove the target
const h=db.hello();printjson({me:h.me,isWritablePrimary:h.isWritablePrimary,setName:h.setName});if (!h.isWritablePrimary) throw new Error("Use the current primary; this lab expects member A to start primary.");const c=db.getSiblingDB("atlasmart").failover_probe_ch14_l5;c.drop();printjson(c.insertOne({_id:"baseline",counter:0,updatedAt:new Date()},{writeConcern:{w:"majority",wtimeout:10000}}));const s=rs.status();printjson({term:s.term,lastCommittedOpTime:s.optimes.lastCommittedOpTime,members:s.members.map(m=>({name:m.name,stateStr:m.stateStr,optimeDate:m.optimeDate}))});
3. Optional replica-set-aware PyMongo probe
Because the set advertises Docker DNS names, run the probe from the same Docker network. If package download is unavailable, use an existing environment with PyMongo 4.17.0 and equivalent DNS visibility.
import timefrom datetime import datetime, timezonefrom pymongo import MongoClientfrom pymongo.errors import AutoReconnect, ConnectionFailure, ServerSelectionTimeoutError, PyMongoErroruri = "mongodb://atlasmart-mongo-ch14-l5-a:27017,atlasmart-mongo-ch14-l5-b:27017,atlasmart-mongo-ch14-l5-c:27017/?replicaSet=atlasmart-rs14-l5&retryWrites=true"client = MongoClient(uri, serverSelectionTimeoutMS=3000, connectTimeoutMS=2000)coll = client.atlasmart.failover_probe_ch14_l5for i in range(40): op_id = f"probe-{i:02d}" started = time.monotonic() try: r = coll.update_one({"_id": op_id},{"$set": {"observedAt": datetime.now(timezone.utc), "sequence": i}},upsert=True) primary = client.admin.command("hello").get("primary") print(f"ok i={i:02d} primary={primary} matched={r.matched_count} upserted={r.upserted_id} elapsed_ms={(time.monotonic()-started)*1000:.1f}") except (AutoReconnect, ConnectionFailure, ServerSelectionTimeoutError) as exc: print(f"topology-window i={i:02d} type={type(exc).__name__} elapsed_ms={(time.monotonic()-started)*1000:.1f}") except PyMongoError as exc: print(f"database-error i={i:02d} type={type(exc).__name__}") time.sleep(1)print("final_count", coll.count_documents({"_id": {"$regex": "^probe-"}}))client.close()
cat > /tmp/ch14_failover_probe.py <<'PY'# Paste the Python block above here exactly.PYdocker run --rm --name atlasmart-client-ch14-l5 --network atlasmart-ch14-l5-net -v /tmp/ch14_failover_probe.py:/app/probe.py:ro python:3.13-slim-bookworm sh -lc 'pip install --no-cache-dir pymongo==4.17.0 >/tmp/pip.log && python /app/probe.py'
Expected logs: success before failure, a bounded topology-selection/reconnect interval, then success on the new primary. Exact duration is environment-dependent.
4. Kill A, observe election, restart A
Blast radius: one disposable container. Verify A is writable primary immediately before the stop. If not, identify and stop the actual primary instead of pretending the test covered failover.
date -Isecondsmongosh "mongodb://127.0.0.1:27097/admin?directConnection=true" --quiet --eval 'const h=db.hello(); printjson(h); quit(h.isWritablePrimary?0:2)'docker stop atlasmart-mongo-ch14-l5-afor i in $(seq 1 30); do printf "sample=%02d time=" "$i"; date -Iseconds for port in 27098 27099; do mongosh "mongodb://127.0.0.1:${port}/admin?directConnection=true" --quiet --eval ' const h=db.hello(); printjson({me:h.me,primary:h.primary,writable:h.isWritablePrimary,secondary:h.secondary});' || true done sleep 1done# Use the port of whichever member actually won.mongosh "mongodb://127.0.0.1:27098/atlasmart?directConnection=true" --quiet --eval 'if (!db.hello().isWritablePrimary) throw new Error("Use elected primary port, possibly 27099");printjson(db.failover_probe_ch14_l5.insertOne({_id:"after-failover",at:new Date()},{writeConcern:{w:"majority",wtimeout:10000}}));printjson(rs.status().optimes);'docker start atlasmart-mongo-ch14-l5-asleep 10for port in 27097 27098 27099; do mongosh "mongodb://127.0.0.1:${port}/admin?directConnection=true" --quiet --eval 'printjson({hello:db.hello(),term:rs.status().term})' || true; done
5. Deliberately wrong: rerun checkout after every network exception
Retryable writes can retry eligible MongoDB operations once after qualifying failures. That guarantee does not wrap a payment capture, email, webhook, or queue publication around the database call. Replaying the whole business workflow can duplicate external effects.
Use a business idempotency key, idempotent/conditional
database transitions, durable outbox/intent state, bounded
retries, error classification, and reconciliation. The probe
uses deterministic _id values so replay
converges.
6. Failure-drill acceptance criteria
| Dimension | Record | Pass condition |
|---|---|---|
| Election | Old/new primary, term, timestamps | Surviving majority elects a new primary without force. |
| Write safety | Majority commit before/after | Baseline survives and post-failover majority write succeeds. |
| Client | Reconnect/server-selection log | Replica-set-aware client resumes within timeout budget. |
| Idempotency | Operation IDs/final counts | Retries do not create duplicate business effects. |
| Rejoin | Old primary state/lag | Restarted A rejoins and catches up before being considered healthy. |
| Observability | Server/driver logs + latency distribution | Incident timeline can be reconstructed. |
Production judgment. Repeat drills across real failure domains, TLS/auth, load, pools, storage pressure, and zone/network scenarios. Measure p50/p95/p99 latency and error rates; avoid retry storms. Test backups separately because replication does not protect against logical corruption.
Bridge to Chapter 15. Replica sets provide
the mechanism; applications still choose acknowledgement and
visibility semantics. Next: w, j,
wtimeout, read concern, read preference, and
causal consistency.
docker rm -f atlasmart-mongo-ch14-l5-a atlasmart-mongo-ch14-l5-b atlasmart-mongo-ch14-l5-c atlasmart-mongo-ch14-l5-d 2>/dev/null || truedocker volume rm atlasmart-mongo-ch14-l5-a-data atlasmart-mongo-ch14-l5-b-data atlasmart-mongo-ch14-l5-c-data atlasmart-mongo-ch14-l5-d-data 2>/dev/null || truedocker network rm atlasmart-ch14-l5-net 2>/dev/null || true
Check your understanding
- Why identify the actual primary before the kill?
- Why can writes fail temporarily with two healthy secondaries?
- What does retryWrites=true not guarantee?
- Why deterministic operation IDs?
- Why restart the old primary?
Review the answers
1. Stopping a random secondary does not test primary failover.
2. Primary loss must be detected, an election completed, and clients must discover/select the new primary.
3. Exactly-once behavior for arbitrary application side effects.
4. They make replay converge and expose missing/duplicate logical work.
5. A complete drill verifies redundancy is restored and the old primary catches up as a healthy member.
Authoritative references
- MongoDB 8.3 release notes — Current server release line and patch-sensitive replication behavior.
- Replication — Replica-set purpose, asynchronous replication, failover, and topology concepts.
- Replica set oplog — Oplog semantics, rolling history, and majority-commit retention behavior.
- replSetGetStatus — Member states, optimes, terms, and majority commit point evidence.
- Replica-set configuration — version, term, members, votes, priorities, heartbeats, and election settings.
- Replica-set data synchronization — Initial sync and ongoing oplog application.
- Rollbacks during failover — Divergent former-primary writes and rollback protection.
- Troubleshoot replica sets — Replication lag, oplog window, and operational diagnostics.
- replSetStepDown — Safe primary stepdown and secondary catch-up behavior.
- replSetReconfig — Reconfiguration commitment, elections, and force-reconfiguration risks.
- Retryable writes — Driver retry semantics on replica sets and sharded clusters.
- PyMongo replica-set connections — Seed lists, discovery, failover, and AutoReconnect behavior.
- PyMongo release notes — Current 4.17 driver baseline.
- mongosh release notes — Current 2.10.0 shell baseline.