Turn failover into an election timeline: heartbeats, terms, votes, priorities, catch-up, stepdown, and the no-primary interval.
Elections, Terms, Heartbeats, Priorities, Votes, Stepdown, and Failover Windows
Observe election terms and failover windows without assuming a fixed failover SLA or permanent primary.
Learning objectives
Explain heartbeat, election timeout, vote, priority, term, and catch-up without promising exact failover time.
Observe term/member-state changes across a controlled stepdown.
Use priority to influence eligibility without treating it as permanent assignment.
Distinguish transient stale primary perception from ability to complete majority writes.
Plan client timeout/retry behavior around a temporary no-primary interval.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo
4.17.0 where the driver is used. The mandatory
topology is a disposable three-member replica set on one Docker
host, with loopback-published diagnostic ports 27090–27092.
Members communicate over a lesson-specific Docker bridge network
by container DNS names. Authentication and TLS are disabled only
for this isolated local learning topology. Feature Compatibility
Version (FCV) is observed and never changed. Unless explicitly
overridden, writes use the deployment default and reads target
the primary; examples that need stronger durability state
w:"majority" explicitly. Run only one Chapter 14
topology at a time and budget roughly 3–4 GB of free RAM plus
disk headroom. Atlas, Search, Vector Search, KMS, and Enterprise
Advanced are not mandatory. Member A gets priority 2 only to
make the starting lab easier to observe; priority is preference,
not permanent ownership. Product commands were not executed in
this generation environment because Docker, mongod, mongosh, and
PyMongo are unavailable here; expected state is
documentation-derived and measured values must be recorded on
the learner’s machine.
1. AtlasMart problem: survive leader loss without a split-brain slogan
Members exchange heartbeats for health/configuration. Eligible voting secondaries can call an election after primary loss. A successful election advances the replica-set term. Priority influences election preference/eligibility and votes participate in quorum. Default heartbeat/election values are configuration inputs, not application SLAs.
docker rm -f atlasmart-mongo-ch14-l3-a atlasmart-mongo-ch14-l3-b atlasmart-mongo-ch14-l3-c 2>/dev/null || truedocker network rm atlasmart-ch14-l3-net 2>/dev/null || truedocker volume rm atlasmart-mongo-ch14-l3-a-data atlasmart-mongo-ch14-l3-b-data atlasmart-mongo-ch14-l3-c-data 2>/dev/null || truedocker network create atlasmart-ch14-l3-netdocker run -d --name atlasmart-mongo-ch14-l3-a --network atlasmart-ch14-l3-net -p 127.0.0.1:27090:27017 -v atlasmart-mongo-ch14-l3-a-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l3 --bind_ip_alldocker run -d --name atlasmart-mongo-ch14-l3-b --network atlasmart-ch14-l3-net -p 127.0.0.1:27091:27017 -v atlasmart-mongo-ch14-l3-b-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l3 --bind_ip_alldocker run -d --name atlasmart-mongo-ch14-l3-c --network atlasmart-ch14-l3-net -p 127.0.0.1:27092:27017 -v atlasmart-mongo-ch14-l3-c-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l3 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-rs14-l3",members:[ {_id:0,host:"atlasmart-mongo-ch14-l3-a:27017",priority:2}, {_id:1,host:"atlasmart-mongo-ch14-l3-b:27017",priority:1}, {_id:2,host:"atlasmart-mongo-ch14-l3-c:27017",priority:1}]})'until mongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'quit(db.hello().isWritablePrimary?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'printjson({server:db.version(),hello:db.hello(),fcv:db.runCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion})'
2. Capture the election baseline
const cfg=rs.conf();const s=rs.status();printjson({term:s.term,primary:s.members.find(m=>m.stateStr==="PRIMARY")?.name});printjson(cfg.members.map(m=>({_id:m._id,host:m.host,priority:m.priority,votes:m.votes})));printjson({heartbeatIntervalMillis:cfg.settings.heartbeatIntervalMillis,electionTimeoutMillis:cfg.settings.electionTimeoutMillis,catchUpTimeoutMillis:cfg.settings.catchUpTimeoutMillis});
3. Controlled stepdown without force
rs.stepDown(30,10) asks the primary to step down
for 30 seconds and allows up to 10 seconds for an electable
secondary to catch up. Do not use force:true for a
normal drill.
date -Isecondsmongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'try { rs.stepDown(30,10) } catch(e) { print(e.message) }'for i in $(seq 1 20); do printf "sample=%02d time=" "$i"; date -Iseconds for port in 27090 27091 27092; do mongosh "mongodb://127.0.0.1:${port}/admin?directConnection=true" --quiet --eval ' const h=db.hello(); printjson({me:h.me,primary:h.primary,writable:h.isWritablePrimary,secondary:h.secondary});' || true done sleep 1done
Record stepdown request time, first new
isWritablePrimary:true, term before/after, and
client-visible errors/latency. Do not copy a fixed “10 second
failover” number; network, catch-up, host scheduling, and
driver monitoring all matter.
4. Priority affects eligibility, not truth
Priority 0 prevents a member from becoming primary. This can support placement/reporting policy, but it reduces the pool of electable members.
let cfg=rs.conf();cfg.members.find(m=>m.host.startsWith("atlasmart-mongo-ch14-l3-c:")).priority=0;rs.reconfig(cfg);printjson(rs.conf().members.map(m=>({host:m.host,priority:m.priority,votes:m.votes})));cfg=rs.conf();cfg.members.find(m=>m.host.startsWith("atlasmart-mongo-ch14-l3-c:")).priority=1;rs.reconfig(cfg);printjson(rs.conf().members.map(m=>({host:m.host,priority:m.priority,votes:m.votes})));
5. Deliberately wrong: “two nodes say primary, so both can commit majority history”
During partitions a former primary can transiently have stale self-perception. That does not mean both branches can complete majority writes. At most one side owns the voting majority; weaker divergent writes on the other side may later roll back.
Reason in terms of term, voting majority, and majority commit point—not one instant of a role label. Use replica-set discovery, bounded timeouts, eligible driver retries, and idempotent application semantics. Never repair a partition by force-reconfiguring a convenient node.
Production judgment. Aggressive election tuning can trade a smaller detection window for election churn during network/host pauses. Priorities should reflect failure domains and policy. Track election frequency, terms, heartbeats, lag, server-selection errors, and p95/p99 write latency.
Bridge. Elections depend on healthy, sufficiently current secondaries. Lesson 4 examines sync, initial sync, rollback, oplog window, and lag.
docker rm -f atlasmart-mongo-ch14-l3-a atlasmart-mongo-ch14-l3-b atlasmart-mongo-ch14-l3-c atlasmart-mongo-ch14-l3-d 2>/dev/null || truedocker volume rm atlasmart-mongo-ch14-l3-a-data atlasmart-mongo-ch14-l3-b-data atlasmart-mongo-ch14-l3-c-data atlasmart-mongo-ch14-l3-d-data 2>/dev/null || truedocker network rm atlasmart-ch14-l3-net 2>/dev/null || true
Check your understanding
- What does term represent?
- Does higher priority guarantee permanent primary status?
- Why avoid forced stepdown?
- Why can a former primary label be misleading during partition?
- Why measure failover instead of assuming electionTimeoutMillis?
Review the answers
1. An election generation that advances when a new primary wins an election.
2. No. Health, freshness, votes, and elections determine the actual primary.
3. It can promote stale candidates and increase rollback/data-loss risk.
4. Only the side with voting majority can complete majority writes as authoritative current primary.
5. Election detection, catch-up, network, scheduling, and client server selection all contribute.
Authoritative references
- MongoDB 8.3 release notes — Current server release line and patch-sensitive replication behavior.
- Replication — Replica-set purpose, asynchronous replication, failover, and topology concepts.
- Replica set oplog — Oplog semantics, rolling history, and majority-commit retention behavior.
- replSetGetStatus — Member states, optimes, terms, and majority commit point evidence.
- Replica-set configuration — version, term, members, votes, priorities, heartbeats, and election settings.
- Replica-set data synchronization — Initial sync and ongoing oplog application.
- Rollbacks during failover — Divergent former-primary writes and rollback protection.
- Troubleshoot replica sets — Replication lag, oplog window, and operational diagnostics.
- replSetStepDown — Safe primary stepdown and secondary catch-up behavior.
- replSetReconfig — Reconfiguration commitment, elections, and force-reconfiguration risks.
- Retryable writes — Driver retry semantics on replica sets and sharded clusters.
- PyMongo replica-set connections — Seed lists, discovery, failover, and AutoReconnect behavior.
- PyMongo release notes — Current 4.17 driver baseline.
- mongosh release notes — Current 2.10.0 shell baseline.