Turn failover into an election timeline: heartbeats, terms, votes, priorities, catch-up, stepdown, and the no-primary interval.

Elections, Terms, Heartbeats, Priorities, Votes, Stepdown, and Failover Windows

Observe election terms and failover windows without assuming a fixed failover SLA or permanent primary.

Intermediate110–170 minutesReplica-set replication/failure labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Explain heartbeat, election timeout, vote, priority, term, and catch-up without promising exact failover time.

02

Observe term/member-state changes across a controlled stepdown.

03

Use priority to influence eligibility without treating it as permanent assignment.

04

Distinguish transient stale primary perception from ability to complete majority writes.

05

Plan client timeout/retry behavior around a temporary no-primary interval.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where the driver is used. The mandatory topology is a disposable three-member replica set on one Docker host, with loopback-published diagnostic ports 27090–27092. Members communicate over a lesson-specific Docker bridge network by container DNS names. Authentication and TLS are disabled only for this isolated local learning topology. Feature Compatibility Version (FCV) is observed and never changed. Unless explicitly overridden, writes use the deployment default and reads target the primary; examples that need stronger durability state w:"majority" explicitly. Run only one Chapter 14 topology at a time and budget roughly 3–4 GB of free RAM plus disk headroom. Atlas, Search, Vector Search, KMS, and Enterprise Advanced are not mandatory. Member A gets priority 2 only to make the starting lab easier to observe; priority is preference, not permanent ownership. Product commands were not executed in this generation environment because Docker, mongod, mongosh, and PyMongo are unavailable here; expected state is documentation-derived and measured values must be recorded on the learner’s machine.

1. AtlasMart problem: survive leader loss without a split-brain slogan

Members exchange heartbeats for health/configuration. Eligible voting secondaries can call an election after primary loss. A successful election advances the replica-set term. Priority influences election preference/eligibility and votes participate in quorum. Default heartbeat/election values are configuration inputs, not application SLAs.

isolated three-member replica-set setup (l3)
docker rm -f atlasmart-mongo-ch14-l3-a atlasmart-mongo-ch14-l3-b atlasmart-mongo-ch14-l3-c 2>/dev/null || truedocker network rm atlasmart-ch14-l3-net 2>/dev/null || truedocker volume rm atlasmart-mongo-ch14-l3-a-data atlasmart-mongo-ch14-l3-b-data atlasmart-mongo-ch14-l3-c-data 2>/dev/null || truedocker network create atlasmart-ch14-l3-netdocker run -d --name atlasmart-mongo-ch14-l3-a --network atlasmart-ch14-l3-net -p 127.0.0.1:27090:27017 -v atlasmart-mongo-ch14-l3-a-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l3 --bind_ip_alldocker run -d --name atlasmart-mongo-ch14-l3-b --network atlasmart-ch14-l3-net -p 127.0.0.1:27091:27017 -v atlasmart-mongo-ch14-l3-b-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l3 --bind_ip_alldocker run -d --name atlasmart-mongo-ch14-l3-c --network atlasmart-ch14-l3-net -p 127.0.0.1:27092:27017 -v atlasmart-mongo-ch14-l3-c-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l3 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-rs14-l3",members:[  {_id:0,host:"atlasmart-mongo-ch14-l3-a:27017",priority:2},  {_id:1,host:"atlasmart-mongo-ch14-l3-b:27017",priority:1},  {_id:2,host:"atlasmart-mongo-ch14-l3-c:27017",priority:1}]})'until mongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'quit(db.hello().isWritablePrimary?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'printjson({server:db.version(),hello:db.hello(),fcv:db.runCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion})' 

2. Capture the election baseline

term, election settings, priorities, and votes
const cfg=rs.conf();const s=rs.status();printjson({term:s.term,primary:s.members.find(m=>m.stateStr==="PRIMARY")?.name});printjson(cfg.members.map(m=>({_id:m._id,host:m.host,priority:m.priority,votes:m.votes})));printjson({heartbeatIntervalMillis:cfg.settings.heartbeatIntervalMillis,electionTimeoutMillis:cfg.settings.electionTimeoutMillis,catchUpTimeoutMillis:cfg.settings.catchUpTimeoutMillis});

3. Controlled stepdown without force

rs.stepDown(30,10) asks the primary to step down for 30 seconds and allows up to 10 seconds for an electable secondary to catch up. Do not use force:true for a normal drill.

record an election timeline
date -Isecondsmongosh "mongodb://127.0.0.1:27090/admin?directConnection=true" --quiet --eval 'try { rs.stepDown(30,10) } catch(e) { print(e.message) }'for i in $(seq 1 20); do  printf "sample=%02d time=" "$i"; date -Iseconds  for port in 27090 27091 27092; do    mongosh "mongodb://127.0.0.1:${port}/admin?directConnection=true" --quiet --eval '      const h=db.hello(); printjson({me:h.me,primary:h.primary,writable:h.isWritablePrimary,secondary:h.secondary});' || true  done  sleep 1done
What to measure

Record stepdown request time, first new isWritablePrimary:true, term before/after, and client-visible errors/latency. Do not copy a fixed “10 second failover” number; network, catch-up, host scheduling, and driver monitoring all matter.

4. Priority affects eligibility, not truth

Priority 0 prevents a member from becoming primary. This can support placement/reporting policy, but it reduces the pool of electable members.

temporarily make member C non-electable, then restore
let cfg=rs.conf();cfg.members.find(m=>m.host.startsWith("atlasmart-mongo-ch14-l3-c:")).priority=0;rs.reconfig(cfg);printjson(rs.conf().members.map(m=>({host:m.host,priority:m.priority,votes:m.votes})));cfg=rs.conf();cfg.members.find(m=>m.host.startsWith("atlasmart-mongo-ch14-l3-c:")).priority=1;rs.reconfig(cfg);printjson(rs.conf().members.map(m=>({host:m.host,priority:m.priority,votes:m.votes})));

5. Deliberately wrong: “two nodes say primary, so both can commit majority history”

During partitions a former primary can transiently have stale self-perception. That does not mean both branches can complete majority writes. At most one side owns the voting majority; weaker divergent writes on the other side may later roll back.

Repair the mental model

Reason in terms of term, voting majority, and majority commit point—not one instant of a role label. Use replica-set discovery, bounded timeouts, eligible driver retries, and idempotent application semantics. Never repair a partition by force-reconfiguring a convenient node.

Production judgment. Aggressive election tuning can trade a smaller detection window for election churn during network/host pauses. Priorities should reflect failure domains and policy. Track election frequency, terms, heartbeats, lag, server-selection errors, and p95/p99 write latency.

Bridge. Elections depend on healthy, sufficiently current secondaries. Lesson 4 examines sync, initial sync, rollback, oplog window, and lag.

cleanup / full reset
docker rm -f atlasmart-mongo-ch14-l3-a atlasmart-mongo-ch14-l3-b atlasmart-mongo-ch14-l3-c atlasmart-mongo-ch14-l3-d 2>/dev/null || truedocker volume rm atlasmart-mongo-ch14-l3-a-data atlasmart-mongo-ch14-l3-b-data atlasmart-mongo-ch14-l3-c-data atlasmart-mongo-ch14-l3-d-data 2>/dev/null || truedocker network rm atlasmart-ch14-l3-net 2>/dev/null || true

Check your understanding

  1. What does term represent?
  2. Does higher priority guarantee permanent primary status?
  3. Why avoid forced stepdown?
  4. Why can a former primary label be misleading during partition?
  5. Why measure failover instead of assuming electionTimeoutMillis?
Review the answers

1. An election generation that advances when a new primary wins an election.

2. No. Health, freshness, votes, and elections determine the actual primary.

3. It can promote stale candidates and increase rollback/data-loss risk.

4. Only the side with voting majority can complete majority writes as authoritative current primary.

5. Election detection, catch-up, network, scheduling, and client server selection all contribute.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.