Measure how secondaries catch up, how initial sync depends on retained history, and when divergent former-primary writes roll back.
Secondary Sync, Initial Sync, Rollback, Oplog Sizing, and Replication Lag
Use lag, oplog-window, initial-sync, and rollback evidence to diagnose replica health and recovery risk.
Learning objectives
Distinguish ongoing secondary synchronization from initial sync and identify STARTUP2 evidence.
Measure replication lag and oplog window rather than size by folklore.
Explain rollback conditions and majority acknowledgement implications.
Run an isolated w:1 divergence/rollback drill without forced reconfiguration or host firewall changes.
Explain why oplog/replication health still does not replace backup.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo
4.17.0 where the driver is used. The mandatory
topology is a disposable three-member replica set on one Docker
host, with loopback-published diagnostic ports 27093–27096.
Members communicate over a lesson-specific Docker bridge network
by container DNS names. Authentication and TLS are disabled only
for this isolated local learning topology. Feature Compatibility
Version (FCV) is observed and never changed. Unless explicitly
overridden, writes use the deployment default and reads target
the primary; examples that need stronger durability state
w:"majority" explicitly. Run only one Chapter 14
topology at a time and budget roughly 3–4 GB of free RAM plus
disk headroom. Atlas, Search, Vector Search, KMS, and Enterprise
Advanced are not mandatory. Three members form the steady-state
set and a fourth empty member is added as non-voting priority-0
to demonstrate initial sync. The rollback probe touches only
this lesson’s synthetic collection. Product commands were not
executed in this generation environment because Docker, mongod,
mongosh, and PyMongo are unavailable here; expected state is
documentation-derived and measured values must be recorded on
the learner’s machine.
1. AtlasMart problem: can a secondary still catch up after maintenance?
After initial sync, secondaries continuously copy/apply oplog history. Replication lag is delay between primary history and secondary application. The oplog window is the retained time range. A member that falls behind beyond retained history may require initial synchronization again.
docker rm -f atlasmart-mongo-ch14-l4-a atlasmart-mongo-ch14-l4-b atlasmart-mongo-ch14-l4-c 2>/dev/null || truedocker network rm atlasmart-ch14-l4-net 2>/dev/null || truedocker volume rm atlasmart-mongo-ch14-l4-a-data atlasmart-mongo-ch14-l4-b-data atlasmart-mongo-ch14-l4-c-data 2>/dev/null || truedocker network create atlasmart-ch14-l4-netdocker run -d --name atlasmart-mongo-ch14-l4-a --network atlasmart-ch14-l4-net -p 127.0.0.1:27093:27017 -v atlasmart-mongo-ch14-l4-a-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l4 --bind_ip_alldocker run -d --name atlasmart-mongo-ch14-l4-b --network atlasmart-ch14-l4-net -p 127.0.0.1:27094:27017 -v atlasmart-mongo-ch14-l4-b-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l4 --bind_ip_alldocker run -d --name atlasmart-mongo-ch14-l4-c --network atlasmart-ch14-l4-net -p 127.0.0.1:27095:27017 -v atlasmart-mongo-ch14-l4-c-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l4 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27093/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27093/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-rs14-l4",members:[ {_id:0,host:"atlasmart-mongo-ch14-l4-a:27017",priority:1}, {_id:1,host:"atlasmart-mongo-ch14-l4-b:27017",priority:1}, {_id:2,host:"atlasmart-mongo-ch14-l4-c:27017",priority:1}]})'until mongosh "mongodb://127.0.0.1:27093/admin?directConnection=true" --quiet --eval 'quit(db.hello().isWritablePrimary?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27093/admin?directConnection=true" --quiet --eval 'printjson({server:db.version(),hello:db.hello(),fcv:db.runCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion})'
2. Measure lag and retention first
print("--- oplog window ---");rs.printReplicationInfo();print("--- secondary lag ---");rs.printSecondaryReplicationInfo();const s=rs.status();printjson({term:s.term,optimes:s.optimes,members:s.members.map(m=>({name:m.name,stateStr:m.stateStr,optimeDate:m.optimeDate,lastHeartbeatRecv:m.lastHeartbeatRecv}))});printjson({oplogMaxBytes:db.getSiblingDB("local").getCollection("oplog.rs").stats().maxSize});printjson({flowControl:db.adminCommand({serverStatus:1}).flowControl});
Plan oplog capacity from write rate and the longest expected member outage/initial-sync duration, with margin. The oplog can exceed configured size to avoid removing the majority commit point, so raw max bytes are not the whole story.
3. Add a fresh fourth member and watch initial sync
docker rm -f atlasmart-mongo-ch14-l4-d 2>/dev/null || truedocker volume rm atlasmart-mongo-ch14-l4-d-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch14-l4-d --network atlasmart-ch14-l4-net -p 127.0.0.1:27096:27017 -v atlasmart-mongo-ch14-l4-d-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs14-l4 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27096/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27093/admin?directConnection=true" --quiet --eval 'printjson(rs.add({_id:3,host:"atlasmart-mongo-ch14-l4-d:27017",priority:0,votes:0}))'for i in $(seq 1 60); do mongosh "mongodb://127.0.0.1:27096/admin?directConnection=true" --quiet --eval ' const s=rs.status(); printjson({myState:s.myState,state:s.members.find(m=>m.self)?.stateStr,initialSyncStatus:s.initialSyncStatus||null});' || true sleep 1done
Logical initial sync clones databases (except local), builds indexes while copying, buffers concurrent oplog history, then applies retained changes to catch up. The oplog window must cover this process. File-copy initial sync is Enterprise-specific and not mandatory here.
4. Controlled rollback drill: one w:1 write exists only on the former primary
Stop both secondaries, immediately attempt one synthetic
w:1 write while A still sees itself as primary,
then stop A and restart B/C so they elect from the branch
lacking that write. Restart A and observe reconciliation. If A
steps down before the probe write, reset and repeat—do not force
topology.
mongosh "mongodb://127.0.0.1:27093/admin?directConnection=true" --quiet --eval 'printjson(db.hello())'docker stop atlasmart-mongo-ch14-l4-b atlasmart-mongo-ch14-l4-cmongosh "mongodb://127.0.0.1:27093/atlasmart?directConnection=true" --quiet --eval 'try { printjson(db.rollback_probe_ch14_l4.insertOne({_id:"DIVERGENT-W1",note:"should disappear after rollback"},{writeConcern:{w:1}})); }catch(e) { printjson({writeDidNotRun:e.codeName||e.name,message:e.message}); }'docker stop atlasmart-mongo-ch14-l4-adocker start atlasmart-mongo-ch14-l4-b atlasmart-mongo-ch14-l4-csleep 15for port in 27094 27095; do mongosh "mongodb://127.0.0.1:${port}/admin?directConnection=true" --quiet --eval 'printjson(db.hello())' || true; donedocker start atlasmart-mongo-ch14-l4-asleep 15mongosh "mongodb://127.0.0.1:27093/atlasmart?directConnection=true" --quiet --eval 'printjson(db.rollback_probe_ch14_l4.findOne({_id:"DIVERGENT-W1"})); printjson(rs.status().members.map(m=>({name:m.name,stateStr:m.stateStr,optimeDate:m.optimeDate})));'docker logs atlasmart-mongo-ch14-l4-a 2>&1 | grep -i rollback | tail -30 || true
If the w:1 write succeeded only on A and B/C later formed the authoritative branch, A must abandon its divergent write when rejoining. The document should be absent after reconciliation. If the write failed because A had stepped down, no divergence occurred; do not fabricate rollback evidence.
5. Health signals and recovery boundaries
| Signal | Question | Escalate when |
|---|---|---|
| Secondary lag | Can members meet failover/read freshness goals? | Lag approaches operational tolerance. |
| Oplog window | Does retained history cover outage/sync duration? | Window shrinks below recovery need. |
| Initial sync | Does STARTUP2 progress to SECONDARY? | Repeated restarts/source failures/gaps. |
| Rollback logs | Did divergent former-primary writes exist? | Unexpected rollback needs reconciliation. |
| Flow control | Is majority lag throttling writes? | Sustained pressure raises tail latency. |
Production judgment. Oplog sizing is capacity planning, not a universal number. Measure write bytes/second, burst rate, maintenance windows, initial-sync duration, large transactions, and disk margin. Rollback of weaker acknowledged writes can require application reconciliation. Never run this rollback drill on production.
Bridge. Lesson 5 turns primary loss into an application-visible failure drill.
docker rm -f atlasmart-mongo-ch14-l4-a atlasmart-mongo-ch14-l4-b atlasmart-mongo-ch14-l4-c atlasmart-mongo-ch14-l4-d 2>/dev/null || truedocker volume rm atlasmart-mongo-ch14-l4-a-data atlasmart-mongo-ch14-l4-b-data atlasmart-mongo-ch14-l4-c-data atlasmart-mongo-ch14-l4-d-data 2>/dev/null || truedocker network rm atlasmart-ch14-l4-net 2>/dev/null || true
Check your understanding
- Initial sync vs ongoing replication?
- Why plan oplog in time as well as bytes?
- What creates rollback?
- Why use w:1 in this drill?
- Why is oplog not backup?
Review the answers
1. Initial sync builds a full new copy then catches up; ongoing replication continuously applies retained oplog history.
2. Recovery depends on how long history remains available under actual write rates.
3. A former primary rejoins with writes absent from the authoritative branch and must revert divergence.
4. To intentionally expose acknowledged-write rollback risk when history exists only on the former primary.
5. It is rolling replication history, not independent durable point-in-time recovery storage.
Authoritative references
- MongoDB 8.3 release notes — Current server release line and patch-sensitive replication behavior.
- Replication — Replica-set purpose, asynchronous replication, failover, and topology concepts.
- Replica set oplog — Oplog semantics, rolling history, and majority-commit retention behavior.
- replSetGetStatus — Member states, optimes, terms, and majority commit point evidence.
- Replica-set configuration — version, term, members, votes, priorities, heartbeats, and election settings.
- Replica-set data synchronization — Initial sync and ongoing oplog application.
- Rollbacks during failover — Divergent former-primary writes and rollback protection.
- Troubleshoot replica sets — Replication lag, oplog window, and operational diagnostics.
- replSetStepDown — Safe primary stepdown and secondary catch-up behavior.
- replSetReconfig — Reconfiguration commitment, elections, and force-reconfiguration risks.
- Retryable writes — Driver retry semantics on replica sets and sharded clusters.
- PyMongo replica-set connections — Seed lists, discovery, failover, and AutoReconnect behavior.
- PyMongo release notes — Current 4.17 driver baseline.
- mongosh release notes — Current 2.10.0 shell baseline.