Connect a physical snapshot to an exact replica-set/oplog boundary and prove the restored files represent the intended point.
Replica-Set Consistent Snapshots, Filesystem/Volume Snapshots, and Oplog-Aware Recovery
Make filesystem snapshot consistency observable with a disposable replica set, fsync locking, physical restore, oplog boundaries, and encrypted-key dependencies.
Learning objectives
Explain the consistency preconditions for filesystem/block snapshots of WiredTiger data and journal state.
Capture a replica-set commit/oplog boundary and connect it to a snapshot recovery point.
Use db.fsyncLock() only inside a disposable lab
and distinguish a locked file copy from a real atomic volume
snapshot.
Explain why future oplog history must be archived independently if recovery must advance beyond the snapshot.
Test a physical restore and identify external encryption/key dependencies.
This chapter pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0,
PyMongo 4.17.0 where application checks are useful,
and MongoDB Database Tools 100.18.0 for
mongodump, mongorestore, and
bsondump. Lesson 2 uses a disposable one-member
replica set and a Docker-volume copy as a portable teaching
substitute for an infrastructure-specific LVM/EBS/SAN snapshot.
Host exposure is loopback-only on 127.0.0.1:27193.
Authentication and TLS are disabled only for disposable local
labs; production security remains the Chapter 22 prerequisite.
Default read/write concern and primary read preference are used
unless a step states otherwise.
FCV is observed and never changed. Atlas,
Enterprise Advanced, Search, Vector Search, and KMS are optional
unless a lesson explicitly labels them. A production replica set
would normally take snapshots from an appropriately configured
secondary/hidden member and use the storage platform’s atomic
snapshot facility. The lab file copy is protected by
fsyncLock precisely because it is not an atomic
volume snapshot. Runtime backup/restore labs were not executed
in the generation environment, so artifact sizes, restore
durations, RPO/RTO measurements, snapshot times, and checksums
must be recorded locally rather than copied as invented output.
1. AtlasMart problem: the logical dump no longer fits the backup window
At hundreds of gigabytes, reading every document through
mongodump can pressure the working set and take
longer than the maintenance window. A filesystem or block
snapshot works below MongoDB: the storage platform records a
point-in-time view of blocks and later copies them to durable
backup storage.
The word “snapshot” does not itself guarantee database consistency. WiredTiger data must represent a valid recoverable state, and any volumes that hold data and journal state must be captured coherently. MongoDB documents that accepted writes need to be on disk in the journal or data files when the snapshot is taken. If your storage cannot provide one coordinated atomic snapshot, temporarily flushing and write-locking a disposable node is a conservative way to make a file copy coherent.
| Evidence | Meaning | Does not prove |
|---|---|---|
rs.status().optimes.lastCommittedOpTime
|
Replica-set commit position around the backup boundary | That future oplog entries are preserved elsewhere |
db.fsyncLock() |
Flushes and blocks writes on the node while locked | That a multi-volume storage platform snapshot is atomic |
| Snapshot/file checksum | Artifact bytes did not change in transit | That the application can start and satisfy invariants |
Restored mongod starts |
WiredTiger recovery found a valid storage state | That RPO/RTO or application cutover objectives are met |
2. Build a replica-set source and record the recovery boundary
docker rm -f atlasmart-ch25-l2 2>/dev/null || truedocker volume rm atlasmart-ch25-l2-db 2>/dev/null || truedocker run -d --name atlasmart-ch25-l2 \ -p 127.0.0.1:27193:27017 \ -v atlasmart-ch25-l2-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --replSet rs25l2 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27193/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27193/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"rs25l2",members:[{_id:0,host:"atlasmart-ch25-l2:27017"}]})'until mongosh "mongodb://127.0.0.1:27193/admin?directConnection=true" --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; done
const d=db.getSiblingDB("atlasmart");d.inventory_ch25_l2.drop();for (let b=0;b<20;b++) { const batch=[]; for (let i=0;i<500;i++) batch.push({sku:`SKU-${b*500+i}`,qty:10+(i%40),updatedAt:new Date()}); d.inventory_ch25_l2.insertMany(batch);}const st=db.getSiblingDB("admin").runCommand({replSetGetStatus:1});const first=db.getSiblingDB("local").oplog.rs.find().sort({$natural:1}).limit(1).next();const last=db.getSiblingDB("local").oplog.rs.find().sort({$natural:-1}).limit(1).next();printjson({count:d.inventory_ch25_l2.countDocuments(),lastCommitted:st.optimes.lastCommittedOpTime,oplogFirst:first.ts,oplogLast:last.ts});
3. Lock only the disposable node, copy the dbPath, then unlock
A real LVM/EBS/SAN snapshot should be nearly instantaneous at the storage layer. This portable lab instead creates a tar copy while writes are locked so the learner can practice the consistency and restore checks without privileged LVM commands. Do not generalize the tar command into a production snapshot mechanism.
rm -f /tmp/atlasmart-ch25-l2-snapshot.tgz /tmp/atlasmart-ch25-l2-snapshot.sha256mongosh "mongodb://127.0.0.1:27193/admin?directConnection=true" --quiet --eval 'printjson(db.fsyncLock())'# No application writes are allowed while this copy runs.docker exec atlasmart-ch25-l2 sh -lc 'tar czf - -C /data/db .' > /tmp/atlasmart-ch25-l2-snapshot.tgzsha256sum /tmp/atlasmart-ch25-l2-snapshot.tgz | tee /tmp/atlasmart-ch25-l2-snapshot.sha256mongosh "mongodb://127.0.0.1:27193/admin?directConnection=true" --quiet --eval 'printjson(db.fsyncUnlock())'
fsyncLock blocks writes on this disposable node.
Never run the lab URI against a production system. On a
production replica set, design the backup member and client
routing so snapshot operations do not unexpectedly stop the
application’s only writable primary.
4. Prove that the snapshot is a point, not the future
const d=db.getSiblingDB("atlasmart");d.inventory_ch25_l2.insertOne({sku:"POST-SNAPSHOT",qty:999,updatedAt:new Date(),note:"must not appear in restored snapshot"});printjson({liveCount:d.inventory_ch25_l2.countDocuments(),postMarker:d.inventory_ch25_l2.countDocuments({sku:"POST-SNAPSHOT"})});
A physical snapshot contains the state at its capture boundary.
To recover to a later point, the recovery system needs an
independently retained operation history that starts no later
than the snapshot’s oplog/commit boundary and continues through
the desired target. The local.oplog.rs stored
inside the snapshot cannot contain operations that happened
after the snapshot was taken.
5. Restore the physical artifact on a disposable replacement
docker rm -f atlasmart-ch25-l2# Preserve the old volume only until the drill succeeds; this is disposable training data.docker volume create atlasmart-ch25-l2-restore-db >/dev/nulldocker run --rm -i -v atlasmart-ch25-l2-restore-db:/restore alpine:3.22.5 sh -lc 'cd /restore && tar xzf - && rm -f mongod.lock' < /tmp/atlasmart-ch25-l2-snapshot.tgzdocker run -d --name atlasmart-ch25-l2 \ -p 127.0.0.1:27193:27017 \ -v atlasmart-ch25-l2-restore-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --replSet rs25l2 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27193/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27193/atlasmart?directConnection=true" --quiet --eval ' printjson({count:db.inventory_ch25_l2.countDocuments(),postMarker:db.inventory_ch25_l2.countDocuments({sku:"POST-SNAPSHOT"})});'
The expected postMarker count is zero. Exact
startup/recovery log messages are version-dependent; record them
locally instead of copying a canned log. For Enterprise
encrypted storage, the restored files still depend on the
external master key/KMIP material; Chapter 23’s key-recovery
drill remains part of this recovery plan. AES256-GCM also has
specific hot/cold backup key-rollover considerations in the
current encryption-at-rest documentation.
Check your understanding
-
Why is a raw copy of a live
dbPathunsafe without a consistency mechanism? - Why record the replica-set optime near the snapshot boundary?
- Does the oplog inside the snapshot contain operations performed after the snapshot?
- Why is the lab tar copy not called an atomic volume snapshot?
- What extra dependency exists for encrypted storage?
Review the answers
1. The copy can observe files from different moments or miss accepted writes that are not represented in a coherent recoverable storage state.
2. It identifies the database replication position associated with the physical recovery point and is needed to reason about later oplog/PIT catch-up.
3. No. Later recovery requires separately retained history.
4. It is a file copy protected by a database lock; real snapshots are implemented by the storage layer and have different atomicity/volume semantics.
5. The required external master/KMS key material and its recovery/rotation procedure must survive independently of the data snapshot.
6. Production judgment
Physical snapshots are attractive when data volume makes logical scanning expensive, but the backup design belongs to the storage topology: which volumes contain data/journal, whether the platform can snapshot them atomically, which replica-set member is safe to pause, where snapshot copies are stored, and how oplog history extends the recovery point. Keep snapshots off the same failure domain as the source. The next lesson moves from self-managed mechanics to Atlas-managed snapshot and continuous PIT policy.
Authoritative references
Backup behavior is topology-, tool-, storage-, and service-version sensitive. Re-check the current server, Database Tools, Atlas, encryption, and restore documentation before adopting a production procedure.
- Backup Methods for Self-Managed Deployments
- Back Up and Restore with MongoDB Tools
- mongodump 100.18.0
- mongorestore 100.18.0
- mongodump Examples
- mongorestore Behavior, Access, and Usage
- Database Tools Release Notes
- Filesystem Snapshots
- db.fsyncLock()
- Replication Oplog
- Backup and Restore Self-Managed Sharded Clusters
- Back Up Sharded Cluster with Database Dumps
- Restore Sharded Cluster from Database Dumps
- Config Servers
- Atlas Backup Architecture Guidance
- Atlas Continuous Cloud Backup Restore
- Atlas Backup Policy
- Atlas Disaster Recovery Guidance
- Encryption at Rest
- Queryable Encryption Key Management
- MongoDB 8.3 Release Notes
- mongosh Changelog
- PyMongo Release Notes