Treat logical export as one recovery mechanism—not as a substitute for physical snapshots, PIT history, key recovery, or restore testing.
Logical Export/Import vs Physical Backup: mongodump / mongorestore Boundaries and Use Cases
Compare logical MongoDB dumps with physical backup, create a consistent replica-set dump with oplog coverage, restore it, and validate application invariants.
Learning objectives
Separate logical BSON export/import from physical storage backup and explain when each is appropriate.
Use Database Tools 100.18.0 with a pinned MongoDB 8.3.8 replica-set lab and capture a backup manifest.
Explain exactly what
mongodump --oplog captures—and why it is not a
continuous PIT archive.
Restore into a clean target, verify counts/indexes/business invariants, and detect a partial or inconsistent recovery.
Carry encryption/key dependencies, topology, FCV, and restore-testing requirements into the backup design.
This chapter pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0,
PyMongo 4.17.0 where application checks are useful,
and MongoDB Database Tools 100.18.0 for
mongodump, mongorestore, and
bsondump. Lesson 1 uses a disposable one-member
replica set as the source so an oplog exists, plus a clean
standalone restore target. Host exposure is loopback-only on
127.0.0.1:27191–27192. Authentication and TLS are
disabled only for disposable local labs; production security
remains the Chapter 22 prerequisite. Default read/write concern
and primary read preference are used unless a step states
otherwise.
FCV is observed and never changed. Atlas,
Enterprise Advanced, Search, Vector Search, and KMS are optional
unless a lesson explicitly labels them. Database Tools run on
the host and are verified with mongodump --version;
they release independently from the server. Runtime
backup/restore labs were not executed in the generation
environment, so artifact sizes, restore durations, RPO/RTO
measurements, snapshot times, and checksums must be recorded
locally rather than copied as invented output.
1. AtlasMart problem: a successful dump is not yet a recovery system
AtlasMart can copy its orders collection every
night and still fail recovery. A
logical backup reads documents, collection
options, and index definitions through MongoDB interfaces and
serializes them as BSON. A
physical backup captures the storage files or
block device containing WiredTiger state. The two solve
different operational problems: logical dumps are portable and
selective, while physical snapshots can be much faster for large
data sets and preserve storage-engine state as a unit.
Neither method is automatically a disaster-recovery system. Recovery also needs a defined Recovery Point Objective (RPO)—how much recent data loss the business can tolerate—and Recovery Time Objective (RTO)—how long service restoration may take. A backup that cannot be restored within the RTO, lacks encryption keys, or restores documents that violate application invariants is not sufficient.
| Mechanism | Captures | Strength | Boundary |
|---|---|---|---|
mongodump |
BSON documents, metadata/options, index definitions | Portable and selective | Reads through a running server; can disturb cache and is not a physical image |
mongodump --oplog |
Full dump plus writes that occur during that dump | Creates a consistent replica-set dump while writes continue | Not a continuous history after the dump finishes |
| Filesystem / block snapshot | Data/journal blocks at a coordinated instant | Fast for large deployments when storage supports it | Requires snapshot consistency, storage coordination, key material, and off-system retention |
| Atlas continuous backup | Managed snapshots plus retained oplog history | PIT recovery inside configured window | Paid managed capability with tier/version/region constraints |
2. Verify the toolchain before creating evidence
MongoDB Database Tools use an independent release train. This
lesson pins 100.18.0, released in August 2026 and
compatible with MongoDB 8.3. Do not assume a server package
automatically installed the matching tools.
mongodump --versionmongorestore --versionbsondump --versiondocker rm -f atlasmart-ch25-l1-src atlasmart-ch25-l1-restore 2>/dev/null || truedocker volume rm atlasmart-ch25-l1-src-db atlasmart-ch25-l1-restore-db 2>/dev/null || truedocker run -d --name atlasmart-ch25-l1-src \ -p 127.0.0.1:27191:27017 \ -v atlasmart-ch25-l1-src-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --replSet rs25l1 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27191/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27191/admin?directConnection=true" --quiet --eval ' rs.initiate({_id:"rs25l1",members:[{_id:0,host:"atlasmart-ch25-l1-src:27017"}]});'until mongosh "mongodb://127.0.0.1:27191/admin?directConnection=true" --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; donemongosh "mongodb://127.0.0.1:27191/admin?directConnection=true" --quiet --eval ' printjson(db.adminCommand({buildInfo:1}).version); printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}));'
The one-member replica set exists only to expose oplog semantics. It has no high availability and does not model majority durability across independent failure domains.
3. Seed a cross-collection invariant before the dump
AtlasMart will validate more than counts. Every paid order must have exactly one payment with the same amount. That business invariant makes a partial restore visible even when individual collections look healthy.
const d=db.getSiblingDB("atlasmart");d.orders_ch25_l1.drop(); d.payments_ch25_l1.drop();d.orders_ch25_l1.createIndex({tenantId:1,createdAt:-1},{name:"idx_tenant_created"});d.payments_ch25_l1.createIndex({orderId:1},{name:"uq_payment_order",unique:true});const orders=[]; const payments=[];for (let i=0;i<5000;i++) { const orderId=`O-${String(i).padStart(6,"0")}`; const totalCents=1000+(i%19000); orders.push({orderId,tenantId:`tenant-${i%8}`,status:"paid",totalCents,createdAt:new Date(Date.UTC(2026,8,3,6,0,i%60))}); payments.push({paymentId:`P-${String(i).padStart(6,"0")}`,orderId,amountCents:totalCents,status:"captured"});}d.orders_ch25_l1.insertMany(orders);d.payments_ch25_l1.insertMany(payments);printjson({orders:d.orders_ch25_l1.countDocuments(),payments:d.payments_ch25_l1.countDocuments()});printjson(d.orders_ch25_l1.getIndexes().map(x=>({name:x.name,key:x.key})));
from pymongo import MongoClientimport hashlib, jsonc=MongoClient("mongodb://127.0.0.1:27191/?directConnection=true")db=c.atlasmartdef digest(name): h=hashlib.sha256() for doc in db[name].find({}, {"_id":0}).sort([(k,1) for k in (["orderId"] if name.startswith("orders") else ["paymentId"]) ]): h.update(json.dumps(doc, sort_keys=True, default=str, separators=(",", ":")).encode()) return h.hexdigest()print({"orders":digest("orders_ch25_l1"),"payments":digest("payments_ch25_l1")})
4. Take a consistent logical dump and inventory the artifact
Without --oplog, writes that occur while a full
replica-set dump is running can leave collections representing
different logical moments. With --oplog, the tools
capture the oplog entries generated during the dump and
mongorestore --oplogReplay applies them after
restoring the BSON files. The captured history ends when the
dump ends; it does not include later writes and therefore is not
continuous point-in-time recovery.
rm -rf /tmp/atlasmart-ch25-l1-dumpmkdir -p /tmp/atlasmart-ch25-l1-dumpSTART_UTC="$(date -u +%Y-%m-%dT%H:%M:%SZ)"mongodump \ --uri="mongodb://127.0.0.1:27191/?directConnection=true" \ --oplog \ --out=/tmp/atlasmart-ch25-l1-dumpEND_UTC="$(date -u +%Y-%m-%dT%H:%M:%SZ)"find /tmp/atlasmart-ch25-l1-dump -type f -printf '%P %s bytes\n' | sortsha256sum /tmp/atlasmart-ch25-l1-dump/oplog.bsonprintf 'dump_start=%s\ndump_end=%s\n' "$START_UTC" "$END_UTC" | tee /tmp/atlasmart-ch25-l1-dump/MANIFEST.times
--oplogReplay requires the full dump. Do not
combine it with namespace-limiting restore options. Also,
Database Tools 100.18.0 reject an oplog replay that spans the
MongoDB 8.x→9.0 time-series format conversion. This course
keeps FCV fixed during backup.
5. Restore into a clean target and prove application integrity
docker run -d --name atlasmart-ch25-l1-restore \ -p 127.0.0.1:27192:27017 \ -v atlasmart-ch25-l1-restore-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27192/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; doneRESTORE_START="$(date +%s)"mongorestore --uri="mongodb://127.0.0.1:27192/" \ --oplogReplay \ /tmp/atlasmart-ch25-l1-dumpRESTORE_END="$(date +%s)"echo "restore_seconds=$((RESTORE_END-RESTORE_START))"
const d=db.getSiblingDB("atlasmart");const bad=d.orders_ch25_l1.aggregate([ {$match:{status:"paid"}}, {$lookup:{from:"payments_ch25_l1",localField:"orderId",foreignField:"orderId",as:"p"}}, {$match:{$expr:{$or:[{$ne:[{$size:"$p"},1]},{$ne:["$totalCents",{$arrayElemAt:["$p.amountCents",0]}]}]}}}, {$limit:10}]).toArray();printjson({orders:d.orders_ch25_l1.countDocuments(),payments:d.payments_ch25_l1.countDocuments(),badInvariantRows:bad.length});printjson(d.orders_ch25_l1.getIndexes().map(x=>({name:x.name,key:x.key})));
A restore log ending successfully is only transport evidence. The acceptance criteria are the restored record counts, index definitions, deterministic checksums, and business invariants. If the restore crashes midway, do not continue from a partially restored target as if it were known-good; recreate the disposable target and rerun the recovery from the beginning.
6. Failure case: a selective dump can look healthy and still be unusable
A common “backup” is a dump of only the most visible collection.
Restore only orders_ch25_l1 into a disposable
database and the order count can be perfect while every
paid-order/payment invariant fails. This is why backup scope
must be derived from application recovery boundaries rather than
from storage size or collection popularity.
Queryable Encryption introduces an additional boundary: current
mongorestore cannot restore a Queryable Encryption
collection. Client-Side Field Level Encryption ciphertext also
remains useless without the external key material needed to
decrypt it. The data artifact and its key-management recovery
path must be tested together.
Check your understanding
- What problem does
mongodump --oplogsolve? -
Does the resulting
oplog.bsonprovide continuous PIT recovery after the dump finishes? - Why are collection counts insufficient restore evidence?
- When is a physical snapshot often preferable to a logical dump?
- What must be versioned with the backup besides data?
Review the answers
1. It captures writes that occur during a full replica-set dump so replay can make the dump consistent through the dump completion boundary.
2. No. It ends with the dump; later writes require a separate continuous oplog/PIT mechanism.
3. Counts can match while cross-collection relationships, totals, uniqueness, indexes, or encryption/key dependencies are wrong.
4. For large deployments where block-level capture can meet the recovery objective with lower read/cache disruption, assuming snapshot consistency is engineered.
5. Tool/server/FCV/topology assumptions, manifests, restore procedure, application invariant tests, and required encryption/key material.
7. Production judgment
Use logical dumps when portability, selective recovery, or a small deployment makes their read cost acceptable. Use coordinated filesystem/block snapshots or managed backup for larger recovery systems. Do not equate replication with backup: an accidental delete is replicated. Treat every backup as untrusted until a clean-target restore proves schema, indexes, business invariants, security dependencies, RPO, and RTO. The next lesson moves below logical BSON and shows what a physically consistent snapshot requires.
Authoritative references
Backup behavior is topology-, tool-, storage-, and service-version sensitive. Re-check the current server, Database Tools, Atlas, encryption, and restore documentation before adopting a production procedure.
- Backup Methods for Self-Managed Deployments
- Back Up and Restore with MongoDB Tools
- mongodump 100.18.0
- mongorestore 100.18.0
- mongodump Examples
- mongorestore Behavior, Access, and Usage
- Database Tools Release Notes
- Filesystem Snapshots
- db.fsyncLock()
- Replication Oplog
- Backup and Restore Self-Managed Sharded Clusters
- Back Up Sharded Cluster with Database Dumps
- Restore Sharded Cluster from Database Dumps
- Config Servers
- Atlas Backup Architecture Guidance
- Atlas Continuous Cloud Backup Restore
- Atlas Backup Policy
- Atlas Disaster Recovery Guidance
- Encryption at Rest
- Queryable Encryption Key Management
- MongoDB 8.3 Release Notes
- mongosh Changelog
- PyMongo Release Notes