Treat logical export as one recovery mechanism—not as a substitute for physical snapshots, PIT history, key recovery, or restore testing.

Logical Export/Import vs Physical Backup: mongodump / mongorestore Boundaries and Use Cases

Compare logical MongoDB dumps with physical backup, create a consistent replica-set dump with oplog coverage, restore it, and validate application invariants.

Advanced120–220 minutesBackup artifact and logical-restore labMongoDB 8.3.8 · Database Tools 100.18.0 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Separate logical BSON export/import from physical storage backup and explain when each is appropriate.

02

Use Database Tools 100.18.0 with a pinned MongoDB 8.3.8 replica-set lab and capture a backup manifest.

03

Explain exactly what mongodump --oplog captures—and why it is not a continuous PIT archive.

04

Restore into a clean target, verify counts/indexes/business invariants, and detect a partial or inconsistent recovery.

05

Carry encryption/key dependencies, topology, FCV, and restore-testing requirements into the backup design.

Reproducible lab baseline

This chapter pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, PyMongo 4.17.0 where application checks are useful, and MongoDB Database Tools 100.18.0 for mongodump, mongorestore, and bsondump. Lesson 1 uses a disposable one-member replica set as the source so an oplog exists, plus a clean standalone restore target. Host exposure is loopback-only on 127.0.0.1:27191–27192. Authentication and TLS are disabled only for disposable local labs; production security remains the Chapter 22 prerequisite. Default read/write concern and primary read preference are used unless a step states otherwise. FCV is observed and never changed. Atlas, Enterprise Advanced, Search, Vector Search, and KMS are optional unless a lesson explicitly labels them. Database Tools run on the host and are verified with mongodump --version; they release independently from the server. Runtime backup/restore labs were not executed in the generation environment, so artifact sizes, restore durations, RPO/RTO measurements, snapshot times, and checksums must be recorded locally rather than copied as invented output.

1. AtlasMart problem: a successful dump is not yet a recovery system

AtlasMart can copy its orders collection every night and still fail recovery. A logical backup reads documents, collection options, and index definitions through MongoDB interfaces and serializes them as BSON. A physical backup captures the storage files or block device containing WiredTiger state. The two solve different operational problems: logical dumps are portable and selective, while physical snapshots can be much faster for large data sets and preserve storage-engine state as a unit.

Neither method is automatically a disaster-recovery system. Recovery also needs a defined Recovery Point Objective (RPO)—how much recent data loss the business can tolerate—and Recovery Time Objective (RTO)—how long service restoration may take. A backup that cannot be restored within the RTO, lacks encryption keys, or restores documents that violate application invariants is not sufficient.

Mechanism Captures Strength Boundary
mongodump BSON documents, metadata/options, index definitions Portable and selective Reads through a running server; can disturb cache and is not a physical image
mongodump --oplog Full dump plus writes that occur during that dump Creates a consistent replica-set dump while writes continue Not a continuous history after the dump finishes
Filesystem / block snapshot Data/journal blocks at a coordinated instant Fast for large deployments when storage supports it Requires snapshot consistency, storage coordination, key material, and off-system retention
Atlas continuous backup Managed snapshots plus retained oplog history PIT recovery inside configured window Paid managed capability with tier/version/region constraints

2. Verify the toolchain before creating evidence

MongoDB Database Tools use an independent release train. This lesson pins 100.18.0, released in August 2026 and compatible with MongoDB 8.3. Do not assume a server package automatically installed the matching tools.

verify server and Database Tools versions
mongodump --versionmongorestore --versionbsondump --versiondocker rm -f atlasmart-ch25-l1-src atlasmart-ch25-l1-restore 2>/dev/null || truedocker volume rm atlasmart-ch25-l1-src-db atlasmart-ch25-l1-restore-db 2>/dev/null || truedocker run -d --name atlasmart-ch25-l1-src \  -p 127.0.0.1:27191:27017 \  -v atlasmart-ch25-l1-src-db:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --replSet rs25l1 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27191/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27191/admin?directConnection=true" --quiet --eval '  rs.initiate({_id:"rs25l1",members:[{_id:0,host:"atlasmart-ch25-l1-src:27017"}]});'until mongosh "mongodb://127.0.0.1:27191/admin?directConnection=true" --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; donemongosh "mongodb://127.0.0.1:27191/admin?directConnection=true" --quiet --eval '  printjson(db.adminCommand({buildInfo:1}).version);  printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}));'

The one-member replica set exists only to expose oplog semantics. It has no high availability and does not model majority durability across independent failure domains.

3. Seed a cross-collection invariant before the dump

AtlasMart will validate more than counts. Every paid order must have exactly one payment with the same amount. That business invariant makes a partial restore visible even when individual collections look healthy.

seed orders, payments, and indexes
const d=db.getSiblingDB("atlasmart");d.orders_ch25_l1.drop(); d.payments_ch25_l1.drop();d.orders_ch25_l1.createIndex({tenantId:1,createdAt:-1},{name:"idx_tenant_created"});d.payments_ch25_l1.createIndex({orderId:1},{name:"uq_payment_order",unique:true});const orders=[]; const payments=[];for (let i=0;i<5000;i++) {  const orderId=`O-${String(i).padStart(6,"0")}`;  const totalCents=1000+(i%19000);  orders.push({orderId,tenantId:`tenant-${i%8}`,status:"paid",totalCents,createdAt:new Date(Date.UTC(2026,8,3,6,0,i%60))});  payments.push({paymentId:`P-${String(i).padStart(6,"0")}`,orderId,amountCents:totalCents,status:"captured"});}d.orders_ch25_l1.insertMany(orders);d.payments_ch25_l1.insertMany(payments);printjson({orders:d.orders_ch25_l1.countDocuments(),payments:d.payments_ch25_l1.countDocuments()});printjson(d.orders_ch25_l1.getIndexes().map(x=>({name:x.name,key:x.key})));
calculate a deterministic application-level checksum
from pymongo import MongoClientimport hashlib, jsonc=MongoClient("mongodb://127.0.0.1:27191/?directConnection=true")db=c.atlasmartdef digest(name):    h=hashlib.sha256()    for doc in db[name].find({}, {"_id":0}).sort([(k,1) for k in (["orderId"] if name.startswith("orders") else ["paymentId"]) ]):        h.update(json.dumps(doc, sort_keys=True, default=str, separators=(",", ":")).encode())    return h.hexdigest()print({"orders":digest("orders_ch25_l1"),"payments":digest("payments_ch25_l1")})

4. Take a consistent logical dump and inventory the artifact

Without --oplog, writes that occur while a full replica-set dump is running can leave collections representing different logical moments. With --oplog, the tools capture the oplog entries generated during the dump and mongorestore --oplogReplay applies them after restoring the BSON files. The captured history ends when the dump ends; it does not include later writes and therefore is not continuous point-in-time recovery.

create a full replica-set dump with oplog coverage
rm -rf /tmp/atlasmart-ch25-l1-dumpmkdir -p /tmp/atlasmart-ch25-l1-dumpSTART_UTC="$(date -u +%Y-%m-%dT%H:%M:%SZ)"mongodump \  --uri="mongodb://127.0.0.1:27191/?directConnection=true" \  --oplog \  --out=/tmp/atlasmart-ch25-l1-dumpEND_UTC="$(date -u +%Y-%m-%dT%H:%M:%SZ)"find /tmp/atlasmart-ch25-l1-dump -type f -printf '%P %s bytes\n' | sortsha256sum /tmp/atlasmart-ch25-l1-dump/oplog.bsonprintf 'dump_start=%s\ndump_end=%s\n' "$START_UTC" "$END_UTC" | tee /tmp/atlasmart-ch25-l1-dump/MANIFEST.times
Boundary:

--oplogReplay requires the full dump. Do not combine it with namespace-limiting restore options. Also, Database Tools 100.18.0 reject an oplog replay that spans the MongoDB 8.x→9.0 time-series format conversion. This course keeps FCV fixed during backup.

5. Restore into a clean target and prove application integrity

restore to a disposable clean target
docker run -d --name atlasmart-ch25-l1-restore \  -p 127.0.0.1:27192:27017 \  -v atlasmart-ch25-l1-restore-db:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27192/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; doneRESTORE_START="$(date +%s)"mongorestore --uri="mongodb://127.0.0.1:27192/" \  --oplogReplay \  /tmp/atlasmart-ch25-l1-dumpRESTORE_END="$(date +%s)"echo "restore_seconds=$((RESTORE_END-RESTORE_START))"
verify counts, indexes, and the paid-order/payment invariant
const d=db.getSiblingDB("atlasmart");const bad=d.orders_ch25_l1.aggregate([  {$match:{status:"paid"}},  {$lookup:{from:"payments_ch25_l1",localField:"orderId",foreignField:"orderId",as:"p"}},  {$match:{$expr:{$or:[{$ne:[{$size:"$p"},1]},{$ne:["$totalCents",{$arrayElemAt:["$p.amountCents",0]}]}]}}},  {$limit:10}]).toArray();printjson({orders:d.orders_ch25_l1.countDocuments(),payments:d.payments_ch25_l1.countDocuments(),badInvariantRows:bad.length});printjson(d.orders_ch25_l1.getIndexes().map(x=>({name:x.name,key:x.key})));

A restore log ending successfully is only transport evidence. The acceptance criteria are the restored record counts, index definitions, deterministic checksums, and business invariants. If the restore crashes midway, do not continue from a partially restored target as if it were known-good; recreate the disposable target and rerun the recovery from the beginning.

6. Failure case: a selective dump can look healthy and still be unusable

A common “backup” is a dump of only the most visible collection. Restore only orders_ch25_l1 into a disposable database and the order count can be perfect while every paid-order/payment invariant fails. This is why backup scope must be derived from application recovery boundaries rather than from storage size or collection popularity.

Queryable Encryption introduces an additional boundary: current mongorestore cannot restore a Queryable Encryption collection. Client-Side Field Level Encryption ciphertext also remains useless without the external key material needed to decrypt it. The data artifact and its key-management recovery path must be tested together.

Check your understanding

  1. What problem does mongodump --oplog solve?
  2. Does the resulting oplog.bson provide continuous PIT recovery after the dump finishes?
  3. Why are collection counts insufficient restore evidence?
  4. When is a physical snapshot often preferable to a logical dump?
  5. What must be versioned with the backup besides data?
Review the answers

1. It captures writes that occur during a full replica-set dump so replay can make the dump consistent through the dump completion boundary.

2. No. It ends with the dump; later writes require a separate continuous oplog/PIT mechanism.

3. Counts can match while cross-collection relationships, totals, uniqueness, indexes, or encryption/key dependencies are wrong.

4. For large deployments where block-level capture can meet the recovery objective with lower read/cache disruption, assuming snapshot consistency is engineered.

5. Tool/server/FCV/topology assumptions, manifests, restore procedure, application invariant tests, and required encryption/key material.

7. Production judgment

Use logical dumps when portability, selective recovery, or a small deployment makes their read cost acceptable. Use coordinated filesystem/block snapshots or managed backup for larger recovery systems. Do not equate replication with backup: an accidental delete is replicated. Treat every backup as untrusted until a clean-target restore proves schema, indexes, business invariants, security dependencies, RPO, and RTO. The next lesson moves below logical BSON and shows what a physically consistent snapshot requires.

Authoritative references

Backup behavior is topology-, tool-, storage-, and service-version sensitive. Re-check the current server, Database Tools, Atlas, encryption, and restore documentation before adopting a production procedure.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.