Measure recovery as a service objective: newest valid recovered event, clean-target readiness, application invariants, and safe cutover.
Perform a Restore Drill and Measure RPO, RTO, Integrity, Index Rebuild, and Application Cutover
Run a measured MongoDB restore drill with clean infrastructure, explicit RPO/RTO, checksums, business invariants, index readiness, and reversible application cutover.
Learning objectives
Run a clean-target restore drill rather than assuming a successful backup job is recoverable.
Measure RPO from the incident boundary to the newest recovered business event and RTO from recovery start through validated readiness.
Validate documents, indexes, checksums, and cross-collection invariants before cutover.
Measure ordinary-index restore/build readiness separately from Search or external derived-index catch-up.
Use canary reads/writes and reconciliation evidence to make application cutover reversible.
This chapter pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0,
PyMongo 4.17.0 where application checks are useful,
and MongoDB Database Tools 100.18.0 for
mongodump, mongorestore, and
bsondump. Lesson 5 uses a disposable one-member
replica-set source and a clean standalone restore target. Host
exposure is loopback-only on 127.0.0.1:27200–27201.
Authentication and TLS are disabled only for disposable local
labs; production security remains the Chapter 22 prerequisite.
Default read/write concern and primary read preference are used
unless a step states otherwise.
FCV is observed and never changed. Atlas,
Enterprise Advanced, Search, Vector Search, and KMS are optional
unless a lesson explicitly labels them. The drill uses logical
backup because it is portable on a learner machine; the same
acceptance framework applies to physical/Atlas restores. All
timers are learner-measured wall-clock evidence, not benchmark
claims. Runtime backup/restore labs were not executed in the
generation environment, so artifact sizes, restore durations,
RPO/RTO measurements, snapshot times, and checksums must be
recorded locally rather than copied as invented output.
1. AtlasMart incident: backup success is not the acceptance criterion
At 09:42 an application bug begins corrupting paid orders. Operations notices at 09:51. The last tested backup finished at 09:35. The question is not “did backup run?” but “what clean point can be recovered, how much valid work will be lost, and how long until the application can safely serve traffic?”
For this drill, RPO is measured from the incident cut line to the newest business event present after recovery. RTO starts when the recovery procedure begins and ends only after restore, index readiness, integrity checks, canary traffic, and cutover readiness all pass.
2. Create a source with deterministic business events
docker rm -f atlasmart-ch25-l5-src atlasmart-ch25-l5-restore 2>/dev/null || truedocker volume rm atlasmart-ch25-l5-src-db atlasmart-ch25-l5-restore-db 2>/dev/null || truedocker run -d --name atlasmart-ch25-l5-src \ -p 127.0.0.1:27200:27017 -v atlasmart-ch25-l5-src-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --replSet rs25l5 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27200/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27200/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"rs25l5",members:[{_id:0,host:"atlasmart-ch25-l5-src:27017"}]})'until mongosh "mongodb://127.0.0.1:27200/admin?directConnection=true" --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; done
const d=db.getSiblingDB("atlasmart");d.orders_ch25_l5.drop(); d.payments_ch25_l5.drop();d.orders_ch25_l5.createIndex({tenantId:1,createdAt:-1},{name:"idx_tenant_created"});d.payments_ch25_l5.createIndex({orderId:1},{name:"uq_payment_order",unique:true});for (let i=0;i<3000;i++) { const t=new Date(Date.UTC(2026,8,3,9,0,0)+i*500); const id=`O-${String(i).padStart(6,"0")}`; const cents=1000+(i%10000); d.orders_ch25_l5.insertOne({orderId:id,tenantId:`tenant-${i%6}`,status:"paid",totalCents:cents,createdAt:t}); d.payments_ch25_l5.insertOne({paymentId:`P-${String(i).padStart(6,"0")}`,orderId:id,amountCents:cents,status:"captured",capturedAt:t});}printjson({orders:d.orders_ch25_l5.countDocuments(),payments:d.payments_ch25_l5.countDocuments(),latest:d.orders_ch25_l5.find().sort({createdAt:-1}).limit(1).next().createdAt});
3. Take the recovery artifact and record a manifest
rm -rf /tmp/atlasmart-ch25-l5-dump /tmp/atlasmart-ch25-l5-manifestmkdir -p /tmp/atlasmart-ch25-l5-manifestBACKUP_START_EPOCH="$(date +%s)"mongodump --uri="mongodb://127.0.0.1:27200/?directConnection=true" --oplog --out=/tmp/atlasmart-ch25-l5-dumpBACKUP_END_EPOCH="$(date +%s)"find /tmp/atlasmart-ch25-l5-dump -type f -printf '%P %s bytes\n' | sort | tee /tmp/atlasmart-ch25-l5-manifest/files.txtfind /tmp/atlasmart-ch25-l5-dump -type f -print0 | sort -z | xargs -0 sha256sum > /tmp/atlasmart-ch25-l5-manifest/sha256.txtprintf 'backup_start_epoch=%s\nbackup_end_epoch=%s\n' "$BACKUP_START_EPOCH" "$BACKUP_END_EPOCH" > /tmp/atlasmart-ch25-l5-manifest/times.txt
After backup completion, insert additional synthetic orders and then mark a chosen incident timestamp. Those later records let the drill measure data that the backup cannot contain instead of assuming RPO from schedule alone.
const d=db.getSiblingDB("atlasmart");for (let i=3000;i<3030;i++) { const t=new Date(Date.UTC(2026,8,3,9,40,0)+(i-3000)*1000); const id=`O-${String(i).padStart(6,"0")}`; const cents=5000+i; d.orders_ch25_l5.insertOne({orderId:id,tenantId:"tenant-drill",status:"paid",totalCents:cents,createdAt:t}); d.payments_ch25_l5.insertOne({paymentId:`P-${String(i).padStart(6,"0")}`,orderId:id,amountCents:cents,status:"captured",capturedAt:t});}d.dr_events_ch25_l5.insertOne({_id:"incident",incidentAt:new Date(Date.UTC(2026,8,3,9,42,0))});printjson(d.orders_ch25_l5.find().sort({createdAt:-1}).limit(3).toArray());
4. Restore to clean infrastructure and measure RTO stages
docker run -d --name atlasmart-ch25-l5-restore \ -p 127.0.0.1:27201:27017 -v atlasmart-ch25-l5-restore-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27201/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; doneDR_START_EPOCH="$(date +%s)"RESTORE_START_EPOCH="$(date +%s)"mongorestore --uri="mongodb://127.0.0.1:27201/" --oplogReplay /tmp/atlasmart-ch25-l5-dumpRESTORE_END_EPOCH="$(date +%s)"printf 'dr_start=%s\nrestore_start=%s\nrestore_end=%s\n' "$DR_START_EPOCH" "$RESTORE_START_EPOCH" "$RESTORE_END_EPOCH"
from pymongo import MongoClientimport hashlib, json, timec=MongoClient("mongodb://127.0.0.1:27201/")db=c.atlasmartdef digest(coll,key): h=hashlib.sha256() for doc in db[coll].find({}, {"_id":0}).sort(key,1): h.update(json.dumps(doc, sort_keys=True, default=str, separators=(",", ":")).encode()) return h.hexdigest()bad=list(db.orders_ch25_l5.aggregate([ {"$match":{"status":"paid"}}, {"$lookup":{"from":"payments_ch25_l5","localField":"orderId","foreignField":"orderId","as":"p"}}, {"$match":{"$expr":{"$or":[{"$ne":[{"$size":"$p"},1]},{"$ne":["$totalCents",{"$arrayElemAt":["$p.amountCents",0]}]}]}}}, {"$limit":10}]))latest=db.orders_ch25_l5.find_one(sort=[("createdAt",-1)])report={ "orders":db.orders_ch25_l5.count_documents({}), "payments":db.payments_ch25_l5.count_documents({}), "latestRecoveredEvent":latest["createdAt"].isoformat(), "ordersHash":digest("orders_ch25_l5","orderId"), "paymentsHash":digest("payments_ch25_l5","paymentId"), "orderIndexes":[x["name"] for x in db.orders_ch25_l5.list_indexes()], "badInvariantRows":len(bad),}print(report)assert report["badInvariantRows"] == 0assert "idx_tenant_created" in report["orderIndexes"]
Record ordinary index readiness as part of the restore time. If Atlas Search/Vector Search is part of the application, its derived index data has a separate rebuild/synchronization readiness stage. Database-document recovery alone must not end the RTO timer for a search-dependent service.
5. Calculate measured RPO and complete a reversible cutover
Use the newest recovered business timestamp and the incident cut
line. Do not use the backup job schedule as a substitute for
evidence. In this synthetic data set the incident cut is
2026-09-03T09:42:00Z; the recovered newest order
should be from the pre-backup set, so the local script can
calculate the actual gap.
from datetime import datetime, timezonefrom pymongo import MongoClientc=MongoClient("mongodb://127.0.0.1:27201/")latest=c.atlasmart.orders_ch25_l5.find_one(sort=[("createdAt",-1)])["createdAt"]incident=datetime(2026,9,3,9,42,0,tzinfo=timezone.utc)if latest.tzinfo is None: latest=latest.replace(tzinfo=timezone.utc)print({"incident":incident.isoformat(),"latestRecovered":latest.isoformat(),"rpoSeconds":(incident-latest).total_seconds()})
const d=db.getSiblingDB("atlasmart");const health={ orders:d.orders_ch25_l5.countDocuments(), payments:d.payments_ch25_l5.countDocuments(), canReadPaid:d.orders_ch25_l5.findOne({status:"paid"}) !== null};printjson(health);d.cutover_canary_ch25_l5.replaceOne({_id:"drill"},{_id:"drill",checkedAt:new Date(),target:"restore"},{upsert:true});printjson(d.cutover_canary_ch25_l5.findOne({_id:"drill"}));
In production, cutover usually changes a service discovery record, secret, load-balancer target, or connection configuration—not application documents. Keep the old environment isolated but available for forensic comparison until reconciliation is complete. Make the cutover reversible, and prevent dual writers unless the recovery design explicitly supports them.
6. Failure drill: “restore succeeded” but application integrity fails
Create a separate bad artifact containing only orders, restore it to another disposable database, and run the same invariant check. The restore process can return success while every order has no payment. This separates infrastructure success from business recovery success and is exactly why restore drills need application-owned validation.
Also test missing key material as a separate exercise from Chapter 23. Never delete a real KMS key to test disaster recovery; simulate denied/wrong disposable key material and prove the runbook can distinguish “data restored” from “data decryptable.”
Check your understanding
- When does the RTO timer stop?
- How is measured RPO calculated?
- Why keep application invariants in the restore test suite?
- Why can Search increase RTO after an Atlas restore?
- What makes cutover safer?
Review the answers
1. Only after restore plus required index/derived-system readiness, integrity checks, canary traffic, and cutover readiness—not when mongorestore exits.
2. From the incident/recovery target to the newest valid business event actually present after recovery.
3. Infrastructure-level success can still restore logically inconsistent or incomplete application state.
4. Search index definitions may be restored while mongot still needs to rebuild/synchronize index data.
5. A clean target, explicit health/invariant gates, canary reads/writes, no unintended dual writers, and a reversible routing/configuration change.
7. Production judgment
Backup quality is measured at restore time. Track backup age, restore-test age, artifact integrity, key availability, RPO, RTO, index/search readiness, and application reconciliation as operational signals. Run drills on clean infrastructure, deliberately break one dependency at a time, and keep evidence. This chapter’s final lesson turns recovery from a file-retention task into a tested service capability; Chapter 26 now uses those recovery assumptions as part of monitoring, capacity, incident response, and upgrade planning.
Authoritative references
Backup behavior is topology-, tool-, storage-, and service-version sensitive. Re-check the current server, Database Tools, Atlas, encryption, and restore documentation before adopting a production procedure.
- Backup Methods for Self-Managed Deployments
- Back Up and Restore with MongoDB Tools
- mongodump 100.18.0
- mongorestore 100.18.0
- mongodump Examples
- mongorestore Behavior, Access, and Usage
- Database Tools Release Notes
- Filesystem Snapshots
- db.fsyncLock()
- Replication Oplog
- Backup and Restore Self-Managed Sharded Clusters
- Back Up Sharded Cluster with Database Dumps
- Restore Sharded Cluster from Database Dumps
- Config Servers
- Atlas Backup Architecture Guidance
- Atlas Continuous Cloud Backup Restore
- Atlas Backup Policy
- Atlas Disaster Recovery Guidance
- Encryption at Rest
- Queryable Encryption Key Management
- MongoDB 8.3 Release Notes
- mongosh Changelog
- PyMongo Release Notes