Measure recovery as a service objective: newest valid recovered event, clean-target readiness, application invariants, and safe cutover.

Perform a Restore Drill and Measure RPO, RTO, Integrity, Index Rebuild, and Application Cutover

Run a measured MongoDB restore drill with clean infrastructure, explicit RPO/RTO, checksums, business invariants, index readiness, and reversible application cutover.

Advanced120–220 minutesRPO/RTO restore acceptance drillMongoDB 8.3.8 · Database Tools 100.18.0 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Run a clean-target restore drill rather than assuming a successful backup job is recoverable.

02

Measure RPO from the incident boundary to the newest recovered business event and RTO from recovery start through validated readiness.

03

Validate documents, indexes, checksums, and cross-collection invariants before cutover.

04

Measure ordinary-index restore/build readiness separately from Search or external derived-index catch-up.

05

Use canary reads/writes and reconciliation evidence to make application cutover reversible.

Reproducible lab baseline

This chapter pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, PyMongo 4.17.0 where application checks are useful, and MongoDB Database Tools 100.18.0 for mongodump, mongorestore, and bsondump. Lesson 5 uses a disposable one-member replica-set source and a clean standalone restore target. Host exposure is loopback-only on 127.0.0.1:27200–27201. Authentication and TLS are disabled only for disposable local labs; production security remains the Chapter 22 prerequisite. Default read/write concern and primary read preference are used unless a step states otherwise. FCV is observed and never changed. Atlas, Enterprise Advanced, Search, Vector Search, and KMS are optional unless a lesson explicitly labels them. The drill uses logical backup because it is portable on a learner machine; the same acceptance framework applies to physical/Atlas restores. All timers are learner-measured wall-clock evidence, not benchmark claims. Runtime backup/restore labs were not executed in the generation environment, so artifact sizes, restore durations, RPO/RTO measurements, snapshot times, and checksums must be recorded locally rather than copied as invented output.

1. AtlasMart incident: backup success is not the acceptance criterion

At 09:42 an application bug begins corrupting paid orders. Operations notices at 09:51. The last tested backup finished at 09:35. The question is not “did backup run?” but “what clean point can be recovered, how much valid work will be lost, and how long until the application can safely serve traffic?”

For this drill, RPO is measured from the incident cut line to the newest business event present after recovery. RTO starts when the recovery procedure begins and ends only after restore, index readiness, integrity checks, canary traffic, and cutover readiness all pass.

2. Create a source with deterministic business events

start source and restore target
docker rm -f atlasmart-ch25-l5-src atlasmart-ch25-l5-restore 2>/dev/null || truedocker volume rm atlasmart-ch25-l5-src-db atlasmart-ch25-l5-restore-db 2>/dev/null || truedocker run -d --name atlasmart-ch25-l5-src \  -p 127.0.0.1:27200:27017 -v atlasmart-ch25-l5-src-db:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --replSet rs25l5 --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27200/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27200/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"rs25l5",members:[{_id:0,host:"atlasmart-ch25-l5-src:27017"}]})'until mongosh "mongodb://127.0.0.1:27200/admin?directConnection=true" --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; done
seed business events with an invariant and indexes
const d=db.getSiblingDB("atlasmart");d.orders_ch25_l5.drop(); d.payments_ch25_l5.drop();d.orders_ch25_l5.createIndex({tenantId:1,createdAt:-1},{name:"idx_tenant_created"});d.payments_ch25_l5.createIndex({orderId:1},{name:"uq_payment_order",unique:true});for (let i=0;i<3000;i++) {  const t=new Date(Date.UTC(2026,8,3,9,0,0)+i*500);  const id=`O-${String(i).padStart(6,"0")}`;  const cents=1000+(i%10000);  d.orders_ch25_l5.insertOne({orderId:id,tenantId:`tenant-${i%6}`,status:"paid",totalCents:cents,createdAt:t});  d.payments_ch25_l5.insertOne({paymentId:`P-${String(i).padStart(6,"0")}`,orderId:id,amountCents:cents,status:"captured",capturedAt:t});}printjson({orders:d.orders_ch25_l5.countDocuments(),payments:d.payments_ch25_l5.countDocuments(),latest:d.orders_ch25_l5.find().sort({createdAt:-1}).limit(1).next().createdAt});

3. Take the recovery artifact and record a manifest

create backup and manifest timing
rm -rf /tmp/atlasmart-ch25-l5-dump /tmp/atlasmart-ch25-l5-manifestmkdir -p /tmp/atlasmart-ch25-l5-manifestBACKUP_START_EPOCH="$(date +%s)"mongodump --uri="mongodb://127.0.0.1:27200/?directConnection=true" --oplog --out=/tmp/atlasmart-ch25-l5-dumpBACKUP_END_EPOCH="$(date +%s)"find /tmp/atlasmart-ch25-l5-dump -type f -printf '%P %s bytes\n' | sort | tee /tmp/atlasmart-ch25-l5-manifest/files.txtfind /tmp/atlasmart-ch25-l5-dump -type f -print0 | sort -z | xargs -0 sha256sum > /tmp/atlasmart-ch25-l5-manifest/sha256.txtprintf 'backup_start_epoch=%s\nbackup_end_epoch=%s\n' "$BACKUP_START_EPOCH" "$BACKUP_END_EPOCH" > /tmp/atlasmart-ch25-l5-manifest/times.txt

After backup completion, insert additional synthetic orders and then mark a chosen incident timestamp. Those later records let the drill measure data that the backup cannot contain instead of assuming RPO from schedule alone.

create post-backup business activity and an incident marker
const d=db.getSiblingDB("atlasmart");for (let i=3000;i<3030;i++) {  const t=new Date(Date.UTC(2026,8,3,9,40,0)+(i-3000)*1000);  const id=`O-${String(i).padStart(6,"0")}`; const cents=5000+i;  d.orders_ch25_l5.insertOne({orderId:id,tenantId:"tenant-drill",status:"paid",totalCents:cents,createdAt:t});  d.payments_ch25_l5.insertOne({paymentId:`P-${String(i).padStart(6,"0")}`,orderId:id,amountCents:cents,status:"captured",capturedAt:t});}d.dr_events_ch25_l5.insertOne({_id:"incident",incidentAt:new Date(Date.UTC(2026,8,3,9,42,0))});printjson(d.orders_ch25_l5.find().sort({createdAt:-1}).limit(3).toArray());

4. Restore to clean infrastructure and measure RTO stages

time restore separately from application validation
docker run -d --name atlasmart-ch25-l5-restore \  -p 127.0.0.1:27201:27017 -v atlasmart-ch25-l5-restore-db:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --bind_ip_all --port 27017until mongosh "mongodb://127.0.0.1:27201/admin" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; doneDR_START_EPOCH="$(date +%s)"RESTORE_START_EPOCH="$(date +%s)"mongorestore --uri="mongodb://127.0.0.1:27201/" --oplogReplay /tmp/atlasmart-ch25-l5-dumpRESTORE_END_EPOCH="$(date +%s)"printf 'dr_start=%s\nrestore_start=%s\nrestore_end=%s\n' "$DR_START_EPOCH" "$RESTORE_START_EPOCH" "$RESTORE_END_EPOCH"
validate checksum-like state, indexes, and business invariants
from pymongo import MongoClientimport hashlib, json, timec=MongoClient("mongodb://127.0.0.1:27201/")db=c.atlasmartdef digest(coll,key):    h=hashlib.sha256()    for doc in db[coll].find({}, {"_id":0}).sort(key,1):        h.update(json.dumps(doc, sort_keys=True, default=str, separators=(",", ":")).encode())    return h.hexdigest()bad=list(db.orders_ch25_l5.aggregate([  {"$match":{"status":"paid"}},  {"$lookup":{"from":"payments_ch25_l5","localField":"orderId","foreignField":"orderId","as":"p"}},  {"$match":{"$expr":{"$or":[{"$ne":[{"$size":"$p"},1]},{"$ne":["$totalCents",{"$arrayElemAt":["$p.amountCents",0]}]}]}}},  {"$limit":10}]))latest=db.orders_ch25_l5.find_one(sort=[("createdAt",-1)])report={  "orders":db.orders_ch25_l5.count_documents({}),  "payments":db.payments_ch25_l5.count_documents({}),  "latestRecoveredEvent":latest["createdAt"].isoformat(),  "ordersHash":digest("orders_ch25_l5","orderId"),  "paymentsHash":digest("payments_ch25_l5","paymentId"),  "orderIndexes":[x["name"] for x in db.orders_ch25_l5.list_indexes()],  "badInvariantRows":len(bad),}print(report)assert report["badInvariantRows"] == 0assert "idx_tenant_created" in report["orderIndexes"]

Record ordinary index readiness as part of the restore time. If Atlas Search/Vector Search is part of the application, its derived index data has a separate rebuild/synchronization readiness stage. Database-document recovery alone must not end the RTO timer for a search-dependent service.

5. Calculate measured RPO and complete a reversible cutover

Use the newest recovered business timestamp and the incident cut line. Do not use the backup job schedule as a substitute for evidence. In this synthetic data set the incident cut is 2026-09-03T09:42:00Z; the recovered newest order should be from the pre-backup set, so the local script can calculate the actual gap.

calculate RPO from recovered business data
from datetime import datetime, timezonefrom pymongo import MongoClientc=MongoClient("mongodb://127.0.0.1:27201/")latest=c.atlasmart.orders_ch25_l5.find_one(sort=[("createdAt",-1)])["createdAt"]incident=datetime(2026,9,3,9,42,0,tzinfo=timezone.utc)if latest.tzinfo is None:    latest=latest.replace(tzinfo=timezone.utc)print({"incident":incident.isoformat(),"latestRecovered":latest.isoformat(),"rpoSeconds":(incident-latest).total_seconds()})
application readiness and cutover canary
const d=db.getSiblingDB("atlasmart");const health={  orders:d.orders_ch25_l5.countDocuments(),  payments:d.payments_ch25_l5.countDocuments(),  canReadPaid:d.orders_ch25_l5.findOne({status:"paid"}) !== null};printjson(health);d.cutover_canary_ch25_l5.replaceOne({_id:"drill"},{_id:"drill",checkedAt:new Date(),target:"restore"},{upsert:true});printjson(d.cutover_canary_ch25_l5.findOne({_id:"drill"}));

In production, cutover usually changes a service discovery record, secret, load-balancer target, or connection configuration—not application documents. Keep the old environment isolated but available for forensic comparison until reconciliation is complete. Make the cutover reversible, and prevent dual writers unless the recovery design explicitly supports them.

6. Failure drill: “restore succeeded” but application integrity fails

Create a separate bad artifact containing only orders, restore it to another disposable database, and run the same invariant check. The restore process can return success while every order has no payment. This separates infrastructure success from business recovery success and is exactly why restore drills need application-owned validation.

Also test missing key material as a separate exercise from Chapter 23. Never delete a real KMS key to test disaster recovery; simulate denied/wrong disposable key material and prove the runbook can distinguish “data restored” from “data decryptable.”

Check your understanding

  1. When does the RTO timer stop?
  2. How is measured RPO calculated?
  3. Why keep application invariants in the restore test suite?
  4. Why can Search increase RTO after an Atlas restore?
  5. What makes cutover safer?
Review the answers

1. Only after restore plus required index/derived-system readiness, integrity checks, canary traffic, and cutover readiness—not when mongorestore exits.

2. From the incident/recovery target to the newest valid business event actually present after recovery.

3. Infrastructure-level success can still restore logically inconsistent or incomplete application state.

4. Search index definitions may be restored while mongot still needs to rebuild/synchronize index data.

5. A clean target, explicit health/invariant gates, canary reads/writes, no unintended dual writers, and a reversible routing/configuration change.

7. Production judgment

Backup quality is measured at restore time. Track backup age, restore-test age, artifact integrity, key availability, RPO, RTO, index/search readiness, and application reconciliation as operational signals. Run drills on clean infrastructure, deliberately break one dependency at a time, and keep evidence. This chapter’s final lesson turns recovery from a file-retention task into a tested service capability; Chapter 26 now uses those recovery assumptions as part of monitoring, capacity, incident response, and upgrade planning.

Authoritative references

Backup behavior is topology-, tool-, storage-, and service-version sensitive. Re-check the current server, Database Tools, Atlas, encryption, and restore documentation before adopting a production procedure.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.