Chapter 22 · Backups, Point-in-Time Recovery, Restore, and Disaster Recovery

Run a Recovery Drill and Measure Data Completeness, RPO, RTO, Rules / Indexes / Configuration, and Application Cutover

Run an end-to-end AtlasMart recovery drill with objective gates for data completeness, RPO/RTO, security/config parity, application cutover, performance, and rollback.

Advanced · 160–220 minutesrecovery drill · RPO · RTO · validation · cutover · rollbackNode 22+ · Firebase CLI 15.30.0 · Firestore emulator 127.0.0.1:8080Mandatory recovery drill local/no-cost · managed backup/PITR optional/billedLast reviewed: 17 September 2026

1. AtlasMart recovery drill: prove the runbook before the incident

A recovery plan that has never been executed is an assumption. AtlasMart therefore runs a deterministic drill: create a baseline, snapshot the evidence, inject an accidental deletion, restore into an isolated target, deploy missing configuration, validate business behavior, measure RPO/RTO, cut application traffic to the recovery target, then prove a rollback path.

Learning outcomes
  • Execute a full recovery rehearsal from detection through validated cutover.
  • Measure counts, canonical hashes, business-query equivalence, configuration completeness, and timing.
  • Separate data restoration from Rules, indexes, TTL, IAM, App Check, Functions, and secrets recovery.
  • Use objective go/no-go gates for cutover and rollback.
  • Produce an auditable recovery report that identifies emulator evidence versus optional managed-service evidence.
Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 22 reproducibility baseline · reviewed 17 September 2026

AtlasMart keeps the course-wide identity demo-atlasmart-firestore. Mandatory work is local and no-cost: Node.js 22+, Firebase CLI 15.30.0, Firebase Admin Node SDK 14.4.0, Firestore emulator 127.0.0.1:8080, Auth emulator 127.0.0.1:9099, and Emulator UI 127.0.0.1:4000. The normal application database is Standard-edition Native mode, database (default). A second emulator project namespace, demo-atlasmart-recovery, acts as the isolated restore target. Managed scheduled backups, PITR, clone, managed export/import, backup storage, and managed restores require a real billing-enabled project and are therefore optional production-verification exercises, not mandatory lab prerequisites.

Drill rule

The source is never destroyed merely to make the exercise realistic. Mandatory work uses the emulator and a separate recovery project namespace. The optional cloud drill also restores/clones to a separate target first. Destructive same-name restoration is discussed, not required.

2. Recovery acceptance gates

Gate Evidence Fail condition
G0 incident declared Start timestamp, affected scope, writer frozen Unknown ongoing corruption
G1 recovery point chosen Backup/PITR timestamp + rationale Outside retention window or after corruption
G2 target created Distinct database/project identity Accidental overwrite of source
G3 data complete Counts + canonical hashes + sampled field/type checks Unexpected count/hash/type drift
G4 config complete Rules/index/TTL/IAM/App Check/deploy inventory Recovered data served with missing controls
G5 application healthy Smoke tests + critical queries + authz tests Functional/security regression
G6 cutover measured Traffic switch timestamp + p50/p95/p99 during bounded test Tail latency/errors exceed agreed SLO
G7 rollback ready Old target preserved/read-only + reconciliation plan No safe route back after writes resume

3. Step 1 — seed and capture the baseline

seed-drill.mjs
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";import { initializeApp, deleteApp } from "firebase-admin/app";import { getFirestore, Timestamp } from "firebase-admin/firestore";const app = initializeApp({ projectId: "demo-atlasmart-firestore" }, "drill-source");const db = getFirestore(app);const fixture = [  ["catalogItems/p-1001", {name:"Trail Camera",price:99,stock:8,tenantId:"atlasmart",updatedAt:Timestamp.fromDate(new Date("2026-09-17T02:00:00Z"))}],  ["catalogItems/p-1002", {name:"USB-C Hub",price:49,stock:3,tenantId:"atlasmart",updatedAt:Timestamp.fromDate(new Date("2026-09-17T02:00:00Z"))}],  ["orders/o-dr-001", {customerId:"user-a",total:148,status:"PAID",tenantId:"atlasmart"}],  ["profiles/user-a", {displayName:"Atlas Buyer",tenantId:"atlasmart",schemaVersion:1}]];for (const [path,data] of fixture) await db.doc(path).set(data);await deleteApp(app);console.log("seeded", fixture.length);

Run the deterministic snapshot script from Lesson 1 and store data.json, manifest.json, firestore.rules, firestore.indexes.json, and recovery-config.json as the baseline evidence bundle.

4. Step 2 — inject the incident and start the RTO clock

incident.mjs
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";initializeApp({ projectId: process.env.GCLOUD_PROJECT });const db = getFirestore();const incidentStarted = process.hrtime.bigint();await db.doc("catalogItems/p-1001").delete();await db.doc("orders/o-dr-001").update({ total: 14800, incident: "DRILL-22" });console.log(JSON.stringify({ incident:"DRILL-22", declaredAt:new Date().toISOString(), monotonicStart:String(incidentStarted) }));
Why one delete and one corrupt update?

The drill forces validation to catch both absence and wrong-but-present data. A recovery test that checks only document count could miss the corrupted order total.

5. Step 3 — restore into a separate emulator project namespace

restore-fixture.mjs
import fs from "node:fs/promises";import { initializeApp, deleteApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";const rows = JSON.parse(await fs.readFile("recovery-fixture/data.json", "utf8"));const app = initializeApp({ projectId: "demo-atlasmart-recovery" }, "drill-target");const db = getFirestore(app);for (const row of rows) await db.doc(row.path).set(row.data);console.log({ restored: rows.length, target: "demo-atlasmart-recovery" });await deleteApp(app);
Format boundary

This script restores the course’s deterministic JSON fixture. It is not a replacement for firebase firestore:databases:restore or gcloud firestore import, and the file is intentionally not represented as a managed-export artifact.

6. Step 4 — validate counts, hashes, types, and business queries

validate-recovery.mjs
import crypto from "node:crypto";import fs from "node:fs/promises";import { initializeApp, deleteApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";const app = initializeApp({ projectId: "demo-atlasmart-recovery" }, "validate-target");const db = getFirestore(app);const paths = ["catalogItems/p-1001","catalogItems/p-1002","orders/o-dr-001","profiles/user-a"];const rows=[];for (const path of paths) { const s=await db.doc(path).get(); if(!s.exists) throw new Error(`missing ${path}`); rows.push({path,data:s.data()}); }rows.sort((a,b)=>a.path.localeCompare(b.path));const hash=crypto.createHash("sha256").update(JSON.stringify(rows)).digest("hex");const order=rows.find(x=>x.path==="orders/o-dr-001").data;if (order.total !== 148) throw new Error(`business regression: order total=${order.total}`);console.log({count:rows.length, hash, businessQuery:"PASS"});await deleteApp(app);

Because Firestore values such as Timestamp require canonical serialization, production drills should define a stable hash encoding rather than relying on JavaScript object iteration or SDK-specific JSON output. The simple lab hash is illustrative; the recovery report should record the canonicalization contract.

7. Step 5 — configuration parity

Before cutover, compare a configuration manifest. A managed backup restore includes indexes but not TTL policies; an export/import workflow does not carry index definitions at all. Neither data mechanism should be expected to reconstruct application Rules, App Check settings, IAM intent, Functions, secrets, or network controls.

recovery-config.json
{  "firestoreRulesSha256": "RECORD_FROM_REPO",  "indexConfigSha256": "RECORD_FROM_REPO",  "ttlPolicies": ["sessions.expireAt", "events.expireAt"],  "requiredIamRoles": ["application-runtime-role", "restore-operator-role"],  "appCheck": "production-only verification required",  "functionsRelease": "RECORD_DEPLOYMENT_ID",  "secretsVersionSet": "RECORD_SECRET_VERSIONS",  "networkPolicyVersion": "RECORD_POLICY_ID"}

8. Step 6 — cutover and rollback boundary

In the emulator, the “application endpoint” is a configuration file that names the project ID. Switch it from demo-atlasmart-firestore to demo-atlasmart-recovery, restart the test client, and rerun the critical query/authz suite. In production, the equivalent may be a database ID/config rollout, backend deployment, DNS/service-routing change, or client release depending on architecture.

Before target writes resume After target writes resume
Rollback can often be a traffic/config reversal because source remains authoritative/read-only evidence. Rollback becomes a data reconciliation problem because writes may exist only on the recovery target.
Keep source frozen and preserve incident evidence. Record every target-side write or establish reverse replication/replay strategy.
Validate security/config parity first. Require explicit incident commander approval for reversal.

9. Measure RPO/RTO and tail latency without inventing numbers

The drill script should emit timestamps for incident declaration, recovery point, restore start/end, validation pass, configuration pass, cutover start, and first successful production-equivalent request. Compute RPO from the age of the chosen clean data; compute RTO from incident declaration to validated service restoration. If you benchmark requests after cutover, report p50/p95/p99 from measured samples. Emulator percentiles are useful for regression detection but not production SLO proof.

percentiles.mjs
export function percentile(xs, p) {  const a=[...xs].sort((x,y)=>x-y);  if (!a.length) return null;  const i=Math.min(a.length-1, Math.ceil((p/100)*a.length)-1);  return a[i];}// Feed real measured request durations; never insert invented values.console.log({p50:percentile(samples,50),p95:percentile(samples,95),p99:percentile(samples,99)});

10. Optional managed-service drill

In an isolated Blaze/billing-enabled project, choose one mechanism and collect real evidence: create a backup schedule and restore a retained backup to a new database; or enable PITR and clone a clean timestamp. Record backup/clone resource names, locations, timestamps, operation duration, backup/PITR storage cost evidence, effective IAM principal, restored index configuration, missing TTL policy, validation hashes, and cutover timing. Do not run destructive same-name restore merely for practice.

11. Recovery report template

Field Required evidence
Incident ID, start/end, detector, affected collections
Recovery point Mechanism, timestamp, earliestVersionTime/backup resource if relevant
RPO Measured age/data loss relative to incident
RTO Detection/declaration → validated application service
Data integrity Counts, canonical hashes, sampled types, business query regression
Configuration Rules, indexes, TTL, IAM, App Check, Functions, secrets, network
Performance Bounded post-cutover p50/p95/p99 and errors
Cost Backup/PITR/export/import/storage/read/write evidence if managed drill
Rollback Source state, write-resumption boundary, reconciliation path
Follow-up Runbook fixes, missing automation, ownership and next drill date

Verification checklist and cleanup

  • Separate recovery target validated before cutover.
  • Both missing-document and corrupted-value failure modes are detected.
  • Counts are supplemented with hashes/types/business-query regression.
  • Configuration parity includes Rules/indexes/TTL/IAM/App Check/Functions/secrets/network.
  • RPO/RTO use measured incident and service timestamps.
  • Rollback behavior changes explicitly once target writes resume.
  • Emulator evidence and managed-production evidence are labeled separately.
  • All temporary recovery databases, exports, and billed resources have a cleanup owner.

Production judgment and bridge to Chapter 23

Backups and PITR reduce the cost of logical failure only when the organization can discover the right recovery point, authorize the operation, reconstruct missing configuration, validate the result, and move application traffic safely. A multi-region database without a practiced recovery runbook can be highly available and still operationally unrecoverable from human mistakes.

Chapter 23 extends the recovery/control plane into encryption, CMEK, private connectivity, data residency, and governance. Those controls affect not only normal requests but also clone/restore destinations, key availability, exports, backups, and disaster-recovery access paths.

Knowledge check

  1. When does RTO end in this drill?
  2. Why combine counts with hashes and business queries?
  3. What configuration must be checked even after a backup restore that contains indexes?
  4. Why is rollback harder after target writes resume?
  5. What does an emulator recovery drill prove?
Review the answers

1. When validated application service is restored on the recovery target, not merely when the storage restore operation finishes.

2. Counts miss wrong-but-present data; hashes/types/business queries catch additional corruption classes.

3. At minimum TTL policies, Rules, IAM intent, App Check, Functions, secrets, networking, and application configuration.

4. Because new writes may exist only on the recovery target, so reversal requires reconciliation/replay rather than a simple routing change.

5. The runbook logic, validation gates, configuration discipline, and application cutover mechanics—not managed backup/PITR performance, IAM, or billing behavior.

Summary and next step

This lesson established the working contract for Run a Recovery Drill and Measure Data Completeness, RPO, RTO, Rules/Indexes/Configuration, and Application Cutover. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to Encryption at Rest/In Transit, Google-Managed Keys, CMEK Concepts, Key Availability, Rotation, and Failure Impact.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.