Chapter 22 · Backups, Point-in-Time Recovery, Restore, and Disaster Recovery
Run a Recovery Drill and Measure Data Completeness, RPO, RTO, Rules / Indexes / Configuration, and Application Cutover
Run an end-to-end AtlasMart recovery drill with objective gates for data completeness, RPO/RTO, security/config parity, application cutover, performance, and rollback.
1. AtlasMart recovery drill: prove the runbook before the incident
A recovery plan that has never been executed is an assumption. AtlasMart therefore runs a deterministic drill: create a baseline, snapshot the evidence, inject an accidental deletion, restore into an isolated target, deploy missing configuration, validate business behavior, measure RPO/RTO, cut application traffic to the recovery target, then prove a rollback path.
- Execute a full recovery rehearsal from detection through validated cutover.
- Measure counts, canonical hashes, business-query equivalence, configuration completeness, and timing.
- Separate data restoration from Rules, indexes, TTL, IAM, App Check, Functions, and secrets recovery.
- Use objective go/no-go gates for cutover and rollback.
- Produce an auditable recovery report that identifies emulator evidence versus optional managed-service evidence.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps the course-wide identity
demo-atlasmart-firestore. Mandatory work is local
and no-cost: Node.js 22+, Firebase CLI 15.30.0,
Firebase Admin Node SDK 14.4.0, Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. The normal application database
is Standard-edition Native mode, database
(default). A second emulator project namespace,
demo-atlasmart-recovery, acts as the isolated
restore target. Managed scheduled backups, PITR, clone,
managed export/import, backup storage, and managed restores
require a real billing-enabled project and are therefore
optional production-verification exercises, not mandatory lab
prerequisites.
The source is never destroyed merely to make the exercise realistic. Mandatory work uses the emulator and a separate recovery project namespace. The optional cloud drill also restores/clones to a separate target first. Destructive same-name restoration is discussed, not required.
2. Recovery acceptance gates
| Gate | Evidence | Fail condition |
|---|---|---|
| G0 incident declared | Start timestamp, affected scope, writer frozen | Unknown ongoing corruption |
| G1 recovery point chosen | Backup/PITR timestamp + rationale | Outside retention window or after corruption |
| G2 target created | Distinct database/project identity | Accidental overwrite of source |
| G3 data complete | Counts + canonical hashes + sampled field/type checks | Unexpected count/hash/type drift |
| G4 config complete | Rules/index/TTL/IAM/App Check/deploy inventory | Recovered data served with missing controls |
| G5 application healthy | Smoke tests + critical queries + authz tests | Functional/security regression |
| G6 cutover measured | Traffic switch timestamp + p50/p95/p99 during bounded test | Tail latency/errors exceed agreed SLO |
| G7 rollback ready | Old target preserved/read-only + reconciliation plan | No safe route back after writes resume |
3. Step 1 — seed and capture the baseline
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";import { initializeApp, deleteApp } from "firebase-admin/app";import { getFirestore, Timestamp } from "firebase-admin/firestore";const app = initializeApp({ projectId: "demo-atlasmart-firestore" }, "drill-source");const db = getFirestore(app);const fixture = [ ["catalogItems/p-1001", {name:"Trail Camera",price:99,stock:8,tenantId:"atlasmart",updatedAt:Timestamp.fromDate(new Date("2026-09-17T02:00:00Z"))}], ["catalogItems/p-1002", {name:"USB-C Hub",price:49,stock:3,tenantId:"atlasmart",updatedAt:Timestamp.fromDate(new Date("2026-09-17T02:00:00Z"))}], ["orders/o-dr-001", {customerId:"user-a",total:148,status:"PAID",tenantId:"atlasmart"}], ["profiles/user-a", {displayName:"Atlas Buyer",tenantId:"atlasmart",schemaVersion:1}]];for (const [path,data] of fixture) await db.doc(path).set(data);await deleteApp(app);console.log("seeded", fixture.length);
Run the deterministic snapshot script from Lesson 1 and store
data.json, manifest.json,
firestore.rules,
firestore.indexes.json, and
recovery-config.json as the baseline evidence
bundle.
4. Step 2 — inject the incident and start the RTO clock
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";initializeApp({ projectId: process.env.GCLOUD_PROJECT });const db = getFirestore();const incidentStarted = process.hrtime.bigint();await db.doc("catalogItems/p-1001").delete();await db.doc("orders/o-dr-001").update({ total: 14800, incident: "DRILL-22" });console.log(JSON.stringify({ incident:"DRILL-22", declaredAt:new Date().toISOString(), monotonicStart:String(incidentStarted) }));
The drill forces validation to catch both absence and wrong-but-present data. A recovery test that checks only document count could miss the corrupted order total.
5. Step 3 — restore into a separate emulator project namespace
import fs from "node:fs/promises";import { initializeApp, deleteApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";const rows = JSON.parse(await fs.readFile("recovery-fixture/data.json", "utf8"));const app = initializeApp({ projectId: "demo-atlasmart-recovery" }, "drill-target");const db = getFirestore(app);for (const row of rows) await db.doc(row.path).set(row.data);console.log({ restored: rows.length, target: "demo-atlasmart-recovery" });await deleteApp(app);
This script restores the course’s deterministic JSON fixture.
It is not a replacement for
firebase firestore:databases:restore or
gcloud firestore import, and the file is
intentionally not represented as a managed-export artifact.
6. Step 4 — validate counts, hashes, types, and business queries
import crypto from "node:crypto";import fs from "node:fs/promises";import { initializeApp, deleteApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";const app = initializeApp({ projectId: "demo-atlasmart-recovery" }, "validate-target");const db = getFirestore(app);const paths = ["catalogItems/p-1001","catalogItems/p-1002","orders/o-dr-001","profiles/user-a"];const rows=[];for (const path of paths) { const s=await db.doc(path).get(); if(!s.exists) throw new Error(`missing ${path}`); rows.push({path,data:s.data()}); }rows.sort((a,b)=>a.path.localeCompare(b.path));const hash=crypto.createHash("sha256").update(JSON.stringify(rows)).digest("hex");const order=rows.find(x=>x.path==="orders/o-dr-001").data;if (order.total !== 148) throw new Error(`business regression: order total=${order.total}`);console.log({count:rows.length, hash, businessQuery:"PASS"});await deleteApp(app);
Because Firestore values such as Timestamp require
canonical serialization, production drills should define a
stable hash encoding rather than relying on JavaScript object
iteration or SDK-specific JSON output. The simple lab hash is
illustrative; the recovery report should record the
canonicalization contract.
7. Step 5 — configuration parity
Before cutover, compare a configuration manifest. A managed backup restore includes indexes but not TTL policies; an export/import workflow does not carry index definitions at all. Neither data mechanism should be expected to reconstruct application Rules, App Check settings, IAM intent, Functions, secrets, or network controls.
{ "firestoreRulesSha256": "RECORD_FROM_REPO", "indexConfigSha256": "RECORD_FROM_REPO", "ttlPolicies": ["sessions.expireAt", "events.expireAt"], "requiredIamRoles": ["application-runtime-role", "restore-operator-role"], "appCheck": "production-only verification required", "functionsRelease": "RECORD_DEPLOYMENT_ID", "secretsVersionSet": "RECORD_SECRET_VERSIONS", "networkPolicyVersion": "RECORD_POLICY_ID"}
8. Step 6 — cutover and rollback boundary
In the emulator, the “application endpoint” is a configuration
file that names the project ID. Switch it from
demo-atlasmart-firestore to
demo-atlasmart-recovery, restart the test client,
and rerun the critical query/authz suite. In production, the
equivalent may be a database ID/config rollout, backend
deployment, DNS/service-routing change, or client release
depending on architecture.
| Before target writes resume | After target writes resume |
|---|---|
| Rollback can often be a traffic/config reversal because source remains authoritative/read-only evidence. | Rollback becomes a data reconciliation problem because writes may exist only on the recovery target. |
| Keep source frozen and preserve incident evidence. | Record every target-side write or establish reverse replication/replay strategy. |
| Validate security/config parity first. | Require explicit incident commander approval for reversal. |
9. Measure RPO/RTO and tail latency without inventing numbers
The drill script should emit timestamps for incident declaration, recovery point, restore start/end, validation pass, configuration pass, cutover start, and first successful production-equivalent request. Compute RPO from the age of the chosen clean data; compute RTO from incident declaration to validated service restoration. If you benchmark requests after cutover, report p50/p95/p99 from measured samples. Emulator percentiles are useful for regression detection but not production SLO proof.
export function percentile(xs, p) { const a=[...xs].sort((x,y)=>x-y); if (!a.length) return null; const i=Math.min(a.length-1, Math.ceil((p/100)*a.length)-1); return a[i];}// Feed real measured request durations; never insert invented values.console.log({p50:percentile(samples,50),p95:percentile(samples,95),p99:percentile(samples,99)});
10. Optional managed-service drill
In an isolated Blaze/billing-enabled project, choose one mechanism and collect real evidence: create a backup schedule and restore a retained backup to a new database; or enable PITR and clone a clean timestamp. Record backup/clone resource names, locations, timestamps, operation duration, backup/PITR storage cost evidence, effective IAM principal, restored index configuration, missing TTL policy, validation hashes, and cutover timing. Do not run destructive same-name restore merely for practice.
11. Recovery report template
| Field | Required evidence |
|---|---|
| Incident | ID, start/end, detector, affected collections |
| Recovery point |
Mechanism, timestamp,
earliestVersionTime/backup resource if
relevant
|
| RPO | Measured age/data loss relative to incident |
| RTO | Detection/declaration → validated application service |
| Data integrity | Counts, canonical hashes, sampled types, business query regression |
| Configuration | Rules, indexes, TTL, IAM, App Check, Functions, secrets, network |
| Performance | Bounded post-cutover p50/p95/p99 and errors |
| Cost | Backup/PITR/export/import/storage/read/write evidence if managed drill |
| Rollback | Source state, write-resumption boundary, reconciliation path |
| Follow-up | Runbook fixes, missing automation, ownership and next drill date |
Verification checklist and cleanup
- Separate recovery target validated before cutover.
- Both missing-document and corrupted-value failure modes are detected.
- Counts are supplemented with hashes/types/business-query regression.
- Configuration parity includes Rules/indexes/TTL/IAM/App Check/Functions/secrets/network.
- RPO/RTO use measured incident and service timestamps.
- Rollback behavior changes explicitly once target writes resume.
- Emulator evidence and managed-production evidence are labeled separately.
- All temporary recovery databases, exports, and billed resources have a cleanup owner.
Production judgment and bridge to Chapter 23
Backups and PITR reduce the cost of logical failure only when the organization can discover the right recovery point, authorize the operation, reconstruct missing configuration, validate the result, and move application traffic safely. A multi-region database without a practiced recovery runbook can be highly available and still operationally unrecoverable from human mistakes.
Chapter 23 extends the recovery/control plane into encryption, CMEK, private connectivity, data residency, and governance. Those controls affect not only normal requests but also clone/restore destinations, key availability, exports, backups, and disaster-recovery access paths.
Knowledge check
- When does RTO end in this drill?
- Why combine counts with hashes and business queries?
- What configuration must be checked even after a backup restore that contains indexes?
- Why is rollback harder after target writes resume?
- What does an emulator recovery drill prove?
Review the answers
1. When validated application service is restored on the recovery target, not merely when the storage restore operation finishes.
2. Counts miss wrong-but-present data; hashes/types/business queries catch additional corruption classes.
3. At minimum TTL policies, Rules, IAM intent, App Check, Functions, secrets, networking, and application configuration.
4. Because new writes may exist only on the recovery target, so reversal requires reconciliation/replay rather than a simple routing change.
5. The runbook logic, validation gates, configuration discipline, and application cutover mechanics—not managed backup/PITR performance, IAM, or billing behavior.
Summary and next step
This lesson established the working contract for Run a Recovery Drill and Measure Data Completeness, RPO, RTO, Rules/Indexes/Configuration, and Application Cutover. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Encryption at Rest/In Transit, Google-Managed Keys, CMEK Concepts, Key Availability, Rotation, and Failure Impact.
Authoritative references
- Firebase: Back up and restore data — scheduled backup semantics, retention, roles, restore behavior, and post-restore checks.
- Firebase: Point-in-time recovery (PITR) — historical-version window, read granularity, clone/export recovery paths, and billing.
- Firebase: Manage databases — clone semantics, destination identity, location, encryption, and permissions.
- Firebase: Export and import data — managed export/import behavior, billing, index handling, and operation caveats.
- Firebase: Disaster recovery planning — availability versus recoverability and recovery mechanism selection.