Chapter 22 · Backups, Point-in-Time Recovery, Restore, and Disaster Recovery
Multi-Region Location Does Not Replace Backup: Failure Modes, Logical Corruption, and Operator Mistakes
Demonstrate why healthy multi-region replication cannot undo logical corruption, privileged mistakes, or malicious writes, and integrate historical recovery with least privilege and business integrity signals.
1. AtlasMart problem: the database stayed available while the data became wrong everywhere
AtlasMart selected a multi-region location for high availability. A compromised privileged backend writes malformed tax rates to every order. The database remains healthy and replicated; the bad state is now durably available in multiple replicas. This is the exact failure mode that demonstrates why availability, durability, and recoverability are separate properties.
- Classify infrastructure outage, logical corruption, malicious writes, accidental deletes, and configuration failures by recovery mechanism.
- Explain why multi-region replication cannot provide historical rewind by itself.
- Model operator/credential mistakes that survive healthy replication.
- Define defense in depth: least privilege, PITR/backups, immutable evidence, configuration source control, and recovery drills.
- Include clients/caches/listeners and downstream systems in cutover/rollback planning.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps the course-wide identity
demo-atlasmart-firestore. Mandatory work is local
and no-cost: Node.js 22+, Firebase CLI 15.30.0,
Firebase Admin Node SDK 14.4.0, Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. The normal application database
is Standard-edition Native mode, database
(default). A second emulator project namespace,
demo-atlasmart-recovery, acts as the isolated
restore target. Managed scheduled backups, PITR, clone,
managed export/import, backup storage, and managed restores
require a real billing-enabled project and are therefore
optional production-verification exercises, not mandatory lab
prerequisites.
A highly available database can be perfectly healthy while business state is catastrophically wrong. Replication protects service availability and data durability against infrastructure failure; it does not know which authorized write was a mistake. Historical mechanisms and operational controls are what create recoverability.
2. Failure-mode matrix
| Failure | Multi-region helps? | Recovery control |
|---|---|---|
| Regional infrastructure disruption | Yes, according to product availability architecture/SLA | Multi-region availability + application retry/failover design |
| Accidental delete by authorized admin | No historical rewind | PITR, backup, restore/clone, audit evidence |
| Bug writes wrong values everywhere | No | PITR selective repair or historical clone/restore |
| Stolen privileged credential | No; malicious writes can replicate | IAM containment + token/key revocation + historical recovery + audit |
| Bad Security Rules deployment | Replication is irrelevant | Version-controlled rules rollback + emulator/CI tests |
| Lost TTL/IAM/App Check configuration after data restore | No | Configuration inventory and deployment runbook |
3. Operator mistakes are application failures, not storage outages
A destructive script can succeed with HTTP 200 responses, low latency, and zero Firestore service errors. Monitoring only infrastructure health will miss the incident. AtlasMart therefore adds business integrity signals: sudden drop in cart count, impossible negative/oversized price deltas, unusual delete rates, privileged principal activity, and divergence between live counts and expected operational baselines.
{ "signal": "catalog-price-mass-change", "window": "5m", "conditions": { "documentsChanged": "> 20% of active catalog", "medianAbsolutePriceChange": "> 50%", "principalClass": "privileged-backend" }, "response": [ "freeze privileged writer", "record incident start/end candidates", "capture audit evidence", "select PITR/backup recovery path", "require two-person approval for rewrite/cutover" ]}
4. Mandatory local failure injection: corruption replicated logically
Use the emulator to create two logical consumers of the same canonical source state: the web read path and an analytics snapshot. Apply one privileged corruption script to the underlying data, then show that both consumers see the wrong data. The exercise demonstrates logical propagation, not physical multi-region replication.
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";initializeApp({ projectId: process.env.GCLOUD_PROJECT });const db = getFirestore();const snap = await db.collection("catalogItems").get();const batch = db.batch();for (const d of snap.docs) batch.update(d.ref, { price: d.get("price") * 100, incident: "DR-LOGICAL-001" });await batch.commit();console.log(`corrupted=${snap.size}; Firestore itself can still be perfectly available`);
Run this only against the emulator fixture. The point is to teach detection and recovery from valid-but-wrong writes, not to test a production blast radius.
5. Recovery must include clients and derived systems
After a database cutover, long-lived listeners, offline caches, backend connection pools, scheduled jobs, search indexes, analytics pipelines, and webhook processors may still point at old state or replay stale assumptions. A “database restored successfully” status is not application recovery. The runbook needs explicit client reconnect/version markers and derived-system reconciliation.
| Layer | Recovery question |
|---|---|
| Browser/mobile client | Will persisted cache surface stale data after endpoint/database switch? |
| Realtime listener | Does the listener reconnect to the recovered database and receive an authoritative snapshot? |
| Backend pool | Has every worker rotated to the new database identity/config? |
| Functions/events | Could replay after restore duplicate side effects? |
| Analytics/search | Must materialized/derived state be rebuilt from recovered data? |
| Billing/observability | Can the team distinguish restore reads/writes from normal traffic? |
6. Security and recovery reinforce each other
Least privilege reduces incident probability and blast radius, but it does not eliminate recovery requirements. Conversely, backups are not a reason to grant broad administrator access. Keep backup deletion, restore initiation, production write authority, and application deploy rights separable where practical. Audit who changed recovery settings and who executed restore/cutover actions.
Verification checklist and cleanup
- At least one logical-corruption alert is independent of Firestore service health.
- Privileged bulk writers have bounded scope and auditable identities.
- PITR/backups exist for authorized-but-wrong writes.
- Rules/IAM/App Check/configuration are versioned separately from data recovery artifacts.
- Client cache/listener behavior is in the cutover checklist.
- Derived systems have a rebuild/reconciliation plan after recovery.
Bridge to Lesson 5
The chapter ends by running the entire recovery process as an operational drill. The success criteria are not “restore command completed”; they are data completeness, measured RPO/RTO, configuration parity, security, application cutover, and a tested rollback point.
Knowledge check
- Why can multi-region replication make logical corruption more durable?
- What detects a destructive script when Firestore reports no service errors?
- Does data restore automatically fix client caches and listeners?
- Should one identity be able to write production data, delete backups, and perform restore by default?
- What is the central distinction of this lesson?
Review the answers
1. Because it faithfully replicates authorized committed writes, including wrong ones.
2. Business-integrity and privileged-activity signals, not only infrastructure health metrics.
3. No. Client/runtime convergence must be tested as part of application cutover.
4. Prefer separation and least privilege where practical to reduce blast radius.
5. Availability keeps service reachable; recoverability lets you return to a known-good historical state after logical failure.
Summary and next step
This lesson established the working contract for Multi-Region Location Does Not Replace Backup: Failure Modes, Logical Corruption, and Operator Mistakes. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Run a Recovery Drill and Measure Data Completeness, RPO, RTO, Rules/Indexes/Configuration, and Application Cutover.
Authoritative references
- Firebase: Back up and restore data — scheduled backup semantics, retention, roles, restore behavior, and post-restore checks.
- Firebase: Point-in-time recovery (PITR) — historical-version window, read granularity, clone/export recovery paths, and billing.
- Firebase: Manage databases — clone semantics, destination identity, location, encryption, and permissions.
- Firebase: Export and import data — managed export/import behavior, billing, index handling, and operation caveats.
- Firebase: Disaster recovery planning — availability versus recoverability and recovery mechanism selection.