Chapter 22 · Backups, Point-in-Time Recovery, Restore, and Disaster Recovery
Scheduled Backups, Backup Retention, Database Restore / Clone Concepts, and Access Control
Protect AtlasMart with scheduled backups and isolated restore targets while separating data/index recovery from TTL, IAM, Rules, App Check, Functions, secrets, and application configuration.
1. AtlasMart problem: high availability did not undo a bad deployment
AtlasMart runs in a multi-region Firestore location. A bad administrative script deletes every current shopping cart and overwrites several catalog flags. Replication does exactly what an availability system should do: the destructive writes propagate consistently. Nothing about multi-region replication creates a historical copy that the team can rewind. Recovery therefore needs a separate mechanism and a runbook that distinguishes availability from recoverability.
- Explain what a scheduled Firestore backup contains, where it lives, and what it deliberately does not contain.
- Separate backup schedules, retained backup objects, restore-to-new-database, PITR clone, and managed export/import.
- Choose least-privilege IAM roles for backup viewing, schedule administration, and restore initiation.
- Prove a recovery target with counts, canonical hashes, index/config inventory, and application-level acceptance checks.
- Measure RPO and RTO from incident detection through validated cutover rather than from a console button click.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps the course-wide identity
demo-atlasmart-firestore. Mandatory work is local
and no-cost: Node.js 22+, Firebase CLI 15.30.0,
Firebase Admin Node SDK 14.4.0, Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. The normal application database
is Standard-edition Native mode, database
(default). A second emulator project namespace,
demo-atlasmart-recovery, acts as the isolated
restore target. Managed scheduled backups, PITR, clone,
managed export/import, backup storage, and managed restores
require a real billing-enabled project and are therefore
optional production-verification exercises, not mandatory lab
prerequisites.
Scheduled backups are relevant to both Firestore Standard and Enterprise editions. A backup is a consistent point-in-time copy containing database data and index configurations. It does not contain TTL policies. It remains in the same Firestore location as the source, and deleting the source database does not automatically delete its backups. Restores normally create a new database identity, which is why this chapter validates a separate recovery target before any cutover.
2. Backup, PITR, clone, and export are different recovery primitives
The word “backup” is often used too loosely. Firestore exposes several distinct mechanisms, each with a different consistency point, retention horizon, restore shape, configuration payload, cost model, and operational risk. Design the recovery plan from failure modes and business RPO/RTO, not from whichever button is most visible in the console.
| Mechanism | What it captures | Typical use | Important boundary |
|---|---|---|---|
| Scheduled backup | Consistent point-in-time data + index configurations | Periodic database recovery and corruption protection | Daily/weekly schedules; retention up to 14 weeks; no TTL policies |
| PITR stale read | Historical document versions | Surgical repair of a subset | Minute-granularity versions up to 7 days when enabled; live rewrite is your responsibility |
| PITR clone | New database from a historical timestamp | Isolated full-database recovery/testing | New database, same location; clone includes data and indexes |
| PITR export/import | Historical export to Cloud Storage, then import | Archive, project move, selective collection-group workflows | Export/import semantics differ from backup restore; export files do not carry index definitions |
| Managed export/import | Cloud Storage export of current data | Data movement, archive, offline processing | Export is not an exact snapshot at its start time; imports use target index definitions |
3. Schedule shape and retention are part of the RPO
For each database you can configure at most one daily backup schedule and one weekly backup schedule. The retention period can be set up to 14 weeks. A weekly schedule lets you choose the weekday; the exact time of day is not a customer-controlled guarantee. That matters: a “daily backup” does not mean “00:00 UTC snapshot.” Your actual recoverable point is the timestamp of a completed retained backup.
# OPTIONAL: billing-enabled isolated project onlyfirebase firestore:backups:schedules:create \ --database '(default)' \ --recurrence DAILY \ --retention 14dfirebase firestore:backups:schedules:list --database '(default)'firebase firestore:backups:list
Deleting a backup schedule stops future backups but does not delete already-created backups. Those retained backup objects expire according to their retention policy or can be deleted explicitly. A deleted backup cannot be recovered.
4. Access control: recovery credentials should be narrower than production administration
Firestore exposes purpose-specific IAM roles: backup
administrators/viewers, backup-schedule administrators/viewers,
and restore administrators, in addition to broad
roles/datastore.owner. A recovery operator who only
needs to enumerate backups should not receive Owner. Likewise,
an automation identity that creates schedules does not
automatically need permission to restore databases.
| Task | Representative predefined role | Blast-radius question |
|---|---|---|
| View backup metadata | roles/datastore.backupsViewer |
Can the identity read inventory without deleting backups? |
| Create/delete backup objects | roles/datastore.backupsAdmin |
Should automation be able to destroy recovery points? |
| Manage schedules | roles/datastore.backupSchedulesAdmin |
Can it alter future retention policy? |
| Initiate restore | roles/datastore.restoreAdmin |
Can it create a new recovered database? |
| Full Firestore administration | roles/datastore.owner |
Is this breadth actually necessary? |
Backup/PITR administration is an IAM-controlled Google Cloud operation. Firebase mobile/web Security Rules and App Check do not authorize these administrative operations. Conversely, restoring application data does not reconstruct your Auth users, App Check configuration, Cloud Functions deployment, secrets, or every surrounding project policy.
5. What a backup does not reconstruct
A database backup is not an application backup. Firestore documents state that backup data includes indexes but not TTL policies. The safe runbook also inventories Security Rules, IAM bindings, App Check settings, application configuration, Cloud Functions/Eventarc deployments, secrets, scheduled jobs, network controls, and any external search/warehouse state. Treat those as versioned infrastructure artifacts with separate recovery evidence.
| Asset | In Firestore backup? | Recovery source |
|---|---|---|
| Documents | Yes | Managed backup |
| Index configurations | Yes | Managed backup; also keep source-controlled index definitions |
| TTL policies | No | Infrastructure/configuration source of truth |
| Security Rules | No database payload guarantee | Version control + CI-tested deployment |
| IAM policy | No | Infrastructure policy / project configuration |
| App Check configuration | No | Firebase project configuration |
| Functions/secrets | No | Deployment artifacts + secret manager/process |
| Client cache | Not a recovery source | Discard/reconnect against validated backend state |
6. Mandatory local lab: create a recoverable evidence bundle
The Local Emulator Suite does not emulate managed backup
scheduling or PITR. Instead, create an application-level
deterministic recovery fixture that teaches the validation
discipline without pretending to produce a Google-managed
backup. The fixture is intentionally a different format from
gcloud firestore export output.
import crypto from "node:crypto";import fs from "node:fs/promises";import { initializeApp, deleteApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";const app = initializeApp({ projectId: "demo-atlasmart-firestore" }, "source");const db = getFirestore(app);const collections = ["catalogItems", "profiles", "orders", "recoveryMarkers"];const rows = [];for (const name of collections) { const snap = await db.collection(name).get(); for (const doc of snap.docs) rows.push({ path: doc.ref.path, data: doc.data() });}rows.sort((a,b) => a.path.localeCompare(b.path));const canonical = JSON.stringify(rows);const sha256 = crypto.createHash("sha256").update(canonical).digest("hex");await fs.mkdir("recovery-fixture", { recursive: true });await fs.writeFile("recovery-fixture/data.json", JSON.stringify(rows, null, 2));await fs.writeFile("recovery-fixture/manifest.json", JSON.stringify({ sourceProject: "demo-atlasmart-firestore", database: "(default)", rowCount: rows.length, sha256, createdAt: "2026-09-17T02:30:00Z", format: "AtlasMart deterministic emulator fixture — NOT managed Firestore export format"}, null, 2));console.log({ rowCount: rows.length, sha256 });await deleteApp(app);
The manifest gives a stable count and SHA-256 over canonicalized application records. It proves the fixture you later restore is byte-for-byte equivalent under this canonicalization. It does not prove production backup consistency, managed service latency, billing, or IAM behavior.
7. Failure injection: restoring “over” the live database
A common unsafe reaction to corruption is to overwrite the live database immediately. That destroys evidence, complicates rollback, and can mix recovered and post-incident writes. The safer default is: freeze or isolate writes, restore/clone to a separate database, validate, deploy missing configuration, test application behavior, and only then redirect traffic.
Firestore also documents an in-place-style recovery path that deletes the source database and restores a backup under the same database name. Treat that as a destructive operation: once the original database is deleted, you have crossed a much more severe rollback boundary. This course does not use it as the default drill.
8. RPO and RTO must include application cutover
RPO is the maximum acceptable data loss relative to the incident. RTO is the maximum acceptable time to restore service. Neither is equal to “backup schedule frequency” or “restore operation duration.” Measured RTO begins when the organization detects/declares the incident and ends when validated application traffic is safely serving from the recovered state. It includes data validation, rules/index/config deployment, smoke tests, client cutover, and rollback readiness.
Verification checklist and cleanup
- A deterministic recovery fixture and manifest exist outside the live emulator namespace.
- The recovery target is a different project/database identity during validation.
- Backup scope is not confused with Security Rules, IAM, App Check, Functions, secrets, or TTL configuration.
- Least-privilege recovery roles are documented for production.
- RPO/RTO measurement starts from business incident timestamps, not only a storage operation timer.
- Optional managed tests are performed only in a billed isolated project and are cleaned up.
Bridge to Lesson 2
Scheduled backups give discrete retained recovery points. Lesson 2 adds a much finer historical mechanism: PITR, its minute-granularity seven-day window, historical reads, and safe repair after accidental writes or deletes.
Knowledge check
- Does multi-region replication provide a rewind point after a destructive write?
- What does a Firestore scheduled backup include?
- Does a scheduled backup include TTL policies?
- Why restore to a separate target first?
- Can a daily backup schedule guarantee an exact clock time?
Review the answers
1. No. Replication improves availability and durability of the current state; a logical mistake can be replicated correctly. Historical recovery needs backups, PITR, exports, or an equivalent independent mechanism.
2. Database data and index configurations at a consistent point in time.
3. No. TTL policies must be reapplied after restore if required.
4. It preserves the damaged source for evidence/rollback and allows validation before traffic moves.
5. No. The exact time of day is not customer-configurable; use the actual backup timestamp when reasoning about RPO.
Summary and next step
This lesson established the working contract for Scheduled Backups, Backup Retention, Database Restore/Clone Concepts, and Access Control. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to PITR Window Semantics, Historical Reads/Recovery, and Accidental Write/Delete Scenarios.
Authoritative references
- Firebase: Back up and restore data — scheduled backup semantics, retention, roles, restore behavior, and post-restore checks.
- Firebase: Point-in-time recovery (PITR) — historical-version window, read granularity, clone/export recovery paths, and billing.
- Firebase: Manage databases — clone semantics, destination identity, location, encryption, and permissions.
- Firebase: Export and import data — managed export/import behavior, billing, index handling, and operation caveats.
- Firebase: Disaster recovery planning — availability versus recoverability and recovery mechanism selection.