Chapter 22 · Backups, Point-in-Time Recovery, Restore, and Disaster Recovery
PITR Window Semantics, Historical Reads / Recovery, and Accidental Write / Delete Scenarios
Use PITR historical versions for bounded accidental-write/delete recovery, selecting a clean timestamp and repairing only the damaged scope without erasing legitimate later changes.
1. AtlasMart problem: one admin update corrupted prices, but only for 12 minutes
An administrator accidentally multiplies prices by 100 for a
subset of catalogItems, notices the error 12
minutes later, and stops the script. Restoring yesterday’s
backup would lose an entire day of valid orders and profile
changes. This is a
surgical historical recovery problem: identify
a clean timestamp, read the earlier versions, validate them, and
write only the damaged records back.
-
Explain PITR enablement, the seven-day retention window,
earliestVersionTime, and minute-granularity historical versions. - Distinguish stale-read repair from database clone and PITR export/import.
- Choose a recovery timestamp that is demonstrably before corruption and within the permissible window.
- Measure logical RPO from the selected historical version rather than from detection time.
- Design repairs that are idempotent, auditable, and safe against concurrent post-incident changes.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps the course-wide identity
demo-atlasmart-firestore. Mandatory work is local
and no-cost: Node.js 22+, Firebase CLI 15.30.0,
Firebase Admin Node SDK 14.4.0, Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. The normal application database
is Standard-edition Native mode, database
(default). A second emulator project namespace,
demo-atlasmart-recovery, acts as the isolated
restore target. Managed scheduled backups, PITR, clone,
managed export/import, backup storage, and managed restores
require a real billing-enabled project and are therefore
optional production-verification exercises, not mandatory lab
prerequisites.
When PITR is enabled, Firestore retains historical document
versions for up to seven days. A single version per minute is
retained in that PITR window, and clone/export recovery
timestamps use whole-minute granularity. You cannot
immediately read seven days into the past right after enabling
PITR; earliestVersionTime defines the oldest
permissible version currently available.
2. The historical timeline: last hour versus seven-day PITR
Firestore retains short historical read capability even when
PITR is disabled: current documentation allows historical reads
within the past hour at microsecond granularity, bounded by
earliestVersionTime. PITR extends historical
retention to seven days, but the retained PITR versions are
minute-granularity. Clone and PITR-export timestamps must be
whole minutes.
| State | Historical reach | Granularity/use |
|---|---|---|
| PITR disabled |
Up to roughly the past hour, subject to
earliestVersionTime
|
Historical reads only; not the seven-day archive |
| PITR newly enabled | From the available earliest version; not instantly seven days | Window grows with time |
| PITR enabled > 7 days | Up to seven days | One retained version per minute |
| Clone/export from PITR | Within valid PITR window | Whole-minute snapshot timestamp |
3. Three recovery shapes
| Recovery shape | Best fit | Write risk |
|---|---|---|
| Stale read + selective rewrite | Known subset of corrupted/deleted documents | Must avoid overwriting legitimate later changes |
| Clone to new database | Full-state validation, isolated testing, broad corruption | Requires application/config cutover if adopted |
| PITR export + import | Archive/data movement or collection-group recovery workflows | Import semantics can overwrite matching IDs and leave unaffected target docs |
Do not turn “PITR exists” into “recovery is automatic.” Firestore retains versions; your runbook still needs incident timestamps, query scope, validation, authorization, and application cutover.
4. Choosing the clean timestamp
Suppose logs show the destructive script began at
2026-09-17T01:42:17Z and finished at
01:54:09Z. A whole-minute PITR clone/export point
such as 01:41:00Z is clearly before corruption.
Choosing 01:42:00Z may or may not include writes
from the corruption minute depending on ordering within that
minute, so the runbook should select a timestamp with an
explicit safety margin and record the rationale.
Within the PITR window Firestore retains one version per minute. Multiple writes within a minute collapse to the retained version for that minute. PITR is therefore excellent for bounded operator mistakes, but it is not an infinite per-write audit log.
5. Mandatory local simulation: deterministic historical versions
The emulator does not provide managed PITR history. Simulate the
mechanism explicitly so the recovery algorithm can be tested
without claiming service fidelity. Keep three canonical
snapshots: T0-clean, T1-corrupted, and
T2-post-incident-valid. The repair logic must
recover only the fields/documents declared damaged and preserve
later valid changes.
{ "incident": "INC-2026-09-17-PRICE-01", "corruptionStart": "2026-09-17T01:42:17Z", "cleanRecoveryPoint": "2026-09-17T01:41:00Z", "scope": ["catalogItems/p-1001", "catalogItems/p-1004"], "fields": ["price"], "preserveCurrentFields": ["stock", "updatedAt", "promotion"], "precondition": "current price must still match known corrupted value", "approval": "two-person review before production writeback"}
// Deterministic emulator exercise: T0-clean is a fixture, not real PITR.import fs from "node:fs/promises";import { initializeApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";const app = initializeApp({ projectId: "demo-atlasmart-firestore" }, "repair");const db = getFirestore(app);const clean = JSON.parse(await fs.readFile("snapshots/T0-clean.json", "utf8"));for (const item of clean.filter(x => ["catalogItems/p-1001","catalogItems/p-1004"].includes(x.path))) { const ref = db.doc(item.path); await db.runTransaction(async tx => { const now = await tx.get(ref); if (!now.exists) throw new Error(`missing current ${item.path}`); const current = now.data(); if (![9900,14900].includes(current.price)) throw new Error(`precondition failed ${item.path}`); tx.update(ref, { price: item.data.price, recoveryIncident: "INC-2026-09-17-PRICE-01" }); });}console.log("Selective repair applied; unrelated current fields preserved.");
A blind overwrite from an old snapshot can erase legitimate writes that happened after the incident. Surgical repair should constrain exactly what it is restoring and fail closed if the current document no longer matches the known damaged state.
6. Optional real-project PITR evidence
In an isolated billing-enabled project, enable PITR, wait until
the desired historical version exists, create a bounded fixture,
mutate/delete it, then either issue supported historical reads
or clone/export from a whole-minute timestamp. Record
earliestVersionTime, selected recovery timestamp,
returned hashes, read counts, and end-to-end repair time. PITR
has no free tier and historical reads/exports incur read-related
charges.
# Illustrative control-plane evidence; run only in an isolated billed project.gcloud firestore databases describe --database='(default)' --format=json# Clone at a validated whole-minute timestamp:gcloud firestore databases clone \ --source-database='projects/YOUR_PROJECT/databases/(default)' \ --snapshot-time='2026-09-17T01:41:00Z' \ --destination-database='recovery-drill-20260917'
7. RPO/RTO interpretation
If the clean historical version is 01:41 and the incident begins at 01:42:17, the recovery point may sacrifice legitimate writes between the chosen clean point and the corruption boundary unless you repair selectively. A lower nominal RPO is not automatically a better outcome if the restoration method overwrites more good data. For surgical PITR repair, measure the number and age of records intentionally reverted, not just the timestamp gap.
Likewise, PITR’s storage capability does not make application RTO zero in practice. Detection, scoping, approval, repair execution, validation, cache convergence, and customer communication are part of operational recovery time.
Verification checklist and cleanup
- Incident start/end timestamps are evidence-backed.
-
Selected PITR point is within
earliestVersionTimeand uses valid granularity. - Repair scope is explicit and narrower than the whole database when possible.
- Preconditions prevent stale recovery data from overwriting legitimate later changes.
- Counts/hashes and business queries are compared before and after repair.
- Optional PITR tests record cost-bearing reads and clean up clone/export artifacts.
Bridge to Lesson 3
PITR, backup restore, and export/import can all recover data, but they are not interchangeable. Lesson 3 compares their snapshot semantics, configuration payloads, RPO/RTO shape, billing, and operational failure modes.
Knowledge check
- How much PITR history is retained after PITR has been enabled long enough?
- What granularity does the seven-day PITR window retain?
- Should a 12-minute localized corruption automatically trigger whole-database restore?
- Why use a precondition during selective repair?
- Does the emulator simulation prove PITR performance or billing?
Review the answers
1. Up to seven days.
2. One document version per minute; clone/export recovery timestamps use whole-minute values.
3. Not necessarily. A selective stale-read repair can preserve valid post-incident changes if the damaged scope is known.
4. To avoid overwriting a document that has legitimately changed since the incident.
5. No. It tests the recovery algorithm and evidence discipline only.
Summary and next step
This lesson established the working contract for PITR Window Semantics, Historical Reads/Recovery, and Accidental Write/Delete Scenarios. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Export/Import vs Managed Backups/PITR: Different Purposes, RPO/RTO, and Operational Workflows.
Authoritative references
- Firebase: Back up and restore data — scheduled backup semantics, retention, roles, restore behavior, and post-restore checks.
- Firebase: Point-in-time recovery (PITR) — historical-version window, read granularity, clone/export recovery paths, and billing.
- Firebase: Manage databases — clone semantics, destination identity, location, encryption, and permissions.
- Firebase: Export and import data — managed export/import behavior, billing, index handling, and operation caveats.
- Firebase: Disaster recovery planning — availability versus recoverability and recovery mechanism selection.