Chapter 22 · Backups, Point-in-Time Recovery, Restore, and Disaster Recovery

PITR Window Semantics, Historical Reads / Recovery, and Accidental Write / Delete Scenarios

Use PITR historical versions for bounded accidental-write/delete recovery, selecting a clean timestamp and repairing only the damaged scope without erasing legitimate later changes.

Advanced · 160–220 minutesPITR · historical reads · repair · RPO · recoveryNode 22+ · Firebase CLI 15.30.0 · Firestore emulator 127.0.0.1:8080Mandatory recovery drill local/no-cost · managed backup/PITR optional/billedLast reviewed: 17 September 2026

1. AtlasMart problem: one admin update corrupted prices, but only for 12 minutes

An administrator accidentally multiplies prices by 100 for a subset of catalogItems, notices the error 12 minutes later, and stops the script. Restoring yesterday’s backup would lose an entire day of valid orders and profile changes. This is a surgical historical recovery problem: identify a clean timestamp, read the earlier versions, validate them, and write only the damaged records back.

Learning outcomes
  • Explain PITR enablement, the seven-day retention window, earliestVersionTime, and minute-granularity historical versions.
  • Distinguish stale-read repair from database clone and PITR export/import.
  • Choose a recovery timestamp that is demonstrably before corruption and within the permissible window.
  • Measure logical RPO from the selected historical version rather than from detection time.
  • Design repairs that are idempotent, auditable, and safe against concurrent post-incident changes.
Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 22 reproducibility baseline · reviewed 17 September 2026

AtlasMart keeps the course-wide identity demo-atlasmart-firestore. Mandatory work is local and no-cost: Node.js 22+, Firebase CLI 15.30.0, Firebase Admin Node SDK 14.4.0, Firestore emulator 127.0.0.1:8080, Auth emulator 127.0.0.1:9099, and Emulator UI 127.0.0.1:4000. The normal application database is Standard-edition Native mode, database (default). A second emulator project namespace, demo-atlasmart-recovery, acts as the isolated restore target. Managed scheduled backups, PITR, clone, managed export/import, backup storage, and managed restores require a real billing-enabled project and are therefore optional production-verification exercises, not mandatory lab prerequisites.

PITR is disabled by default

When PITR is enabled, Firestore retains historical document versions for up to seven days. A single version per minute is retained in that PITR window, and clone/export recovery timestamps use whole-minute granularity. You cannot immediately read seven days into the past right after enabling PITR; earliestVersionTime defines the oldest permissible version currently available.

2. The historical timeline: last hour versus seven-day PITR

Firestore retains short historical read capability even when PITR is disabled: current documentation allows historical reads within the past hour at microsecond granularity, bounded by earliestVersionTime. PITR extends historical retention to seven days, but the retained PITR versions are minute-granularity. Clone and PITR-export timestamps must be whole minutes.

State Historical reach Granularity/use
PITR disabled Up to roughly the past hour, subject to earliestVersionTime Historical reads only; not the seven-day archive
PITR newly enabled From the available earliest version; not instantly seven days Window grows with time
PITR enabled > 7 days Up to seven days One retained version per minute
Clone/export from PITR Within valid PITR window Whole-minute snapshot timestamp

3. Three recovery shapes

Recovery shape Best fit Write risk
Stale read + selective rewrite Known subset of corrupted/deleted documents Must avoid overwriting legitimate later changes
Clone to new database Full-state validation, isolated testing, broad corruption Requires application/config cutover if adopted
PITR export + import Archive/data movement or collection-group recovery workflows Import semantics can overwrite matching IDs and leave unaffected target docs

Do not turn “PITR exists” into “recovery is automatic.” Firestore retains versions; your runbook still needs incident timestamps, query scope, validation, authorization, and application cutover.

4. Choosing the clean timestamp

Suppose logs show the destructive script began at 2026-09-17T01:42:17Z and finished at 01:54:09Z. A whole-minute PITR clone/export point such as 01:41:00Z is clearly before corruption. Choosing 01:42:00Z may or may not include writes from the corruption minute depending on ordering within that minute, so the runbook should select a timestamp with an explicit safety margin and record the rationale.

Minute granularity is a recovery design constraint

Within the PITR window Firestore retains one version per minute. Multiple writes within a minute collapse to the retained version for that minute. PITR is therefore excellent for bounded operator mistakes, but it is not an infinite per-write audit log.

5. Mandatory local simulation: deterministic historical versions

The emulator does not provide managed PITR history. Simulate the mechanism explicitly so the recovery algorithm can be tested without claiming service fidelity. Keep three canonical snapshots: T0-clean, T1-corrupted, and T2-post-incident-valid. The repair logic must recover only the fields/documents declared damaged and preserve later valid changes.

repair-plan.json
{  "incident": "INC-2026-09-17-PRICE-01",  "corruptionStart": "2026-09-17T01:42:17Z",  "cleanRecoveryPoint": "2026-09-17T01:41:00Z",  "scope": ["catalogItems/p-1001", "catalogItems/p-1004"],  "fields": ["price"],  "preserveCurrentFields": ["stock", "updatedAt", "promotion"],  "precondition": "current price must still match known corrupted value",  "approval": "two-person review before production writeback"}
selective-repair.mjs
// Deterministic emulator exercise: T0-clean is a fixture, not real PITR.import fs from "node:fs/promises";import { initializeApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";const app = initializeApp({ projectId: "demo-atlasmart-firestore" }, "repair");const db = getFirestore(app);const clean = JSON.parse(await fs.readFile("snapshots/T0-clean.json", "utf8"));for (const item of clean.filter(x => ["catalogItems/p-1001","catalogItems/p-1004"].includes(x.path))) {  const ref = db.doc(item.path);  await db.runTransaction(async tx => {    const now = await tx.get(ref);    if (!now.exists) throw new Error(`missing current ${item.path}`);    const current = now.data();    if (![9900,14900].includes(current.price)) throw new Error(`precondition failed ${item.path}`);    tx.update(ref, { price: item.data.price, recoveryIncident: "INC-2026-09-17-PRICE-01" });  });}console.log("Selective repair applied; unrelated current fields preserved.");
Why the transaction/precondition matters

A blind overwrite from an old snapshot can erase legitimate writes that happened after the incident. Surgical repair should constrain exactly what it is restoring and fail closed if the current document no longer matches the known damaged state.

6. Optional real-project PITR evidence

In an isolated billing-enabled project, enable PITR, wait until the desired historical version exists, create a bounded fixture, mutate/delete it, then either issue supported historical reads or clone/export from a whole-minute timestamp. Record earliestVersionTime, selected recovery timestamp, returned hashes, read counts, and end-to-end repair time. PITR has no free tier and historical reads/exports incur read-related charges.

optional-pitr.txt
# Illustrative control-plane evidence; run only in an isolated billed project.gcloud firestore databases describe --database='(default)' --format=json# Clone at a validated whole-minute timestamp:gcloud firestore databases clone \  --source-database='projects/YOUR_PROJECT/databases/(default)' \  --snapshot-time='2026-09-17T01:41:00Z' \  --destination-database='recovery-drill-20260917' 

7. RPO/RTO interpretation

If the clean historical version is 01:41 and the incident begins at 01:42:17, the recovery point may sacrifice legitimate writes between the chosen clean point and the corruption boundary unless you repair selectively. A lower nominal RPO is not automatically a better outcome if the restoration method overwrites more good data. For surgical PITR repair, measure the number and age of records intentionally reverted, not just the timestamp gap.

Likewise, PITR’s storage capability does not make application RTO zero in practice. Detection, scoping, approval, repair execution, validation, cache convergence, and customer communication are part of operational recovery time.

Verification checklist and cleanup

  • Incident start/end timestamps are evidence-backed.
  • Selected PITR point is within earliestVersionTime and uses valid granularity.
  • Repair scope is explicit and narrower than the whole database when possible.
  • Preconditions prevent stale recovery data from overwriting legitimate later changes.
  • Counts/hashes and business queries are compared before and after repair.
  • Optional PITR tests record cost-bearing reads and clean up clone/export artifacts.

Bridge to Lesson 3

PITR, backup restore, and export/import can all recover data, but they are not interchangeable. Lesson 3 compares their snapshot semantics, configuration payloads, RPO/RTO shape, billing, and operational failure modes.

Knowledge check

  1. How much PITR history is retained after PITR has been enabled long enough?
  2. What granularity does the seven-day PITR window retain?
  3. Should a 12-minute localized corruption automatically trigger whole-database restore?
  4. Why use a precondition during selective repair?
  5. Does the emulator simulation prove PITR performance or billing?
Review the answers

1. Up to seven days.

2. One document version per minute; clone/export recovery timestamps use whole-minute values.

3. Not necessarily. A selective stale-read repair can preserve valid post-incident changes if the damaged scope is known.

4. To avoid overwriting a document that has legitimately changed since the incident.

5. No. It tests the recovery algorithm and evidence discipline only.

Summary and next step

This lesson established the working contract for PITR Window Semantics, Historical Reads/Recovery, and Accidental Write/Delete Scenarios. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to Export/Import vs Managed Backups/PITR: Different Purposes, RPO/RTO, and Operational Workflows.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.