Chapter 27 · Production Capstone: Design, Secure, Scale, Search, Recover, and Operate a Firestore Application
Load-Test Hotspots, Test Offline / Conflict Paths, Run Security Tests, Enable Backup / PITR, Monitor Queries, and Rehearse Recovery
Attack the architecture with hotspot, contention, offline, security, deployment and recovery game days while preserving emulator-versus-production evidence boundaries.
1. AtlasMart pre-launch game day: prove the design under failure
A green happy-path test suite is not a production argument. The capstone now attacks the architecture with the failures most likely to produce expensive incidents: hotspot keys, transaction contention, offline edits, stale cache, cross-tenant access, leaked listeners, duplicate retries, missing production indexes, logical deletion, and a recovery cutover. Each test records what local evidence proves and what still needs a managed canary.
- Load-test distribution and distinguish local regression data from production capacity evidence.
- Test offline/cache/conflict paths without claiming offline correctness for checkout.
- Run Security Rules and trusted-backend authorization failure cases.
- Connect Query Explain/Insights/hotspot evidence to a bounded managed verification plan.
- Rehearse recovery and deployment rollback with RPO/RTO and configuration completeness.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart keeps project ID
demo-atlasmart-firestore, Standard edition /
Native mode / Core operations, database
(default), Node.js 22+, Firebase CLI course
baseline 15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node SDK
14.4.0,
@firebase/rules-unit-testing 5.0.2, Firestore
emulator 127.0.0.1:8080, Auth emulator
127.0.0.1:9099, and Emulator UI
127.0.0.1:4000. The mandatory capstone uses only
a demo- project and local tooling. Managed
location, production IAM/App Check, real composite/vector
indexes, Query Explain/Insights, billing, quotas, Key
Visualizer, scheduled backups/PITR, CMEK/network controls,
Enterprise/Pipeline, and MongoDB-compatibility validation are
optional managed evidence and are never inferred from emulator
success. Where a managed Standard example needs a concrete
location, this chapter uses us-central1 only as
an explicit example—not a universal recommendation.
2. Load test the shape, not a magic throughput number
The local benchmark compares key distributions and contention patterns using the same dataset/concurrency. It reports distributions (p50/p95/p99/max), error counts and operation mix. The absolute emulator throughput is not a Firestore production capacity number.
const stages = [ { name: "scatter", id: i => crypto.randomUUID() }, { name: "sequential", id: i => `event_${String(i).padStart(8,"0")}` }, { name: "hot-doc", id: i => "global_counter" }];for (const stage of stages) { const latencies = []; let errors = 0; await runConcurrent(32, 1000, async i => { const t0 = performance.now(); try { await writeFixture(stage.id(i), i); } catch { errors++; } finally { latencies.push(performance.now() - t0); } }); console.log(summarize(stage.name, latencies, errors));}
Production follows hotspot physics even with automatic scaling. Scatter IDs reduce sequential key-range concentration, but a single hot document still cannot be split into parallel independent documents. For a new key range, use gradual traffic ramping such as the documented 500/50/5 best practice rather than treating any one number as a universal hard capacity limit.
3. Offline/conflict game day
| Scenario | Expected product behavior | Invariant |
|---|---|---|
| Cart edit while offline | UI marks pending/local state; write may synchronize later. | Server recalculates authoritative checkout price/inventory. |
| Two offline devices edit cart quantity | Last synchronized write may win unless app adds stronger merge semantics. | Cart convenience state must not represent inventory reservation. |
| Checkout while offline | Blocked with explicit online-required state. | No offline transaction promise for inventory/order invariant. |
| Order listener loses network | Last cached order may render as stale/cache state. | UI does not claim server-confirmed payment/shipment from cache alone. |
function orderBadge(snapshot) { if (!snapshot.exists()) return "missing"; if (snapshot.metadata.fromCache) return "last known — reconnecting"; if (snapshot.metadata.hasPendingWrites) return "local update pending"; return `server confirmed: ${snapshot.data().status}`;}
4. Security game day: prove denies, not only allows
- Unauthenticated cart read → DENY.
- Cross-user cart read/write → DENY.
- Client attempts order status mutation → DENY.
- Client attempts product price mutation → DENY.
- Server checkout with valid service identity but unauthorized user/business scope → application authorization DENY.
- Replay checkout idempotency key → existing order, no repeated inventory decrement.
- Overbroad product query incompatible with Rules → DENY.
App Check, if later enforced, adds an app-attestation gate; it does not change the requirement for user authentication, Rules/IAM and server business authorization.
5. Query and hotspot evidence in a managed canary
The emulator cannot establish index readiness, Query Explain scan counts, Query Insights normalized patterns, Key Visualizer heatmaps, managed p95/p99 or billing. The release checklist therefore names a small isolated managed canary:
| Check | Evidence | Stop condition |
|---|---|---|
| Composite query | Index state READY + successful query + Query Explain when supported. | Missing index/building/error → feature flag stays off. |
| High-frequency browse | p50/p95/p99 + returned/scanned work + error codes. | Regression beyond approved SLO/error budget. |
| Hot write path | Latency/contention metrics + key distribution evidence. | Narrow hot key/document signature or deadline errors. |
| Cost | Measured reads/index entries/writes/listener updates + billing/usage metrics. | Worst-case slope exceeds approved envelope. |
A production canary is bounded by traffic, duration, budget, synthetic data and stop conditions. It is not permission to discover capacity by uncontrolled load on customer data.
6. Recovery game day: availability is not recoverability
AtlasMart simulates accidental deletion locally by exporting deterministic fixtures into a separate recovery namespace, deleting a subset, restoring into an isolated target, validating counts/hashes/config inventory, and switching a test client. This teaches the runbook without pretending the emulator implements managed scheduled backups/PITR.
{ "incident": "simulated accidental order deletion", "declaredAt": "2026-09-17T10:00:00Z", "recoveryPoint": "2026-09-17T09:57:00Z", "target": "demo-atlasmart-recovery", "validation": { "documentCountMatch": true, "sampleHashMatch": true, "rulesShaMatch": true, "indexesShaMatch": true, "ttlConfig": "VERIFY_MANAGED_SEPARATELY" }, "cutoverValidatedAt": "2026-09-17T10:08:30Z", "measuredRpoMinutes": 3, "measuredRtoMinutes": 8.5}
For the managed optional drill, scheduled backup, PITR, restore/clone permissions, target identity, data validation, Rules/index/config inventory, application cutover and rollback are separate checkpoints. Multi-region replication does not undo a logically valid destructive write.
7. Deployment rollback game day
| Fault | Immediate response | Data/config response |
|---|---|---|
| New query missing managed index | Disable feature flag. | Wait/build/verify index; do not remove unrelated indexes. |
| Rules update causes deny spike | Restore reviewed compatible ruleset if safe. | Confirm mixed schema/client versions before rollback. |
| Backfill introduces mismatch | Pause resumable backfill; keep dual reader. | Reconcile from checkpoint; do not blindly reverse all data. |
| Checkout canary correctness failure | Stop canary immediately. | Preserve evidence; repair invariant before re-enable. |
| Cost spike from listener fan-out | Disable/reduce affected feature/query scope. | Use logs/metrics to find reconnect/query amplification; retain security/recovery controls. |
8. Mandatory local game-day sequence
- Run Rules deny matrix and checkout replay/contention tests.
- Run scatter/sequential/hot-document local benchmark and save p50/p95/p99/error JSON.
- Disconnect the client, edit cart, assert pending/cache indicators, then reconnect and verify synchronization.
- Attempt offline checkout and assert it is blocked.
- Inject one materialized-summary mismatch and reconcile it.
- Simulate accidental deletion, restore to a separate local target/fixture, compare count/hash/config evidence and measure local drill RPO/RTO.
- Simulate deployment canary failure and assert the feature flag/rollback policy trips.
-
Generate
capstone/managed-verification.jsonlisting index readiness, IAM/App Check, Query Explain/Insights, Key Visualizer, billing/quotas, backup/PITR, regional p95/p99 asVERIFY_MANAGED.
The architecture survives controlled failures locally, every local measurement is labeled with its scope, and production-only unknowns are visible as release gates rather than hidden assumptions.
Production judgment and bridge
Passing the game day does not make Firestore universally correct. It makes the current AtlasMart decision auditable. Lesson 5 packages the evidence into an architecture review, identifies known limits and vendor dependencies, and states concrete criteria for choosing another database.
Knowledge check
- Why is emulator p99 not production p99?
- What should happen if a new composite query works locally but fails managed staging?
- Why is cached order state labeled differently?
- What is the difference between multi-region availability and PITR/backup recovery?
- Why must rollback criteria exist before a canary starts?
Review the answers
1. The emulator runs locally with different network/backend/index/limit/contention/billing behavior.
2. Keep the feature disabled, verify/build the production index, then rerun the managed query evidence.
3. Cache can be stale and cannot prove current server state.
4. Replication keeps the service/data available through infrastructure failures; recovery mechanisms address historical/logical corruption and operator mistakes.
5. Because once the fault appears, the team needs a deterministic stop condition rather than debating acceptable damage during the incident.
Summary and next step
This lesson established the working contract for Load-Test Hotspots, Test Offline/Conflict Paths, Run Security Tests, Enable Backup/PITR, Monitor Queries, and Rehearse Recovery. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.
Next, continue to Present the Architecture with Measured Cost/Latency, Data-Governance Controls, Known Limits, and Criteria for Choosing Another Database.
Authoritative references
- Firebase: Cloud Firestore documentation — Standard Native/Core application semantics.
- Firebase: Security Rules conditions — Rules are not filters and server libraries bypass Rules in favor of IAM/ADC.
- Firebase: Connect to the Cloud Firestore emulator — demo projects and emulator/production differences.
- Firebase: Transactions and batched writes — retries, atomicity and offline constraints.
- Firebase: Firestore best practices — hotspot avoidance and gradual traffic ramping.
- Firebase: Firestore pricing — document/index-entry/listener/aggregation billing in Standard.
- Firebase: Vector search — vector indexes, dimensions, query limits and supported server SDKs.
- Google Cloud: Firestore release notes — current Enterprise/Pipeline launch status and product changes.
- Google Cloud: Core and Pipeline query interfaces — Standard/Enterprise indexing and query-interface differences.
- Google Cloud: Firestore with MongoDB compatibility overview — serverless compatibility surface, not MongoDB itself.
- Google Cloud: Firestore backups and restore — managed recovery mechanisms.
- Google Cloud: Point-in-time recovery — historical recovery window and managed semantics.