Chapter 27 · Production Capstone: Design, Secure, Scale, Search, Recover, and Operate a Firestore Application

Load-Test Hotspots, Test Offline / Conflict Paths, Run Security Tests, Enable Backup / PITR, Monitor Queries, and Rehearse Recovery

Attack the architecture with hotspot, contention, offline, security, deployment and recovery game days while preserving emulator-versus-production evidence boundaries.

Advanced · 180–240 minuteshotspots · offline/conflicts · security tests · observability · backup/PITR · recoveryNode 22+ · Firebase CLI course baseline 15.30.0 · JS SDK 12.19.0 · Admin SDK 14.4.0 · rules-unit-testing 5.0.2Mandatory capstone demo project + Emulator Suite/no-cost · managed production verification explicitly separatedLast reviewed: 17 September 2026

1. AtlasMart pre-launch game day: prove the design under failure

A green happy-path test suite is not a production argument. The capstone now attacks the architecture with the failures most likely to produce expensive incidents: hotspot keys, transaction contention, offline edits, stale cache, cross-tenant access, leaked listeners, duplicate retries, missing production indexes, logical deletion, and a recovery cutover. Each test records what local evidence proves and what still needs a managed canary.

Learning outcomes
  • Load-test distribution and distinguish local regression data from production capacity evidence.
  • Test offline/cache/conflict paths without claiming offline correctness for checkout.
  • Run Security Rules and trusted-backend authorization failure cases.
  • Connect Query Explain/Insights/hotspot evidence to a bounded managed verification plan.
  • Rehearse recovery and deployment rollback with RPO/RTO and configuration completeness.
Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 27 reproducibility baseline · reviewed 17 September 2026

AtlasMart keeps project ID demo-atlasmart-firestore, Standard edition / Native mode / Core operations, database (default), Node.js 22+, Firebase CLI course baseline 15.30.0, Firebase JavaScript SDK 12.19.0, Firebase Admin Node SDK 14.4.0, @firebase/rules-unit-testing 5.0.2, Firestore emulator 127.0.0.1:8080, Auth emulator 127.0.0.1:9099, and Emulator UI 127.0.0.1:4000. The mandatory capstone uses only a demo- project and local tooling. Managed location, production IAM/App Check, real composite/vector indexes, Query Explain/Insights, billing, quotas, Key Visualizer, scheduled backups/PITR, CMEK/network controls, Enterprise/Pipeline, and MongoDB-compatibility validation are optional managed evidence and are never inferred from emulator success. Where a managed Standard example needs a concrete location, this chapter uses us-central1 only as an explicit example—not a universal recommendation.

2. Load test the shape, not a magic throughput number

The local benchmark compares key distributions and contention patterns using the same dataset/concurrency. It reports distributions (p50/p95/p99/max), error counts and operation mix. The absolute emulator throughput is not a Firestore production capacity number.

Local distribution benchmark sketch
const stages = [  { name: "scatter", id: i => crypto.randomUUID() },  { name: "sequential", id: i => `event_${String(i).padStart(8,"0")}` },  { name: "hot-doc", id: i => "global_counter" }];for (const stage of stages) {  const latencies = [];  let errors = 0;  await runConcurrent(32, 1000, async i => {    const t0 = performance.now();    try { await writeFixture(stage.id(i), i); }    catch { errors++; }    finally { latencies.push(performance.now() - t0); }  });  console.log(summarize(stage.name, latencies, errors));}

Production follows hotspot physics even with automatic scaling. Scatter IDs reduce sequential key-range concentration, but a single hot document still cannot be split into parallel independent documents. For a new key range, use gradual traffic ramping such as the documented 500/50/5 best practice rather than treating any one number as a universal hard capacity limit.

3. Offline/conflict game day

Scenario Expected product behavior Invariant
Cart edit while offline UI marks pending/local state; write may synchronize later. Server recalculates authoritative checkout price/inventory.
Two offline devices edit cart quantity Last synchronized write may win unless app adds stronger merge semantics. Cart convenience state must not represent inventory reservation.
Checkout while offline Blocked with explicit online-required state. No offline transaction promise for inventory/order invariant.
Order listener loses network Last cached order may render as stale/cache state. UI does not claim server-confirmed payment/shipment from cache alone.
Offline product-state assertion
function orderBadge(snapshot) {  if (!snapshot.exists()) return "missing";  if (snapshot.metadata.fromCache) return "last known — reconnecting";  if (snapshot.metadata.hasPendingWrites) return "local update pending";  return `server confirmed: ${snapshot.data().status}`;}

4. Security game day: prove denies, not only allows

  • Unauthenticated cart read → DENY.
  • Cross-user cart read/write → DENY.
  • Client attempts order status mutation → DENY.
  • Client attempts product price mutation → DENY.
  • Server checkout with valid service identity but unauthorized user/business scope → application authorization DENY.
  • Replay checkout idempotency key → existing order, no repeated inventory decrement.
  • Overbroad product query incompatible with Rules → DENY.

App Check, if later enforced, adds an app-attestation gate; it does not change the requirement for user authentication, Rules/IAM and server business authorization.

5. Query and hotspot evidence in a managed canary

The emulator cannot establish index readiness, Query Explain scan counts, Query Insights normalized patterns, Key Visualizer heatmaps, managed p95/p99 or billing. The release checklist therefore names a small isolated managed canary:

Check Evidence Stop condition
Composite query Index state READY + successful query + Query Explain when supported. Missing index/building/error → feature flag stays off.
High-frequency browse p50/p95/p99 + returned/scanned work + error codes. Regression beyond approved SLO/error budget.
Hot write path Latency/contention metrics + key distribution evidence. Narrow hot key/document signature or deadline errors.
Cost Measured reads/index entries/writes/listener updates + billing/usage metrics. Worst-case slope exceeds approved envelope.
Do not benchmark until it breaks

A production canary is bounded by traffic, duration, budget, synthetic data and stop conditions. It is not permission to discover capacity by uncontrolled load on customer data.

6. Recovery game day: availability is not recoverability

AtlasMart simulates accidental deletion locally by exporting deterministic fixtures into a separate recovery namespace, deleting a subset, restoring into an isolated target, validating counts/hashes/config inventory, and switching a test client. This teaches the runbook without pretending the emulator implements managed scheduled backups/PITR.

Recovery evidence record
{  "incident": "simulated accidental order deletion",  "declaredAt": "2026-09-17T10:00:00Z",  "recoveryPoint": "2026-09-17T09:57:00Z",  "target": "demo-atlasmart-recovery",  "validation": {    "documentCountMatch": true,    "sampleHashMatch": true,    "rulesShaMatch": true,    "indexesShaMatch": true,    "ttlConfig": "VERIFY_MANAGED_SEPARATELY"  },  "cutoverValidatedAt": "2026-09-17T10:08:30Z",  "measuredRpoMinutes": 3,  "measuredRtoMinutes": 8.5}

For the managed optional drill, scheduled backup, PITR, restore/clone permissions, target identity, data validation, Rules/index/config inventory, application cutover and rollback are separate checkpoints. Multi-region replication does not undo a logically valid destructive write.

7. Deployment rollback game day

Fault Immediate response Data/config response
New query missing managed index Disable feature flag. Wait/build/verify index; do not remove unrelated indexes.
Rules update causes deny spike Restore reviewed compatible ruleset if safe. Confirm mixed schema/client versions before rollback.
Backfill introduces mismatch Pause resumable backfill; keep dual reader. Reconcile from checkpoint; do not blindly reverse all data.
Checkout canary correctness failure Stop canary immediately. Preserve evidence; repair invariant before re-enable.
Cost spike from listener fan-out Disable/reduce affected feature/query scope. Use logs/metrics to find reconnect/query amplification; retain security/recovery controls.

8. Mandatory local game-day sequence

  1. Run Rules deny matrix and checkout replay/contention tests.
  2. Run scatter/sequential/hot-document local benchmark and save p50/p95/p99/error JSON.
  3. Disconnect the client, edit cart, assert pending/cache indicators, then reconnect and verify synchronization.
  4. Attempt offline checkout and assert it is blocked.
  5. Inject one materialized-summary mismatch and reconcile it.
  6. Simulate accidental deletion, restore to a separate local target/fixture, compare count/hash/config evidence and measure local drill RPO/RTO.
  7. Simulate deployment canary failure and assert the feature flag/rollback policy trips.
  8. Generate capstone/managed-verification.json listing index readiness, IAM/App Check, Query Explain/Insights, Key Visualizer, billing/quotas, backup/PITR, regional p95/p99 as VERIFY_MANAGED.
Expected state

The architecture survives controlled failures locally, every local measurement is labeled with its scope, and production-only unknowns are visible as release gates rather than hidden assumptions.

Production judgment and bridge

Passing the game day does not make Firestore universally correct. It makes the current AtlasMart decision auditable. Lesson 5 packages the evidence into an architecture review, identifies known limits and vendor dependencies, and states concrete criteria for choosing another database.

Knowledge check

  1. Why is emulator p99 not production p99?
  2. What should happen if a new composite query works locally but fails managed staging?
  3. Why is cached order state labeled differently?
  4. What is the difference between multi-region availability and PITR/backup recovery?
  5. Why must rollback criteria exist before a canary starts?
Review the answers

1. The emulator runs locally with different network/backend/index/limit/contention/billing behavior.

2. Keep the feature disabled, verify/build the production index, then rerun the managed query evidence.

3. Cache can be stale and cannot prove current server state.

4. Replication keeps the service/data available through infrastructure failures; recovery mechanisms address historical/logical corruption and operator mistakes.

5. Because once the fault appears, the team needs a deterministic stop condition rather than debating acceptable damage during the incident.

Summary and next step

This lesson established the working contract for Load-Test Hotspots, Test Offline/Conflict Paths, Run Security Tests, Enable Backup/PITR, Monitor Queries, and Rehearse Recovery. Keep its edition/mode assumptions, trust boundary, verification evidence, and operational constraints explicit when reusing the pattern.

Next, continue to Present the Architecture with Measured Cost/Latency, Data-Governance Controls, Known Limits, and Criteria for Choosing Another Database.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.