Chapter 15 · Scaling and Hotspots: Key Distribution, Index Fan-Out, Sequential Values, and Ramp-Up
Load-Test a High-Write Model, Identify Hotspots, and Redesign Keys, Indexes, or Data Distribution
Build a repeatable AtlasMart write-load harness, compare key/index shapes with p50/p95/p99 evidence, diagnose hotspots and document a production-safe redesign decision.
1. From “load test” to a diagnostic experiment
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
A useful Firestore load test is not “send as many writes as possible.” It starts with a hypothesis: random keys should distribute document writes better than sequential keys; a hot global document should exhibit contention; removing unused index fan-out should reduce managed write work; gradual ramp should be safer than an instant burst on a cold range. The experiment records enough context to accept or reject each hypothesis without turning one laptop result into a production guarantee.
AtlasMart continues the same mandatory environment used in
Chapters 01–14: project ID
demo-atlasmart-firestore, Standard edition /
Native mode / (default) database, Firestore
emulator 127.0.0.1:8080, Authentication emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase CLI
15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node.js SDK
14.4.0 carrying
@google-cloud/firestore 9.1.0,
@firebase/rules-unit-testing 5.0.2, and Node.js
22+. Mandatory benchmarks remain emulator-only and no-cost.
They teach measurement mechanics and relative shapes; they do
not certify production throughput, split behavior,
Key Visualizer patterns, billing, or regional latency.
Firebase JavaScript SDK 12.19.0 was released 9
September 2026. Firebase Admin Node.js 14.4.0 was
released 10 September 2026 and uses
@google-cloud/firestore 9.1.0. Firebase CLI
15.30.0 was released 9 September 2026. Standard
Native documentation still describes 500/50/5 gradual warm-up
and the 500 writes/s constraint for a collection with a
monotonically changing indexed field. Enterprise
Native indexes are optional rather than automatic; unindexed
queries can scan, so index and cost reasoning is different
even though document/key-range hotspots still exist.
Learning outcomes
Construct a repeatable high-write test plan with controlled variables, raw evidence, percentiles, error taxonomy and cleanup.
Compare random-key, sequential-key and hot-document shapes without conflating emulator observations with production split behavior.
Diagnose whether a production symptom points to document contention, document-key hotspot, index-key hotspot, fan-out or cold-range ramp.
Redesign keys, indexes or state partitioning and define an explicit rollback/verification plan.
Produce a decision record that can be reproduced later instead of a one-off benchmark screenshot.
2. Define the experiment before writing the generator
| Control | Record explicitly | Why it matters |
|---|---|---|
| Environment | emulator vs managed; project/database/edition/mode/location | Determines which backend mechanisms the result can represent |
| Dataset | document count, field sizes, array/map shape, existing indexes | Changes storage/index work |
| Key strategy | auto/random, sequential, shard prefix, one hot document | Defines key distribution and contention |
| Load shape | concurrency, target rate, stage duration, warmup, burst/ramp | Changes split/queue/contention behavior |
| Metrics | p50/p95/p99/max, success/error rate, retries, ops/s | Tail behavior is more informative than average alone |
| Security path | Admin/server vs client Rules path | Adds authorization/network work and changes failure classes |
| Cleanup | isolated collection prefix + recursive deletion | Prevents stale state from contaminating later runs |
3. Canonical local harness
{ "name": "atlasmart-firestore-ch15", "private": true, "type": "module", "engines": { "node": ">=22" }, "dependencies": { "firebase-admin": "14.4.0" }, "devDependencies": { "firebase-tools": "15.30.0" }, "scripts": { "emulators": "firebase emulators:start --only firestore --project demo-atlasmart-firestore", "bench": "node bench-writes.mjs", "reset": "node reset-ch15.mjs" }}
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore, FieldValue } from "firebase-admin/firestore";import { randomUUID } from "node:crypto";initializeApp({ projectId: "demo-atlasmart-firestore" });const db = getFirestore();const percentile = (sorted, p) => { if (!sorted.length) return null; const index = Math.min(sorted.length - 1, Math.ceil((p / 100) * sorted.length) - 1); return sorted[index];};async function runStage({ mode, writes, concurrency }) { const latencies = []; let next = 0; let errors = 0; async function worker(workerId) { while (true) { const i = next++; if (i >= writes) return; const started = performance.now(); try { if (mode === "hot-document") { await db.doc("benchmarks/ch15-hot").set({ n: FieldValue.increment(1), updatedAt: FieldValue.serverTimestamp() }, { merge: true }); } else { const id = mode === "sequential" ? `evt-${String(i).padStart(10, "0")}` : `evt-${randomUUID()}`; await db.doc(`benchmarks/ch15-${mode}/events/${id}`).set({ mode, ordinal: i, observedAt: new Date().toISOString(), workerId }); } } catch (error) { errors++; console.error(JSON.stringify({ mode, i, code: error.code, message: error.message })); } finally { latencies.push(performance.now() - started); } } } const t0 = performance.now(); await Promise.all(Array.from({ length: concurrency }, (_, i) => worker(i))); const elapsedMs = performance.now() - t0; latencies.sort((a, b) => a - b); return { mode, writes, concurrency, elapsedMs, observedOpsPerSecond: writes / (elapsedMs / 1000), errors, latencyMs: { p50: percentile(latencies, 50), p95: percentile(latencies, 95), p99: percentile(latencies, 99), max: latencies.at(-1) } };}for (const stage of [ { mode: "random", writes: 600, concurrency: 12 }, { mode: "sequential", writes: 600, concurrency: 12 }, { mode: "hot-document", writes: 250, concurrency: 12 }]) { console.log(JSON.stringify(await runStage(stage)));}
import { appendFile } from "node:fs/promises";export async function recordStage({ name, run }) { const startedAt = new Date().toISOString(); const result = await run(); const record = { schemaVersion: 1, startedAt, finishedAt: new Date().toISOString(), environment: "firestore-emulator", projectId: "demo-atlasmart-firestore", ...result, disclaimer: "Local emulator result; not production Firestore capacity evidence." }; await appendFile("ch15-results.jsonl", JSON.stringify(record) + "\n"); return record;}
Wrap each benchmark stage with recordStage() so the
result file becomes the source of truth. Add Git commit SHA and
exact firestore.indexes.json hash if running from a
repository. Never overwrite prior runs; append or timestamp them
so regression analysis has history.
4. Diagnosis matrix: symptoms are clues, not proofs
| Observed managed symptom | Likely mechanism to investigate | Evidence to seek | Typical redesign |
|---|---|---|---|
| High latency/errors on one document path | Single-document contention | path-level logs, transaction retries, Key Visualizer narrow hot key | Shard/partition state; remove global counter/lock |
| Diagonal hot band across document keys | Sequential IDs | Key Visualizer key pattern + ID distribution | Scatter/random IDs |
| Index-key hotspot while document keys look healthy | Monotonic indexed field | index-key heatmap, index manifest, field distribution | Exempt unused field or shard required index |
| Broad write latency increase after schema/index change | Index fan-out | before/after index manifest, doc shape, index write activity | Remove unused indexes; split payload; reduce arrays/maps |
| Tail spike immediately after cutover to new collection | Cold-range/sudden ramp | migration timeline, ops/s stages, Key Visualizer/latency | staged cohorts; 500/50/5-style warm-up in Standard |
| Enterprise query slow but writes healthy | Unindexed scan | Query Explain/scan evidence where supported | create targeted index or redesign query |
Do not diagnose from one metric. A p99 spike plus a diagonal Key Visualizer band is much stronger evidence of sequential keys than a p99 spike alone, which could also come from network or downstream application work.
5. Deliberately wrong approach: benchmark production until it breaks
This can create customer impact, unexpected billing, autoscaling state changes and noisy incident signals. It also lacks repeatability because live traffic is uncontrolled. The repair is a dedicated pre-production project with production-like region/index/schema, explicit billing budget/alerts, synthetic isolated data, declared stop conditions and an owner who can abort the test. The mandatory academy path remains emulator-only; production-scale testing is optional and must be deliberately provisioned.
If no eligible managed scan was generated, state that. If the load test ran only on the emulator, label every metric accordingly. “Expected” should describe the evidence pattern to look for, not invented numbers.
6. Redesign example: AtlasMart event ingestion
Before: IDs are event-000001...,
ingestedAt is automatically indexed despite no
query needing it, a 2,000-value debug array is indexed, and all
migration traffic switches at once.
After: auto/scatter IDs,
ingestedAt and debug payload exemptions, only
query-backed composites, bounded arrays or separate child
documents, sticky-cohort migration and staged ramp. If a
required chronological index itself becomes limiting, add
timestamp sharding and accept the read merge cost.
The redesign is not “faster by definition.” It is a set of hypotheses with specific expected evidence: broader key distribution, lower index-write fan-out, stable tails under staged growth and unchanged query-contract correctness.
7. Benchmark report template
{ "change": "events-v2 key/index redesign", "environment": { "kind": "emulator | preproduction-managed", "edition": "Standard", "mode": "Native", "location": "record when managed", "sdk": "firebase-admin 14.4.0 / @google-cloud/firestore 9.1.0" }, "workload": { "datasetDocuments": "record actual", "concurrency": "record actual", "stages": "record actual", "keyStrategy": "record actual", "indexManifestSha256": "record actual" }, "evidence": { "rawResultsFile": "ch15-results.jsonl", "p50Ms": "measured", "p95Ms": "measured", "p99Ms": "measured", "errorRate": "measured", "keyVisualizerScan": "URL or NOT_AVAILABLE" }, "decision": "accept | hold | rollback", "reason": "tie directly to evidence", "knownNonGuarantees": [ "emulator results are not production capacity evidence" ]}
Verification checklist and cleanup
- All benchmark stages are bounded and isolated from course fixtures.
- Raw results include environment and load parameters.
- Reports show p50/p95/p99 and errors, not average alone.
- The diagnosis names document, document-key, index-key, fan-out or ramp mechanism explicitly.
- Every index/key redesign has query-contract and security regression tests.
- Managed load tests, if performed, use a dedicated environment, billing controls and stop conditions.
- Rollback retains enough evidence to explain the failure.
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore } from "firebase-admin/firestore";initializeApp({ projectId: "demo-atlasmart-firestore" });const db = getFirestore();await db.recursiveDelete(db.collection("benchmarks"));console.log("Chapter 15 benchmark data removed from emulator.");
Production judgment and bridge to Chapter 16
A scalable Firestore model combines distributable keys, bounded
contention, query-backed indexes, controlled fan-out and gradual
traffic changes. Performance evidence is meaningful only when
the environment and workload are recorded. Chapter 16 changes
the read side of the equation: server-side count(),
sum() and avg() can avoid downloading
every document, but aggregation cost and latency still depend on
the query/index work they execute.
Knowledge check
- Why is maximum achieved ops/s a poor standalone benchmark result?
- What distinguishes a diagnosis from a symptom?
- When may Key Visualizer be absent?
- What must follow a key/index redesign?
- What is the mandatory academy load-test environment?
Review the answers
1. It hides tail latency, errors, retries, workload shape and whether the test used the same key/index pattern as production.
2. A diagnosis ties multiple pieces of evidence to a specific mechanism such as one hot document, sequential keys, sequential index values, fan-out or sudden ramp.
3. When managed traffic did not meet scan eligibility or no managed project was used; the report should say NOT_AVAILABLE rather than invent a scan.
4. Query-contract, security and correctness regression tests plus staged performance verification and rollback readiness.
5. The local Firestore emulator; managed production-like load tests are optional and require explicit project/billing/safety controls.
Summary
Chapter 15 ends with a reusable performance engineering method: define the mechanism hypothesis, control the workload, preserve raw evidence, diagnose with multiple signals, redesign the right layer, and verify correctness before scaling traffic.
Authoritative references
- Understand reads and writes at scale
- Firestore best practices
- Standard index overview and indexing best practices
- Sharded timestamps
- Key Visualizer overview
- Key Visualizer document-key patterns
- Enterprise Native Core/Pipeline overview
- Enterprise Native index overview
- Enterprise latency troubleshooting
- Firestore quotas and limits
- Firebase current releases
- Firebase Admin Node.js SDK release notes