Chapter 15 · Scaling and Hotspots: Key Distribution, Index Fan-Out, Sequential Values, and Ramp-Up

Index Fan-Out, Large Arrays / Maps, High-Cardinality Fields, and Write Latency / Cost

Trace how one document write expands into index mutations, quantify fan-out risk from arrays/maps and remove unnecessary AtlasMart indexes without breaking query contracts.

Advanced · 165–195 minutesindex fan-out · arrays/maps · write amplificationFirebase JS 12.19.0 · Admin 14.4.0 · CLI 15.30.0Standard Native canonical lab · Enterprise differences explicitLast reviewed: September 2026

1. Why a scattered key can still be an expensive write

Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

AtlasMart’s product telemetry uses random IDs and shows no obvious document-key hotspot, yet write tail latency grows after a schema change. The new document contains a 1,200-element tags array, an 800-element audiences array, a deeply nested debug map, and several automatic/composite indexes. Every document write can produce many index-entry mutations. This is index fan-out: one logical write expands into many physical index updates.

Chapter 15 reproducibility baseline · reviewed 17 September 2026

AtlasMart continues the same mandatory environment used in Chapters 01–14: project ID demo-atlasmart-firestore, Standard edition / Native mode / (default) database, Firestore emulator 127.0.0.1:8080, Authentication emulator 127.0.0.1:9099, Emulator UI 127.0.0.1:4000, Firebase CLI 15.30.0, Firebase JavaScript SDK 12.19.0, Firebase Admin Node.js SDK 14.4.0 carrying @google-cloud/firestore 9.1.0, @firebase/rules-unit-testing 5.0.2, and Node.js 22+. Mandatory benchmarks remain emulator-only and no-cost. They teach measurement mechanics and relative shapes; they do not certify production throughput, split behavior, Key Visualizer patterns, billing, or regional latency.

Current documentation check

Firebase JavaScript SDK 12.19.0 was released 9 September 2026. Firebase Admin Node.js 14.4.0 was released 10 September 2026 and uses @google-cloud/firestore 9.1.0. Firebase CLI 15.30.0 was released 9 September 2026. Standard Native documentation still describes 500/50/5 gradual warm-up and the 500 writes/s constraint for a collection with a monotonically changing indexed field. Enterprise Native indexes are optional rather than automatic; unindexed queries can scan, so index and cost reasoning is different even though document/key-range hotspots still exist.

Learning outcomes

01

Trace a Firestore write into document-row and index-entry mutations so index fan-out becomes an explicit part of write design.

02

Identify large strings, arrays, maps, TTL/sequential fields and unused descending/array indexes that can create avoidable Standard write/storage cost.

03

Use the 40,000 index-entry-per-document and 8 MiB total index-entry-size limits as hard design boundaries, not optimization targets.

04

Remove an index only after a query-contract dependency audit and rollback plan.

05

Contrast Standard automatic indexes with Enterprise optional indexes and unindexed-scan behavior.

2. Fan-out is the number of index structures one write must maintain

Document feature Why fan-out can grow Design question
Scalar fields Standard automatically creates ascending/descending single-field entries by default Do queries need both directions for every field?
Arrays Array membership indexing can create one entry per array value and interact with composites Is this really a queryable set, or payload/debug data?
Maps/embedded fields Nested leaf fields can be indexed recursively in Standard Which nested paths are actual query keys?
Composite indexes Each matching document contributes entries according to composite definition Is each composite tied to a named query contract?
Large strings Indexed values consume index storage; values over 1,500 bytes are truncated for indexing If not queried, exempt the field
TTL/sequential timestamps Automatic indexing can add write work and a sequential index edge Exempt when no query depends on it

Firestore documents have a hard maximum of 40,000 index entries per document and 8 MiB total index-entry size. These are failure boundaries. Good designs leave operating margin; they do not deliberately approach them.

3. Use query contracts to justify every index

firestore.indexes.json · keep only query-backed indexes
{  "indexes": [    {      "collectionGroup": "events",      "queryScope": "COLLECTION",      "fields": [        { "fieldPath": "tenantId", "order": "ASCENDING" },        { "fieldPath": "kind", "order": "ASCENDING" },        { "fieldPath": "occurredAt", "order": "DESCENDING" }      ]    }  ],  "fieldOverrides": [    {      "collectionGroup": "events",      "fieldPath": "ingestedAt",      "indexes": []    },    {      "collectionGroup": "events",      "fieldPath": "rawPayload",      "indexes": []    },    {      "collectionGroup": "events",      "fieldPath": "debugAttributes",      "indexes": []    }  ]}

In Chapter 6 AtlasMart retained an index manifest. Chapter 15 adds a rule: every composite index must point to at least one named query contract and every exemption must point to a statement that the field is not queried—or to the replacement access path. Index configuration is code, so review it with the same dependency discipline as an API change.

Deleting a composite is not a “cleanup” until dependency tests prove it.

A composite that appears unused in one service may be required by a mobile release still in circulation. Before deletion, search code, query-contract fixtures, dashboards/jobs and release support windows. Stage the removal in a test environment and preserve rollback instructions.

4. Design-time estimator: catch suspicious arrays/maps before deployment

estimate-fanout.mjs · deterministic design-time warning, not backend billing
function estimateArrayEntries(document, indexedArrayFields) {  return indexedArrayFields.reduce((sum, field) => {    const value = document[field];    return sum + (Array.isArray(value) ? value.length : 0);  }, 0);}const candidate = {  tags: Array.from({ length: 1200 }, (_, i) => `tag-${i}`),  audiences: Array.from({ length: 800 }, (_, i) => `aud-${i}`),  rawPayload: "x".repeat(20_000)};console.log({  approximateArrayMembershipValues: estimateArrayEntries(candidate, ["tags", "audiences"]),  warning: "Actual Firestore index-entry count depends on every single-field and composite index. Verify the real index configuration; do not treat this script as billing evidence."});

This small script is intentionally conservative and incomplete. It counts obvious array membership values so code review can flag an extreme document shape. It cannot compute Firestore’s exact index-entry count because that depends on the full deployed single-field and composite index configuration. Exact backend behavior belongs to official index-entry calculation rules and, where appropriate, production/pre-production measurements.

5. Deliberately wrong approach: disable every index to make writes fast

In Standard Native, that breaks required queries. In Enterprise Native, an unindexed query may still execute, which can hide the mistake until data growth turns a prototype scan into expensive latency. “Fewer indexes” is not the objective; “only indexes that buy a required query/performance outcome” is.

The repair sequence is query inventory → index dependency mapping → candidate removal/exemption → contract tests → bounded write/query measurement → staged deployment → observe → retain rollback. Performance changes without query correctness evidence are incomplete.

6. Bounded before/after experiment

bench-writes.mjs · measure observed latency, never invent it
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore, FieldValue } from "firebase-admin/firestore";import { randomUUID } from "node:crypto";initializeApp({ projectId: "demo-atlasmart-firestore" });const db = getFirestore();const percentile = (sorted, p) => {  if (!sorted.length) return null;  const index = Math.min(sorted.length - 1, Math.ceil((p / 100) * sorted.length) - 1);  return sorted[index];};async function runStage({ mode, writes, concurrency }) {  const latencies = [];  let next = 0;  let errors = 0;  async function worker(workerId) {    while (true) {      const i = next++;      if (i >= writes) return;      const started = performance.now();      try {        if (mode === "hot-document") {          await db.doc("benchmarks/ch15-hot").set({            n: FieldValue.increment(1),            updatedAt: FieldValue.serverTimestamp()          }, { merge: true });        } else {          const id = mode === "sequential"            ? `evt-${String(i).padStart(10, "0")}`            : `evt-${randomUUID()}`;          await db.doc(`benchmarks/ch15-${mode}/events/${id}`).set({            mode,            ordinal: i,            observedAt: new Date().toISOString(),            workerId          });        }      } catch (error) {        errors++;        console.error(JSON.stringify({ mode, i, code: error.code, message: error.message }));      } finally {        latencies.push(performance.now() - started);      }    }  }  const t0 = performance.now();  await Promise.all(Array.from({ length: concurrency }, (_, i) => worker(i)));  const elapsedMs = performance.now() - t0;  latencies.sort((a, b) => a - b);  return {    mode, writes, concurrency, elapsedMs,    observedOpsPerSecond: writes / (elapsedMs / 1000),    errors,    latencyMs: {      p50: percentile(latencies, 50),      p95: percentile(latencies, 95),      p99: percentile(latencies, 99),      max: latencies.at(-1)    }  };}for (const stage of [  { mode: "random", writes: 600, concurrency: 12 },  { mode: "sequential", writes: 600, concurrency: 12 },  { mode: "hot-document", writes: 250, concurrency: 12 }]) {  console.log(JSON.stringify(await runStage(stage)));}

Run the same local write harness against two document shapes: a lean event and a deliberately bloated event with large arrays/maps. Record raw latency distributions and serialized document sizes. Then modify the local firestore.indexes.json manifest and rerun. The emulator does not reproduce managed index write cost faithfully enough to claim a production percentage improvement, but it validates the experiment workflow and catches application-level regressions.

7. Standard vs Enterprise index-cost reasoning

Question Standard Native Enterprise Native
What exists by default? Automatic single-field indexes plus explicitly created composites No indexes by default; create them deliberately
Can a query run without its needed index? Core query normally requires matching index and may return an index-creation error Query can scan without index, potentially increasing latency/cost
What raises write work? Automatic + composite entries touched by changed fields Only indexes you created, but broad/complex indexes can still add write work
Primary failure if you remove too much Query becomes unavailable until index exists Query may continue but scan too much, hiding the regression

Verification checklist

  • Every retained composite maps to a named query contract.
  • Every exemption names the field’s non-query purpose or replacement query path.
  • Large arrays/maps have an explicit reason to be in one document and an index policy.
  • Index limits are treated as hard ceilings, not desired operating points.
  • Before/after results preserve environment, document shape, index manifest and raw latency samples.
  • Enterprise unindexed scans are not described as free.

Production judgment and bridge to Lesson 5

Index fan-out is part of the write path. Reducing it can improve write efficiency, but only when query capability remains intentional. Lesson 5 combines document-key distribution, sequential-index behavior, fan-out and traffic ramp into one repeatable load-test and redesign workflow.

Knowledge check

  1. What is index fan-out?
  2. Why is 40,000 index entries not a tuning target?
  3. Can Enterprise simply omit every index?
  4. Why map indexes to named query contracts?
  5. What must accompany a before/after write-latency claim?
Review the answers

1. The multiplication of physical index-entry mutations caused by one logical document write across applicable single-field and composite indexes.

2. It is a hard maximum per document; production designs should keep margin and avoid shapes that can unpredictably approach it.

3. Queries can run without indexes, but they may scan large collections and become slow/costly; create indexes for important access patterns.

4. It gives each index a verifiable consumer and lets tests detect whether removal changes supported behavior.

5. Environment/edition/mode, dataset and document shape, index manifest, concurrency, raw samples, percentiles, errors and the measurement limitations.

Summary

AtlasMart now treats indexes as write-path structures with measurable amplification, not merely query configuration. Large arrays/maps and redundant indexes are reviewed against query contracts and hard index-entry limits.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.