Chapter 06 · Indexes: Automatic/Composite/Collection-Group/Vector, Exemptions, and Index Cost

Index Exemptions for Large Strings / Arrays / Maps, Sequential Fields, TTL Fields, and Write / Storage Cost

Design Firestore index exemptions for large strings, arrays, maps, sequential and TTL fields while reasoning about index-entry limits, storage, writes, and vector indexes.

Intermediate120–145 minutesExemptions + index-entry limitsFirebase JS 12.19.0 · CLI 15.30.0Last reviewed: September 2026

Learning outcomes

AtlasMart has a 5–20 KiB product description, nested vendor metadata, user-generated tag arrays, event timestamps, and future TTL fields. Default indexing every surface would be easy—but it can waste storage, amplify writes, approach per-document index-entry limits, and create a hotspot on sequential fields that the application never queries.

01

Design single-field exemptions for large unqueried strings, arrays, maps, sequential fields, and TTL timestamps.

02

Explain Standard index-entry limits and why array/map cardinality can be more dangerous than document count.

03

Distinguish document-write billing from index storage/index-entry read billing in Standard, and avoid importing Enterprise unit pricing into Standard reasoning.

04

Use a deterministic comparison model and optional managed evidence without claiming emulator timings predict production p95/p99.

05

Treat vector indexes as explicit manual structures with dimensions and query ownership, not as automatic indexing of arbitrary arrays.

Execution and safety note

Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.

Chapter 06 reproducibility baseline · reviewed 15 September 2026

AtlasMart continues the same environment used in Chapters 01–05: project ID demo-atlasmart-firestore, Standard edition / Native mode / (default) database for the main labs, Firestore emulator 127.0.0.1:8080, Authentication emulator 127.0.0.1:9099, Emulator UI 127.0.0.1:4000, Firebase CLI 15.30.0, Firebase JavaScript SDK 12.19.0, Firebase Admin Node SDK 14.4.0 (with @google-cloud/firestore 9.1.0), @firebase/rules-unit-testing 5.0.2, and Node.js 22+. The Enterprise-only exercise in Lesson 4 uses an isolated emulator configuration rather than mutating the Standard lab.

Evidence boundary for this generated lesson

The authoring environment did not execute Firebase emulators or a billed Firestore project. Local commands below are deterministic exercises to run on your machine; any shown output is labeled as an expected invariant, not captured benchmark evidence. The Firestore emulator does not reproduce production composite-index enforcement, managed index build/backfill state, billing, or production Query Explain metrics. Those observations are separated into optional managed-project checks.

1. Exemption means “this field is not a query key here”

An automatic-index exemption is a schema contract: AtlasMart agrees that a field will not participate in the corresponding automatic query/order modes. The benefit is reduced index storage and write-side fan-out. The cost is capability: a future screen cannot suddenly filter/order on the exempted field without revisiting the index design.

Field Why default indexing is questionable Safe decision when not queried
description Long text; only an indexed prefix participates in Standard query ordering/comparison and storage grows Exempt; use dedicated text search later if requirements demand it.
attributes map Recursive subfields can multiply automatic surfaces as vendors add keys Exempt the map if its children are display-only; promote intentionally queryable attributes to governed fields.
searchTokens large array Membership entries grow with array cardinality and composites can multiply the effect Do not use unbounded token arrays as improvised full-text search; exempt/remodel.
ingestedAt sequential timestamp Indexed monotonic values can hotspot a high-write collection Exempt if not queried; keep a separate queryable/sharded access pattern if needed.
expiresAt TTL timestamp Sequential lifecycle field; TTL does not require it to remain queryable Keep TTL policy but exempt indexing when the app does not query expiration.

2. Hard index limits make unbounded maps/arrays a correctness problem

In Standard, one document may have at most 40,000 total index entries across automatic and manual indexes; one index entry is limited to 7.5 KiB; the combined index-entry size per document is limited to 8 MiB; and indexed field values beyond 1,500 bytes are truncated for index purposes. These are not merely cost warnings: crossing index limits can make a write or index build fail.

Arrays are especially easy to underestimate. An array membership index creates entries per array value, and a manual index containing an array field can produce an entry for each member combined with the other indexed fields. A growing product field with thousands of tags is therefore not just bad taxonomy—it is an index-shape risk.

Wrong approach: use a giant searchTokens array as a search engine.

It couples document size, index entry count, write amplification, ranking quality, and query limitations into one brittle field. For true text search, use a capability designed for search or evaluate current Enterprise text-search support where appropriate; do not force every token into Standard array indexing.

3. Express exemptions and TTL policy together

firestore.indexes.json · exemptions + TTL lifecycle field
{  "indexes": [    {      "collectionGroup": "catalogItems",      "queryScope": "COLLECTION",      "fields": [        { "fieldPath": "category", "order": "ASCENDING" },        { "fieldPath": "price", "order": "ASCENDING" }      ]    },    {      "collectionGroup": "orders",      "queryScope": "COLLECTION_GROUP",      "fields": [        { "fieldPath": "tenantId", "order": "ASCENDING" },        { "fieldPath": "createdAt", "order": "DESCENDING" }      ]    },    {      "collectionGroup": "catalogItems",      "queryScope": "COLLECTION",      "fields": [        { "fieldPath": "tags", "arrayConfig": "CONTAINS" },        { "fieldPath": "rating", "order": "DESCENDING" }      ]    }  ],  "fieldOverrides": [    { "collectionGroup": "catalogItems", "fieldPath": "description", "indexes": [] },    { "collectionGroup": "catalogItems", "fieldPath": "attributes", "indexes": [] },    { "collectionGroup": "catalogEvents", "fieldPath": "expiresAt", "ttl": true, "indexes": [] }  ]}

The expiresAt example intentionally combines ttl: true with an empty indexes array. Current index-definition syntax permits a field-level configuration to carry TTL plus indexing configuration. The mandatory local lab does not rely on TTL deletion—the emulator is not a proof of managed asynchronous TTL behavior. Chapter 21 will treat TTL timing, billing, events, and legal/lifecycle consequences in depth.

4. Sequential fields: understand the specific 500-write/s documented case

Firestore documentation calls out a Standard pattern where an indexed field that increases or decreases sequentially across documents, such as a timestamp in a high-write collection, can cap writes to that collection at 500 writes per second. Exempting the field can bypass that particular index constraint when you do not query it.

This is not a blanket statement that “Firestore only handles 500 writes/s.” Document IDs, key ranges, other indexes, contention, traffic ramp, and schema distribution all matter. Chapter 15 will cover hotspot physics and the 500/50/5 ramp concept. Here, the actionable rule is narrower: do not maintain a sequential index you have no query owner for.

5. Compare index surfaces without inventing backend cost

fanout-model.mjs · deterministic relative comparison
const products = [  {id:"p-1001", tags:["outdoor","camera"], descriptionBytes:9000, attributeLeaves:4},  {id:"p-1002", tags:["studio","camera"], descriptionBytes:12000, attributeLeaves:6},  {id:"p-1003", tags:["outdoor","iot"], descriptionBytes:7000, attributeLeaves:5}];function relativeSurface(p,{indexDescription,indexAttributes}){  // Relative design metric only: scalar directions + array members + optional nested leaves.  return 2*5 + p.tags.length + (indexDescription?2:0) + (indexAttributes?2*p.attributeLeaves:0);}for (const p of products) console.log(p.id, {  defaultLike: relativeSurface(p,{indexDescription:true,indexAttributes:true}),  exempted: relativeSurface(p,{indexDescription:false,indexAttributes:false})});console.log("Do not convert this model directly into Firestore billing or latency.");

The model demonstrates direction: exempting large, unqueried surfaces lowers index maintenance. It deliberately does not reproduce Firestore's exact internal entry-size accounting. For exact index-entry size, quotas, storage, or billable reads, use the official sizing/pricing formulas and managed metrics for your actual schema.

6. Standard cost semantics: separate storage, writes, and query index reads

In Standard edition, a document set/update is billed as a document write; automatic index-entry writes are not billed as separate document writes, but indexes consume storage and influence write work/latency. Query billing can also include index entries read: current pricing exempts queries with up to one range field from index-entry read charges, while queries with multiple range fields can incur index-entry read charges. KNN vector queries have a different batching rule for vector index entries.

Do not copy those semantics into Enterprise. Enterprise uses read/write units based on byte tranches and index-entry writes can consume write units. Lesson 4 isolates that model explicitly.

7. Vector indexes are manual, dimensional contracts

A vector embedding may look like an array of numbers, but Firestore vector search uses a dedicated vectorConfig index mode. Current vector indexes use the flat index type and support dimensions up to 2048. AtlasMart will not generate embeddings in this chapter; Chapter 17 owns embedding model/version/quality. Here, the goal is to recognize that a vector index has explicit dimensions and a nearest-neighbor query owner.

firestore.indexes.vector-example.json · configuration only
{  "indexes": [    {      "collectionGroup": "catalogItems",      "queryScope": "COLLECTION",      "fields": [        { "fieldPath": "embedding", "vectorConfig": { "dimension": 8, "flat": {} } }      ]    }  ],  "fieldOverrides": []}
Do not deploy this just because it exists.

The toy dimension is a deterministic teaching placeholder, not a production embedding contract. A real vector index belongs only after the model/provider/version/dimension and retrieval test suite are chosen.

Hands-on lab and verification checklist

  1. Extend the Chapter 05 seed with large description and nested attributes exactly as shown in Lesson 1.
  2. Run the Chapter 05 query contract suite before and after adding exemptions; expected result IDs must remain identical because those fields are intentionally not query keys.
  3. Run fanout-model.mjs and record the relative surface delta for the fixed fixtures.
  4. Add a deliberately huge synthetic searchTokens array in a throwaway in-memory object and show the model growing; do not write an unbounded stress object into production.
  5. Keep the vector index example out of the mandatory deployment. Record it as “unowned / Chapter 17 pending.”
  6. Optional managed project: compare index storage/Query Explain/usage evidence before and after an exemption only with bounded synthetic data and a rollback copy of the old config.

Production judgment

Index exemptions are safest when they encode a stable product decision: “this field is display/lifecycle payload, not a query key.” If a future feature changes that decision, treat re-indexing as a migration with build time and cost, not a one-line code edit. Never use emulator timings to claim write-throughput improvements.

Knowledge check

  1. Why can a large map create unexpected index work?
  2. What does the 1,500-byte indexed-field limit imply for long text?
  3. Why might you exempt a TTL timestamp?
  4. Does exempting an indexed sequential timestamp prove unlimited write throughput?
  5. Why is a Firestore vector index not equivalent to automatic indexing of a number array?
Review the answers

1. Standard automatic indexing recursively considers map subfields, so adding vendor keys can expand the indexed surface.

2. Only the indexed prefix participates in the index; queries involving truncated values can be inconsistent, so long unqueried text is a strong exemption candidate.

3. TTL needs a timestamp field, but if the application does not query it, keeping its sequential automatic index can add unnecessary performance/storage cost.

4. No. It removes one documented index hotspot; other key-range, document, index, contention, and ramp constraints remain.

5. It is a dedicated vector index mode with fixed dimensions and nearest-neighbor semantics, created manually for a vector-search contract.

Summary and next step

Exemptions reduce work by declaring what is not queryable. Lesson 4 flips the edition: Enterprise Native makes indexes optional by default, so the engineering question becomes when a successful table scan must be replaced by an explicit index.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.