Chapter 06 · Indexes: Automatic/Composite/Collection-Group/Vector, Exemptions, and Index Cost
Index Exemptions for Large Strings / Arrays / Maps, Sequential Fields, TTL Fields, and Write / Storage Cost
Design Firestore index exemptions for large strings, arrays, maps, sequential and TTL fields while reasoning about index-entry limits, storage, writes, and vector indexes.
Learning outcomes
AtlasMart has a 5–20 KiB product description, nested vendor metadata, user-generated tag arrays, event timestamps, and future TTL fields. Default indexing every surface would be easy—but it can waste storage, amplify writes, approach per-document index-entry limits, and create a hotspot on sequential fields that the application never queries.
Design single-field exemptions for large unqueried strings, arrays, maps, sequential fields, and TTL timestamps.
Explain Standard index-entry limits and why array/map cardinality can be more dangerous than document count.
Distinguish document-write billing from index storage/index-entry read billing in Standard, and avoid importing Enterprise unit pricing into Standard reasoning.
Use a deterministic comparison model and optional managed evidence without claiming emulator timings predict production p95/p99.
Treat vector indexes as explicit manual structures with dimensions and query ownership, not as automatic indexing of arbitrary arrays.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart continues the same environment used in Chapters
01–05: project ID demo-atlasmart-firestore,
Standard edition / Native mode /
(default) database for the main labs, Firestore
emulator 127.0.0.1:8080, Authentication emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase CLI
15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node SDK
14.4.0 (with
@google-cloud/firestore 9.1.0),
@firebase/rules-unit-testing 5.0.2,
and Node.js 22+. The Enterprise-only exercise in Lesson 4 uses
an isolated emulator configuration rather than mutating the
Standard lab.
The authoring environment did not execute Firebase emulators or a billed Firestore project. Local commands below are deterministic exercises to run on your machine; any shown output is labeled as an expected invariant, not captured benchmark evidence. The Firestore emulator does not reproduce production composite-index enforcement, managed index build/backfill state, billing, or production Query Explain metrics. Those observations are separated into optional managed-project checks.
1. Exemption means “this field is not a query key here”
An automatic-index exemption is a schema contract: AtlasMart agrees that a field will not participate in the corresponding automatic query/order modes. The benefit is reduced index storage and write-side fan-out. The cost is capability: a future screen cannot suddenly filter/order on the exempted field without revisiting the index design.
| Field | Why default indexing is questionable | Safe decision when not queried |
|---|---|---|
description |
Long text; only an indexed prefix participates in Standard query ordering/comparison and storage grows | Exempt; use dedicated text search later if requirements demand it. |
attributes map |
Recursive subfields can multiply automatic surfaces as vendors add keys | Exempt the map if its children are display-only; promote intentionally queryable attributes to governed fields. |
searchTokens large array |
Membership entries grow with array cardinality and composites can multiply the effect | Do not use unbounded token arrays as improvised full-text search; exempt/remodel. |
ingestedAt sequential timestamp |
Indexed monotonic values can hotspot a high-write collection | Exempt if not queried; keep a separate queryable/sharded access pattern if needed. |
expiresAt TTL timestamp |
Sequential lifecycle field; TTL does not require it to remain queryable | Keep TTL policy but exempt indexing when the app does not query expiration. |
2. Hard index limits make unbounded maps/arrays a correctness problem
In Standard, one document may have at most 40,000 total index entries across automatic and manual indexes; one index entry is limited to 7.5 KiB; the combined index-entry size per document is limited to 8 MiB; and indexed field values beyond 1,500 bytes are truncated for index purposes. These are not merely cost warnings: crossing index limits can make a write or index build fail.
Arrays are especially easy to underestimate. An array membership index creates entries per array value, and a manual index containing an array field can produce an entry for each member combined with the other indexed fields. A growing product field with thousands of tags is therefore not just bad taxonomy—it is an index-shape risk.
searchTokens array
as a search engine.
It couples document size, index entry count, write amplification, ranking quality, and query limitations into one brittle field. For true text search, use a capability designed for search or evaluate current Enterprise text-search support where appropriate; do not force every token into Standard array indexing.
3. Express exemptions and TTL policy together
{ "indexes": [ { "collectionGroup": "catalogItems", "queryScope": "COLLECTION", "fields": [ { "fieldPath": "category", "order": "ASCENDING" }, { "fieldPath": "price", "order": "ASCENDING" } ] }, { "collectionGroup": "orders", "queryScope": "COLLECTION_GROUP", "fields": [ { "fieldPath": "tenantId", "order": "ASCENDING" }, { "fieldPath": "createdAt", "order": "DESCENDING" } ] }, { "collectionGroup": "catalogItems", "queryScope": "COLLECTION", "fields": [ { "fieldPath": "tags", "arrayConfig": "CONTAINS" }, { "fieldPath": "rating", "order": "DESCENDING" } ] } ], "fieldOverrides": [ { "collectionGroup": "catalogItems", "fieldPath": "description", "indexes": [] }, { "collectionGroup": "catalogItems", "fieldPath": "attributes", "indexes": [] }, { "collectionGroup": "catalogEvents", "fieldPath": "expiresAt", "ttl": true, "indexes": [] } ]}
The expiresAt example intentionally combines
ttl: true with an empty indexes array.
Current index-definition syntax permits a field-level
configuration to carry TTL plus indexing configuration. The
mandatory local lab does not rely on TTL deletion—the emulator
is not a proof of managed asynchronous TTL behavior. Chapter 21
will treat TTL timing, billing, events, and legal/lifecycle
consequences in depth.
4. Sequential fields: understand the specific 500-write/s documented case
Firestore documentation calls out a Standard pattern where an indexed field that increases or decreases sequentially across documents, such as a timestamp in a high-write collection, can cap writes to that collection at 500 writes per second. Exempting the field can bypass that particular index constraint when you do not query it.
This is not a blanket statement that “Firestore only handles 500 writes/s.” Document IDs, key ranges, other indexes, contention, traffic ramp, and schema distribution all matter. Chapter 15 will cover hotspot physics and the 500/50/5 ramp concept. Here, the actionable rule is narrower: do not maintain a sequential index you have no query owner for.
5. Compare index surfaces without inventing backend cost
const products = [ {id:"p-1001", tags:["outdoor","camera"], descriptionBytes:9000, attributeLeaves:4}, {id:"p-1002", tags:["studio","camera"], descriptionBytes:12000, attributeLeaves:6}, {id:"p-1003", tags:["outdoor","iot"], descriptionBytes:7000, attributeLeaves:5}];function relativeSurface(p,{indexDescription,indexAttributes}){ // Relative design metric only: scalar directions + array members + optional nested leaves. return 2*5 + p.tags.length + (indexDescription?2:0) + (indexAttributes?2*p.attributeLeaves:0);}for (const p of products) console.log(p.id, { defaultLike: relativeSurface(p,{indexDescription:true,indexAttributes:true}), exempted: relativeSurface(p,{indexDescription:false,indexAttributes:false})});console.log("Do not convert this model directly into Firestore billing or latency.");
The model demonstrates direction: exempting large, unqueried surfaces lowers index maintenance. It deliberately does not reproduce Firestore's exact internal entry-size accounting. For exact index-entry size, quotas, storage, or billable reads, use the official sizing/pricing formulas and managed metrics for your actual schema.
6. Standard cost semantics: separate storage, writes, and query index reads
In Standard edition, a document set/update
is billed as a document write; automatic index-entry writes are
not billed as separate document writes, but indexes consume
storage and influence write work/latency. Query billing can also
include index entries read: current pricing exempts queries with
up to one range field from index-entry read charges, while
queries with multiple range fields can incur index-entry read
charges. KNN vector queries have a different batching rule for
vector index entries.
Do not copy those semantics into Enterprise. Enterprise uses read/write units based on byte tranches and index-entry writes can consume write units. Lesson 4 isolates that model explicitly.
7. Vector indexes are manual, dimensional contracts
A vector embedding may look like an array of numbers, but
Firestore vector search uses a dedicated
vectorConfig index mode. Current vector indexes use
the flat index type and support dimensions up to 2048. AtlasMart
will not generate embeddings in this chapter; Chapter 17 owns
embedding model/version/quality. Here, the goal is to recognize
that a vector index has explicit dimensions and a
nearest-neighbor query owner.
{ "indexes": [ { "collectionGroup": "catalogItems", "queryScope": "COLLECTION", "fields": [ { "fieldPath": "embedding", "vectorConfig": { "dimension": 8, "flat": {} } } ] } ], "fieldOverrides": []}
The toy dimension is a deterministic teaching placeholder, not a production embedding contract. A real vector index belongs only after the model/provider/version/dimension and retrieval test suite are chosen.
Hands-on lab and verification checklist
-
Extend the Chapter 05 seed with large
descriptionand nestedattributesexactly as shown in Lesson 1. - Run the Chapter 05 query contract suite before and after adding exemptions; expected result IDs must remain identical because those fields are intentionally not query keys.
-
Run
fanout-model.mjsand record the relative surface delta for the fixed fixtures. -
Add a deliberately huge synthetic
searchTokensarray in a throwaway in-memory object and show the model growing; do not write an unbounded stress object into production. - Keep the vector index example out of the mandatory deployment. Record it as “unowned / Chapter 17 pending.”
- Optional managed project: compare index storage/Query Explain/usage evidence before and after an exemption only with bounded synthetic data and a rollback copy of the old config.
Production judgment
Index exemptions are safest when they encode a stable product decision: “this field is display/lifecycle payload, not a query key.” If a future feature changes that decision, treat re-indexing as a migration with build time and cost, not a one-line code edit. Never use emulator timings to claim write-throughput improvements.
Knowledge check
- Why can a large map create unexpected index work?
- What does the 1,500-byte indexed-field limit imply for long text?
- Why might you exempt a TTL timestamp?
- Does exempting an indexed sequential timestamp prove unlimited write throughput?
- Why is a Firestore vector index not equivalent to automatic indexing of a number array?
Review the answers
1. Standard automatic indexing recursively considers map subfields, so adding vendor keys can expand the indexed surface.
2. Only the indexed prefix participates in the index; queries involving truncated values can be inconsistent, so long unqueried text is a strong exemption candidate.
3. TTL needs a timestamp field, but if the application does not query it, keeping its sequential automatic index can add unnecessary performance/storage cost.
4. No. It removes one documented index hotspot; other key-range, document, index, contention, and ramp constraints remain.
5. It is a dedicated vector index mode with fixed dimensions and nearest-neighbor semantics, created manually for a vector-search contract.
Summary and next step
Exemptions reduce work by declaring what is not queryable. Lesson 4 flips the edition: Enterprise Native makes indexes optional by default, so the engineering question becomes when a successful table scan must be replaced by an explicit index.
Authoritative references
- Index types in Cloud Firestore — Standard automatic/manual index modes, scopes, entry limits, exemptions, and index fan-out guidance.
- Manage indexes in Cloud Firestore — Missing-index workflow, roles, build state, CLI/console management, and vector indexes.
-
Cloud Firestore Index Definition Reference
— Current
firestore.indexes.jsonschema, vector configuration, field overrides, and TTL configuration. - Best practices for Cloud Firestore — Index fan-out, sequential-field, TTL, large string/array/map exemption guidance.
- Understand query performance using Query Explain — Planner versus analyze evidence and billing/scan statistics for managed Firestore.
- Enterprise edition index overview — Optional indexing, sparse/dense behavior, and query-performance reasoning in Enterprise Native mode.
- Firestore Native mode Core/Pipeline overview — Standard/Enterprise indexing requirements and interface differences.
- Search with vector embeddings — Vector index management, flat index type, supported dimensions, and vector-search limitations.
- Firestore pricing — Current document/index-entry billing semantics; re-check region and edition before budgeting.
- Firebase release notes — Current CLI/SDK versions used by the pinned lab.
- Firestore release notes — Enterprise Native/Pipeline launch-stage changes and emulator support.