Chapter 06 · Indexes: Automatic/Composite/Collection-Group/Vector, Exemptions, and Index Cost
Single-Field Indexing in Standard Native Mode, Ascending / Descending / Array Entries, and Automatic Behavior
Understand Standard Firestore automatic indexing, ascending/descending/array modes, field exemptions, query scope, and observable index fan-out.
Learning outcomes
AtlasMart's Chapter 05 query suite works locally, but the team
now wants to know what Firestore must maintain on every product
write. A product has scalar fields such as price,
an array such as tags, a large
description, and a nested
attributes map. In Standard Native mode, these are
not just stored values: default automatic indexing creates
additional ordered structures unless you deliberately exempt
fields.
Explain what Standard automatic indexes are and how ascending, descending, and array-contains modes map stored fields to query capabilities.
Distinguish collection-scope automatic indexes from manually configured collection-group/composite/vector indexes.
Trace how one document write fans out into document storage plus index maintenance without pretending the emulator exposes production index I/O.
Read and version firestore.indexes.json as
database configuration, including justified field
exemptions.
Identify when default indexing is safe convenience and when large arrays/maps/strings or sequential fields make it an operational design decision.
Use the Emulator Suite, a Firebase demo project, or an isolated test project for destructive, security-sensitive, billing-sensitive, migration, backup/restore, or write-heavy exercises unless the lesson explicitly marks managed verification as required. Treat shown output as expected evidence unless it is explicitly identified as captured output, and re-check current Firebase/Google Cloud edition, mode, quota, pricing, and security documentation before production execution.
AtlasMart continues the same environment used in Chapters
01–05: project ID demo-atlasmart-firestore,
Standard edition / Native mode /
(default) database for the main labs, Firestore
emulator 127.0.0.1:8080, Authentication emulator
127.0.0.1:9099, Emulator UI
127.0.0.1:4000, Firebase CLI
15.30.0, Firebase JavaScript SDK
12.19.0, Firebase Admin Node SDK
14.4.0 (with
@google-cloud/firestore 9.1.0),
@firebase/rules-unit-testing 5.0.2,
and Node.js 22+. The Enterprise-only exercise in Lesson 4 uses
an isolated emulator configuration rather than mutating the
Standard lab.
The authoring environment did not execute Firebase emulators or a billed Firestore project. Local commands below are deterministic exercises to run on your machine; any shown output is labeled as an expected invariant, not captured benchmark evidence. The Firestore emulator does not reproduce production composite-index enforcement, managed index build/backfill state, billing, or production Query Explain metrics. Those observations are separated into optional managed-project checks.
1. The Standard mental model: every supported query walks an index
In Firestore Standard edition using Native/Core operations, indexes are required for queries. The service automatically maintains single-field indexes for common query shapes and requires manual indexes when one ordered structure must combine fields. Think of the document as the source record and indexes as derived, sorted lookup structures that are updated as the document changes. A successful write therefore has two conceptual effects: persist the document and maintain every index entry affected by the changed fields.
The term automatic index is current Firestore
documentation terminology for the structures historically called
single-field indexes. A manual index is an
explicitly managed index, historically called a composite index.
A query scope says whether an index serves one
collection at a path (COLLECTION) or all
collections with the same ID anywhere in the database
(COLLECTION_GROUP).
| Index mode | What it orders/matches | Typical Core clauses | Standard default |
|---|---|---|---|
ASCENDING |
One non-array field by value, then document name tie-break semantics |
==, range, in,
not-in, ascending sort
|
Automatic collection-scope index |
DESCENDING |
Same field in reverse value order | Same comparisons plus descending sort | Automatic collection-scope index |
CONTAINS |
Array membership entries |
array-contains,
array-contains-any
|
Automatic array membership support |
vectorConfig |
Embedding vector search structure | FindNearest / nearest-neighbor APIs |
Manual; not an automatic scalar index |
Standard automatic indexing recursively considers nested map subfields. That is convenient for small queryable maps, but it is precisely why a large unqueried metadata map can create index work you did not intend. Exempting an unqueried map is a schema/index decision, not cosmetic cleanup.
2. Collection scope is the default; collection-group scope is deliberate
Automatic Standard indexes are collection scoped. AtlasMart can
query catalogItems by
price immediately because every document in that
collection participates in automatic field indexes. But Chapter
05's cross-user orders query uses
collectionGroup("orders"); collection-group support
is not automatically maintained in the same blanket way. When a
compound collection-group query needs a manual index, its scope
must be COLLECTION_GROUP.
This is also why “the field is indexed” is incomplete language.
A usable index contract includes collection ID, query scope,
field order, field mode, edition, and database. An index for
orders at collection scope is not interchangeable
with a collection-group index over every
orders subcollection.
3. Turn Chapter 05 query contracts into configuration-as-code
Start with only indexes justified by a named query contract. The following baseline supports the catalog query “category equality + price order” and the seller order-history collection-group query. It also exempts two product fields that AtlasMart intentionally never filters or orders by.
{ "indexes": [ { "collectionGroup": "catalogItems", "queryScope": "COLLECTION", "fields": [ { "fieldPath": "category", "order": "ASCENDING" }, { "fieldPath": "price", "order": "ASCENDING" } ] }, { "collectionGroup": "orders", "queryScope": "COLLECTION_GROUP", "fields": [ { "fieldPath": "tenantId", "order": "ASCENDING" }, { "fieldPath": "createdAt", "order": "DESCENDING" } ] } ], "fieldOverrides": [ { "collectionGroup": "catalogItems", "fieldPath": "description", "indexes": [] }, { "collectionGroup": "catalogItems", "fieldPath": "attributes", "indexes": [] } ]}
The file is not a performance guarantee. It is a versioned declaration of intended manual indexes and field-level overrides. Production still has an index lifecycle: a newly created index enters a build/backfill operation and becomes usable when ready. Deleting a definition changes future query support only after the managed service applies the change.
{ "firestore": { "rules": "firestore.rules", "indexes": "firestore.indexes.json", "edition": "standard" }, "emulators": { "firestore": { "port": 8080, "edition": "standard" }, "auth": { "port": 9099 }, "ui": { "enabled": true, "port": 4000 } }}
4. Observe data and query correctness locally—do not fake production index evidence
Use the same awkward Chapter 05 fixtures: tied prices, a missing
rank, a null rank, arrays, nested
orders, and deterministic timestamps. Chapter 06 adds a large
description and nested attributes so index exemptions have
something real to govern without changing earlier query
behavior.
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore, Timestamp } from "firebase-admin/firestore";initializeApp({ projectId: "demo-atlasmart-firestore" });const db = getFirestore();const fixed = Timestamp.fromMillis(1760000000000);const catalog = { "p-1001": { category:"camera", price:99, rating:4.8, tags:["outdoor","camera"], published:true, rank:10, sellerId:"seller-a" }, "p-1002": { category:"camera", price:99, rating:4.5, tags:["studio","camera"], published:true, rank:20, sellerId:"seller-a" }, "p-1003": { category:"sensor", price:49, rating:4.8, tags:["outdoor","iot"], published:true, sellerId:"seller-b" }, "p-1004": { category:"sensor", price:149, rating:4.1, tags:["industrial","iot"], published:false, rank:null, sellerId:"seller-b" }, "p-1005": { category:"camera", price:199, rating:4.9, tags:["outdoor","camera"], published:true, rank:30, sellerId:"seller-c" }, "p-1006": { category:"gateway", price:99, rating:4.2, tags:["iot"], published:true, rank:40, sellerId:"seller-c" }};for (const [id, data] of Object.entries(catalog)) { await db.collection("catalogItems").doc(id).set({ ...data, description: `AtlasMart ${id} product description `.repeat(40), attributes: { source:"chapter06", color:id.endsWith("1")?"black":"gray", dimensions:{ unit:"cm", width:10 } }, updatedAt: fixed, schemaVersion: 3 });}const orders = [ ["u-alice","o-1001",{ ownerUid:"u-alice", tenantId:"seller-a", status:"paid", total:198, createdAt:Timestamp.fromMillis(1760000100000) }], ["u-alice","o-1002",{ ownerUid:"u-alice", tenantId:"seller-b", status:"shipped", total:49, createdAt:Timestamp.fromMillis(1760000200000) }], ["u-bob","o-1003",{ ownerUid:"u-bob", tenantId:"seller-a", status:"paid", total:99, createdAt:Timestamp.fromMillis(1760000300000) }]];for (const [uid,id,data] of orders) await db.doc(`users/${uid}/orders/${id}`).set(data);console.log("seeded", Object.keys(catalog).length, "catalog items and", orders.length, "orders");
mkdir atlasmart-firestore-index-lab && cd atlasmart-firestore-index-labnpm init -ynpm i firebase@12.19.0 firebase-admin@14.4.0 @firebase/rules-unit-testing@5.0.2npm pkg set type=modulenpx firebase-tools@15.30.0 --version# copy the chapter's firebase.json, firestore.rules and firestore.indexes.json herenpx firebase-tools@15.30.0 emulators:exec --project demo-atlasmart-firestore --only firestore,auth "node seed.mjs && node query-contracts.mjs"
On the emulator, assert the query results from Chapter 05 still match. That proves your data/schema changes did not break query semantics. It does not prove that a production composite index exists, that an index build completed, or that a production query scanned a certain number of entries.
5. Make fan-out visible with a deterministic model
You can still reason about index fan-out locally. The script below is intentionally a model, not a reverse-engineered billing calculator. It counts the field surfaces AtlasMart chose to make queryable and highlights arrays/maps whose cardinality grows with data. The point is to compare schema alternatives before deployment.
const doc = { category:"camera", price:99, rating:4.8, tags:["outdoor","camera"], description:"x".repeat(5000), attributes:{ color:"black", dimensions:{ unit:"cm", width:10 } }};const exempt = new Set(["description","attributes"]);const scalar = ["category","price","rating"];const arrays = { tags: doc.tags.length };console.table({ automaticScalarDirections: scalar.length * 2, arrayMembersForContains: arrays.tags, exemptLargeFields: exempt.size, note: "Design-time comparison only; not backend index-entry billing"});if (!exempt.has("description")) throw new Error("large unqueried description should be reviewed");
That ignores write amplification, index storage, field-value truncation/limits, array/map entry growth, and hotspot behavior for indexed sequential fields. The repair is not “disable all indexes.” Keep the smallest index surface that satisfies named query contracts, then prove those queries in tests and managed evidence.
6. Hard boundaries that turn schema into index engineering
Current Standard limits include a maximum of 40,000 index entries per document, 100 fields in one composite index, 7.5 KiB per index entry, 8 MiB total index-entry size per document, and 1,500 bytes considered for an indexed field value before truncation. These are hard limits, not tuning targets. Large arrays/maps can approach the entry limit quickly; long unqueried text wastes storage and can create surprising query behavior if you rely on content beyond the indexed prefix.
A separate boundary is throughput: an indexed field whose values increase or decrease sequentially across a high-write collection—such as an ingestion timestamp—can constrain a collection to 500 writes per second in the documented Standard case. If AtlasMart does not query that sequential field, an exemption removes that index hotspot. Do not generalize 500 writes/s to the whole database; it is specifically the documented indexed-sequential-field pattern.
Hands-on lab and verification checklist
- Create the isolated directory and install the pinned packages.
-
Use the Standard
firebase.json, the baselinefirestore.indexes.json, and the deny-by-default Rules from earlier chapters. -
Start Firestore/Auth emulators with project ID
demo-atlasmart-firestore. - Seed the six catalog documents and three nested orders, then rerun Chapter 05 filter/pagination contracts.
-
Run
index-surface-report.mjstwice: once with the exemptions and once with them removed. Record only the model delta, not a claim about production billing. - Save the baseline index file in source control and annotate each manual index with its owning query-test name in a neighboring manifest document.
In a disposable real project/database only, run
firebase firestore:indexes --database="(default)", deploy with
firebase deploy --only firestore:indexes, wait
for the managed index state to become ready, and use Query
Explain on a server-authenticated query. This can consume
quota/billing; record project, database, edition, region,
fixture size, query, index state, and Explain output. Do not
paste a production service-account key into the lab.
Production judgment
Default automatic indexing is a productivity feature, not an excuse to ignore schema shape. Keep query ownership explicit, exempt unqueried high-fan-out fields, and treat every new manual index as a write/storage/deploy-time dependency. p95/p99 latency cannot be inferred from six emulator documents; measure realistic production-like data with cache/network state and region disclosed.
Knowledge check
- Why does a Standard query normally need an index even if the collection is small?
- What extra information is missing from the statement “price is indexed”?
- Why is an emulator query success not proof that a production composite index exists?
- When is exempting a large description field reasonable?
- What does the 40,000 entry limit constrain?
Review the answers
1. Standard Core query execution is index-backed by design; collection size does not turn an unsupported query into a table scan.
2. Edition/database, collection or collection-group scope, index mode/direction, and—if manual—the other fields and their order.
3. The emulator does not enforce production composite-index behavior or managed index lifecycle the same way as the service.
4. When no supported query, ordering, or other feature requires that field to be indexed; the decision should be tied to query contracts and tests.
5. The sum of a document’s automatic/single-field and manual/composite index entries; it is not a per-collection query limit.
Summary and next step
You now have a field-to-index mental model and a configuration-as-code baseline. Lesson 2 derives manual composite indexes from exact query shapes, including collection-group scope, field order, missing-index failures, and safe error-driven creation.
Authoritative references
- Index types in Cloud Firestore — Standard automatic/manual index modes, scopes, entry limits, exemptions, and index fan-out guidance.
- Manage indexes in Cloud Firestore — Missing-index workflow, roles, build state, CLI/console management, and vector indexes.
-
Cloud Firestore Index Definition Reference
— Current
firestore.indexes.jsonschema, vector configuration, field overrides, and TTL configuration. - Best practices for Cloud Firestore — Index fan-out, sequential-field, TTL, large string/array/map exemption guidance.
- Understand query performance using Query Explain — Planner versus analyze evidence and billing/scan statistics for managed Firestore.
- Enterprise edition index overview — Optional indexing, sparse/dense behavior, and query-performance reasoning in Enterprise Native mode.
- Firestore Native mode Core/Pipeline overview — Standard/Enterprise indexing requirements and interface differences.
- Search with vector embeddings — Vector index management, flat index type, supported dimensions, and vector-search limitations.
- Firestore pricing — Current document/index-entry billing semantics; re-check region and edition before budgeting.
- Firebase release notes — Current CLI/SDK versions used by the pinned lab.
- Firestore release notes — Enterprise Native/Pipeline launch-stage changes and emulator support.