Chapter 02 · Documents, Collections, Fields, References, Maps, Arrays, Timestamps, GeoPoints, and Limits
Document Size, Field / Path / Nesting Constraints, Index Entry Implications, and Schema Governance
Treat Firestore limits and indexing as part of the schema contract: reason about document growth, field paths, nesting depth, index-entry fan-out, exemptions, and migration-safe governance before production.
Learning outcomes
A Firestore document can be syntactically valid and still be a production liability. Large documents increase transfer cost, deeply nested maps complicate updates and rules, and automatic indexing can turn one logical write into many index-entry updates. Chapter 02 therefore treats service limits and index fan-out as part of schema design rather than cleanup work for later.
Apply current document, ID, field-name/path, nesting, subcollection, and request limits to an AtlasMart schema.
Explain how automatic indexing creates entries for fields, arrays, and map subfields in Standard Native mode.
Estimate which large strings, arrays, maps, sequential fields, or raw payloads should be exempt from indexing.
Separate hard service limits from internal team guardrails that should trigger much earlier.
Version the schema contract so backfills and client rollouts can evolve safely rather than mutate document meaning silently.
The lab continues Chapter 01 exactly: project
demo-atlasmart-firestore; Firestore emulator
127.0.0.1:8080; Auth emulator
127.0.0.1:9099; Emulator UI
127.0.0.1:4000; Firebase CLI
15.30.0; Firebase JavaScript SDK
12.19.0; Firebase Admin Node.js SDK
14.4.0; Node.js 22 or newer. Standard Native
semantics are the default. Enterprise-only differences are
labeled explicitly.
These lessons are authored against current official documentation and the deterministic emulator design from Chapter 01. This generation environment does not run the Firebase Emulator Suite or install the pinned npm dependencies, so command output is described by invariant and expected shape rather than presented as captured execution. The emulator is evidence for local application behavior, not proof of production quotas, regional latency, backend index topology, or billing.
1. Hard limits are correctness boundaries, not optimization hints
Current Standard Native limits include a 1 MiB maximum document size, 6 KiB maximum document name, 1,500-byte document IDs, 1,500-byte field names and field paths, 20 levels of map/array field nesting, and 100 levels of subcollections. A field value can be at most 1 MiB minus 89 bytes. These are service contracts; applications should normally adopt lower internal guardrails so they fail predictably before a user request approaches the edge.
| Constraint | Current documented boundary | AtlasMart guardrail reasoning |
|---|---|---|
| Document | 1 MiB | Keep normal product/order docs far below the limit; unbounded histories belong elsewhere. |
| Document name | 6 KiB | Do not encode payloads or compound business state in a path. |
| Document ID | ≤ 1,500 bytes; additional naming restrictions |
Use opaque stable IDs; avoid slash, .,
.., reserved patterns.
|
| Field name / path | ≤ 1,500 bytes | Choose simple stable field names; avoid escaping-heavy punctuation. |
| Nested map/array depth | 20 | Flatten or split before rule/update complexity becomes unmanageable. |
| Subcollection depth | 100 | A hard maximum, not an invitation to model 100-level ownership trees. |
| API request | 10 MiB | Bulk import/write tooling needs bounded chunks, independent of per-document size. |
2. Index entries are a second size system attached to every document
In Standard Native mode, automatic indexing commonly creates ascending and descending entries for non-array/non-map fields, recursively indexes map subfields, and creates array-related entries for arrays. Manual/composite indexes add more entries. The current per-document index-entry count limit is 40,000; individual index entries have a 7.5 KiB maximum, the sum of index-entry sizes for one document is capped at 8 MiB, and indexed field values above 1,500 bytes are truncated for the index in Standard.
This is why a 400 KiB product document is not automatically "safe." If it contains a huge array or map and every element/subfield is indexed, one write may have much larger index work than its logical document size suggests. Conversely, exempting a field from indexing reduces index work but removes query/sort capabilities that depended on that index.
| Field | Query need | Default risk | Chapter 02 decision |
|---|---|---|---|
name |
Yes | Normal index entries | Keep indexed. |
category |
Yes | Normal index entries | Keep indexed. |
description |
Not with Core exact/range queries in this lab | Long string storage/index cost | Exempt from automatic indexing. |
attributes |
No dynamic subfield queries yet | Map subfields can multiply index entries | Exempt whole map until a query requirement appears. |
rawGatewayPayload |
Never queried | Large server-only map/string | Exempt and consider external object storage if growth is unbounded. |
3. Put index intent in source control, not in tribal knowledge
The Chapter 01 repository committed
firestore.indexes.json even though it was empty.
Chapter 02 gives that file a purpose: document which fields
should not receive automatic indexes because the application has
no query contract for them. This makes the choice reviewable and
reproducible across environments.
{ "indexes": [], "fieldOverrides": [ { "collectionGroup": "products", "fieldPath": "description", "indexes": [] }, { "collectionGroup": "products", "fieldPath": "attributes", "indexes": [] }, { "collectionGroup": "orders", "fieldPath": "rawGatewayPayload", "indexes": [] } ]}
An exemption is not a generic performance tweak. Before
disabling an index, identify every current query, planned query,
orderBy, aggregation, and rules interaction that may depend on
that field. If a later feature needs to query
attributes.color, AtlasMart should add a deliberate
index for the stable field rather than restoring indexing for an
uncontrolled dynamic map.
4. Schema governance means bounding growth before the service does
AtlasMart uses schemaVersion because documents will
evolve across mobile/web releases and server backfills. A useful
schema contract records more than field names: ownership, writer
trust level, type, nullable/missing policy, maximum expected
size/count, index intent, lifecycle, and migration behavior.
Internal guardrails should be stricter than service maxima—for
example, a product might allow at most 30 display tags even
though Firestore's array/index limits are much larger.
{ "collection": "products", "schemaVersion": 2, "fields": { "name": { "type": "string", "required": true, "maxUtf8Bytes": 200, "indexed": true }, "description": { "type": "string", "required": true, "maxUtf8Bytes": 20000, "indexed": false }, "tags": { "type": "array<string>", "maxItems": 30, "indexed": true }, "attributes": { "type": "map", "maxKeys": 50, "indexed": false }, "schemaVersion": { "type": "integer", "required": true, "indexed": true } }}
This file is an application contract, not a Firestore-native schema. It exists so CI, backfills, server validation, and client converters can agree on tighter rules than the database's schemaless storage layer enforces.
5. Hands-on lab: estimate size, test boundaries, inspect index configuration
Start with deterministic local guards. JavaScript's
Buffer.byteLength(JSON.stringify(...)) is only an
approximation of Firestore wire/storage size because native
types and field-name overhead matter, but it is still useful as
an application guardrail if you clearly label it as such. The
service remains the authority for hard limits.
import assert from "node:assert/strict";function utf8Bytes(s) { return Buffer.byteLength(s, "utf8"); }function checkProduct(p) { assert.equal(p.schemaVersion, 2); assert.ok(utf8Bytes(p.name) <= 200, "name too large for AtlasMart contract"); assert.ok(utf8Bytes(p.description) <= 20_000, "description exceeds AtlasMart contract"); assert.ok(p.tags.length <= 30, "too many tags"); assert.ok(Object.keys(p.attributes).length <= 50, "too many attributes");}const candidate = { name: "Trail Camera", description: "Weatherproof camera".repeat(100), tags: ["outdoor", "camera"], attributes: { resolution: "4K", weatherproof: true }, schemaVersion: 2};checkProduct(candidate);console.log("application guardrails OK");
Then deploy or load the index configuration only in the disposable emulator/demo project workflow. The Firestore emulator does not reproduce every production limit/index behavior, so do not use emulator acceptance of an oversize or unusual document as evidence that production will accept it.
# Validate the file in source control.cat firestore.indexes.json# Emulator-first workflow from Chapter 01:npx firebase emulators:start --project demo-atlasmart-firestore --only auth,firestore# Optional isolated real-project verification only after billing/target review:# npx firebase deploy --only firestore:indexes --project YOUR_DISPOSABLE_PROJECT
6. Deliberately wrong approach: giant document + giant dynamic map
A common anti-pattern is one product document containing
thousands of reviews, every per-store inventory record, raw
payment/log payloads, and a dynamic attributes map
whose keys come directly from sellers. Even before the 1 MiB
document limit, every product read transfers unrelated data,
updates contend on one document, Security Rules become complex,
and automatic indexing can explode across arrays/map subfields.
The repair is requirement-driven decomposition. Keep bounded product summary fields together. Put independently growing reviews in a subcollection. Put operational inventory in a queryable structure designed for its update/query pattern. Exempt long/dynamic fields from indexing when no query requires them. Move genuinely large binary/raw payloads to appropriate object/log storage and keep references/metadata in Firestore.
A service limit is where a request becomes invalid, not a recommended target. Build business-level guardrails that preserve latency, cost, update granularity, security clarity, and migration headroom.
Knowledge check
Check your understanding
- What is the current maximum Firestore document size?
- How deep can map/array fields nest, and is that the same as subcollection depth?
- Why can a document hit an index limit even when the document itself is below 1 MiB?
-
What happens when AtlasMart exempts
attributesfrom automatic indexing? -
Why keep
schemaVersionif Firestore is schemaless?
Review the answers
1. 1 MiB (1,048,576 bytes). Applications should generally enforce lower domain-specific guardrails.
2. Map/array field nesting is limited to 20 levels; subcollection nesting has a separate 100-level maximum.
3. Automatic and manual indexes create separate index entries. Large arrays/maps or many index combinations can approach the 40,000-entry and index-size limits.
4. Queries/sorts that require those indexed subfields are no longer supported by those automatic indexes; the exemption trades queryability for lower index work/storage.
5. Schemaless storage does not eliminate application contracts. schemaVersion lets readers, migrations, backfills, and tests distinguish document meanings over time.
Summary and next step
AtlasMart now treats Firestore's 1 MiB document boundary, path/nesting limits, and index-entry system as part of schema engineering. Field exemptions are version-controlled decisions, and internal guardrails intentionally trigger earlier than service maxima. The next lesson applies this discipline to document IDs, where seemingly harmless sequential or tenant-prefixed keys can create operational hotspots or privacy coupling.
Next: Auto IDs vs Semantic IDs, Hotspot Risks, Tenant Prefixes, and External Identity Mapping.
Authoritative references
- Cloud Firestore data model — Official document/collection hierarchy and subcollection model.
- Choose a data structure — Official guidance for nested data, subcollections, and root-level collections.
- Supported data types — Current Native data types, sort ordering, precision, and edition-specific value behavior.
- Usage and limits — Current document, field, path, nesting, index-entry, and request limits.
- Index types in Cloud Firestore — Automatic/manual indexing, map/array indexing, index entries, and exemptions.
- Best practices for Cloud Firestore — Official document-ID, hotspot, location, and index-fan-out guidance.
- Add data to Cloud Firestore — Current SDK examples for IDs, timestamps, nested fields, and custom object conversion.
- Firebase release notes — Current Firebase SDK and CLI version baseline.
- Connect to the Firestore Emulator — Local emulator connection and production-difference guidance.