Chapter 02 · Documents, Collections, Fields, References, Maps, Arrays, Timestamps, GeoPoints, and Limits

Document Size, Field / Path / Nesting Constraints, Index Entry Implications, and Schema Governance

Treat Firestore limits and indexing as part of the schema contract: reason about document growth, field paths, nesting depth, index-entry fan-out, exemptions, and migration-safe governance before production.

Beginner → Advanced90–120 minutesAtlasMart emulator-first labFirebase CLI 15.30.0 · Web SDK 12.19.0 · Admin Node 14.4.0 · Node.js 22+Firestore Standard Native semantics unless explicitly labeled EnterpriseLast reviewed: September 2026

Learning outcomes

A Firestore document can be syntactically valid and still be a production liability. Large documents increase transfer cost, deeply nested maps complicate updates and rules, and automatic indexing can turn one logical write into many index-entry updates. Chapter 02 therefore treats service limits and index fan-out as part of schema design rather than cleanup work for later.

01

Apply current document, ID, field-name/path, nesting, subcollection, and request limits to an AtlasMart schema.

02

Explain how automatic indexing creates entries for fields, arrays, and map subfields in Standard Native mode.

03

Estimate which large strings, arrays, maps, sequential fields, or raw payloads should be exempt from indexing.

04

Separate hard service limits from internal team guardrails that should trigger much earlier.

05

Version the schema contract so backfills and client rollouts can evolve safely rather than mutate document meaning silently.

Chapter 02 baseline reviewed 14 September 2026

The lab continues Chapter 01 exactly: project demo-atlasmart-firestore; Firestore emulator 127.0.0.1:8080; Auth emulator 127.0.0.1:9099; Emulator UI 127.0.0.1:4000; Firebase CLI 15.30.0; Firebase JavaScript SDK 12.19.0; Firebase Admin Node.js SDK 14.4.0; Node.js 22 or newer. Standard Native semantics are the default. Enterprise-only differences are labeled explicitly.

Execution and evidence note

These lessons are authored against current official documentation and the deterministic emulator design from Chapter 01. This generation environment does not run the Firebase Emulator Suite or install the pinned npm dependencies, so command output is described by invariant and expected shape rather than presented as captured execution. The emulator is evidence for local application behavior, not proof of production quotas, regional latency, backend index topology, or billing.

1. Hard limits are correctness boundaries, not optimization hints

Current Standard Native limits include a 1 MiB maximum document size, 6 KiB maximum document name, 1,500-byte document IDs, 1,500-byte field names and field paths, 20 levels of map/array field nesting, and 100 levels of subcollections. A field value can be at most 1 MiB minus 89 bytes. These are service contracts; applications should normally adopt lower internal guardrails so they fail predictably before a user request approaches the edge.

Constraint Current documented boundary AtlasMart guardrail reasoning
Document 1 MiB Keep normal product/order docs far below the limit; unbounded histories belong elsewhere.
Document name 6 KiB Do not encode payloads or compound business state in a path.
Document ID ≤ 1,500 bytes; additional naming restrictions Use opaque stable IDs; avoid slash, ., .., reserved patterns.
Field name / path ≤ 1,500 bytes Choose simple stable field names; avoid escaping-heavy punctuation.
Nested map/array depth 20 Flatten or split before rule/update complexity becomes unmanageable.
Subcollection depth 100 A hard maximum, not an invitation to model 100-level ownership trees.
API request 10 MiB Bulk import/write tooling needs bounded chunks, independent of per-document size.

2. Index entries are a second size system attached to every document

In Standard Native mode, automatic indexing commonly creates ascending and descending entries for non-array/non-map fields, recursively indexes map subfields, and creates array-related entries for arrays. Manual/composite indexes add more entries. The current per-document index-entry count limit is 40,000; individual index entries have a 7.5 KiB maximum, the sum of index-entry sizes for one document is capped at 8 MiB, and indexed field values above 1,500 bytes are truncated for the index in Standard.

This is why a 400 KiB product document is not automatically "safe." If it contains a huge array or map and every element/subfield is indexed, one write may have much larger index work than its logical document size suggests. Conversely, exempting a field from indexing reduces index work but removes query/sort capabilities that depended on that index.

Field Query need Default risk Chapter 02 decision
name Yes Normal index entries Keep indexed.
category Yes Normal index entries Keep indexed.
description Not with Core exact/range queries in this lab Long string storage/index cost Exempt from automatic indexing.
attributes No dynamic subfield queries yet Map subfields can multiply index entries Exempt whole map until a query requirement appears.
rawGatewayPayload Never queried Large server-only map/string Exempt and consider external object storage if growth is unbounded.

3. Put index intent in source control, not in tribal knowledge

The Chapter 01 repository committed firestore.indexes.json even though it was empty. Chapter 02 gives that file a purpose: document which fields should not receive automatic indexes because the application has no query contract for them. This makes the choice reviewable and reproducible across environments.

firestore.indexes.json · Chapter 02 field exemptions
{  "indexes": [],  "fieldOverrides": [    {      "collectionGroup": "products",      "fieldPath": "description",      "indexes": []    },    {      "collectionGroup": "products",      "fieldPath": "attributes",      "indexes": []    },    {      "collectionGroup": "orders",      "fieldPath": "rawGatewayPayload",      "indexes": []    }  ]}

An exemption is not a generic performance tweak. Before disabling an index, identify every current query, planned query, orderBy, aggregation, and rules interaction that may depend on that field. If a later feature needs to query attributes.color, AtlasMart should add a deliberate index for the stable field rather than restoring indexing for an uncontrolled dynamic map.

4. Schema governance means bounding growth before the service does

AtlasMart uses schemaVersion because documents will evolve across mobile/web releases and server backfills. A useful schema contract records more than field names: ownership, writer trust level, type, nullable/missing policy, maximum expected size/count, index intent, lifecycle, and migration behavior. Internal guardrails should be stricter than service maxima—for example, a product might allow at most 30 display tags even though Firestore's array/index limits are much larger.

schema-contract.json · versioned application guardrails
{  "collection": "products",  "schemaVersion": 2,  "fields": {    "name": { "type": "string", "required": true, "maxUtf8Bytes": 200, "indexed": true },    "description": { "type": "string", "required": true, "maxUtf8Bytes": 20000, "indexed": false },    "tags": { "type": "array<string>", "maxItems": 30, "indexed": true },    "attributes": { "type": "map", "maxKeys": 50, "indexed": false },    "schemaVersion": { "type": "integer", "required": true, "indexed": true }  }}

This file is an application contract, not a Firestore-native schema. It exists so CI, backfills, server validation, and client converters can agree on tighter rules than the database's schemaless storage layer enforces.

5. Hands-on lab: estimate size, test boundaries, inspect index configuration

Start with deterministic local guards. JavaScript's Buffer.byteLength(JSON.stringify(...)) is only an approximation of Firestore wire/storage size because native types and field-name overhead matter, but it is still useful as an application guardrail if you clearly label it as such. The service remains the authority for hard limits.

scripts/ch02-guardrails.mjs · application-level checks
import assert from "node:assert/strict";function utf8Bytes(s) { return Buffer.byteLength(s, "utf8"); }function checkProduct(p) {  assert.equal(p.schemaVersion, 2);  assert.ok(utf8Bytes(p.name) <= 200, "name too large for AtlasMart contract");  assert.ok(utf8Bytes(p.description) <= 20_000, "description exceeds AtlasMart contract");  assert.ok(p.tags.length <= 30, "too many tags");  assert.ok(Object.keys(p.attributes).length <= 50, "too many attributes");}const candidate = {  name: "Trail Camera",  description: "Weatherproof camera".repeat(100),  tags: ["outdoor", "camera"],  attributes: { resolution: "4K", weatherproof: true },  schemaVersion: 2};checkProduct(candidate);console.log("application guardrails OK");

Then deploy or load the index configuration only in the disposable emulator/demo project workflow. The Firestore emulator does not reproduce every production limit/index behavior, so do not use emulator acceptance of an oversize or unusual document as evidence that production will accept it.

inspect/deploy index configuration boundary
# Validate the file in source control.cat firestore.indexes.json# Emulator-first workflow from Chapter 01:npx firebase emulators:start --project demo-atlasmart-firestore --only auth,firestore# Optional isolated real-project verification only after billing/target review:# npx firebase deploy --only firestore:indexes --project YOUR_DISPOSABLE_PROJECT

6. Deliberately wrong approach: giant document + giant dynamic map

A common anti-pattern is one product document containing thousands of reviews, every per-store inventory record, raw payment/log payloads, and a dynamic attributes map whose keys come directly from sellers. Even before the 1 MiB document limit, every product read transfers unrelated data, updates contend on one document, Security Rules become complex, and automatic indexing can explode across arrays/map subfields.

The repair is requirement-driven decomposition. Keep bounded product summary fields together. Put independently growing reviews in a subcollection. Put operational inventory in a queryable structure designed for its update/query pattern. Exempt long/dynamic fields from indexing when no query requires them. Move genuinely large binary/raw payloads to appropriate object/log storage and keep references/metadata in Firestore.

Do not chase the maximum

A service limit is where a request becomes invalid, not a recommended target. Build business-level guardrails that preserve latency, cost, update granularity, security clarity, and migration headroom.

Knowledge check

Check your understanding

  1. What is the current maximum Firestore document size?
  2. How deep can map/array fields nest, and is that the same as subcollection depth?
  3. Why can a document hit an index limit even when the document itself is below 1 MiB?
  4. What happens when AtlasMart exempts attributes from automatic indexing?
  5. Why keep schemaVersion if Firestore is schemaless?
Review the answers

1. 1 MiB (1,048,576 bytes). Applications should generally enforce lower domain-specific guardrails.

2. Map/array field nesting is limited to 20 levels; subcollection nesting has a separate 100-level maximum.

3. Automatic and manual indexes create separate index entries. Large arrays/maps or many index combinations can approach the 40,000-entry and index-size limits.

4. Queries/sorts that require those indexed subfields are no longer supported by those automatic indexes; the exemption trades queryability for lower index work/storage.

5. Schemaless storage does not eliminate application contracts. schemaVersion lets readers, migrations, backfills, and tests distinguish document meanings over time.

Summary and next step

AtlasMart now treats Firestore's 1 MiB document boundary, path/nesting limits, and index-entry system as part of schema engineering. Field exemptions are version-controlled decisions, and internal guardrails intentionally trigger earlier than service maxima. The next lesson applies this discipline to document IDs, where seemingly harmless sequential or tenant-prefixed keys can create operational hotspots or privacy coupling.

Next: Auto IDs vs Semantic IDs, Hotspot Risks, Tenant Prefixes, and External Identity Mapping.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.