Chapter 02 · Documents, Collections, Fields, References, Maps, Arrays, Timestamps, GeoPoints, and Limits
Document / Collection Hierarchy, Paths, IDs, Subcollections, and the Absence of Traditional Server-Side Joins in Core Operations
Model Firestore hierarchy deliberately: understand alternating collection/document paths, subcollections, references, and why Native Core operations do not provide traditional server-side joins.
Learning outcomes
AtlasMart now has a safe emulator environment, but its data model is still only a pair of toy product documents. The next design mistake would be to treat Firestore collections as SQL tables, references as foreign keys, and subcollections as automatic child rows. Firestore gives you a hierarchy of addressable documents and typed fields; the application must still decide which facts live together, which paths express ownership, and which reads can be satisfied without a server-side join.
Read and validate an alternating collection/document path, including nested subcollections and document IDs.
Explain what a collection, document, subcollection, map, and document reference are—and what guarantees they do not create.
Distinguish embedding, referencing, and separate reads without pretending Native Core operations provide SQL-style joins.
Recognize non-existent parent documents and the fact that deleting a parent document does not cascade to subcollections.
Build an AtlasMart hierarchy whose paths, ownership, and future query surfaces are explicit before application code spreads them.
The lab continues Chapter 01 exactly: project
demo-atlasmart-firestore; Firestore emulator
127.0.0.1:8080; Auth emulator
127.0.0.1:9099; Emulator UI
127.0.0.1:4000; Firebase CLI
15.30.0; Firebase JavaScript SDK
12.19.0; Firebase Admin Node.js SDK
14.4.0; Node.js 22 or newer. Standard Native
semantics are the default. Enterprise-only differences are
labeled explicitly.
These lessons are authored against current official documentation and the deterministic emulator design from Chapter 01. This generation environment does not run the Firebase Emulator Suite or install the pinned npm dependencies, so command output is described by invariant and expected shape rather than presented as captured execution. The emulator is evidence for local application behavior, not proof of production quotas, regional latency, backend index topology, or billing.
1. Firestore paths alternate collection and document segments
Cloud Firestore stores documents inside
collections. A collection is a container for
documents; it does not directly hold raw fields. A document is
the storage unit that owns fields. Paths therefore alternate:
collection/document/collection/document. AtlasMart
can address products/p-1001 or a nested review at
products/p-1001/reviews/r-001, but
products/p-1001/reviews names a collection, not a
document, and products/products/p-1001 is a
different path entirely.
Collections are implicit. You do not provision an empty
collection object first: writing the first document makes the
collection visible, and removing all its documents removes the
collection from normal data views. Document IDs are unique only
within their immediate collection. The same ID
r-001 can exist under many different product review
subcollections without collision because the full paths differ.
| Path | Kind | What it establishes |
|---|---|---|
products |
Collection | A namespace of product documents; no fields live directly here. |
products/p-1001 |
Document | One addressable document with fields. |
products/p-1001/reviews |
Subcollection | A collection nested under the product path. |
products/p-1001/reviews/r-001 |
Document | A review document; its existence does not create a relational foreign-key constraint. |
users/u-1001/orders/o-9001/items/i-01 |
Nested document | A deeper hierarchy; depth is a design and lifecycle decision, not free organization. |
A Firestore document reference points to a database/document path. Moving a document to another path is not a metadata rename; application code must create/copy data at the new path and handle the old path deliberately.
2. Subcollections are independent storage, not embedded children
A subcollection is convenient when child data grows
independently of the parent. Reviews can grow without making the
product document approach its size limit, and queries can target
one review collection or later use a collection-group query
across every reviews subcollection. The tradeoff is
lifecycle complexity: deleting products/p-1001 does
not automatically delete products/p-1001/reviews/*.
A subtle but important boundary is that a subcollection document
can exist even if the parent document has no stored fields. The
path products/p-ghost/reviews/r-1 can exist while
products/p-ghost itself does not. The console can
visually surface such non-existent parents, but ordinary
document reads/queries still distinguish the missing parent from
a real stored document. Therefore, "the path contains a product
ID" is not proof that the product document exists.
Ownership encoded in a path does not give you cascade delete, referential integrity, or orphan prevention. AtlasMart must implement and verify those policies explicitly; Chapter 11 will revisit recursive deletion and data lifecycle.
3. Document references are typed values, not joins
Firestore can store a DocumentReference value
such as a reference from an order line to
products/p-1001. That value preserves the target
path as a Firestore-native type. It does not cause the target
document to be fetched when the order is read, it does not
enforce that the target exists, and deleting the target does not
rewrite or null out stored references.
With Native Core operations, the application normally performs another read when it needs the referenced document, or it duplicates the read-critical fields it needs inside the source document. That is why references must be evaluated against read count, latency, offline behavior, consistency requirements, and security rules—not against the intuition that they behave like SQL foreign keys. Enterprise Pipeline operations add different query capabilities and later chapters will cover them; Chapter 02 must not back-port those capabilities into Standard Core semantics.
| Relational intuition | Firestore reality | AtlasMart response |
|---|---|---|
| Foreign key guarantees target existence | Reference is a typed path value | Validate target existence in trusted workflow if the invariant matters. |
| JOIN returns related rows in one query | Core operations do not provide a traditional join | Embed/duplicate bounded read-critical data or issue explicit reads. |
| Cascade delete follows ownership | Parent deletion does not delete subcollections or referenced docs | Create an explicit deletion workflow and verification step. |
| Schema enforced by table | Documents in one collection can differ | Use application/rules/schemaVersion contracts and tests. |
4. AtlasMart hierarchy: make ownership and query surfaces visible
For Chapter 02, AtlasMart will keep product catalog documents at
the root because many screens query products independently of a
user. Reviews belong under a product path for local ownership
examples, while orders live at a root collection because
operational services often need cross-user order workflows. The
order stores a stable userId and selected snapshot
fields instead of assuming a join to a profile at read time.
products/p-1001 name: "Trail Camera" category: "cameras" sellerId: "s-001" description: "..." attributes: { resolution: "4K", weatherproof: true }products/p-1001/reviews/r-001 userId: "u-1001" rating: 5 body: "Works well in rain"users/u-1001 displayName: "A. Shopper" preferredCurrency: "USD"orders/o-9001 userId: "u-1001" status: "placed" itemSnapshots: [{ productId: "p-1001", name: "Trail Camera", unitPrice: 129.90 }]
The duplicated product name and unit price in the order are not accidental corruption: they can be a historical snapshot of what was purchased. Whether that is the right model is an access-pattern decision. Chapter 03 will compare embedding, references, duplication, and fan-out systematically.
5. Hands-on lab: seed exact paths and prove what exists
Continue the Chapter 01 project. The Admin SDK writes deterministic fixtures because the server path is trusted and bypasses client Security Rules. The script then reads the parent product, nested review, and a deliberately missing parent to show that path syntax and document existence are separate facts.
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";import { initializeApp } from "firebase-admin/app";import { getFirestore, Timestamp } from "firebase-admin/firestore";initializeApp({ projectId: "demo-atlasmart-firestore" });const db = getFirestore();await db.doc("products/p-1001").set({ name: "Trail Camera", category: "cameras", sellerId: "s-001", schemaVersion: 2, updatedAt: Timestamp.fromDate(new Date("2026-09-14T09:00:00Z"))}, { merge: true });await db.doc("products/p-1001/reviews/r-001").set({ userId: "u-1001", rating: 5, body: "Works well in rain", schemaVersion: 1});// Deliberately create a child below a parent document that has no fields.await db.doc("products/p-ghost/reviews/r-orphan-demo").set({ userId: "u-1002", rating: 3, schemaVersion: 1});for (const path of [ "products/p-1001", "products/p-1001/reviews/r-001", "products/p-ghost", "products/p-ghost/reviews/r-orphan-demo"]) { const snap = await db.doc(path).get(); console.log(path, { exists: snap.exists, data: snap.exists ? snap.data() : null });}
Expected invariant: products/p-1001, its review,
and the orphan-demo review exist;
products/p-ghost itself does not. Inspect the same
paths in the Emulator UI. That visual inspection is supporting
evidence, not a substitute for the programmatic
exists assertions.
process.env.FIRESTORE_EMULATOR_HOST = "127.0.0.1:8080";process.env.GCLOUD_PROJECT = "demo-atlasmart-firestore";// Use Admin SDK deletes only for the exact fixture paths created above.// Do not translate this into an unscoped production recursive delete.
6. Deliberately wrong approach: build a relational graph out of references
Suppose every order stores only a product reference, every product stores only a seller reference, and the UI expects one query to return order + product + seller. The model looks normalized, but Core operations have no traditional server-side join to materialize that graph. A mobile client now performs a chain of network reads, offline screens become harder to reason about, read cost increases, and Security Rules must authorize each path separately.
The repair is not “never use references.” It is to name the screen/query contract first. If the order receipt must show the purchased name and price forever, copy those immutable snapshot values into the order. If a seller profile is optional and changes independently, a reference plus an explicit read may be appropriate. If the target must exist, enforce that invariant in a trusted transaction/workflow rather than assuming the reference type does it for you.
Hierarchy should make ownership and lifecycle easier to reason about without hiding query cost. Deep nesting, cross-document references, and duplicated fields are all valid tools, but each changes security rules, deletion, migration, offline behavior, and read amplification.
Knowledge check
Check your understanding
-
Why does
products/p-1001/reviewsnot identify a document? - Can a review document exist below a product path whose parent product document does not exist?
- What does storing a DocumentReference guarantee about its target?
- Why might an order duplicate product name and purchase price?
- What must AtlasMart do when deleting a product with nested reviews?
Review the answers
1. Firestore paths alternate collection and document segments. That path ends at a collection; a review document requires one more document-ID segment.
2. Yes. Subcollection documents can exist even when the parent document has no stored fields; path ancestry is not referential integrity.
3. It preserves a typed Firestore document path. It does not guarantee target existence, eager loading, cascade behavior, or join semantics.
4. Those fields can be an intentional historical snapshot needed by the receipt, avoiding a later join/read and preserving what the buyer actually purchased.
5. Use an explicit lifecycle/deletion workflow that finds and deletes or archives child documents, then verifies completion; deleting the parent alone is not enough.
Summary and next step
Firestore gives AtlasMart addressable documents in alternating collection/document paths, not relational tables with foreign keys. Subcollections scale child lists but create lifecycle obligations; references are typed paths rather than joins. With that hierarchy explicit, the next lesson can focus on what values actually live inside a document and how their types sort, serialize, and cross SDK boundaries.
Next: Supported Native Data Types, Null, Numeric Ordering, Timestamps, References, GeoPoints, Bytes, Arrays, and Maps.
Authoritative references
- Cloud Firestore data model — Official document/collection hierarchy and subcollection model.
- Choose a data structure — Official guidance for nested data, subcollections, and root-level collections.
- Supported data types — Current Native data types, sort ordering, precision, and edition-specific value behavior.
- Usage and limits — Current document, field, path, nesting, index-entry, and request limits.
- Index types in Cloud Firestore — Automatic/manual indexing, map/array indexing, index entries, and exemptions.
- Best practices for Cloud Firestore — Official document-ID, hotspot, location, and index-fan-out guidance.
- Add data to Cloud Firestore — Current SDK examples for IDs, timestamps, nested fields, and custom object conversion.
- Firebase release notes — Current Firebase SDK and CLI version baseline.
- Connect to the Firestore Emulator — Local emulator connection and production-difference guidance.