Chapter 03 · Data Modeling by Access Pattern: Embedding, Referencing, Duplication, Fan-Out, and Denormalization

Subcollections vs Root Collections, Collection Groups, Ownership, Lifecycle, and Security Rules

Place data in root collections or subcollections deliberately, then prove collection-group query, lifecycle, ownership, and Security Rules consequences with executable AtlasMart fixtures.

Beginner → Advanced100–130 minutesAtlasMart emulator-first modeling labFirebase CLI 15.30.0 · Web SDK 12.19.0 · Admin Node 14.4.0 · Node.js 22+Firestore Standard Native Core semantics unless explicitly labeled EnterpriseLast reviewed: September 2026

Learning outcomes

AtlasMart can represent reviews, carts, orders, and feeds either as root collections with owner fields or as subcollections beneath owner documents. Both can work. The important question is which shape makes local ownership, cross-parent querying, Security Rules, deletion, indexes, and migrations easiest to reason about. This lesson turns that choice into observable queries and lifecycle tests.

01

Compare root collections and subcollections by query scope, ownership, lifecycle, rules, and migration—not aesthetics.

02

Use collection-group queries across same-named subcollections and understand their index/rule requirements.

03

Prove that deleting a parent document does not delete its subcollection documents.

04

Write Security Rules that constrain client collection-group queries instead of relying on post-query filtering.

05

Design owner/tenant fields redundantly when they are needed for query authorization or lifecycle tooling.

Chapter 03 baseline reviewed 15 September 2026

The lab continues Chapters 01–02 exactly: project demo-atlasmart-firestore; Firestore emulator 127.0.0.1:8080; Auth emulator 127.0.0.1:9099; Emulator UI 127.0.0.1:4000; Firebase CLI 15.30.0; Firebase JavaScript SDK 12.19.0; Firebase Admin Node.js SDK 14.4.0; Node.js 22 or newer. Standard Native Core semantics are the default. Enterprise Native Pipeline or MongoDB-compatibility behavior is mentioned only when the distinction changes a modeling decision.

Execution, pricing, and evidence note

The mandatory lab is emulator-first and free/local. Emulator reads and writes are useful for deterministic operation counting and correctness tests, but they are not billable production operations and do not reproduce regional latency, production index topology, contention, autoscaling, or billing. Where a table estimates production operations, it reports document/query/write counts only; convert them to money only after re-checking the current production pricing contract for the exact edition, region, query shape, listener state, and index work.

1. Subcollections express path ownership but do not enforce lifecycle

products/p-1001/reviews/r-001 communicates that the review is organized beneath a product. The child can grow independently and the parent document stays small. But path ancestry does not provide referential integrity or cascade deletion. A review can survive after the product document is deleted, and the application must decide whether that is an orphan bug, an intentional audit record, or data awaiting asynchronous cleanup.

A subcollection is therefore useful when local child queries and ownership are natural, but only if AtlasMart also defines recursive deletion/retention behavior. Chapter 11 will deepen lifecycle mechanics; here we make the obligation visible in the model.

2. Root collections make cross-parent queries direct

A root collection such as reviews/{reviewId} can include productId, sellerId, tenantId, and authorUid. This often simplifies global moderation or analytics-like operational queries. The tradeoff is that hierarchy is now represented by fields instead of path ancestry, so writes must validate those fields and lifecycle jobs must find children by query rather than by path.

Decision Subcollection shape Root collection shape
Product-local query Natural path products/{id}/reviews Query reviews where productId == ...
Cross-product moderation Collection-group query reviews Root reviews query
Ownership signal Path + optionally redundant owner fields Explicit fields required
Parent deletion Child docs survive unless explicitly removed Independent by design; cleanup by owner-field query
Rules Can use path variables; collection-group rules need recursive wildcard design Rules inspect explicit fields and query constraints

3. Collection-group queries make same-named subcollections globally queryable

A collection group is all collections with the same collection ID regardless of parent path. If every product stores reviews in a subcollection named reviews, AtlasMart can query across all of them with collectionGroup("reviews"). This is powerful, but it means the child document often needs fields such as productId, sellerId, or tenantId if global filters/rules need those values.

Web SDK · seller-scoped collection-group query
import { collectionGroup, query, where, orderBy, limit, getDocs } from "firebase/firestore";const q = query(  collectionGroup(db, "reviews"),  where("sellerId", "==", "s-2001"),  orderBy("createdAt", "desc"),  limit(20));const snap = await getDocs(q);console.log(snap.docs.map(d => ({ path: d.ref.path, ...d.data() })));

Current Firestore guidance requires indexes that support collection-group queries, and mobile/web access must also be permitted by Security Rules. The emulator does not reproduce every production index enforcement detail, so keep production index definitions in source control and retain a production-only verification layer before launch.

4. Security Rules are not post-query filters

Suppose clients may read only reviews for their tenant. A rule such as "allow when resource.data.tenantId == request.auth.token.tenantId" does not mean the client can query every review and let Rules remove unauthorized rows. Firestore evaluates whether the query could return disallowed documents. The query must carry constraints compatible with the rules.

firestore.rules · conceptual tenant-constrained collection-group rule
rules_version = '2';service cloud.firestore {  match /databases/{database}/documents {    function signedIn() { return request.auth != null; }    function tenant() { return request.auth.token.tenantId; }    match /{path=**}/reviews/{{reviewId}} {      allow read: if signedIn() && resource.data.tenantId == tenant();      allow create: if signedIn()        && request.resource.data.tenantId == tenant()        && request.resource.data.authorUid == request.auth.uid;    }  }}

This is a teaching baseline, not a complete production rule set. Production rules must validate allowed fields, immutable owner identifiers, types, update transitions, and query shapes. Server/Admin code bypasses Firestore Security Rules and must apply application authorization under IAM-authenticated trusted execution.

5. Hands-on lab: prove query scope, rules surface, and orphan behavior

Use the same local workspace and pinned package baseline from Chapter 02. The examples assume firebase emulators:start --only auth,firestore is running and that Admin SDK code points at FIRESTORE_EMULATOR_HOST=127.0.0.1:8080 with GCLOUD_PROJECT=demo-atlasmart-firestore. Browser examples connect the modular Web SDK to the Firestore/Auth emulators. Keep the fixtures synthetic and delete only the Chapter 03 paths you create.

Admin seed · two products with same-named review subcollections
const fixtures = [  ["products/p-1001/reviews/r-001", { productId:"p-1001", sellerId:"s-2001", tenantId:"t-01", authorUid:"u-1001", rating:5, createdAt:new Date("2026-09-15T10:10:00Z") }],  ["products/p-1002/reviews/r-002", { productId:"p-1002", sellerId:"s-2001", tenantId:"t-01", authorUid:"u-1002", rating:4, createdAt:new Date("2026-09-15T10:11:00Z") }],  ["products/p-2001/reviews/r-003", { productId:"p-2001", sellerId:"s-3001", tenantId:"t-02", authorUid:"u-2001", rating:5, createdAt:new Date("2026-09-15T10:12:00Z") }],];for (const [path, data] of fixtures) await db.doc(path).set({ ...data, schemaVersion: 1 });

Run a collection-group query for sellerId == s-2001 from Admin code and assert exactly the two t-01 paths. Then run the equivalent authorized web query under the Auth emulator using a synthetic user whose token/claims match the intended test setup from Chapter 01. Also include a deliberately overbroad client query in the rules test suite and assert that it fails rather than being filtered.

lifecycle failure injection · parent delete does not remove review
await db.doc("products/p-1002").set({ name: "Action Camera", sellerId: "s-2001" });await db.doc("products/p-1002").delete();const parent = await db.doc("products/p-1002").get();const child = await db.doc("products/p-1002/reviews/r-002").get();console.log({ parentExists: parent.exists, childExists: child.exists });// Expected invariant: false, true.

That output proves a lifecycle obligation. It does not mean orphans are always wrong: an audit or legal-retention design may intentionally keep them. The schema decision record must say which policy applies and how cleanup/retention is verified.

Verification checklist

  • Local product review query and cross-product collection-group query are both tested.
  • Overbroad client query is denied under Security Rules rather than filtered.
  • Admin query is recognized as an IAM/trusted path that bypasses Rules.
  • Deleting a parent leaves the child review observable.
  • Schema contract records owner/tenant fields needed for global query/rules/lifecycle tasks.

6. Deliberately wrong approach: deep nesting without a deletion/query plan

A team may encode every ownership relationship into paths—tenants/{t}/users/{u}/orders/{o}/items/{i}/events/{e}—because the hierarchy looks organized. Later it needs cross-tenant support queries, right-to-delete workflows, moderation, and migrations. Each task now needs collection-group indexes/rules or recursive traversal, while parent deletion still does not clean descendants automatically.

The repair is to use nesting only where local ownership and independent child growth justify it, duplicate owner/tenant identifiers when global operations need them, and document recursive lifecycle operations. Root collections are not "less NoSQL"; they are sometimes the clearer access-pattern choice.

Knowledge check

Check your understanding

  1. What does a subcollection path communicate, and what does it not guarantee?
  2. When might a root reviews collection be preferable?
  3. What is a collection-group query?
  4. Why can an overbroad client query fail even if each returned document would later be checked?
  5. What happens to subcollection documents when the parent document is deleted?
Review the answers

1. It communicates hierarchical organization/ownership context, but does not guarantee parent existence, referential integrity, or cascade deletion.

2. When global/cross-parent queries, moderation, lifecycle jobs, or security constraints are simpler with explicit owner fields in one root collection.

3. A query across all collections with the same collection ID, regardless of parent path.

4. Security Rules are not filters; Firestore must be able to prove the query cannot return disallowed documents from its potential result set.

5. They remain unless explicitly deleted or otherwise handled by a lifecycle workflow.

Summary and next step

Root collections and subcollections are query/lifecycle/security choices, not stylistic nesting. AtlasMart can now query same-named review subcollections globally, test Rules-compatible constraints, and prove orphan behavior after parent deletion. Next we deliberately create derived query shapes—feeds, dashboard summaries, and counters—and make their fan-out, idempotency, and repair costs explicit.

Next: Fan-Out Writes, Materialized Views, Counters, Feeds, and Precomputed Query Shapes.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.