Chapter 03 · Data Modeling by Access Pattern: Embedding, Referencing, Duplication, Fan-Out, and Denormalization
Start from Screens / Queries / Transactions—not Entity Diagrams—When Designing Firestore Models
Design Firestore from concrete screens, queries, listeners, and atomic workflows so document boundaries and duplicated views are consequences of access patterns rather than copied relational entities.
Learning outcomes
AtlasMart's team has five concrete journeys: product detail, cart, order history, seller dashboard, and a customer activity feed. A relational instinct would begin by drawing Product, Seller, User, Order, OrderItem, Review, and Event entities and normalizing relationships. In Firestore that is only a vocabulary list. The database is most useful when document boundaries and duplicated projections are chosen from the exact reads, listeners, writes, transactions, authorization constraints, and lifecycle operations the application must perform.
Convert screens and workflows into explicit read/query/write/transaction contracts before proposing collections.
Distinguish a source-of-truth document from a read-optimized projection or historical snapshot.
Estimate read and write amplification for competing schemas without pretending emulator counts are a production bill.
Identify invariants that require a transaction or trusted workflow instead of eventual repair.
Produce a schema decision record that includes rules, indexes, lifecycle, repair, migration, and rollback—not just field names.
The lab continues Chapters 01–02 exactly: project
demo-atlasmart-firestore; Firestore emulator
127.0.0.1:8080; Auth emulator
127.0.0.1:9099; Emulator UI
127.0.0.1:4000; Firebase CLI
15.30.0; Firebase JavaScript SDK
12.19.0; Firebase Admin Node.js SDK
14.4.0; Node.js 22 or newer. Standard Native Core
semantics are the default. Enterprise Native Pipeline or
MongoDB-compatibility behavior is mentioned only when the
distinction changes a modeling decision.
The mandatory lab is emulator-first and free/local. Emulator reads and writes are useful for deterministic operation counting and correctness tests, but they are not billable production operations and do not reproduce regional latency, production index topology, contention, autoscaling, or billing. Where a table estimates production operations, it reports document/query/write counts only; convert them to money only after re-checking the current production pricing contract for the exact edition, region, query shape, listener state, and index work.
1. Start with an access-pattern ledger, not an entity diagram
Firestore Core reads are document- and query-oriented. A model that looks elegant on an ER diagram can be expensive or awkward if one screen must chase several references, if a query cannot be expressed with available indexes, or if Security Rules cannot prove the query's result set is safe. An access pattern is a named application operation with its input key, result shape, freshness requirement, authorization context, cardinality, frequency, and atomicity needs.
| AtlasMart journey | Required read shape | Write/invariant | Freshness/lifecycle |
|---|---|---|---|
| Product detail | One product + seller display snapshot + recent review summary | Catalog edits controlled by seller/admin | Seconds/minutes acceptable for duplicated seller display name; product price current |
| Cart | User cart items, current availability, current price decision | Quantity edits by owner; checkout revalidates server-side | Cart can be stale; checkout invariant cannot |
| Order history | User's orders newest-first with immutable purchase snapshots | Order state changes on trusted backend | Historical line-item name/price must not change when catalog changes |
| Seller dashboard | Seller-scoped product/order summaries | Summary may be materialized from operational writes | Small bounded lag acceptable if observable/reparable |
| Activity feed | User-scoped newest events | Fan-out or append projection | Eventual delivery acceptable if idempotent and backfillable |
Notice how this ledger already suggests different storage semantics. An order line wants an immutable snapshot; a cart item may keep an ID plus a few display fields but must revalidate at checkout; a dashboard can tolerate a materialized projection; a business invariant such as "do not oversell reserved stock" belongs in a transaction or another trusted consistency mechanism rather than in a duplicated display document.
2. Derive two competing Firestore schemas deliberately
Schema A is intentionally reference-heavy. It minimizes duplication: order lines keep product IDs, product documents keep seller IDs, feed events keep object IDs, and the UI follows references. Schema B is access-pattern-first: operational sources remain authoritative, while order lines embed purchase snapshots, feed entries duplicate compact display fields, and seller dashboards maintain bounded summary documents.
Schema A — reference-heavyproducts/{productId} -> sellerIdorders/{orderId} -> userId, lines:[{ productId, quantity }]users/{{uid}}/feed/{{eventId}} -> objectType, objectIdSchema B — access-pattern-firstproducts/{productId} -> sellerId, sellerNameSnapshot, price, reviewSummaryorders/{orderId} -> userId, lines:[{ productId, nameSnapshot, unitPriceSnapshot, quantity }]sellers/{sellerId}/dashboard/{bucketId} -> materialized metricsusers/{{uid}}/feed/{{eventId}} -> objectId + compact display snapshot + sourceVersion
Neither schema is universally correct. Schema B increases write amplification and repair obligations but can reduce read chains and preserve history. Schema A reduces duplicated mutable fields but can multiply reads and rules checks. The design decision must state which side of that trade is intentional for each access pattern.
3. Make operation counts observable before arguing about cost
For a deterministic fixture, instrument the application repository layer instead of timing UI clicks. Count each document read, query execution, and document write that your code requests. The emulator can then verify the logical operation plan. Do not call those counters "billable reads" because production billing can depend on query/index/listener semantics that the local counter cannot prove.
const ops = { docReads: 0, queries: 0, writes: 0 };async function readDoc(ref) { ops.docReads++; return ref.get(); }async function runQuery(query) { ops.queries++; return query.get(); }async function writeDoc(ref, data, options) { ops.writes++; return options ? ref.set(data, options) : ref.set(data);}function report(label) { console.log(label, structuredClone(ops)); }
| Journey | Reference-heavy expected application work | Access-pattern-first expected work |
|---|---|---|
| Order receipt with 4 distinct products | Order read + up to 4 product reads if display fields are not copied | Order read if immutable display snapshot is embedded |
| Activity feed page | Feed query + follow-up reads for referenced objects | Feed query if compact display projection is sufficient |
| Seller rename | One source write, later reads see new name | Source write + projection repair/backfill if duplicated name must converge |
| Checkout stock invariant | Must use an explicit trusted atomic workflow; display projections are not the authority. | |
When the access-pattern-first model wins a read, ask what write or repair it added. That symmetry prevents the common mistake of labeling denormalization "faster" without acknowledging consistency and operational cost.
4. Transaction boundaries define where denormalization must stop
Some facts may be duplicated safely because the business accepts a consistency window. A seller display name copied into product cards can be repaired asynchronously. Other facts define invariants: reservation quantity, payment state, or a one-time idempotency marker. Those must not be inferred from a stale projection. If the invariant spans multiple documents and fits Firestore transaction semantics, the trusted workflow should read the authoritative documents and commit dependent writes atomically; external side effects must remain outside retryable transaction callbacks or be durably idempotent.
Mobile/web clients are authorized by Firebase Authentication plus Security Rules. Admin/server libraries use IAM/ADC and bypass Firestore Security Rules. A model that is safe only because "the UI hides the field" is not a security model. Put client-writable projections behind rules that constrain ownership and fields, and put privileged fan-out/repair work on a trusted server path with application authorization.
Enterprise Native Pipeline operations can express query patterns that Standard/Core cannot, but that does not erase lifecycle, authorization, write amplification, or historical-snapshot decisions. MongoDB compatibility is a separate Enterprise mode. Chapter 03 uses Standard Native Core behavior unless explicitly labeled otherwise so the model remains portable across the course's default lab.
5. Hands-on lab: compare two schemas from the five AtlasMart journeys
Use the same local workspace and pinned package baseline from
Chapter 02. The examples assume
firebase emulators:start --only auth,firestore is
running and that Admin SDK code points at
FIRESTORE_EMULATOR_HOST=127.0.0.1:8080 with
GCLOUD_PROJECT=demo-atlasmart-firestore. Browser
examples connect the modular Web SDK to the Firestore/Auth
emulators. Keep the fixtures synthetic and delete only the
Chapter 03 paths you create.
import { getFirestore } from "firebase-admin/firestore";const db = getFirestore();const now = new Date("2026-09-15T10:00:00Z");await db.doc("sellers/s-2001").set({ displayName: "Northwind Outdoor", schemaVersion: 1 });await db.doc("products/p-1001").set({ sellerId: "s-2001", name: "Trail Camera", priceCents: 12990, schemaVersion: 2 });await db.doc("users/u-1001").set({ displayName: "Asha", schemaVersion: 1 });await db.doc("orders/o-9001").set({ userId: "u-1001", createdAt: now, lines: [{ productId: "p-1001", nameSnapshot: "Trail Camera", unitPriceCentsSnapshot: 12990, quantity: 1 }], schemaVersion: 2});await db.doc("users/u-1001/feed/e-001").set({ type: "order_created", objectId: "o-9001", titleSnapshot: "Order o-9001 placed", sourceVersion: 1, createdAt: now});
Create a second fixture namespace such as
modelComparisons/schemaA/* and
modelComparisons/schemaB/* if you want both shapes
present simultaneously. Run the product-detail, order-history,
seller-dashboard, and feed repository functions through the
operation counter. Record counts in a small JSON result file.
Then change sellers/s-2001.displayName to
Northwind Field Gear and intentionally leave one
duplicated product/feed snapshot stale.
const seller = await db.doc("sellers/s-2001").get();const product = await db.doc("products/p-1001").get();console.log({ authoritative: seller.get("displayName"), duplicated: product.get("sellerNameSnapshot") ?? null, diverged: product.get("sellerNameSnapshot") !== seller.get("displayName")});
Your expected evidence is not a fabricated latency number. It is a table containing each access pattern, requested document/query/write counts, duplicated fields, atomic boundary, Security Rules surface, and observed divergence after the injected rename. Finish by writing a repair plan: source path, projection paths, idempotency key/version, retry policy, verification query, and rollback/reset.
Verification checklist
- Every schema object exists because a named screen/query/workflow needs it.
- Historical snapshots are labeled snapshots and never mistaken for current catalog authority.
- Operation counters distinguish document reads, queries, and writes.
- The stale seller-name injection is detected programmatically.
- Invariant-bearing writes identify their transaction/trusted-workflow boundary.
- Cleanup targets only Chapter 03 fixture paths in the demo project/emulator.
6. Deliberately wrong approach: normalize by reflex
A team can reproduce a relational schema almost literally: every relationship becomes an ID/reference, every mutable attribute lives in one document, and every screen assembles its view by following references. The model looks tidy, but product detail and order history become client-side joins. Each extra read adds latency and authorization surface; offline behavior becomes more complex; historical receipts can accidentally display current product names/prices instead of purchase-time facts.
The repair is not indiscriminate duplication. Return to the access-pattern ledger. Duplicate only fields whose read value justifies the write/repair cost, mark snapshots versus current projections, version repairable projections, and keep invariants attached to authoritative documents and atomic workflows. A denormalized field without a source-of-truth and repair strategy is technical debt, not a performance feature.
Knowledge check
Check your understanding
- Why is an ER diagram insufficient as the starting point for a Firestore model?
- When is a duplicated field an intentional projection instead of accidental drift?
- Why should emulator operation counters not be labeled a production bill?
- Which AtlasMart facts can tolerate eventual repair, and which require an atomic trusted workflow?
- What extra obligation appears when a screen saves reads by copying seller/product fields?
Review the answers
1. Firestore design depends on executable query, listener, authorization, lifecycle, and atomicity constraints. Entity relationships alone do not tell you how many documents a screen must read or whether a query/rule shape is possible.
2. It is intentional when it has a named authoritative source, a documented freshness contract, a version/repair path, and tests that can detect divergence.
3. The emulator proves local logical behavior but not production billing dimensions, index work, regional latency, or listener/reconnect semantics.
4. Display projections such as seller names or dashboard summaries can often tolerate bounded lag; reservation/payment/stock invariants require authoritative checks and atomic or otherwise explicit consistency control.
5. Write amplification, consistency-window definition, repair/backfill logic, extra rules/indexes, and migration/rollback work.
Summary and next step
Firestore modeling starts from access patterns and consistency contracts. AtlasMart now has a ledger that makes reads, writes, rules, lifecycle, and transaction boundaries visible, plus two competing schemas whose amplification can be measured. Next, we make the embed/reference/duplicate choice explicit field by field and quantify the document-growth and repair consequences.
Next: Embed vs Reference vs Duplicate: Read Cost, Write Amplification, Consistency, and Document Limits.
Authoritative references
- Choose a data structure — Official guidance for nested data, subcollections, and root-level collections.
- Cloud Firestore data model — Documents, collections, subcollections, paths, and parent-deletion behavior.
- Perform simple and compound queries — Collection and collection-group query semantics.
- Securely query data — Security Rules and query constraints; rules are not filters.
- Transactions and batched writes — Atomicity and retry boundaries used by bounded fan-out workflows.
- Distributed counters — Official sharded-counter pattern for higher update rates.
- Best practices for Cloud Firestore — IDs, hotspotting, index fan-out, location, and write-shape guidance.
- Delete data — Parent/subcollection lifecycle and bulk-delete considerations.
- Cloud Firestore pricing — Living production billing contract; re-check before converting operation counts into currency.
- Firebase release notes — Current Firebase SDK and CLI version baseline.
- Connect to the Firestore Emulator — Local development path and production-difference guidance.