Chapter 03 · Data Modeling by Access Pattern: Embedding, Referencing, Duplication, Fan-Out, and Denormalization

Start from Screens / Queries / Transactions—not Entity Diagrams—When Designing Firestore Models

Design Firestore from concrete screens, queries, listeners, and atomic workflows so document boundaries and duplicated views are consequences of access patterns rather than copied relational entities.

Beginner → Advanced100–130 minutesAtlasMart emulator-first modeling labFirebase CLI 15.30.0 · Web SDK 12.19.0 · Admin Node 14.4.0 · Node.js 22+Firestore Standard Native Core semantics unless explicitly labeled EnterpriseLast reviewed: September 2026

Learning outcomes

AtlasMart's team has five concrete journeys: product detail, cart, order history, seller dashboard, and a customer activity feed. A relational instinct would begin by drawing Product, Seller, User, Order, OrderItem, Review, and Event entities and normalizing relationships. In Firestore that is only a vocabulary list. The database is most useful when document boundaries and duplicated projections are chosen from the exact reads, listeners, writes, transactions, authorization constraints, and lifecycle operations the application must perform.

01

Convert screens and workflows into explicit read/query/write/transaction contracts before proposing collections.

02

Distinguish a source-of-truth document from a read-optimized projection or historical snapshot.

03

Estimate read and write amplification for competing schemas without pretending emulator counts are a production bill.

04

Identify invariants that require a transaction or trusted workflow instead of eventual repair.

05

Produce a schema decision record that includes rules, indexes, lifecycle, repair, migration, and rollback—not just field names.

Chapter 03 baseline reviewed 15 September 2026

The lab continues Chapters 01–02 exactly: project demo-atlasmart-firestore; Firestore emulator 127.0.0.1:8080; Auth emulator 127.0.0.1:9099; Emulator UI 127.0.0.1:4000; Firebase CLI 15.30.0; Firebase JavaScript SDK 12.19.0; Firebase Admin Node.js SDK 14.4.0; Node.js 22 or newer. Standard Native Core semantics are the default. Enterprise Native Pipeline or MongoDB-compatibility behavior is mentioned only when the distinction changes a modeling decision.

Execution, pricing, and evidence note

The mandatory lab is emulator-first and free/local. Emulator reads and writes are useful for deterministic operation counting and correctness tests, but they are not billable production operations and do not reproduce regional latency, production index topology, contention, autoscaling, or billing. Where a table estimates production operations, it reports document/query/write counts only; convert them to money only after re-checking the current production pricing contract for the exact edition, region, query shape, listener state, and index work.

1. Start with an access-pattern ledger, not an entity diagram

Firestore Core reads are document- and query-oriented. A model that looks elegant on an ER diagram can be expensive or awkward if one screen must chase several references, if a query cannot be expressed with available indexes, or if Security Rules cannot prove the query's result set is safe. An access pattern is a named application operation with its input key, result shape, freshness requirement, authorization context, cardinality, frequency, and atomicity needs.

AtlasMart journey Required read shape Write/invariant Freshness/lifecycle
Product detail One product + seller display snapshot + recent review summary Catalog edits controlled by seller/admin Seconds/minutes acceptable for duplicated seller display name; product price current
Cart User cart items, current availability, current price decision Quantity edits by owner; checkout revalidates server-side Cart can be stale; checkout invariant cannot
Order history User's orders newest-first with immutable purchase snapshots Order state changes on trusted backend Historical line-item name/price must not change when catalog changes
Seller dashboard Seller-scoped product/order summaries Summary may be materialized from operational writes Small bounded lag acceptable if observable/reparable
Activity feed User-scoped newest events Fan-out or append projection Eventual delivery acceptable if idempotent and backfillable

Notice how this ledger already suggests different storage semantics. An order line wants an immutable snapshot; a cart item may keep an ID plus a few display fields but must revalidate at checkout; a dashboard can tolerate a materialized projection; a business invariant such as "do not oversell reserved stock" belongs in a transaction or another trusted consistency mechanism rather than in a duplicated display document.

2. Derive two competing Firestore schemas deliberately

Schema A is intentionally reference-heavy. It minimizes duplication: order lines keep product IDs, product documents keep seller IDs, feed events keep object IDs, and the UI follows references. Schema B is access-pattern-first: operational sources remain authoritative, while order lines embed purchase snapshots, feed entries duplicate compact display fields, and seller dashboards maintain bounded summary documents.

design sketch · competing AtlasMart shapes
Schema A — reference-heavyproducts/{productId} -> sellerIdorders/{orderId} -> userId, lines:[{ productId, quantity }]users/{{uid}}/feed/{{eventId}} -> objectType, objectIdSchema B — access-pattern-firstproducts/{productId} -> sellerId, sellerNameSnapshot, price, reviewSummaryorders/{orderId} -> userId, lines:[{ productId, nameSnapshot, unitPriceSnapshot, quantity }]sellers/{sellerId}/dashboard/{bucketId} -> materialized metricsusers/{{uid}}/feed/{{eventId}} -> objectId + compact display snapshot + sourceVersion

Neither schema is universally correct. Schema B increases write amplification and repair obligations but can reduce read chains and preserve history. Schema A reduces duplicated mutable fields but can multiply reads and rules checks. The design decision must state which side of that trade is intentional for each access pattern.

3. Make operation counts observable before arguing about cost

For a deterministic fixture, instrument the application repository layer instead of timing UI clicks. Count each document read, query execution, and document write that your code requests. The emulator can then verify the logical operation plan. Do not call those counters "billable reads" because production billing can depend on query/index/listener semantics that the local counter cannot prove.

Node.js · tiny operation counter for schema comparison
const ops = { docReads: 0, queries: 0, writes: 0 };async function readDoc(ref) { ops.docReads++; return ref.get(); }async function runQuery(query) { ops.queries++; return query.get(); }async function writeDoc(ref, data, options) {  ops.writes++;  return options ? ref.set(data, options) : ref.set(data);}function report(label) { console.log(label, structuredClone(ops)); }
Journey Reference-heavy expected application work Access-pattern-first expected work
Order receipt with 4 distinct products Order read + up to 4 product reads if display fields are not copied Order read if immutable display snapshot is embedded
Activity feed page Feed query + follow-up reads for referenced objects Feed query if compact display projection is sufficient
Seller rename One source write, later reads see new name Source write + projection repair/backfill if duplicated name must converge
Checkout stock invariant Must use an explicit trusted atomic workflow; display projections are not the authority.

When the access-pattern-first model wins a read, ask what write or repair it added. That symmetry prevents the common mistake of labeling denormalization "faster" without acknowledging consistency and operational cost.

4. Transaction boundaries define where denormalization must stop

Some facts may be duplicated safely because the business accepts a consistency window. A seller display name copied into product cards can be repaired asynchronously. Other facts define invariants: reservation quantity, payment state, or a one-time idempotency marker. Those must not be inferred from a stale projection. If the invariant spans multiple documents and fits Firestore transaction semantics, the trusted workflow should read the authoritative documents and commit dependent writes atomically; external side effects must remain outside retryable transaction callbacks or be durably idempotent.

Client/server boundary

Mobile/web clients are authorized by Firebase Authentication plus Security Rules. Admin/server libraries use IAM/ADC and bypass Firestore Security Rules. A model that is safe only because "the UI hides the field" is not a security model. Put client-writable projections behind rules that constrain ownership and fields, and put privileged fan-out/repair work on a trusted server path with application authorization.

Enterprise Native Pipeline operations can express query patterns that Standard/Core cannot, but that does not erase lifecycle, authorization, write amplification, or historical-snapshot decisions. MongoDB compatibility is a separate Enterprise mode. Chapter 03 uses Standard Native Core behavior unless explicitly labeled otherwise so the model remains portable across the course's default lab.

5. Hands-on lab: compare two schemas from the five AtlasMart journeys

Use the same local workspace and pinned package baseline from Chapter 02. The examples assume firebase emulators:start --only auth,firestore is running and that Admin SDK code points at FIRESTORE_EMULATOR_HOST=127.0.0.1:8080 with GCLOUD_PROJECT=demo-atlasmart-firestore. Browser examples connect the modular Web SDK to the Firestore/Auth emulators. Keep the fixtures synthetic and delete only the Chapter 03 paths you create.

seed · authoritative product, seller, order, dashboard, and feed fixtures
import { getFirestore } from "firebase-admin/firestore";const db = getFirestore();const now = new Date("2026-09-15T10:00:00Z");await db.doc("sellers/s-2001").set({ displayName: "Northwind Outdoor", schemaVersion: 1 });await db.doc("products/p-1001").set({ sellerId: "s-2001", name: "Trail Camera", priceCents: 12990, schemaVersion: 2 });await db.doc("users/u-1001").set({ displayName: "Asha", schemaVersion: 1 });await db.doc("orders/o-9001").set({  userId: "u-1001", createdAt: now,  lines: [{ productId: "p-1001", nameSnapshot: "Trail Camera", unitPriceCentsSnapshot: 12990, quantity: 1 }],  schemaVersion: 2});await db.doc("users/u-1001/feed/e-001").set({  type: "order_created", objectId: "o-9001", titleSnapshot: "Order o-9001 placed", sourceVersion: 1, createdAt: now});

Create a second fixture namespace such as modelComparisons/schemaA/* and modelComparisons/schemaB/* if you want both shapes present simultaneously. Run the product-detail, order-history, seller-dashboard, and feed repository functions through the operation counter. Record counts in a small JSON result file. Then change sellers/s-2001.displayName to Northwind Field Gear and intentionally leave one duplicated product/feed snapshot stale.

assertion · stale duplication must be visible, not silently accepted
const seller = await db.doc("sellers/s-2001").get();const product = await db.doc("products/p-1001").get();console.log({  authoritative: seller.get("displayName"),  duplicated: product.get("sellerNameSnapshot") ?? null,  diverged: product.get("sellerNameSnapshot") !== seller.get("displayName")});

Your expected evidence is not a fabricated latency number. It is a table containing each access pattern, requested document/query/write counts, duplicated fields, atomic boundary, Security Rules surface, and observed divergence after the injected rename. Finish by writing a repair plan: source path, projection paths, idempotency key/version, retry policy, verification query, and rollback/reset.

Verification checklist

  • Every schema object exists because a named screen/query/workflow needs it.
  • Historical snapshots are labeled snapshots and never mistaken for current catalog authority.
  • Operation counters distinguish document reads, queries, and writes.
  • The stale seller-name injection is detected programmatically.
  • Invariant-bearing writes identify their transaction/trusted-workflow boundary.
  • Cleanup targets only Chapter 03 fixture paths in the demo project/emulator.

6. Deliberately wrong approach: normalize by reflex

A team can reproduce a relational schema almost literally: every relationship becomes an ID/reference, every mutable attribute lives in one document, and every screen assembles its view by following references. The model looks tidy, but product detail and order history become client-side joins. Each extra read adds latency and authorization surface; offline behavior becomes more complex; historical receipts can accidentally display current product names/prices instead of purchase-time facts.

The repair is not indiscriminate duplication. Return to the access-pattern ledger. Duplicate only fields whose read value justifies the write/repair cost, mark snapshots versus current projections, version repairable projections, and keep invariants attached to authoritative documents and atomic workflows. A denormalized field without a source-of-truth and repair strategy is technical debt, not a performance feature.

Knowledge check

Check your understanding

  1. Why is an ER diagram insufficient as the starting point for a Firestore model?
  2. When is a duplicated field an intentional projection instead of accidental drift?
  3. Why should emulator operation counters not be labeled a production bill?
  4. Which AtlasMart facts can tolerate eventual repair, and which require an atomic trusted workflow?
  5. What extra obligation appears when a screen saves reads by copying seller/product fields?
Review the answers

1. Firestore design depends on executable query, listener, authorization, lifecycle, and atomicity constraints. Entity relationships alone do not tell you how many documents a screen must read or whether a query/rule shape is possible.

2. It is intentional when it has a named authoritative source, a documented freshness contract, a version/repair path, and tests that can detect divergence.

3. The emulator proves local logical behavior but not production billing dimensions, index work, regional latency, or listener/reconnect semantics.

4. Display projections such as seller names or dashboard summaries can often tolerate bounded lag; reservation/payment/stock invariants require authoritative checks and atomic or otherwise explicit consistency control.

5. Write amplification, consistency-window definition, repair/backfill logic, extra rules/indexes, and migration/rollback work.

Summary and next step

Firestore modeling starts from access patterns and consistency contracts. AtlasMart now has a ledger that makes reads, writes, rules, lifecycle, and transaction boundaries visible, plus two competing schemas whose amplification can be measured. Next, we make the embed/reference/duplicate choice explicit field by field and quantify the document-growth and repair consequences.

Next: Embed vs Reference vs Duplicate: Read Cost, Write Amplification, Consistency, and Document Limits.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.