Chapter 03 · Data Modeling by Access Pattern: Embedding, Referencing, Duplication, Fan-Out, and Denormalization

Embed vs Reference vs Duplicate: Read Cost, Write Amplification, Consistency, and Document Limits

Choose embedding, references, or deliberate duplication by measuring read amplification, write fan-out, consistency windows, document growth, and repair obligations in AtlasMart.

Beginner → Advanced100–130 minutesAtlasMart emulator-first modeling labFirebase CLI 15.30.0 · Web SDK 12.19.0 · Admin Node 14.4.0 · Node.js 22+Firestore Standard Native Core semantics unless explicitly labeled EnterpriseLast reviewed: September 2026

Learning outcomes

AtlasMart now knows which screens and workflows matter. The next question is where each fact lives. A product description may belong in one authoritative product document; a purchase-time unit price belongs in the order history even after the catalog changes; seller display data may be a repairable projection; an ever-growing review list does not belong in one product array. The choice among embedding, referencing, and duplication is therefore a contract about growth, reads, writes, and consistency.

01

Choose embedding when data is bounded, co-read, and shares lifecycle rather than simply because nesting is convenient.

02

Choose references when independent lifecycle/current authority matters and the extra read/query is acceptable.

03

Use duplication deliberately for immutable snapshots or repairable projections with explicit source/version semantics.

04

Quantify read amplification, write amplification, document growth, index fan-out, and rules complexity for each choice.

05

Inject and repair a stale projection without corrupting the authoritative source.

Chapter 03 baseline reviewed 15 September 2026

The lab continues Chapters 01–02 exactly: project demo-atlasmart-firestore; Firestore emulator 127.0.0.1:8080; Auth emulator 127.0.0.1:9099; Emulator UI 127.0.0.1:4000; Firebase CLI 15.30.0; Firebase JavaScript SDK 12.19.0; Firebase Admin Node.js SDK 14.4.0; Node.js 22 or newer. Standard Native Core semantics are the default. Enterprise Native Pipeline or MongoDB-compatibility behavior is mentioned only when the distinction changes a modeling decision.

Execution, pricing, and evidence note

The mandatory lab is emulator-first and free/local. Emulator reads and writes are useful for deterministic operation counting and correctness tests, but they are not billable production operations and do not reproduce regional latency, production index topology, contention, autoscaling, or billing. Where a table estimates production operations, it reports document/query/write counts only; convert them to money only after re-checking the current production pricing contract for the exact edition, region, query shape, listener state, and index work.

1. Embedding: one read, shared lifecycle, bounded growth

Embedding places child fields inside the same Firestore document. It is attractive when the fields are almost always read together, updated under the same authorization/lifecycle, and remain bounded well below Firestore's document/index limits. AtlasMart order line snapshots are a strong example when an order contains a bounded number of purchased lines: the receipt can read one order document and preserve purchase-time values.

Embedding becomes dangerous when a list grows without a designed bound. An array of every product review, every feed event, or every order history item forces the entire document to grow and rewrites/indexes that large field on updates. The 1 MiB document boundary and index-entry limits are correctness constraints, but you should design a much earlier application bound based on expected workload and update contention rather than treating the hard limit as a target.

Candidate Embed? Reason
Order line purchase snapshot Often yes Bounded, co-read, immutable history.
Product's latest 3 review snippets Possibly as a materialized summary Bounded projection; source reviews remain separate.
Every review ever written No as one array Unbounded growth and update/index amplification.
Seller legal/compliance profile Usually separate Different trust, lifecycle, and access surface.

2. Referencing: independent authority at the price of another lookup

A reference can be a string ID or Firestore DocumentReference depending on your contract. The important property is not the host-language type; it is that the related object remains independent. A product can keep sellerId as the authoritative relationship while the seller document changes independently. Reading current seller details then requires another read or a separate query unless a duplicated projection is also stored.

References do not enforce existence, cascades, or joins. If AtlasMart requires "every product must point to an active seller" as a write invariant, a trusted workflow must validate that state. If the client writes the relationship directly, Security Rules may need get()/exists() checks subject to current rules-access-call limits; those checks themselves must be designed and tested rather than assumed free or unlimited.

3. Duplication: snapshot versus projection

Duplication has two very different meanings. An immutable snapshot intentionally freezes the value as it was at a business event: order item name, unit price, tax label, or shipping-address summary. It should not be repaired when the source changes. A mutable projection copies current data to make a read shape cheaper: seller display name on product cards or a precomputed dashboard total. A projection needs a convergence strategy.

Duplicated field Kind Source of truth On source change
unitPriceCentsSnapshot in order line Historical snapshot Order business event Do not rewrite to current catalog price.
sellerNameSnapshot in product card Current projection sellers/{sellerId}.displayName Repair/backfill with version/idempotency.
reviewCount on product Materialized aggregate Review set + update workflow Reconcile against authoritative collection.
titleSnapshot in feed event Display projection/historical context Policy-dependent Document whether it should converge or remain historical.

Without this distinction, a repair job can corrupt history by rewriting order receipts, or a team can leave a "current" projection stale indefinitely because nobody knows it was meant to converge.

4. Quantify amplification with a small worksheet

modeling worksheet · counts are application operations, not billing claims
Pattern: order-history page, 10 orders, 3 lines/orderReference-heavy: 1 query + up to distinct product lookupsEmbedded snapshots: 1 query, larger order documents, no product joins for historyPattern: seller renameReference-heavy: 1 seller write; screens read current seller separatelyDuplicated projection: 1 source write + N projection writes/repairsPattern: product reviewsEmbedded unbounded array: 1 doc read, growing write/index/contention surfaceSubcollection: query child docs, stable parent size, separate lifecycle/index/rules surface

For production cost engineering, later chapters will combine these logical counts with current pricing, index-entry reads, listeners, storage, network, and edition-specific billing. Here the modeling objective is earlier: know where amplification comes from and why.

5. Hands-on lab: compare read/write shapes and inject divergence

Use the same local workspace and pinned package baseline from Chapter 02. The examples assume firebase emulators:start --only auth,firestore is running and that Admin SDK code points at FIRESTORE_EMULATOR_HOST=127.0.0.1:8080 with GCLOUD_PROJECT=demo-atlasmart-firestore. Browser examples connect the modular Web SDK to the Firestore/Auth emulators. Keep the fixtures synthetic and delete only the Chapter 03 paths you create.

Node.js · seed source plus immutable and mutable duplicates
const sellerRef = db.doc("sellers/s-2001");const productRef = db.doc("products/p-1001");const orderRef = db.doc("orders/o-9002");await sellerRef.set({ displayName: "Northwind Outdoor", projectionVersion: 3, schemaVersion: 1 });await productRef.set({  sellerId: "s-2001", sellerNameSnapshot: "Northwind Outdoor", sellerProjectionVersion: 3,  name: "Trail Camera", priceCents: 12990, schemaVersion: 3}, { merge: true });await orderRef.set({  userId: "u-1001",  lines: [{ productId: "p-1001", nameSnapshot: "Trail Camera", unitPriceCentsSnapshot: 12990, quantity: 1 }],  schemaVersion: 3});

Now change only the seller source to Northwind Field Gear and increment projectionVersion. The product's duplicated name should become detectably stale while the order's nameSnapshot remains intentionally unchanged.

repair · update mutable projection only when source version is newer
await db.runTransaction(async tx => {  const sellerSnap = await tx.get(sellerRef);  const productSnap = await tx.get(productRef);  const sourceVersion = sellerSnap.get("projectionVersion");  const copiedVersion = productSnap.get("sellerProjectionVersion") ?? 0;  if (sourceVersion > copiedVersion) {    tx.update(productRef, {      sellerNameSnapshot: sellerSnap.get("displayName"),      sellerProjectionVersion: sourceVersion    });  }});

This transaction is small and emulator-testable. It demonstrates a convergence mechanism, not a recommendation to update every duplicated document in one giant transaction. Large fan-out repair needs pagination, idempotency, retries, and progress tracking; Lesson 4 builds that operational model.

Verification checklist

  • Order receipt still shows the original purchase snapshot after catalog/seller changes.
  • Product projection divergence is detected by version comparison.
  • Repair updates only the mutable projection, not historical snapshots.
  • No unbounded review/feed arrays are introduced.
  • Operation worksheet records both saved reads and added writes.

6. Deliberately wrong approach: duplicate mutable fields with no repair contract

Suppose AtlasMart copies sellerName into every product and feed item because it makes reads easy. Months later a seller rebrands. Some products update, others fail during a partial backfill, old feed entries are ambiguous, and nobody can tell which copies are expected to be historical. The problem is not Firestore eventual consistency; it is an undefined application consistency contract.

The repair starts by classifying each duplicate as snapshot or projection, adding source identity/version metadata where convergence matters, making the updater idempotent, and providing a reconciliation query/job. Duplication should reduce a named read cost in exchange for a named write/repair cost.

Knowledge check

Check your understanding

  1. When is embedding a good fit?
  2. Why is a document reference not equivalent to a foreign key?
  3. What is the difference between an immutable snapshot and a mutable projection?
  4. Why should every mutable projection carry a repair/convergence strategy?
  5. Why is an unbounded review array a poor default even if one read can fetch it?
Review the answers

1. When the data is bounded, commonly read together, shares lifecycle/authorization, and the parent document remains comfortably within size/index/update constraints.

2. It identifies a path/value but does not enforce target existence, cascade behavior, or automatic join loading.

3. A snapshot intentionally preserves a historical value and should not converge; a projection represents current source state and should converge according to a documented policy.

4. Because write failures, retries, migrations, or asynchronous fan-out can leave copies stale. A source/version/idempotent repair path makes drift detectable and recoverable.

5. It creates unbounded document growth, write/index amplification, and contention risk; a subcollection keeps child growth independent.

Summary and next step

Embedding, referencing, and duplication are not style preferences. They trade document growth, extra reads, write amplification, historical correctness, and convergence work. AtlasMart now labels snapshots separately from projections and can repair mutable copies without rewriting history. Next, we place those data shapes into root collections or subcollections and prove the query, lifecycle, ownership, and Security Rules consequences.

Next: Subcollections vs Root Collections, Collection Groups, Ownership, Lifecycle, and Security Rules.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.