Chapter 03 · Data Modeling by Access Pattern: Embedding, Referencing, Duplication, Fan-Out, and Denormalization
Embed vs Reference vs Duplicate: Read Cost, Write Amplification, Consistency, and Document Limits
Choose embedding, references, or deliberate duplication by measuring read amplification, write fan-out, consistency windows, document growth, and repair obligations in AtlasMart.
Learning outcomes
AtlasMart now knows which screens and workflows matter. The next question is where each fact lives. A product description may belong in one authoritative product document; a purchase-time unit price belongs in the order history even after the catalog changes; seller display data may be a repairable projection; an ever-growing review list does not belong in one product array. The choice among embedding, referencing, and duplication is therefore a contract about growth, reads, writes, and consistency.
Choose embedding when data is bounded, co-read, and shares lifecycle rather than simply because nesting is convenient.
Choose references when independent lifecycle/current authority matters and the extra read/query is acceptable.
Use duplication deliberately for immutable snapshots or repairable projections with explicit source/version semantics.
Quantify read amplification, write amplification, document growth, index fan-out, and rules complexity for each choice.
Inject and repair a stale projection without corrupting the authoritative source.
The lab continues Chapters 01–02 exactly: project
demo-atlasmart-firestore; Firestore emulator
127.0.0.1:8080; Auth emulator
127.0.0.1:9099; Emulator UI
127.0.0.1:4000; Firebase CLI
15.30.0; Firebase JavaScript SDK
12.19.0; Firebase Admin Node.js SDK
14.4.0; Node.js 22 or newer. Standard Native Core
semantics are the default. Enterprise Native Pipeline or
MongoDB-compatibility behavior is mentioned only when the
distinction changes a modeling decision.
The mandatory lab is emulator-first and free/local. Emulator reads and writes are useful for deterministic operation counting and correctness tests, but they are not billable production operations and do not reproduce regional latency, production index topology, contention, autoscaling, or billing. Where a table estimates production operations, it reports document/query/write counts only; convert them to money only after re-checking the current production pricing contract for the exact edition, region, query shape, listener state, and index work.
1. Embedding: one read, shared lifecycle, bounded growth
Embedding places child fields inside the same Firestore document. It is attractive when the fields are almost always read together, updated under the same authorization/lifecycle, and remain bounded well below Firestore's document/index limits. AtlasMart order line snapshots are a strong example when an order contains a bounded number of purchased lines: the receipt can read one order document and preserve purchase-time values.
Embedding becomes dangerous when a list grows without a designed bound. An array of every product review, every feed event, or every order history item forces the entire document to grow and rewrites/indexes that large field on updates. The 1 MiB document boundary and index-entry limits are correctness constraints, but you should design a much earlier application bound based on expected workload and update contention rather than treating the hard limit as a target.
| Candidate | Embed? | Reason |
|---|---|---|
| Order line purchase snapshot | Often yes | Bounded, co-read, immutable history. |
| Product's latest 3 review snippets | Possibly as a materialized summary | Bounded projection; source reviews remain separate. |
| Every review ever written | No as one array | Unbounded growth and update/index amplification. |
| Seller legal/compliance profile | Usually separate | Different trust, lifecycle, and access surface. |
2. Referencing: independent authority at the price of another lookup
A reference can be a string ID or Firestore
DocumentReference depending on your contract. The
important property is not the host-language type; it is that the
related object remains independent. A product can keep
sellerId as the authoritative relationship while
the seller document changes independently. Reading current
seller details then requires another read or a separate query
unless a duplicated projection is also stored.
References do not enforce existence, cascades, or joins. If
AtlasMart requires "every product must point to an active
seller" as a write invariant, a trusted workflow must validate
that state. If the client writes the relationship directly,
Security Rules may need get()/exists()
checks subject to current rules-access-call limits; those checks
themselves must be designed and tested rather than assumed free
or unlimited.
3. Duplication: snapshot versus projection
Duplication has two very different meanings. An immutable snapshot intentionally freezes the value as it was at a business event: order item name, unit price, tax label, or shipping-address summary. It should not be repaired when the source changes. A mutable projection copies current data to make a read shape cheaper: seller display name on product cards or a precomputed dashboard total. A projection needs a convergence strategy.
| Duplicated field | Kind | Source of truth | On source change |
|---|---|---|---|
unitPriceCentsSnapshot in order line |
Historical snapshot | Order business event | Do not rewrite to current catalog price. |
sellerNameSnapshot in product card |
Current projection | sellers/{sellerId}.displayName |
Repair/backfill with version/idempotency. |
reviewCount on product |
Materialized aggregate | Review set + update workflow | Reconcile against authoritative collection. |
titleSnapshot in feed event |
Display projection/historical context | Policy-dependent | Document whether it should converge or remain historical. |
Without this distinction, a repair job can corrupt history by rewriting order receipts, or a team can leave a "current" projection stale indefinitely because nobody knows it was meant to converge.
4. Quantify amplification with a small worksheet
Pattern: order-history page, 10 orders, 3 lines/orderReference-heavy: 1 query + up to distinct product lookupsEmbedded snapshots: 1 query, larger order documents, no product joins for historyPattern: seller renameReference-heavy: 1 seller write; screens read current seller separatelyDuplicated projection: 1 source write + N projection writes/repairsPattern: product reviewsEmbedded unbounded array: 1 doc read, growing write/index/contention surfaceSubcollection: query child docs, stable parent size, separate lifecycle/index/rules surface
For production cost engineering, later chapters will combine these logical counts with current pricing, index-entry reads, listeners, storage, network, and edition-specific billing. Here the modeling objective is earlier: know where amplification comes from and why.
5. Hands-on lab: compare read/write shapes and inject divergence
Use the same local workspace and pinned package baseline from
Chapter 02. The examples assume
firebase emulators:start --only auth,firestore is
running and that Admin SDK code points at
FIRESTORE_EMULATOR_HOST=127.0.0.1:8080 with
GCLOUD_PROJECT=demo-atlasmart-firestore. Browser
examples connect the modular Web SDK to the Firestore/Auth
emulators. Keep the fixtures synthetic and delete only the
Chapter 03 paths you create.
const sellerRef = db.doc("sellers/s-2001");const productRef = db.doc("products/p-1001");const orderRef = db.doc("orders/o-9002");await sellerRef.set({ displayName: "Northwind Outdoor", projectionVersion: 3, schemaVersion: 1 });await productRef.set({ sellerId: "s-2001", sellerNameSnapshot: "Northwind Outdoor", sellerProjectionVersion: 3, name: "Trail Camera", priceCents: 12990, schemaVersion: 3}, { merge: true });await orderRef.set({ userId: "u-1001", lines: [{ productId: "p-1001", nameSnapshot: "Trail Camera", unitPriceCentsSnapshot: 12990, quantity: 1 }], schemaVersion: 3});
Now change only the seller source to
Northwind Field Gear and increment
projectionVersion. The product's duplicated name
should become detectably stale while the order's
nameSnapshot remains intentionally unchanged.
await db.runTransaction(async tx => { const sellerSnap = await tx.get(sellerRef); const productSnap = await tx.get(productRef); const sourceVersion = sellerSnap.get("projectionVersion"); const copiedVersion = productSnap.get("sellerProjectionVersion") ?? 0; if (sourceVersion > copiedVersion) { tx.update(productRef, { sellerNameSnapshot: sellerSnap.get("displayName"), sellerProjectionVersion: sourceVersion }); }});
This transaction is small and emulator-testable. It demonstrates a convergence mechanism, not a recommendation to update every duplicated document in one giant transaction. Large fan-out repair needs pagination, idempotency, retries, and progress tracking; Lesson 4 builds that operational model.
Verification checklist
- Order receipt still shows the original purchase snapshot after catalog/seller changes.
- Product projection divergence is detected by version comparison.
- Repair updates only the mutable projection, not historical snapshots.
- No unbounded review/feed arrays are introduced.
- Operation worksheet records both saved reads and added writes.
6. Deliberately wrong approach: duplicate mutable fields with no repair contract
Suppose AtlasMart copies sellerName into every
product and feed item because it makes reads easy. Months later
a seller rebrands. Some products update, others fail during a
partial backfill, old feed entries are ambiguous, and nobody can
tell which copies are expected to be historical. The problem is
not Firestore eventual consistency; it is an undefined
application consistency contract.
The repair starts by classifying each duplicate as snapshot or projection, adding source identity/version metadata where convergence matters, making the updater idempotent, and providing a reconciliation query/job. Duplication should reduce a named read cost in exchange for a named write/repair cost.
Knowledge check
Check your understanding
- When is embedding a good fit?
- Why is a document reference not equivalent to a foreign key?
- What is the difference between an immutable snapshot and a mutable projection?
- Why should every mutable projection carry a repair/convergence strategy?
- Why is an unbounded review array a poor default even if one read can fetch it?
Review the answers
1. When the data is bounded, commonly read together, shares lifecycle/authorization, and the parent document remains comfortably within size/index/update constraints.
2. It identifies a path/value but does not enforce target existence, cascade behavior, or automatic join loading.
3. A snapshot intentionally preserves a historical value and should not converge; a projection represents current source state and should converge according to a documented policy.
4. Because write failures, retries, migrations, or asynchronous fan-out can leave copies stale. A source/version/idempotent repair path makes drift detectable and recoverable.
5. It creates unbounded document growth, write/index amplification, and contention risk; a subcollection keeps child growth independent.
Summary and next step
Embedding, referencing, and duplication are not style preferences. They trade document growth, extra reads, write amplification, historical correctness, and convergence work. AtlasMart now labels snapshots separately from projections and can repair mutable copies without rewriting history. Next, we place those data shapes into root collections or subcollections and prove the query, lifecycle, ownership, and Security Rules consequences.
Next: Subcollections vs Root Collections, Collection Groups, Ownership, Lifecycle, and Security Rules.
Authoritative references
- Choose a data structure — Official guidance for nested data, subcollections, and root-level collections.
- Cloud Firestore data model — Documents, collections, subcollections, paths, and parent-deletion behavior.
- Perform simple and compound queries — Collection and collection-group query semantics.
- Securely query data — Security Rules and query constraints; rules are not filters.
- Transactions and batched writes — Atomicity and retry boundaries used by bounded fan-out workflows.
- Distributed counters — Official sharded-counter pattern for higher update rates.
- Best practices for Cloud Firestore — IDs, hotspotting, index fan-out, location, and write-shape guidance.
- Delete data — Parent/subcollection lifecycle and bulk-delete considerations.
- Cloud Firestore pricing — Living production billing contract; re-check before converting operation counts into currency.
- Firebase release notes — Current Firebase SDK and CLI version baseline.
- Connect to the Firestore Emulator — Local development path and production-difference guidance.