Chapter 03 · Data Modeling by Access Pattern: Embedding, Referencing, Duplication, Fan-Out, and Denormalization
Refactor a Relational Model into Firestore and Document Every Consistency and Cost Tradeoff
Translate a normalized AtlasMart relational model into Firestore access-pattern models and document every consistency, cost, security, lifecycle, migration, and rollback tradeoff.
Learning outcomes
AtlasMart begins with a normalized relational design: users, sellers, products, orders, order_items, reviews, carts, and events connected by foreign keys. The goal is not to "convert tables to collections." It is to preserve business meaning while reshaping storage around Firestore access patterns, explicit authorization, bounded documents, query/index constraints, lifecycle, and consistency requirements.
Refactor a relational schema by preserving business invariants while changing read/write shapes.
Compare at least two Firestore target models and document why one is selected per access pattern.
Separate authoritative operational documents, immutable event snapshots, and rebuildable projections.
Plan a versioned migration with dual-read/dual-write or backfill stages, validation, and rollback.
Produce a decision matrix covering reads, writes, rules, indexes, lifecycle, analytics, cost dimensions, and future portability.
The lab continues Chapters 01–02 exactly: project
demo-atlasmart-firestore; Firestore emulator
127.0.0.1:8080; Auth emulator
127.0.0.1:9099; Emulator UI
127.0.0.1:4000; Firebase CLI
15.30.0; Firebase JavaScript SDK
12.19.0; Firebase Admin Node.js SDK
14.4.0; Node.js 22 or newer. Standard Native Core
semantics are the default. Enterprise Native Pipeline or
MongoDB-compatibility behavior is mentioned only when the
distinction changes a modeling decision.
The mandatory lab is emulator-first and free/local. Emulator reads and writes are useful for deterministic operation counting and correctness tests, but they are not billable production operations and do not reproduce regional latency, production index topology, contention, autoscaling, or billing. Where a table estimates production operations, it reports document/query/write counts only; convert them to money only after re-checking the current production pricing contract for the exact edition, region, query shape, listener state, and index work.
1. Preserve invariants, not relational layout
A normalized database might store orders and
order_items separately so product, seller, and
pricing facts are referenced by foreign keys. In Firestore, an
order-history screen often benefits from one order document
containing a bounded array of immutable line snapshots. That is
a deliberate denormalization, but it does not mean every
relationship should be embedded.
Start by listing invariants that must survive the migration: an order belongs to exactly one user; purchase-time totals cannot silently change when the product catalog changes; stock/reservation transitions must be checked atomically on authoritative records; seller/admin access must not rely on client UI state; deleted users/products need documented retention and right-to-delete behavior. Those are business semantics independent of database shape.
2. Map relational queries to Firestore access contracts
| Relational operation | Firestore target shape | Tradeoff introduced |
|---|---|---|
| JOIN order + order_items + product for receipt | Embed bounded immutable line snapshots in order | Larger order write; snapshot intentionally diverges from current catalog. |
| SELECT reviews WHERE product_id=? | products/{productId}/reviews |
Natural local ownership; explicit recursive lifecycle; collection-group for global moderation. |
| JOIN seller + products for product cards | Product keeps sellerId + repairable sellerNameSnapshot | Saved read versus projection repair/write amplification. |
| GROUP BY seller/day for dashboard | Materialized seller/day summary or later aggregation query | Projection lag/replay or query-time cost. |
| ORDER BY user events DESC |
users/{uid}/feed/{eventId} with queryable
createdAt
|
Fan-out/appended projection; retention/indexing decisions. |
Enterprise Native Pipeline operations can support additional query shapes, including capabilities that may look more relational, but this course's Standard/Core target model should not depend on them unless AtlasMart explicitly chooses Enterprise and verifies current launch stage, SDK, location, and billing. MongoDB compatibility is a separate mode and is not a migration shortcut for relational semantics.
3. Choose a target model and label every consistency class
users/{uid} authoritative profile fieldssellers/{sellerId} authoritative seller identity/display sourceproducts/{productId} authoritative catalog + repairable seller display projectionproducts/{{productId}}/reviews/{id} authoritative review records owned by review workflowcarts/{uid}/items/{productId} user-owned mutable intent; checkout revalidates authorityorders/{orderId} authoritative order state + immutable purchase snapshotsusers/{{uid}}/feed/{{eventId}} rebuildable user read projectionsellers/{sellerId}/dashboard/{day} rebuildable/materialized aggregateprojectionEvents/{eventId} idempotency/application marker
Add schemaVersion to documents whose shape will
evolve and use projection/source versions where copied mutable
values must converge. Version fields are not magic migrations:
the reader/writer code and backfill plan must still support
old/new versions during rollout.
4. Migration plan: inventory → backfill → validate → cut over → reconcile
A safe migration from a relational source starts with a compatibility and access-pattern inventory, not a blind export. Decide which system is authoritative during each phase and how writes are handled. A common strategy is: create new Firestore documents with versioned IDs; backfill historical data; validate counts and sampled business invariants; introduce dual-read or shadow-read comparison; optionally dual-write through a trusted service; cut reads over by feature; then retire the old path only after rollback windows and reconciliation are complete.
| Phase | Evidence | Rollback |
|---|---|---|
| Inventory | Query/access ledger, invariants, PII/retention map, expected volume | No data mutation yet. |
| Backfill | Source IDs → Firestore paths, counts, checksums/samples, schemaVersion | Delete only migration namespace/version or rerun idempotently. |
| Shadow read | Compare relational result with Firestore canonicalized result | Keep production read on source. |
| Dual write / event replicate | Lag, failures, idempotency markers, reconciliation | Source remains authority; stop secondary writer. |
| Cutover | Error rate, read shape, authorization/rules, business metrics | Feature flag back to source if contracts permit. |
| Retire | Retention/legal/backup signoff | Only after explicit irreversible-deletion approval. |
5. Hands-on lab: refactor and compare two Firestore models
Use the same local workspace and pinned package baseline from
Chapter 02. The examples assume
firebase emulators:start --only auth,firestore is
running and that Admin SDK code points at
FIRESTORE_EMULATOR_HOST=127.0.0.1:8080 with
GCLOUD_PROJECT=demo-atlasmart-firestore. Browser
examples connect the modular Web SDK to the Firestore/Auth
emulators. Keep the fixtures synthetic and delete only the
Chapter 03 paths you create.
Create two emulator namespaces from the same small synthetic relational fixture. Model A keeps reference-heavy documents; Model B embeds order snapshots and maintains feed/dashboard projections. Write repository functions for the five Chapter 03 journeys and record logical operations, document paths, rules surface, and consistency class.
{ "users": [{"id":"u-1001","display_name":"Asha"}], "sellers": [{"id":"s-2001","display_name":"Northwind Outdoor"}], "products": [{"id":"p-1001","seller_id":"s-2001","name":"Trail Camera","price_cents":12990}], "orders": [{"id":"o-9004","user_id":"u-1001","status":"placed"}], "order_items": [{"order_id":"o-9004","product_id":"p-1001","unit_price_cents":12990,"quantity":1}], "reviews": [{"id":"r-010","product_id":"p-1001","author_uid":"u-1001","rating":5}]}
For each target model, generate a machine-readable decision
record. Include readOpsRequested,
queryOpsRequested, writeDocsTouched,
atomicBoundary, sourceOfTruth,
projectionRepair, rulesNotes,
indexNotes, lifecycleNotes, and
rollback. Deliberately change the seller name and
product price after order creation. The selected model must
preserve the order's purchase snapshot while current product
screens converge to the new seller projection.
const order = await db.doc("orders/o-9004").get();const product = await db.doc("products/p-1001").get();console.assert(order.get("lines.0.unitPriceCentsSnapshot") === 12990);console.assert(order.get("lines.0.nameSnapshot") === "Trail Camera");console.assert(product.get("priceCents") !== undefined);// Add project-specific canonical comparison for seller projection version.
If your SDK accessor does not support an array field path
exactly as shown, retrieve lines and assert in
JavaScript; the lesson's invariant is the important part. Do not
fabricate successful output—run these assertions locally after
extracting the ZIP into the academy repository and installing
the pinned lab dependencies.
Decision matrix acceptance criteria
- Every duplicated mutable field has a source and repair/version contract.
- Every immutable snapshot is explicitly protected from "repair" to current source values.
- Every cross-document invariant names its transaction/trusted workflow.
- Every subcollection has a lifecycle/delete/retention policy.
- Every client query is compatible with its Security Rules constraints.
- Every projection has idempotent replay/reconciliation.
- Operation counts are recorded separately from current production pricing.
- Migration includes rollback until the retirement phase is explicitly approved.
6. Production judgment: denormalization is a portfolio of obligations
The chosen Firestore model should be defended as a set of tradeoffs, not as "NoSQL best practice." Embedding can reduce reads while increasing document size/update scope. References preserve independent authority while increasing read chains. Duplication accelerates read shapes while increasing writes and repair. Subcollections encode local hierarchy but create recursive lifecycle work. Materialized views shift query work into fan-out/reconciliation. These choices also alter Security Rules complexity, indexes, offline behavior, analytics export needs, and future migrations.
Keep exit strategy visible. Stable external IDs, schema versions, source/projection labeling, and canonical validation make it easier to move data later. Avoid burying essential business meaning in undocumented path conventions or one-off client code.
7. Deliberately wrong approach: "convert every SQL table into one collection"
This mechanical migration preserves relational storage boundaries while removing the database join engine. The application then performs client-side joins across many collections, duplicates authorization checks, and often pays more reads without gaining an access-pattern model. At the other extreme, placing the entire relational graph into one giant document creates size, contention, and update-granularity problems.
The repair is access-pattern refactoring: preserve invariants, then choose document boundaries and projections based on exact reads, writes, rules, and lifecycle. Relational normalization and Firestore denormalization are both tools; neither is a universal moral rule.
Knowledge check
Check your understanding
- What should survive a relational-to-Firestore migration even when document boundaries change?
- Why can an order embed product fields while the product remains an independent source?
- What makes a projection safe to rebuild?
- Why is a shadow-read phase useful?
- What is wrong with converting every SQL table directly to a Firestore collection?
Review the answers
1. Business invariants, authorization requirements, historical semantics, retention/lifecycle obligations, and externally meaningful identities.
2. The embedded fields are purchase-time immutable snapshots; current catalog authority stays on the product document.
3. It has authoritative source data, deterministic derivation, stable IDs/idempotency, schema/source versions, and reconciliation tests.
4. It compares Firestore results with the source without immediately changing the user-visible production read path, exposing semantic mismatches before cutover.
5. It preserves storage normalization after removing server-side relational joins, often creating client-side join/read amplification and awkward rules instead of modeling the actual access patterns.
Summary and next step
Chapter 03 turns Firestore data modeling into a measurable contract. AtlasMart can now derive schemas from screens and workflows, choose embedding/reference/duplication deliberately, place data in root/subcollection shapes with lifecycle and rules in mind, build idempotent materialized views, and refactor a relational model without losing business invariants. Chapter 04 can now focus on CRUD semantics against this stable model.
Next: Get Documents, Create/Set/Merge, Update, Delete, and Understand Missing vs Null Fields.
Authoritative references
- Choose a data structure — Official guidance for nested data, subcollections, and root-level collections.
- Cloud Firestore data model — Documents, collections, subcollections, paths, and parent-deletion behavior.
- Perform simple and compound queries — Collection and collection-group query semantics.
- Securely query data — Security Rules and query constraints; rules are not filters.
- Transactions and batched writes — Atomicity and retry boundaries used by bounded fan-out workflows.
- Distributed counters — Official sharded-counter pattern for higher update rates.
- Best practices for Cloud Firestore — IDs, hotspotting, index fan-out, location, and write-shape guidance.
- Delete data — Parent/subcollection lifecycle and bulk-delete considerations.
- Cloud Firestore pricing — Living production billing contract; re-check before converting operation counts into currency.
- Firebase release notes — Current Firebase SDK and CLI version baseline.
- Connect to the Firestore Emulator — Local development path and production-difference guidance.