Chapter 03 · Data Modeling by Access Pattern: Embedding, Referencing, Duplication, Fan-Out, and Denormalization

Refactor a Relational Model into Firestore and Document Every Consistency and Cost Tradeoff

Translate a normalized AtlasMart relational model into Firestore access-pattern models and document every consistency, cost, security, lifecycle, migration, and rollback tradeoff.

Beginner → Advanced100–130 minutesAtlasMart emulator-first modeling labFirebase CLI 15.30.0 · Web SDK 12.19.0 · Admin Node 14.4.0 · Node.js 22+Firestore Standard Native Core semantics unless explicitly labeled EnterpriseLast reviewed: September 2026

Learning outcomes

AtlasMart begins with a normalized relational design: users, sellers, products, orders, order_items, reviews, carts, and events connected by foreign keys. The goal is not to "convert tables to collections." It is to preserve business meaning while reshaping storage around Firestore access patterns, explicit authorization, bounded documents, query/index constraints, lifecycle, and consistency requirements.

01

Refactor a relational schema by preserving business invariants while changing read/write shapes.

02

Compare at least two Firestore target models and document why one is selected per access pattern.

03

Separate authoritative operational documents, immutable event snapshots, and rebuildable projections.

04

Plan a versioned migration with dual-read/dual-write or backfill stages, validation, and rollback.

05

Produce a decision matrix covering reads, writes, rules, indexes, lifecycle, analytics, cost dimensions, and future portability.

Chapter 03 baseline reviewed 15 September 2026

The lab continues Chapters 01–02 exactly: project demo-atlasmart-firestore; Firestore emulator 127.0.0.1:8080; Auth emulator 127.0.0.1:9099; Emulator UI 127.0.0.1:4000; Firebase CLI 15.30.0; Firebase JavaScript SDK 12.19.0; Firebase Admin Node.js SDK 14.4.0; Node.js 22 or newer. Standard Native Core semantics are the default. Enterprise Native Pipeline or MongoDB-compatibility behavior is mentioned only when the distinction changes a modeling decision.

Execution, pricing, and evidence note

The mandatory lab is emulator-first and free/local. Emulator reads and writes are useful for deterministic operation counting and correctness tests, but they are not billable production operations and do not reproduce regional latency, production index topology, contention, autoscaling, or billing. Where a table estimates production operations, it reports document/query/write counts only; convert them to money only after re-checking the current production pricing contract for the exact edition, region, query shape, listener state, and index work.

1. Preserve invariants, not relational layout

A normalized database might store orders and order_items separately so product, seller, and pricing facts are referenced by foreign keys. In Firestore, an order-history screen often benefits from one order document containing a bounded array of immutable line snapshots. That is a deliberate denormalization, but it does not mean every relationship should be embedded.

Start by listing invariants that must survive the migration: an order belongs to exactly one user; purchase-time totals cannot silently change when the product catalog changes; stock/reservation transitions must be checked atomically on authoritative records; seller/admin access must not rely on client UI state; deleted users/products need documented retention and right-to-delete behavior. Those are business semantics independent of database shape.

2. Map relational queries to Firestore access contracts

Relational operation Firestore target shape Tradeoff introduced
JOIN order + order_items + product for receipt Embed bounded immutable line snapshots in order Larger order write; snapshot intentionally diverges from current catalog.
SELECT reviews WHERE product_id=? products/{productId}/reviews Natural local ownership; explicit recursive lifecycle; collection-group for global moderation.
JOIN seller + products for product cards Product keeps sellerId + repairable sellerNameSnapshot Saved read versus projection repair/write amplification.
GROUP BY seller/day for dashboard Materialized seller/day summary or later aggregation query Projection lag/replay or query-time cost.
ORDER BY user events DESC users/{uid}/feed/{eventId} with queryable createdAt Fan-out/appended projection; retention/indexing decisions.

Enterprise Native Pipeline operations can support additional query shapes, including capabilities that may look more relational, but this course's Standard/Core target model should not depend on them unless AtlasMart explicitly chooses Enterprise and verifies current launch stage, SDK, location, and billing. MongoDB compatibility is a separate mode and is not a migration shortcut for relational semantics.

3. Choose a target model and label every consistency class

target model · authoritative, snapshot, and projection labels
users/{uid}                         authoritative profile fieldssellers/{sellerId}                 authoritative seller identity/display sourceproducts/{productId}               authoritative catalog + repairable seller display projectionproducts/{{productId}}/reviews/{id}  authoritative review records owned by review workflowcarts/{uid}/items/{productId}       user-owned mutable intent; checkout revalidates authorityorders/{orderId}                    authoritative order state + immutable purchase snapshotsusers/{{uid}}/feed/{{eventId}}          rebuildable user read projectionsellers/{sellerId}/dashboard/{day} rebuildable/materialized aggregateprojectionEvents/{eventId}          idempotency/application marker

Add schemaVersion to documents whose shape will evolve and use projection/source versions where copied mutable values must converge. Version fields are not magic migrations: the reader/writer code and backfill plan must still support old/new versions during rollout.

4. Migration plan: inventory → backfill → validate → cut over → reconcile

A safe migration from a relational source starts with a compatibility and access-pattern inventory, not a blind export. Decide which system is authoritative during each phase and how writes are handled. A common strategy is: create new Firestore documents with versioned IDs; backfill historical data; validate counts and sampled business invariants; introduce dual-read or shadow-read comparison; optionally dual-write through a trusted service; cut reads over by feature; then retire the old path only after rollback windows and reconciliation are complete.

Phase Evidence Rollback
Inventory Query/access ledger, invariants, PII/retention map, expected volume No data mutation yet.
Backfill Source IDs → Firestore paths, counts, checksums/samples, schemaVersion Delete only migration namespace/version or rerun idempotently.
Shadow read Compare relational result with Firestore canonicalized result Keep production read on source.
Dual write / event replicate Lag, failures, idempotency markers, reconciliation Source remains authority; stop secondary writer.
Cutover Error rate, read shape, authorization/rules, business metrics Feature flag back to source if contracts permit.
Retire Retention/legal/backup signoff Only after explicit irreversible-deletion approval.

5. Hands-on lab: refactor and compare two Firestore models

Use the same local workspace and pinned package baseline from Chapter 02. The examples assume firebase emulators:start --only auth,firestore is running and that Admin SDK code points at FIRESTORE_EMULATOR_HOST=127.0.0.1:8080 with GCLOUD_PROJECT=demo-atlasmart-firestore. Browser examples connect the modular Web SDK to the Firestore/Auth emulators. Keep the fixtures synthetic and delete only the Chapter 03 paths you create.

Create two emulator namespaces from the same small synthetic relational fixture. Model A keeps reference-heavy documents; Model B embeds order snapshots and maintains feed/dashboard projections. Write repository functions for the five Chapter 03 journeys and record logical operations, document paths, rules surface, and consistency class.

fixture · relational source rows represented as deterministic JSON
{  "users": [{"id":"u-1001","display_name":"Asha"}],  "sellers": [{"id":"s-2001","display_name":"Northwind Outdoor"}],  "products": [{"id":"p-1001","seller_id":"s-2001","name":"Trail Camera","price_cents":12990}],  "orders": [{"id":"o-9004","user_id":"u-1001","status":"placed"}],  "order_items": [{"order_id":"o-9004","product_id":"p-1001","unit_price_cents":12990,"quantity":1}],  "reviews": [{"id":"r-010","product_id":"p-1001","author_uid":"u-1001","rating":5}]}

For each target model, generate a machine-readable decision record. Include readOpsRequested, queryOpsRequested, writeDocsTouched, atomicBoundary, sourceOfTruth, projectionRepair, rulesNotes, indexNotes, lifecycleNotes, and rollback. Deliberately change the seller name and product price after order creation. The selected model must preserve the order's purchase snapshot while current product screens converge to the new seller projection.

validation · invariants that survive the refactor
const order = await db.doc("orders/o-9004").get();const product = await db.doc("products/p-1001").get();console.assert(order.get("lines.0.unitPriceCentsSnapshot") === 12990);console.assert(order.get("lines.0.nameSnapshot") === "Trail Camera");console.assert(product.get("priceCents") !== undefined);// Add project-specific canonical comparison for seller projection version.

If your SDK accessor does not support an array field path exactly as shown, retrieve lines and assert in JavaScript; the lesson's invariant is the important part. Do not fabricate successful output—run these assertions locally after extracting the ZIP into the academy repository and installing the pinned lab dependencies.

Decision matrix acceptance criteria

  • Every duplicated mutable field has a source and repair/version contract.
  • Every immutable snapshot is explicitly protected from "repair" to current source values.
  • Every cross-document invariant names its transaction/trusted workflow.
  • Every subcollection has a lifecycle/delete/retention policy.
  • Every client query is compatible with its Security Rules constraints.
  • Every projection has idempotent replay/reconciliation.
  • Operation counts are recorded separately from current production pricing.
  • Migration includes rollback until the retirement phase is explicitly approved.

6. Production judgment: denormalization is a portfolio of obligations

The chosen Firestore model should be defended as a set of tradeoffs, not as "NoSQL best practice." Embedding can reduce reads while increasing document size/update scope. References preserve independent authority while increasing read chains. Duplication accelerates read shapes while increasing writes and repair. Subcollections encode local hierarchy but create recursive lifecycle work. Materialized views shift query work into fan-out/reconciliation. These choices also alter Security Rules complexity, indexes, offline behavior, analytics export needs, and future migrations.

Keep exit strategy visible. Stable external IDs, schema versions, source/projection labeling, and canonical validation make it easier to move data later. Avoid burying essential business meaning in undocumented path conventions or one-off client code.

7. Deliberately wrong approach: "convert every SQL table into one collection"

This mechanical migration preserves relational storage boundaries while removing the database join engine. The application then performs client-side joins across many collections, duplicates authorization checks, and often pays more reads without gaining an access-pattern model. At the other extreme, placing the entire relational graph into one giant document creates size, contention, and update-granularity problems.

The repair is access-pattern refactoring: preserve invariants, then choose document boundaries and projections based on exact reads, writes, rules, and lifecycle. Relational normalization and Firestore denormalization are both tools; neither is a universal moral rule.

Knowledge check

Check your understanding

  1. What should survive a relational-to-Firestore migration even when document boundaries change?
  2. Why can an order embed product fields while the product remains an independent source?
  3. What makes a projection safe to rebuild?
  4. Why is a shadow-read phase useful?
  5. What is wrong with converting every SQL table directly to a Firestore collection?
Review the answers

1. Business invariants, authorization requirements, historical semantics, retention/lifecycle obligations, and externally meaningful identities.

2. The embedded fields are purchase-time immutable snapshots; current catalog authority stays on the product document.

3. It has authoritative source data, deterministic derivation, stable IDs/idempotency, schema/source versions, and reconciliation tests.

4. It compares Firestore results with the source without immediately changing the user-visible production read path, exposing semantic mismatches before cutover.

5. It preserves storage normalization after removing server-side relational joins, often creating client-side join/read amplification and awkward rules instead of modeling the actual access patterns.

Summary and next step

Chapter 03 turns Firestore data modeling into a measurable contract. AtlasMart can now derive schemas from screens and workflows, choose embedding/reference/duplication deliberately, place data in root/subcollection shapes with lifecycle and rules in mind, build idempotent materialized views, and refactor a relational model without losing business invariants. Chapter 04 can now focus on CRUD semantics against this stable model.

Next: Get Documents, Create/Set/Merge, Update, Delete, and Understand Missing vs Null Fields.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.