Chapter 11 · Document Databases and Aggregate-Oriented Modeling
Embed vs Reference: Aggregate Boundaries, Duplication, Update Frequency, and Read Locality
Choose embedding or references from ownership, cardinality, update frequency, read locality, and transactional boundaries, then expose the consistency cost of duplicated mutable facts.
Learning outcomes
AtlasMart now has documents, but the harder question is where one document should end. Embedding and referencing are not style preferences. They move read locality, duplication, update fan-out, and atomicity boundaries.
Define an aggregate as a consistency/ownership boundary, not merely a nested JSON shape.
Choose embedding when owned data is bounded and usually read with its parent.
Choose references when related data has independent lifecycle, high fan-out, or frequent independent access.
Separate immutable snapshots from mutable duplicated truth.
1. Embed for owned locality; reference for independent identity
An aggregate is a group of state that should normally be read and changed as a unit under one owner. Embedding places child state inside the parent document. Referencing stores the related entity separately and keeps an identifier/link. The decision depends on workload evidence, not on whether a relational schema previously had a foreign key.
| Criterion | Embedding pressure | Referencing pressure |
|---|---|---|
| ownership | child exists only with parent | entity has independent identity/lifecycle |
| cardinality | small and predictably bounded | large or unbounded |
| read pattern | usually fetched with parent | often queried independently |
| update frequency | changes with parent or is immutable snapshot | changes frequently across many parents |
| consistency | single-document atomic update is valuable | independent update plus explicit synchronization is acceptable |
2. AtlasMart: orders should snapshot some facts and reference others
An order line’s price paid is historical evidence. If the catalog price changes tomorrow, the old order must not change. Embedding the purchased SKU, display name, quantity, and price-at-purchase is therefore deliberate duplication. A customer’s current marketing preference is different: copying it into every order creates mutable duplicated truth and repair work.
from collections import Counter
orders = [
{"id": "o1", "customer_id": "c1", "items": [{"sku": "A", "name": "Lamp", "price_paid": 30}]},
{"id": "o2", "customer_id": "c1", "items": [{"sku": "A", "name": "Lamp", "price_paid": 32}]},
]
customer = {"id": "c1", "email": "old@example.test", "marketing_opt_in": False}
# Bad model: copy mutable customer email into every order and treat it as current truth.
for o in orders:
o["customer_email"] = customer["email"]
customer["email"] = "new@example.test"
stale = [o["id"] for o in orders if o["customer_email"] != customer["email"]]
print("stale duplicated customer email in orders:", stale)
# Better boundary: immutable purchase snapshot is embedded; mutable customer profile is referenced.
read_cost = Counter()
def load_order(order_id):
read_cost["order_docs"] += 1
o = next(o for o in orders if o["id"] == order_id)
read_cost["customer_docs"] += 1
return {"order": o, "current_customer": customer}
view = load_order("o2")
print("price-at-purchase:", view["order"]["items"][0]["price_paid"])
print("current email:", view["current_customer"]["email"])
print("document reads:", dict(read_cost))
The lab exposes the stale duplicated email, then uses the order as the source of historical purchase facts while resolving mutable customer state separately. The extra document read is a real tradeoff: reference-heavy models can create request fan-out or join-like application work.
3. Wrong approach: mechanically replace every foreign key with embedding
Embedding an unbounded review list inside a product, all orders inside a customer, or a frequently changing supplier record into thousands of products creates growth and write fan-out. The opposite extreme—referencing every tiny owned value—can turn one request into many network round trips. Model the aggregate around invariants and access paths, then document what duplication is a snapshot versus what duplication must converge.
For every duplicated mutable fact, intentionally drop one projection/update in a test. If you cannot detect the divergence and repair it from an authoritative source, the denormalized design has no credible recovery story.
4. Production judgment
Embedding can reduce latency and expand the single-document atomic boundary, but it spends document-size budget and can increase write amplification. References preserve independent lifecycle and reduce duplication but may add network hops, transaction scope, and availability dependencies. In a sharded topology, also ask whether related referenced documents are co-located or whether a request becomes scatter-gather.
Check your understanding
- Why embed order-line price-at-purchase?
- Why reference mutable customer profile state?
- When does a one-to-many relationship argue against embedding?
- What must accompany intentional denormalization?
Review the answers
1. It is an immutable historical snapshot owned by the order and should not follow future catalog price changes.
2. It has independent lifecycle and copying it into many orders would create mutable duplication and repair work.
3. When the many side is unbounded, independently queried, independently updated, or too large for a stable aggregate.
4. An authoritative source plus a synchronization, divergence-detection, and repair plan.
Authoritative references
- MongoDB — Embedded data models — concrete current implementation guidance for aggregate locality and the 16 MiB BSON document limit.
- MongoDB — References — cases where independent lifecycle, many-to-many relationships, or frequent independent access favor references.
- MongoDB — Multikey indexes — array-index behavior and index-entry implications.
- MongoDB — Schema validation — evidence that flexible documents still benefit from enforced structural rules.
- MongoDB — Avoid unbounded arrays — implementation example of bounding growth with subsetting/references.
- MongoDB 8.3 release notes — dated implementation snapshot; 8.3.8 is the latest released patch as of this chapter review.