Chapter 11 · Document Databases and Aggregate-Oriented Modeling
Model One-to-One, One-to-Many, Many-to-Many, and Event Data with Workload Evidence
Redesign one-to-one, one-to-many, many-to-many, and event data from workload evidence, boundedness, ownership, and a concrete consistency/repair contract.
Learning outcomes
Relationship names—one-to-one, one-to-many, many-to-many—describe cardinality but do not choose a document model. AtlasMart will now design each relationship from read paths, write paths, invariants, boundedness, ownership, and recovery evidence, then carry that discipline into the next chapter’s query-first wide-column modeling.
Model one-to-one data from co-read/co-update behavior.
Bound one-to-many collections and preserve independent child access where needed.
Represent many-to-many membership without uncontrolled duplication.
Bucket event/time-series-like data and document repair for duplicated summaries.
1. Relationship cardinality is only one input
| AtlasMart fact | Chosen shape | Why |
|---|---|---|
| shipping address at purchase | embedded immutable snapshot in order | co-read with order; historical value must not follow later profile edits |
| product reviews | separate review documents + bounded product summary | unbounded many side; reviews queried/moderated independently |
| campaign ↔ product | membership/edge documents or bounded references | many-to-many lifecycle and independent campaign operations |
| customer activity events | time/count buckets | append-heavy, retention-aware, bounded growth |
The same cardinality can produce a different answer under another workload. A one-to-one profile image containing large binary content may belong outside the core document. A tiny one-to-many list of fixed shipping dimensions may embed naturally. The model must state what is authoritative and what is derived.
2. A repairable model makes duplicated facts explicit
from collections import defaultdict
from datetime import date
model = {
"order": {
"id":"o-7",
"shipping_address_snapshot":{"city":"Baku","street":"Nizami 1"},
"customer_id":"c-9",
"items":[{"sku":"A","qty":1,"price_paid":30}],
},
"product": {"sku":"A", "review_summary":{"count":2,"avg":4.5}},
"reviews": [
{"id":"r1","sku":"A","customer_id":"c1","stars":5},
{"id":"r2","sku":"A","customer_id":"c2","stars":4},
],
"campaign_membership": [
{"campaign_id":"summer","sku":"A"},
{"campaign_id":"clearance","sku":"A"},
]
}
# Event data is bounded into daily buckets instead of one unbounded customer document.
events = [
{"customer_id":"c-9","day":"2026-08-29","kind":"view","sku":"A"},
{"customer_id":"c-9","day":"2026-08-29","kind":"cart","sku":"A"},
{"customer_id":"c-9","day":"2026-08-30","kind":"purchase","sku":"A"},
]
buckets=defaultdict(list)
for e in events:
buckets[(e["customer_id"],e["day"])].append(e)
print("event buckets:", {str(k):len(v) for k,v in buckets.items()})
# Detect/repair duplicated review summary after a missed projection update.
model["reviews"].append({"id":"r3","sku":"A","customer_id":"c3","stars":1})
actual_count=len(model["reviews"])
summary_count=model["product"]["review_summary"]["count"]
print("summary drift:", summary_count, "vs", actual_count)
model["product"]["review_summary"]={
"count":actual_count,
"avg":sum(r["stars"] for r in model["reviews"])/actual_count,
}
print("repaired summary:", model["product"]["review_summary"])
The product’s review summary is deliberately stale after a new review arrives. Because review documents are authoritative, AtlasMart can recompute the summary. That is a recoverable denormalization. If both copies were allowed to mutate independently, reconciliation would not know which truth to preserve.
3. Wrong approach: put the customer’s entire world in one document
Embedding every order, review, cart event, support message, session, and recommendation under one customer record creates an unbounded hot aggregate. One popular customer or automation account can dominate a partition; unrelated writers contend; retention becomes all-or-nothing; and partial reads still move or parse a giant value. The repair is to keep a small customer aggregate and move independently growing histories to bounded child documents/buckets with explicit keys.
For every duplicated value, write down: source of truth, propagation mechanism, expected lag, idempotency key/version, mismatch detector, repair command/job, and rollback behavior. “Eventually consistent” without these details is not an operational design.
4. Chapter decision notebook
Before selecting a document product or schema, record the access-pattern evidence: dominant reads/writes, maximum cardinality, expected document size distribution, array lengths, update frequency, partition key, indexes, atomicity boundary, staleness tolerance, projection lag SLO, backup/restore implications, tenant isolation, and migration path. Revisit the notebook under realistic skew—not just average traffic.
Check your understanding
- Why can the same one-to-many cardinality lead to embedding in one case and referencing in another?
- Why is a cached/denormalized review summary safe only with an authoritative source?
- What is the main risk of one giant customer document?
- What bridges this chapter to wide-column modeling?
Review the answers
1. Because ownership, boundedness, independent access, update frequency, and atomicity needs differ by workload.
2. The source lets the system detect drift and deterministically rebuild the derived value.
3. Unbounded growth, hot-key/write contention, poor retention granularity, oversized transfers, and a large failure blast radius.
4. The discipline of starting from concrete queries, bounded partitions/aggregates, ordering, and explicit duplication rather than normalizing by habit.
5. Production judgment and bridge to Chapter 12
Document databases are excellent when aggregate locality matches request locality and when bounded records support the required atomic updates. They are weaker when the workload demands unbounded fan-out, frequent global joins, or cross-aggregate invariants that dominate the design. The next chapter takes the same workload-first principle further: wide-column databases make the partition key and clustering order explicit parts of query design.
Authoritative references
- MongoDB — Embedded data models — concrete current implementation guidance for aggregate locality and the 16 MiB BSON document limit.
- MongoDB — References — cases where independent lifecycle, many-to-many relationships, or frequent independent access favor references.
- MongoDB — Multikey indexes — array-index behavior and index-entry implications.
- MongoDB — Schema validation — evidence that flexible documents still benefit from enforced structural rules.
- MongoDB — Avoid unbounded arrays — implementation example of bounding growth with subsetting/references.
- MongoDB 8.3 release notes — dated implementation snapshot; 8.3.8 is the latest released patch as of this chapter review.