Produce and defend a production-oriented AtlasMart MongoDB schema in which every embedding, reference, duplicated field, bound, and transaction boundary follows from a stated workload and invariant.
Design a Schema for a Real Application and Defend Every Duplication Decision
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning outcomes
The chapter ends with a design review. AtlasMart needs catalog, customers, inventory, reviews, orders, and activity data. The goal is not to create the fewest collections or the most embedded schema. The goal is to make every boundary defensible: what is canonical, what is copied, what is bounded, what changes atomically, what can grow independently, and what will be reconciled if derived state diverges.
Turn workload evidence into a coherent multi-collection AtlasMart schema with explicit aggregate owners.
Defend every duplicated field as either historical snapshot or bounded derived state with a freshness contract.
Keep high-cardinality review/activity history outside hot product/account documents while preserving fast common reads.
Keep correctness-critical inventory counters in one atomic document and identify where cross-document workflows remain.
Produce a schema decision register that records bounds, ownership, indexes, retention, failure behavior, and migration rollback.
Mandatory examples use MongoDB Community Server
8.3.8 in the pinned
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim
image, a disposable standalone mongod published
only on 127.0.0.1:27046, and mongosh 2.10.0.
PyMongo examples target the 4.17 line where driver behavior
matters. Authentication and Transport Layer Security (TLS) are
intentionally disabled only inside this isolated loopback lab;
do not copy that posture to a shared or remotely reachable
server. The lab uses standalone default read/write concern
semantics, no replica-set or sharding guarantee is implied,
and cleanup removes atlasmart-mongo-ch06-l5. This
final modeling lab remains a standalone because it validates
document shapes and single-document invariants. It does not
claim atomic checkout across inventory, order, and payment
documents; that workflow is deliberately deferred to the
transactions chapter and would require a replica set or
sharded cluster for a MongoDB multi-document transaction.
1. Final schema: separate aggregates where ownership really differs
| Collection | Aggregate owner | Embedded/duplicated state | Reference/high-cardinality state |
|---|---|---|---|
| customers_final | Customer profile | Bounded settings subdocument | Orders/reviews reference customerId |
| products_final | Current catalog item | 3 recent review projections + computed reviewStats | Full reviews live in reviews_final |
| inventory_final | Sellable stock invariant | onHand + reserved counters together | References productId |
| orders_final | Immutable purchase record | Customer/name/price snapshots + bounded lines | customerId/productId retained for lineage |
| reviews_final | Review lifecycle | Review body fields | References product/customer |
| activity_final | Activity bucket | Bounded events array | Multiple buckets per account over time |
Notice that inventory is not embedded inside the product even though product pages read both. Inventory has a high-contention correctness invariant and may scale/partition differently from catalog text. A second read—or an application service boundary—can be a better trade than turning a catalog document into a hot write target.
2. Defend each duplication decision
AtlasMart duplicates four kinds of values in this design. The
order’s customer display name and line-item product name/price
are historical snapshots; changing them later
would corrupt the purchase record. The product’s recent review
subset and reviewStats are
derived projections; they may be rebuilt from
canonical review documents and therefore require
freshness/reconciliation rules.
from dataclasses import dataclass@dataclass(frozen=True)class Decision: field: str owner: str copy_kind: str bound: str freshness: strregistry = [ Decision("orders.customerSnapshot.displayName", "orders", "historical snapshot", "1 per order", "immutable after purchase"), Decision("orders.lines[].nameAtOrder", "orders", "historical snapshot", "business-limited line count", "immutable after purchase"), Decision("products.recentReviews", "reviews + product projection", "bounded derived subset", "3", "eventual/reconciled"), Decision("products.reviewStats", "reviews + product projection", "computed derived value", "1 summary", "explicit freshness SLA"),]for row in registry: print(row)
A decision register prevents future engineers from “fixing duplication” by deleting intentional snapshots or from assuming a derived copy is always current. It should live beside architecture documentation and migrations, and it should change when the workload changes.
3. Reject the giant-customer-document anti-pattern
A tempting design puts profile, all orders, every review, wishlist history, support tickets, and activity events inside one customer document. It appears to optimize “load customer” but creates an aggregate whose cardinality and retention are the union of unrelated lifecycles. One heavy buyer becomes a hot and continuously growing document; deleting one review or archiving old activity rewrites customer-owned state; indexes on multiple large arrays multiply storage and write cost.
The repair is not “normalize everything.” Keep bounded co-owned settings with the customer, immutable bounded line items with the order, full review history with reviews, and long activity series in bounded buckets. Preserve strategically duplicated snapshots/projections when they remove measured read fan-out.
If the team cannot state the maximum reasonable cardinality, retention window, or spill strategy for an embedded array, the design is incomplete. Outliers must be part of capacity testing, not discovered after one celebrity product or power user creates the first pathological document.
4. Build and inspect the defended AtlasMart schema
docker rm -f atlasmart-mongo-ch06-l5 2>/dev/null || truedocker run --name atlasmart-mongo-ch06-l5 -p 127.0.0.1:27046:27017 -d mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27046/atlasmart?directConnection=true" --quiet --eval 'printjson({version:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1}))'
db.customers_final.drop();db.products_final.drop();db.reviews_final.drop();db.orders_final.drop();db.inventory_final.drop();db.activity_final.drop();db.customers_final.createIndex({tenantId:1,email:1},{unique:true});db.products_final.createIndex({tenantId:1,sku:1},{unique:true});db.reviews_final.createIndex({tenantId:1,productId:1,createdAt:-1});db.orders_final.createIndex({tenantId:1,customerId:1,createdAt:-1});db.activity_final.createIndex({tenantId:1,accountId:1,bucketStart:1},{unique:true});db.customers_final.insertOne({_id:"cust-1",tenantId:"t1",email:"mina@example.test",displayName:"Mina",settings:{locale:"en",marketing:false}});db.products_final.insertOne({_id:"prod-1",tenantId:"t1",sku:"CAM-1",name:"Atlas Camera",priceCents:12900,recentReviews:[],reviewStats:{count:0,sum:0,average:null}});db.inventory_final.insertOne({_id:"inv-prod-1",tenantId:"t1",productId:"prod-1",onHand:10,reserved:0});db.reviews_final.insertMany([ {_id:"rv-1",tenantId:"t1",productId:"prod-1",customerId:"cust-1",rating:5,summary:"Great",createdAt:ISODate("2026-09-01T10:00:00Z")}, {_id:"rv-2",tenantId:"t1",productId:"prod-1",customerId:"cust-1",rating:4,summary:"Good",createdAt:ISODate("2026-09-02T10:00:00Z")}]);db.products_final.updateOne({_id:"prod-1"},{$set:{recentReviews:[ {reviewId:"rv-2",rating:4,summary:"Good",createdAt:ISODate("2026-09-02T10:00:00Z")}, {reviewId:"rv-1",rating:5,summary:"Great",createdAt:ISODate("2026-09-01T10:00:00Z")}],reviewStats:{count:2,sum:9,average:4.5}}});db.orders_final.insertOne({_id:"ord-1",tenantId:"t1",customerId:"cust-1",customerSnapshot:{displayName:"Mina"},createdAt:ISODate("2026-09-02T11:00:00Z"),lines:[{productId:"prod-1",sku:"CAM-1",nameAtOrder:"Atlas Camera",unitPriceCents:12900,qty:1}],totalCents:12900,status:"paid"});db.activity_final.insertOne({_id:"t1:acct-1:2026-09-02T11",tenantId:"t1",accountId:"acct-1",bucketStart:ISODate("2026-09-02T11:00:00Z"),count:2,events:[{seq:1,type:"view"},{seq:2,type:"cart"}]});const sizeReport=["customers_final","products_final","orders_final","inventory_final","activity_final"].map(name=>({ collection:name, docs:db.getCollection(name).countDocuments({}), maxBytes:db.getCollection(name).aggregate([{$project:{b:{$bsonSize:"$$ROOT"}}},{$group:{_id:null,max:{$max:"$b"}}}]).next()?.max||0}));printjson({sizeReport});const product=db.products_final.findOne({_id:"prod-1"});const recent=product.recentReviews;const fullHistory=db.reviews_final.find({tenantId:"t1",productId:"prod-1"}).sort({createdAt:-1}).toArray();const order=db.orders_final.findOne({_id:"ord-1"});printjson({commonRead:{productName:product.name,recentReviewCount:recent.length,queryCount:1}});printjson({historyRead:{reviewCount:fullHistory.length,queryCount:1}});printjson({orderSnapshot:{customerSnapshot:order.customerSnapshot,lineSnapshot:order.lines[0]}});// Single-document inventory invariant check.const inv=db.inventory_final.updateOne({_id:"inv-prod-1",onHand:{$gte:1}},{$inc:{onHand:-1,reserved:1}});printjson({inventoryReservation:{matched:inv.matchedCount,modified:inv.modifiedCount,state:db.inventory_final.findOne({_id:"inv-prod-1"})}});
What the evidence proves
The size report confirms the sample aggregate documents are
small and independently bounded; it does not prove future
production cardinalities are safe. The product common read
returns the recent subset in one query, while full review
history remains independently readable. The paid order carries
purchase-time snapshots plus canonical identifiers. The
inventory reservation changes onHand and
reserved together in one document-level conditional
write.
What it does not prove
This lab is not a latency benchmark, does not test shard placement, does not make the derived review summary strongly consistent with review history, and does not make checkout across order/payment/inventory atomic. Those are separate topology, consistency, transaction, indexing, and performance decisions later in the course.
Verification checklist
- Every collection has a named aggregate owner and lifecycle.
- Every embedded array is bounded by a business/application rule or bucket size.
- Every duplicated field is classified as snapshot or derived state.
- Indexes support demonstrated identity/access patterns without pretending index design is complete before Chapters 10–12.
- Server-measured BSON sizes are recorded for representative aggregate documents.
- The inventory invariant fits one document-level conditional update.
- Cross-document checkout atomicity is explicitly not claimed.
Check your understanding
- Why keep inventory separate from product catalog?
- Which duplicated values should not be updated when canonical data changes?
- Which duplicated values require reconciliation?
- What is the strongest evidence that an array is safely bounded?
- When should this schema be revisited?
Review the answers
Its high-write correctness invariant and scaling/ownership characteristics differ from mostly read-oriented catalog state; embedding could turn the catalog document into a hot aggregate.
Historical snapshots such as purchase-time names and prices, because their purpose is to preserve what the order meant at purchase.
Derived projections such as recentReviews and reviewStats when their contract allows asynchronous or fallible updates.
A documented business/access bound plus production-like cardinality tests and monitoring—not merely current sample size.
When access patterns, cardinalities, retention, latency SLOs, topology, tenant distribution, or invariants change enough that the current locality/growth tradeoff no longer holds.
docker rm -f atlasmart-mongo-ch06-l5
5. Production judgment and migration/rollback plan
Good MongoDB modeling reduces coordination on the common path without hiding coordination that truly exists. Embedding improves locality and single-document atomicity, but it consumes document growth and couples lifecycle. References preserve independent ownership and reduce duplication, but can add network round trips, join/batching logic, and cross-document invariants. Buckets/subsets/computed fields can keep hot documents bounded, but create maintenance and reconciliation work.
Before production, test distributions rather than averages:
largest orders, viral products, highest-activity tenants,
long-retention customers, and bursty concurrent inventory
reservations. Monitor document size/cardinality, update latency,
write conflicts, query fan-out, derived-data lag, index growth,
and tenant skew. Security boundaries remain independent of
schema locality: a tenantId field does not by
itself authorize access.
Migration should be reversible. For an embedding-to-reference change, backfill canonical child documents, dual-read or shadow-read, compare results, cut over writers, then remove old embedded copies only after rollback windows expire. For reference-to-embedded projections, populate the derived copy, measure freshness and fan-out reduction, and retain the canonical source until the projection is proven rebuildable.
Chapter 07 adds the next missing layer: schema validation, evolution, versioning, defaults, and data-quality rules so these modeled contracts remain enforceable as AtlasMart changes.
Authoritative references
- MongoDB release notes — Current stable server series and patch history.
- MongoDB 8.3 release notes — 8.3.8 is the latest released 8.3 patch at review time; 8.3.9 is upcoming.
- mongosh release notes — mongosh 2.10.0 was released August 13, 2026.
- PyMongo release notes — Current PyMongo 4.17 line and driver changes.
- MongoDB schema design process — Official workload-first process: identify workload, map relationships, apply patterns, and create supporting indexes.
- Data modeling best practices — Official embedding-versus-referencing decision guidance.
- Designing your schema — Workload → relationships → patterns → indexes process.
- Handle duplicate data — Consistency and read-performance tradeoffs for duplicated fields.
- Avoid unbounded arrays — Bound large arrays through subsets/references.
- Group data patterns — Bucket/outlier approaches for long or skewed series.
- Computed values — Read-versus-write/freshness tradeoff for stored derived data.
- Atomicity and transactions — Single-document boundary and transaction considerations.