Design MongoDB documents from AtlasMart request patterns and invariants, measuring locality, cardinality, growth, transaction boundaries, and retention before choosing a shape.
Start from Workload: Queries, Updates, Atomicity, Cardinality, Growth, and Retention
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning outcomes
AtlasMart is about to add product detail pages, inventory reservation, customer reviews, and a seven-year order-history policy. A schema copied mechanically from a relational diagram would split every entity into a separate collection; an equally mechanical “MongoDB means embed everything” design would place all reviews, inventory movements, and history into one ever-growing product document. Both ignore the actual work. MongoDB schema design begins with workload—the operations the application performs—and then asks what data must be read, changed, retained, or archived together.
Translate AtlasMart requests and invariants into explicit read, write, atomicity, cardinality, growth, and retention requirements.
Compare embedded, referenced, and hybrid candidate shapes using query count and server-measured BSON document size.
Explain why a single MongoDB document is an atomic write boundary and when that boundary can eliminate cross-document coordination.
Recognize cardinality and retention as growth mechanisms rather than static ER-diagram labels.
Reject both automatic normalization and automatic embedding in favor of measured workload evidence.
Mandatory examples use MongoDB Community Server
8.3.8 in the pinned
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim
image, a disposable standalone mongod published
only on 127.0.0.1:27042, and mongosh 2.10.0.
PyMongo examples target the 4.17 line where driver behavior
matters. Authentication and Transport Layer Security (TLS) are
intentionally disabled only inside this isolated loopback lab;
do not copy that posture to a shared or remotely reachable
server. The lab uses standalone default read/write concern
semantics, no replica-set or sharding guarantee is implied,
and cleanup removes atlasmart-mongo-ch06-l1. This
lesson intentionally uses a standalone deployment because the
key observation is single-document atomicity. Multi-document
transactions require a replica set or sharded cluster and are
discussed only as a design consequence here.
1. Write the workload before the schema
A query is a read request shape, such as “load product card plus two recent reviews.” An update is a mutation shape, such as “reserve one unit if stock remains.” Atomicity means whether a group of changes becomes visible as one indivisible write. Cardinality describes how many related items can exist, while growth asks how that number evolves over time. Retention defines how long data remains operationally available before archival or deletion.
| AtlasMart operation | Frequency / bound | Correctness or latency need | Modeling pressure |
|---|---|---|---|
| Product detail | Very frequent; reads product + 2 recent reviews | One low-latency response | Embed a small recent subset or pay another query |
| Reserve inventory | High write rate | onHand/reserved invariant must change together | Keep inventory counters in one atomic document when ownership permits |
| Review history | Unbounded over product lifetime | Independent moderation/deletion lifecycle | Reference full review history; do not grow one array forever |
| Paid order display | Read-mostly after purchase | Must preserve purchase-time names/prices | Historical snapshots are valid duplication |
| Seven-year retention | Time-based | Delete/archive old reviews independently | Independent collections simplify lifecycle operations |
The table is the beginning of the schema. Relationship notation alone cannot tell you whether a customer name should be copied into an order or whether 500,000 reviews should live inside one product.
2. Compare candidate shapes, not slogans
Candidate A embeds every review in the product. It maximizes read locality but allows the document to grow with review history. Candidate B references every review. The product remains bounded and reviews have an independent lifecycle, but the product-detail path needs another request (or a later aggregation join). Candidate C keeps a bounded subset—for example the two most recent reviews—inside the product and stores the complete review history separately. Candidate C deliberately duplicates a small amount of data to match the common read while bounding growth.
MongoDB limits a BSON document to 16 mebibytes (MiB), but a good model should not aim for 15.99 MiB. Larger documents consume more network bandwidth, cache space, serialization work, and update I/O even before the hard limit is reached. The right bound comes from the workload, not from the maximum.
That confuses business identity with an aggregate boundary. Review history can have much higher cardinality and a different retention/moderation lifecycle than the product. The repair is to keep strongly co-owned bounded state together and move independently growing histories behind references or bounded patterns.
3. Atomicity is a modeling input
MongoDB guarantees atomicity for a write to one document, even
when the update changes several fields or array elements inside
that document. That means an inventory document can update
onHand and reserved together with one
conditional update. If the same invariant is split across
multiple documents, a single ordinary write no longer covers the
whole invariant. MongoDB supports multi-document transactions on
replica sets and sharded clusters, but those introduce
additional coordination and should not be treated as permission
to ignore aggregate design.
Atomicity does not mean “nobody can ever observe stale data,” nor does it define replication durability. Those depend on topology, read concern, write concern, and read preference. This chapter uses atomicity only to decide what state should change as one document-level unit.
4. Run the workload-and-growth lab
docker rm -f atlasmart-mongo-ch06-l1 2>/dev/null || truedocker run --name atlasmart-mongo-ch06-l1 -p 127.0.0.1:27042:27017 -d mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27042/atlasmart?directConnection=true" --quiet --eval 'printjson({version:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1}))'
db.model_products.drop();db.model_reviews.drop();const recent = [ {reviewId:"r-49",rating:5,summary:"Sharp image",createdAt:ISODate("2026-09-01T10:00:00Z")}, {reviewId:"r-50",rating:4,summary:"Good battery",createdAt:ISODate("2026-09-02T09:00:00Z")}];const allReviews = Array.from({length:50},(_,i)=>({ reviewId:`r-${String(i+1).padStart(2,"0")}`, rating:(i%5)+1, summary:`deterministic review ${i+1}`, createdAt:new Date(Date.UTC(2026,7,1+i%30,12,0,0))}));db.model_products.insertMany([ {_id:"sku-camera-1",name:"AtlasCam",priceCents:12900,inventory:{onHand:12,reserved:2},recentReviews:recent,reviewCount:50}, {_id:"sku-camera-1-all",name:"AtlasCam",priceCents:12900,inventory:{onHand:12,reserved:2},reviews:allReviews,reviewCount:50}]);db.model_reviews.insertMany(allReviews.map(r=>({...r,productId:"sku-camera-1"})));print("candidate document sizes (server-measured BSON bytes)");printjson(db.model_products.aggregate([ {$project:{_id:1,bytes:{$bsonSize:"$$ROOT"},arrayCount:{$size:{$ifNull:["$reviews","$recentReviews"]}}}}]).toArray());print("request-shape evidence");printjson({ productDetailEmbeddedQueries:1, productDetailReferencedQueries:2, recentEmbeddedCount:db.model_products.findOne({_id:"sku-camera-1"}).recentReviews.length, referencedReviewCount:db.model_reviews.countDocuments({productId:"sku-camera-1"})});const before=db.model_products.findOne({_id:"sku-camera-1"});const wr=db.model_products.updateOne( {_id:"sku-camera-1","inventory.onHand":{$gte:1}}, {$inc:{"inventory.onHand":-1,"inventory.reserved":1}});const after=db.model_products.findOne({_id:"sku-camera-1"});printjson({singleDocumentAtomicWrite:{matched:wr.matchedCount,modified:wr.modifiedCount,before:before.inventory,after:after.inventory}});print("retention boundary: removing old review does not rewrite product");const del=db.model_reviews.deleteMany({productId:"sku-camera-1",createdAt:{$lt:ISODate("2026-08-15T00:00:00Z")}});printjson({deletedHistoricalReviews:del.deletedCount,productStillExists:!!db.model_products.findOne({_id:"sku-camera-1"})});
Expected evidence
The fully embedded 50-review product is materially larger than the product with two recent embedded reviews. The hybrid product-detail path can satisfy the common “product + recent reviews” request with one document read while the full history remains independently countable and deletable. The conditional inventory update reports one matched/modified document and changes both counters together.
The exact BSON byte counts depend on field values and encoding.
That is why the lab asks MongoDB itself through
$bsonSize rather than estimating from JSON
character count.
Verification checklist
-
The lab runs only against the disposable
atlasmart-mongo-ch06-l1container. - Two candidate products contain identical core fields but different review embedding strategies.
-
$bsonSizereports server-measured BSON bytes for each stored document. - The product-detail query-count tradeoff is stated explicitly rather than assumed.
- The inventory update changes its related counters in one document-level write.
- Deleting old referenced reviews does not require rewriting the product document.
Check your understanding
- Why is cardinality not enough to decide embed versus reference?
- What does the 16 MiB limit tell you?
- Why can inventory counters belong in one document?
- Why might full review history be referenced even when recent reviews are embedded?
- When would a multi-document transaction enter the design?
Review the answers
Because update frequency, read locality, lifecycle ownership, atomicity needs, growth rate, and retention determine whether that cardinality remains a safe aggregate.
It is a hard BSON document ceiling, not a performance target. You should choose a lower workload-specific bound based on memory, bandwidth, update cost, and growth.
If they participate in the same invariant and share ownership, one document lets one atomic write change them together.
The full history grows independently and has its own moderation/retention lifecycle, while a small recent subset can optimize the dominant product-detail read.
When a correctness invariant truly spans multiple documents/collections and cannot be remodeled into one safe atomic aggregate; on MongoDB it requires a replica set or sharded cluster.
docker rm -f atlasmart-mongo-ch06-l1
The next lesson makes the central choice explicit: what exactly you gain and pay for when data is embedded versus referenced.
5. Production judgment: optimize the dominant path without hiding growth
Embed when data is strongly co-owned, normally read or changed together, and remains bounded under realistic high-percentile cardinality. Reference when the child grows independently, needs a different retention/security lifecycle, or mutable duplication would create large update fan-out. The guarantee gained from embedding is one-document atomic mutation; it does not by itself guarantee majority durability, linearizable reads, cross-document consistency, or low tail latency under a hot-document workload.
Measure p95/p99 document size, array cardinality, bytes returned, update conflicts, query count, index-key growth, and derived-data freshness. On replica sets or sharded clusters, re-evaluate placement and transaction needs; on sharded deployments, a poor shard key can still turn references into scatter/gather. Search/vector indexes and encryption can add their own storage/CPU constraints later. Tenant identifiers must be enforced by authorization, not assumed to be a security boundary because they are embedded in a document.
Test outliers, concurrent updates, failed derived-copy refreshes, and archival/deletion paths. Community tooling is sufficient for these modeling labs; Atlas/Enterprise features are not required. Migration should be reversible: backfill new shapes, compare shadow reads, cut over writers, retain old representation through a rollback window, then remove it only after correctness and latency evidence is stable. Lesson 2 now focuses on the cost ledger of embedding versus referencing.
Authoritative references
- MongoDB release notes — Current stable server series and patch history.
- MongoDB 8.3 release notes — 8.3.8 is the latest released 8.3 patch at review time; 8.3.9 is upcoming.
- mongosh release notes — mongosh 2.10.0 was released August 13, 2026.
- PyMongo release notes — Current PyMongo 4.17 line and driver changes.
- MongoDB schema design process — Official workload-first process: identify workload, map relationships, apply patterns, and create supporting indexes.
- Data modeling best practices — Official embedding-versus-referencing decision guidance.
- Embedded data — Benefits, use cases, and document-size boundary for embedded models.
- Referenced data — Use cases for references, including complex relationships and independently changing data.
- MongoDB limits and thresholds — Official 16 MiB BSON document and nesting limits.
- Atomicity and transactions — Single-document atomicity and why transactions do not replace effective schema design.