Keep MongoDB aggregates bounded by measuring growth and applying subset, bucket, computed, and outlier-aware patterns without hiding consistency or maintenance costs.
Document Growth, Large Arrays, Bucket/Subset/Computed Patterns, and Bounded Aggregates
Batch heterogeneous writes safely, interpret partial success, compare ordered and unordered execution, and use modern cross-namespace bulk APIs without assuming all-or-nothing behavior.
Learning outcomes
AtlasMart’s first product document looked healthy when it had three reviews and ten activity events. A year later, one viral product has hundreds of thousands of reviews and one merchant account produces millions of activity records. “It still fits in MongoDB” is not an adequate growth strategy. A bounded aggregate needs an explicit maximum shape, and data that exceeds that bound needs a different storage pattern.
Explain why document growth and large arrays affect more than the 16 MiB hard limit.
Use the subset pattern to keep a small hot working set embedded while moving full history to references.
Use the bucket pattern to partition a long homogeneous series into bounded documents aligned with access units.
Use computed fields only with an explicit freshness and reconciliation contract.
Define a workload-derived aggregate bound and recognize outliers that should not force every normal document to grow.
Mandatory examples use MongoDB Community Server
8.3.8 in the pinned
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim
image, a disposable standalone mongod published
only on 127.0.0.1:27045, and mongosh 2.10.0.
PyMongo examples target the 4.17 line where driver behavior
matters. Authentication and Transport Layer Security (TLS) are
intentionally disabled only inside this isolated loopback lab;
do not copy that posture to a shared or remotely reachable
server. The lab uses standalone default read/write concern
semantics, no replica-set or sharding guarantee is implied,
and cleanup removes atlasmart-mongo-ch06-l4. The
lesson stores only small deterministic fixtures. It does not
intentionally create a near-16-MiB document; the hard size
ceiling is verified from official documentation while the lab
measures realistic growth with $bsonSize.
1. The hard limit is not the first growth problem you hit
MongoDB’s maximum BSON document size is 16 MiB. Long before that ceiling, a growing document can enlarge every network response, occupy more cache, increase serialization/deserialization cost, create a larger multikey index footprint when arrays are indexed, and concentrate writes on one hot document. An unbounded array is therefore an anti-pattern even if today’s fixture contains only 100 elements.
A bounded aggregate declares a maximum cardinality or size justified by the application. “Last three reviews,” “one page of 20 activities,” or “100 line items maximum per order” are business/model bounds. “Whatever fits under 16 MiB” is not.
2. Subset, bucket, and computed patterns solve different pressure
| Pattern | Problem | AtlasMart use | Hidden cost |
|---|---|---|---|
| Subset | Common read needs only small hot portion of large related set | 3 recent reviews embedded; all reviews referenced | Duplicate subset must be refreshed/reconciled |
| Bucket | Long homogeneous series needs bounded grouped documents | 5 activity events per bucket/page | Bucket assignment and concurrency need deterministic logic |
| Computed | Expensive aggregate is read much more often than source changes | review count/sum/average stored on product | Can become stale if update path fails or is asynchronous |
| Outlier-aware split | A few entities grow far beyond normal population | viral product review history moved to separate documents | Read path must recognize outlier storage |
These patterns can be combined. A product may embed three recent reviews (subset), store all reviews separately, maintain a computed rating summary, and bucket high-volume activity logs. Combining patterns is justified only when each one solves a measured access/growth problem.
3. Computed fields trade read work for write/reconciliation work
A computed value such as reviewStats.average avoids
scanning all review documents on every product-card read. If
every review insert updates the computed summary synchronously,
writes become more complex. If a background job updates it
periodically, reads may be stale. Neither is inherently wrong;
the contract must state the allowed freshness window.
The lab deliberately inserts one review without refreshing
reviewStats. That produces visible drift. A
reconciliation aggregation recomputes authoritative count, sum,
and average from review history and repairs the cached value.
This is a safer teaching model than pretending duplicated
computed state cannot diverge.
One busy account becomes a giant hot document, while normal accounts remain small. Repair by defining a bucket size/window, keeping each bucket bounded, and treating the account identifier plus bucket key as the routing/ownership unit.
4. Run the bounded-aggregate lab
docker rm -f atlasmart-mongo-ch06-l4 2>/dev/null || truedocker run --name atlasmart-mongo-ch06-l4 -p 127.0.0.1:27045:27017 -d mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimmongosh "mongodb://127.0.0.1:27045/atlasmart?directConnection=true" --quiet --eval 'printjson({version:db.version(), hello:db.hello().isWritablePrimary}); printjson(db.getSiblingDB("admin").runCommand({getParameter:1,featureCompatibilityVersion:1}))'
db.product_bounded.drop();db.review_history.drop();db.activity_buckets.drop();const recentReviews = [ {reviewId:"r-18",rating:4,createdAt:ISODate("2026-08-31T12:00:00Z")}, {reviewId:"r-19",rating:5,createdAt:ISODate("2026-09-01T12:00:00Z")}, {reviewId:"r-20",rating:3,createdAt:ISODate("2026-09-02T12:00:00Z")}];db.product_bounded.insertOne({_id:"p-1",name:"Atlas Camera",recentReviews,reviewStats:{count:20,sum:60,average:3.0}});db.review_history.insertMany(Array.from({length:20},(_,i)=>({_id:`r-${i+1}`,productId:"p-1",rating:(i%5)+1,createdAt:new Date(Date.UTC(2026,7,14+i,12,0,0))})));// Bound the hot subset to the three most recent reviews.const add=db.product_bounded.updateOne({_id:"p-1"},{$push:{recentReviews:{$each:[{reviewId:"r-21",rating:5,createdAt:ISODate("2026-09-02T13:00:00Z")}],$sort:{createdAt:-1},$slice:3}}});printjson({subsetUpdate:{matched:add.matchedCount,modified:add.modifiedCount,recent:db.product_bounded.findOne({_id:"p-1"}).recentReviews}});// Bucket 12 activity events into deterministic groups of five.const events=Array.from({length:12},(_,i)=>({seq:i+1,type:(i%2?"view":"cart"),at:new Date(Date.UTC(2026,8,2,10,0,i))}));for(let start=0;start<events.length;start+=5){ const slice=events.slice(start,start+5); db.activity_buckets.insertOne({_id:`acct-1-b${Math.floor(start/5)+1}`,accountId:"acct-1",count:slice.length,events:slice});}printjson({bucketCounts:db.activity_buckets.find({accountId:"acct-1"}).sort({_id:1}).map(x=>({id:x._id,count:x.count,bytes:null})).toArray()});printjson({bucketSizes:db.activity_buckets.aggregate([{$project:{_id:1,count:1,bytes:{$bsonSize:"$$ROOT"}}},{$sort:{_id:1}}]).toArray()});// Deliberately create computed-field drift: insert r-21 without updating reviewStats.db.review_history.insertOne({_id:"r-21",productId:"p-1",rating:5,createdAt:ISODate("2026-09-02T13:00:00Z")});const actual=db.review_history.aggregate([{$match:{productId:"p-1"}},{$group:{_id:"$productId",count:{$sum:1},sum:{$sum:"$rating"},avg:{$avg:"$rating"}}}]).next();const cached=db.product_bounded.findOne({_id:"p-1"}).reviewStats;printjson({computedDrift:{cached,actual}});db.product_bounded.updateOne({_id:"p-1"},{$set:{reviewStats:{count:actual.count,sum:actual.sum,average:actual.avg}}});printjson({reconciled:db.product_bounded.findOne({_id:"p-1"}).reviewStats});
Expected evidence
After adding review r-21, the product still
contains exactly three recent reviews because
$push uses $sort plus
$slice. Twelve activity events become three bucket
documents with counts 5, 5, and 2, and MongoDB reports the BSON
bytes of each bucket. Inserting the review into history without
updating reviewStats creates an intentional
discrepancy; the reconciliation query computes the authoritative
values and repairs the product summary.
Verification checklist
- The embedded recent-review array remains at the declared bound.
- The full review history is independently queryable.
- Every activity bucket has no more than five events.
- Bucket BSON sizes are measured by MongoDB.
- Computed drift is visible before reconciliation and eliminated afterward.
- No exercise attempts to allocate a near-limit document merely to prove the documented 16 MiB ceiling.
Check your understanding
- Why is an unbounded array dangerous before 16 MiB?
- What does the subset pattern optimize?
- What does the bucket pattern bound?
- Why can a computed field be stale?
- How should you choose the bound?
Review the answers
It can increase network, cache, serialization, index, update, and contention costs as it grows, and eventually reaches the hard size ceiling.
The hot working set: it embeds only the small frequently read portion while keeping the complete history elsewhere.
A long series is split into multiple documents, each with a defined maximum group/window aligned with reads or writes.
It duplicates a derived value. If its refresh is asynchronous or a write path fails between source and summary updates, the copy diverges until reconciliation.
From application request size, retention, write rate, index/memory budget, concurrency, and growth—not from a universal MongoDB magic number.
docker rm -f atlasmart-mongo-ch06-l4
The final lesson uses all of these mechanisms to produce a defensible AtlasMart schema and a duplication decision register.
5. Production judgment: bounds are capacity contracts
Subset and bucket patterns are appropriate when they align storage with a stable access unit and keep the working set/document size predictable. Computed values are appropriate when reads materially outnumber source changes and the business can state a freshness requirement. Their non-guarantees are equally important: subsets can be stale, bucket assignment can contend under concurrency, and computed values can drift after partial failure.
Monitor array/bucket cardinality, max/p95 BSON size, bucket fill distribution, hot-bucket write latency, multikey index growth, reconciliation lag, and source-versus-derived mismatch counts. Replica/shard topology can change hotspot behavior; a monotonic or tenant-skewed bucket key can concentrate writes. Security and retention should be applied to the canonical history as well as projections, and deletion workflows must remove or rebuild copies where policy requires.
Failure-injection should include a skipped projection update, a duplicate/replayed event, a full bucket race, and an outlier entity. Choose bounds from measured payload and concurrency budgets rather than folklore. These patterns are available with Community/local tooling; later managed features may automate parts of the problem but do not change the contract. Lesson 5 combines these decisions into a schema that can be reviewed and migrated safely.
Authoritative references
- MongoDB release notes — Current stable server series and patch history.
- MongoDB 8.3 release notes — 8.3.8 is the latest released 8.3 patch at review time; 8.3.9 is upcoming.
- mongosh release notes — mongosh 2.10.0 was released August 13, 2026.
- PyMongo release notes — Current PyMongo 4.17 line and driver changes.
- MongoDB schema design process — Official workload-first process: identify workload, map relationships, apply patterns, and create supporting indexes.
- Data modeling best practices — Official embedding-versus-referencing decision guidance.
- Avoid unbounded arrays — Official anti-pattern guidance and subset/reference alternatives.
- Subset pattern — Reduce working set by embedding only frequently accessed related data.
- Bucket pattern — Group long series into bounded documents aligned with access patterns.
- Computed pattern — Store frequently read derived values and understand freshness tradeoffs.
- Group data patterns — Bucket and outlier pattern guidance.