Chapter 11 · Document Databases and Aggregate-Oriented Modeling
Document Growth, Size Limits, Large Arrays, Bucketing, and Schema Versioning
Prevent unbounded document growth with buckets and versioned schemas, distinguish product limits from modeling principles, and practice an online-compatible migration path.
Learning outcomes
Document models often start clean and fail later through growth. AtlasMart’s customer activity history is a perfect trap: an array append looks cheap until one record becomes enormous, indexes multiply, updates move more bytes, and one hot customer becomes one hot document.
Recognize unbounded arrays and document-growth risk before a hard engine limit is reached.
Use bucketing, child collections, and bounded subsets intentionally.
Evolve schemas online with versioned readers/writers.
Separate implementation limits from vendor-neutral modeling principles.
1. Growth is an operational property
A document has a serialized size, update frequency, index footprint, and placement. An unbounded array is one whose maximum element count is not constrained by a business rule or bucket boundary. Even before a hard size limit, it can make reads transfer unnecessary history, updates contend on one aggregate, and array indexes fan out. MongoDB is one concrete example: BSON documents must remain below 16 MiB, and its documentation explicitly warns against unbounded arrays. The vendor-neutral lesson is earlier: design an upper bound.
2. Bucketing turns one unbounded aggregate into bounded units
A bucket groups records by a stable boundary such as time window, count, tenant, or key range. AtlasMart can store customer events in daily or fixed-count buckets. Recent history remains easy to fetch while old buckets become immutable and independently retainable. The exact bucket size is workload-dependent: choose it from write rate, read window, retention, index cost, and maximum object size.
import json
from collections import defaultdict
# Simulate growth; these byte counts are JSON serialization sizes, not MongoDB BSON sizes.
events = [{"ts": i, "kind": "view", "sku": f"sku-{i%20}"} for i in range(5000)]
monolith = {"customer_id": "c1", "events": events}
print("monolith JSON bytes:", len(json.dumps(monolith).encode()))
buckets = []
for start in range(0, len(events), 200):
buckets.append({"customer_id": "c1", "bucket": start//200, "events": events[start:start+200]})
sizes = [len(json.dumps(b).encode()) for b in buckets]
print("bucket count:", len(buckets), "largest bucket bytes:", max(sizes))
v1 = {"schema_version": 1, "name": "Desk Lamp", "price": 30}
v2 = {"schema_version": 2, "name": "Desk Lamp", "pricing": {"amount": 30, "currency": "USD"}}
def read_price(d):
if d.get("schema_version", 1) == 1:
return {"amount": d["price"], "currency": "USD"}
return d["pricing"]
print("v1 normalized:", read_price(v1))
print("v2 normalized:", read_price(v2))
The JSON byte counts are deliberately labeled simulation evidence, not BSON measurements. The important observation is structural: 25 bounded event documents have a known maximum in this workload, while the monolith grows with every event.
3. Schema versioning avoids flag-day migration
Changing price into a nested
pricing object affects stored data, indexes, API
serializers, and downstream consumers. A safe rollout commonly
uses expand → migrate → contract: readers
accept old and new, writers emit the new form, a backfill
converts old records, telemetry proves old forms are gone, and
only then are compatibility paths removed. Do not run an
irreversible mass rewrite before rollback is proven.
“The database is flexible, so we will just start writing v2.” During mixed-version deployment, old readers fail or silently ignore fields. The repair is explicit version compatibility and measurable migration state.
4. Production judgment
Bucket boundaries affect query fan-out, hotspot risk, TTL/retention granularity, and repair cost. Too-small buckets create metadata/index overhead; too-large buckets recreate the original problem. Treat product size limits as guardrails, not target sizes. Leave headroom for future fields and serialization differences.
Check your understanding
- What makes an array dangerous even before a hard document-size limit?
- Why are the lab byte counts not a MongoDB benchmark?
- What is the safe sequence for a breaking field-shape change?
- What should determine bucket size?
Review the answers
1. Unbounded growth increases transfer, contention, rewrite cost, index fan-out, and hotspot concentration.
2. They measure Python JSON serialization only and intentionally do not model BSON encoding, storage compression, or engine I/O.
3. Deploy compatible readers, emit the new form, backfill/verify, then remove old compatibility.
4. Observed write/read windows, retention, object-size headroom, index cost, fan-out, and operational repair behavior.
Authoritative references
- MongoDB — Embedded data models — concrete current implementation guidance for aggregate locality and the 16 MiB BSON document limit.
- MongoDB — References — cases where independent lifecycle, many-to-many relationships, or frequent independent access favor references.
- MongoDB — Multikey indexes — array-index behavior and index-entry implications.
- MongoDB — Schema validation — evidence that flexible documents still benefit from enforced structural rules.
- MongoDB — Avoid unbounded arrays — implementation example of bounding growth with subsetting/references.
- MongoDB 8.3 release notes — dated implementation snapshot; 8.3.8 is the latest released patch as of this chapter review.