Chapter 11 · Document Databases and Aggregate-Oriented Modeling

Document Growth, Size Limits, Large Arrays, Bucketing, and Schema Versioning

Prevent unbounded document growth with buckets and versioned schemas, distinguish product limits from modeling principles, and practice an online-compatible migration path.

Intermediate90–115 minutesGrowth + bucketing + versioning labPython 3.13+ · standard libraryMongoDB 8.3.8 optional referenceLast reviewed: August 2026

Learning outcomes

Document models often start clean and fail later through growth. AtlasMart’s customer activity history is a perfect trap: an array append looks cheap until one record becomes enormous, indexes multiply, updates move more bytes, and one hot customer becomes one hot document.

01

Recognize unbounded arrays and document-growth risk before a hard engine limit is reached.

02

Use bucketing, child collections, and bounded subsets intentionally.

03

Evolve schemas online with versioned readers/writers.

04

Separate implementation limits from vendor-neutral modeling principles.

1. Growth is an operational property

A document has a serialized size, update frequency, index footprint, and placement. An unbounded array is one whose maximum element count is not constrained by a business rule or bucket boundary. Even before a hard size limit, it can make reads transfer unnecessary history, updates contend on one aggregate, and array indexes fan out. MongoDB is one concrete example: BSON documents must remain below 16 MiB, and its documentation explicitly warns against unbounded arrays. The vendor-neutral lesson is earlier: design an upper bound.

2. Bucketing turns one unbounded aggregate into bounded units

A bucket groups records by a stable boundary such as time window, count, tenant, or key range. AtlasMart can store customer events in daily or fixed-count buckets. Recent history remains easy to fetch while old buckets become immutable and independently retainable. The exact bucket size is workload-dependent: choose it from write rate, read window, retention, index cost, and maximum object size.

python · AtlasMart deterministic simulation
import json
from collections import defaultdict

# Simulate growth; these byte counts are JSON serialization sizes, not MongoDB BSON sizes.
events = [{"ts": i, "kind": "view", "sku": f"sku-{i%20}"} for i in range(5000)]
monolith = {"customer_id": "c1", "events": events}
print("monolith JSON bytes:", len(json.dumps(monolith).encode()))

buckets = []
for start in range(0, len(events), 200):
    buckets.append({"customer_id": "c1", "bucket": start//200, "events": events[start:start+200]})
sizes = [len(json.dumps(b).encode()) for b in buckets]
print("bucket count:", len(buckets), "largest bucket bytes:", max(sizes))

v1 = {"schema_version": 1, "name": "Desk Lamp", "price": 30}
v2 = {"schema_version": 2, "name": "Desk Lamp", "pricing": {"amount": 30, "currency": "USD"}}

def read_price(d):
    if d.get("schema_version", 1) == 1:
        return {"amount": d["price"], "currency": "USD"}
    return d["pricing"]

print("v1 normalized:", read_price(v1))
print("v2 normalized:", read_price(v2))

The JSON byte counts are deliberately labeled simulation evidence, not BSON measurements. The important observation is structural: 25 bounded event documents have a known maximum in this workload, while the monolith grows with every event.

3. Schema versioning avoids flag-day migration

Changing price into a nested pricing object affects stored data, indexes, API serializers, and downstream consumers. A safe rollout commonly uses expand → migrate → contract: readers accept old and new, writers emit the new form, a backfill converts old records, telemetry proves old forms are gone, and only then are compatibility paths removed. Do not run an irreversible mass rewrite before rollback is proven.

Wrong approach

“The database is flexible, so we will just start writing v2.” During mixed-version deployment, old readers fail or silently ignore fields. The repair is explicit version compatibility and measurable migration state.

4. Production judgment

Bucket boundaries affect query fan-out, hotspot risk, TTL/retention granularity, and repair cost. Too-small buckets create metadata/index overhead; too-large buckets recreate the original problem. Treat product size limits as guardrails, not target sizes. Leave headroom for future fields and serialization differences.

Check your understanding

  1. What makes an array dangerous even before a hard document-size limit?
  2. Why are the lab byte counts not a MongoDB benchmark?
  3. What is the safe sequence for a breaking field-shape change?
  4. What should determine bucket size?
Review the answers

1. Unbounded growth increases transfer, contention, rewrite cost, index fan-out, and hotspot concentration.

2. They measure Python JSON serialization only and intentionally do not model BSON encoding, storage compression, or engine I/O.

3. Deploy compatible readers, emit the new form, backfill/verify, then remove old compatibility.

4. Observed write/read windows, retention, object-size headroom, index cost, fan-out, and operational repair behavior.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.