Prompt 19 · Lesson 05 · Capacity and operations

Operate High-Volume Time-Series Workloads with Bounded Cardinality and Capacity Planning

Convert cardinality, rate, retention, storage, and latency evidence into an operational capacity model.

Advanced150–220 minutesCapacity/performance labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Build a capacity model from measurement rate, metadata cardinality, retention horizon, observed compression, index bytes, and working-set behavior.

02

Measure local batch-ingest throughput and latency distributions without turning one laptop run into a universal performance claim.

03

Use $collStats and bucket catalog diagnostics to detect sparse buckets, reopen pressure, cache-related closures, and storage growth.

04

Distinguish throughput bottlenecks caused by document shape/metadata/indexes from those caused by hardware, durability, network, or client configuration.

05

Design a production rollout with bounded cardinality, retention tests, failure injection, and explicit scaling/sharding boundaries.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where a driver is used. The mandatory lab is a disposable standalone on loopback port 27164 named atlasmart-ch19-l5; replication or sharding is not required to learn the time-series mechanisms in this chapter. Authentication and TLS are disabled only for the isolated local lab. Feature Compatibility Version (FCV) is inspected. FCV is observed and never changed. Default read/write concern and primary read preference apply. Atlas, Search, Vector Search, KMS, and Enterprise Advanced are not mandatory. Internal system.buckets.* data and time-series diagnostic counters are used only for observation; they are not application APIs. Product commands were not executed in this generation environment because Docker, mongod, mongosh, and PyMongo are unavailable here, so environment-dependent bucket counts, storage ratios, explain plans, and throughput/latency values must be measured on the learner machine rather than copied as invented output. The optional PyMongo harness is deliberately small and reports the learner machine’s measured p50/p95/p99 rather than embedding benchmark numbers in the lesson.

1. Capacity starts with workload arithmetic, then gets corrected by measurement

Suppose AtlasMart has 20,000 devices. If each writes one measurement every 10 seconds, the system receives 2,000 measurements/s before retries, bursts, or backfills. A 30-day retention horizon contains about 5.184 billion measurements. That arithmetic is useful for scale awareness, but it is not a disk-size prediction: bucket compression, indexes, metadata shape, allocation, late data, and replication/sharding all change physical cost.

Capacity planning therefore needs two layers: a deterministic logical workload model, then empirical measurements from a representative staging environment. Never extrapolate from document count alone.

Input How to obtain it Why it matters
measurement rate devices × samples/device/second plus bursts Ingest CPU/network and daily logical growth.
unique stable series distinct metaField identities Open/reopened bucket working set and cardinality.
average logical BSON bytes sample representative documents Uncompressed logical volume.
observed storage/index bytes $collStats on representative load Physical footprint after compression + index cost.
retention horizon business/compliance policy Steady-state retained measurement volume.
late-arrival distribution eventAt → ingestedAt telemetry Bucket reopening and retention surprises.
p50/p95/p99 ingest latency client-side timed batches Tail behavior; averages hide overload.

2. Build a bounded local workload, then inspect bucket behavior

start the disposable MongoDB 8.3.8 lab
docker rm -f atlasmart-ch19-l5 2>/dev/null || truedocker volume rm atlasmart-ch19-l5-data 2>/dev/null || truedocker run -d --name atlasmart-ch19-l5 \  -p 127.0.0.1:27164:27017 \  -v atlasmart-ch19-l5-data:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --bind_ip_alldocker exec atlasmart-ch19-l5 mongosh --quiet --eval 'printjson(db.adminCommand({buildInfo:1}).version);printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion);' 
create the capacity-test collection
const d=db.getSiblingDB("atlasmart");d.telemetry_ch19_l5.drop();d.createCollection("telemetry_ch19_l5",{timeseries:{timeField:"eventAt",metaField:"meta",granularity:"seconds"},expireAfterSeconds:2592000});d.telemetry_ch19_l5.createIndex({"meta.tenantId":1,"meta.sensorId":1,eventAt:1},{name:"tenant_sensor_time"});printjson(d.runCommand({listCollections:1,filter:{name:"telemetry_ch19_l5"}}).cursor.firstBatch[0]);printjson(d.telemetry_ch19_l5.getIndexes());
PyMongo batch-ingest harness with local latency percentiles
from datetime import datetime, timezone, timedeltafrom statistics import medianfrom time import perf_counterfrom pymongo import MongoClientclient = MongoClient("mongodb://127.0.0.1:27164/?directConnection=true")coll = client.atlasmart.telemetry_ch19_l5base = datetime(2026, 9, 3, tzinfo=timezone.utc)latencies_ms=[]inserted=0for batch_no in range(100):    docs=[]    for j in range(200):        n=batch_no*200+j        sensor=n % 200        docs.append({            "eventAt": base + timedelta(seconds=n//200),            "ingestedAt": datetime.now(timezone.utc),            "meta": {"tenantId":"tenant-a","storeId":f"store-{sensor%10:02d}","sensorId":f"sensor-{sensor:03d}"},            "temperatureC": 15.0 + (n % 100)/20.0,            "powerW": 70 + (n % 50),        })    t0=perf_counter()    coll.insert_many(docs, ordered=False)    latencies_ms.append((perf_counter()-t0)*1000)    inserted += len(docs)values=sorted(latencies_ms)def pct(p):    idx=min(len(values)-1, round((len(values)-1)*p))    return values[idx]print({"inserted":inserted,"batch_size":200,"batches":len(values),"p50_ms":pct(0.50),"p95_ms":pct(0.95),"p99_ms":pct(0.99),"max_ms":values[-1]})client.close()

The harness inserts exactly 20,000 measurements, but its timing output is intentionally not pre-filled. Report server version, hardware/VM limits, cache state, durability settings, batch size, client version, indexes, and whether competing workloads were present before comparing results across machines.

3. Inspect storage, bucket catalog, and index cost together

collect the evidence after the workload
const d=db.getSiblingDB("atlasmart");const s=d.telemetry_ch19_l5.aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats;printjson({  count:s.count,  logicalBytes:s.size,  storageBytes:s.storageSize,  totalIndexBytes:s.totalIndexSize,  totalBytes:s.totalSize,  timeseries:s.timeseries});printjson(db.serverStatus().bucketCatalog);printjson(d.telemetry_ch19_l5.explain("executionStats").find({"meta.tenantId":"tenant-a","meta.sensorId":"sensor-042",eventAt:{$gte:ISODate("2026-09-03T00:00:10Z"),$lt:ISODate("2026-09-03T00:01:10Z")}}).sort({eventAt:1}));

Useful diagnostics include measurement commits, bucket count/average size, buckets opened because of metadata, buckets closed because of size/count/cache pressure, and reopen activity. The exact field set is diagnostic and can evolve. Pair these counters with storage/index bytes and query evidence; no one bucket metric is a complete capacity signal.

4. Turn measured ratios into a scenario worksheet—not a promise

capacity worksheet using your measured values
from dataclasses import dataclass@dataclassclass Scenario:    devices: int    seconds_per_sample: float    retention_days: int    measured_storage_bytes_per_measurement: float    measured_index_bytes_per_measurement: float    replication_factor: float = 1.0    def report(self):        rate=self.devices/self.seconds_per_sample        retained=rate*self.retention_days*86400        data=retained*self.measured_storage_bytes_per_measurement*self.replication_factor        index=retained*self.measured_index_bytes_per_measurement*self.replication_factor        return {"measurements_per_second":rate,"retained_measurements":retained,"estimated_data_bytes":data,"estimated_index_bytes":index,"estimated_total_bytes":data+index}# Replace the two measured byte ratios with values from your representative staging run.s=Scenario(devices=20000,seconds_per_sample=10,retention_days=30,measured_storage_bytes_per_measurement=0.0,measured_index_bytes_per_measurement=0.0)print(s.report())

The zeros are deliberate placeholders for measurements you must supply, not unfinished content. Production sizing must also reserve headroom for checkpoints, replication/oplog where applicable, backups, compaction/fragmentation, temporary operations, balancer/resharding if sharded, workload growth, and failure-domain capacity. Never size a cluster to steady-state averages with no margin.

5. Controlled failure: unbounded metadata cardinality

A load test can look healthy at 200 stable sensor series and collapse later if production metadata contains a UUID per reading. Repeat a bounded comparison in a separate disposable collection rather than mutating the good dataset.

inject high-cardinality metadata without risking unrelated data
const d=db.getSiblingDB("atlasmart");d.telemetry_highcard_ch19_l5.drop();d.createCollection("telemetry_highcard_ch19_l5",{timeseries:{timeField:"eventAt",metaField:"meta",granularity:"seconds"}});const base=ISODate("2026-09-03T00:00:00Z");const docs=[];for(let i=0;i<5000;i++) docs.push({eventAt:new Date(base.getTime()+i*1000),meta:{tenantId:"tenant-a",sensorId:"sensor-042",eventId:`unique-${i}`},v:i%100});d.telemetry_highcard_ch19_l5.insertMany(docs);for(const n of ["telemetry_ch19_l5","telemetry_highcard_ch19_l5"]){  const s=d[n].aggregate([{$collStats:{storageStats:{}}}]).toArray()[0].storageStats;  printjson({name:n,count:s.count,bucketCount:s.timeseries.bucketCount,avgBucketSize:s.timeseries.avgBucketSize,storageSize:s.storageSize,totalIndexSize:s.totalIndexSize});}

The collections do not contain identical counts, so this is not a normalized performance benchmark. It is a controlled structural demonstration: unique metadata forces many distinct series. If you need a numeric comparison, run equal-sized variants and report the exact environment and distribution.

6. Production judgment and operating checklist

Operate time-series workloads from explicit budgets: series cardinality, measurement rate, retention horizon, query windows, storage/index growth, and latency SLOs. Monitor bucket closures/reopens/cache pressure, oldest/newest event times, TTL lag, WiredTiger cache, disk IOPS/latency, CPU, network, replication lag when replicated, and query plan regressions. Test burst ingestion, delayed gateways, historical backfills, retention changes, index additions, node restarts, and disk-pressure scenarios in staging.

Scale vertically while the working set and throughput fit comfortably; shard only when capacity/throughput/operational evidence requires it. For sharded time-series collections, prefer metadata-based shard-key design and remember that timeField-containing shard keys are deprecated starting MongoDB 8.0. Zone sharding is not supported for time-series collections. If a workload is not predominantly time-dependent append-like measurements, a normal collection may be simpler.

Bridge. Chapter 20 applies the same evidence-first method to geospatial data: model coordinates correctly, choose 2dsphere versus 2d intentionally, and verify spatial query/index behavior.

cleanup only the chapter-specific lab
docker rm -f atlasmart-ch19-l5 2>/dev/null || truedocker volume rm atlasmart-ch19-l5-data 2>/dev/null || true

Check your understanding

  1. Why is devices × sample rate × retention not enough to predict disk size?
  2. Why record p50/p95/p99 batch latency instead of only an average?
  3. What bucket-catalog signals can indicate pressure?
  4. Why does the capacity worksheet require measured bytes per measurement?
  5. When should AtlasMart consider sharding time-series data?
Review the answers

1. Compression, bucket density, indexes, allocation, metadata cardinality, replication, and operational headroom change physical storage.

2. Tail latency reveals stalls and overload that averages can hide.

3. Examples include closures due to memory/cache, many buckets opened due to metadata, waits, and reopen/fetch activity; interpret them with storage and query evidence.

4. Compression/index cost is workload-specific, so staging evidence is safer than universal ratios.

5. When measured capacity/throughput/working-set or operational requirements exceed comfortable single-replica-set headroom—not merely because the collection is large.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.