Capacity is the margin needed to survive growth, bursts, member loss, maintenance, resync, and recovery—not a single utilization percentage.
Capacity for Working Set, WiredTiger Cache, Indexes, Oplog, Connections, Disk, and Replication Headroom
Turn working set, cache, indexes, oplog, connections, disk, and replication measurements into capacity and failure-recovery headroom.
Learning objectives
Translate working-set, WiredTiger cache, index, oplog, connection, disk, and replication evidence into explicit capacity headroom.
Measure the oplog window and replication lag on a disposable three-member replica set.
Distinguish “current usage” from growth rate, burst allowance, failure headroom, and recovery headroom.
Build a capacity worksheet from observed bytes/rates rather than fixed folklore percentages.
Recognize when a topology change or workload/data-model change is safer than another tuning knob.
This chapter pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and
PyMongo 4.17.0 where client behavior or load
generation matters. Lesson 2 uses a disposable three-member
replica set so replication lag, oplog history, and member-level
headroom are observable. All three members run on one host, so
the lab does not model independent disks, networks, or failure
domains. Host exposure is loopback-only on
127.0.0.1:27203–27205. Authentication and TLS are
disabled only for disposable local labs; production security
remains the Chapter 22 prerequisite. Default read/write concern
and primary read preference are used unless a step states
otherwise.
FCV is observed and never changed in mandatory labs.
Atlas, Enterprise Advanced, Search, Vector Search, and KMS are
optional unless explicitly labeled. Use at least several GB of
free disk and roughly 2–3 GB RAM for the full topology; if that
is unreasonable, the included Python worksheet can be run
without MongoDB. Runtime performance, failover, capacity, and
upgrade labs were not executed in the generation environment, so
metric values, latency distributions, queue depths, oplog
windows, replication lag, incident times, and upgrade durations
must be measured locally rather than copied as invented output.
1. Capacity is a failure-budget problem, not “how full are we?”
AtlasMart can operate normally at 70% disk usage and still be dangerously under-provisioned if growth will exhaust the remaining space before the next maintenance window, or if a failed member has no disk/network headroom to resynchronize. Capacity planning therefore combines current utilization, growth rate, burst load, failure mode, and recovery work.
| Dimension | Measure | Headroom question |
|---|---|---|
| Working set / cache | bytes in cache, pages read, query latency under warm/cold access | Can the active data/index set fit with enough room for bursts and maintenance? |
| Indexes |
indexSize, totalIndexSize,
write cost
|
Are indexes growing faster than useful query value? |
| Oplog | configured/used size, GB/hour, first/last timestamp | Does the history window cover plausible secondary outages and maintenance? |
| Connections | current, active, available, pool counts | Can a failover or traffic burst reconnect without a storm? |
| Disk | free bytes, data/index/journal growth, latency | Can the node absorb compaction/build/recovery and growth before intervention? |
| Replication | lag, headroom, secondary apply rate | Can a slow/failed member catch up before oplog rollover? |
2. Build a three-member replica set and capture member state
docker network rm atlasmart-ch26-l2-net 2>/dev/null || truedocker network create atlasmart-ch26-l2-net >/dev/nullfor n in 1 2 3; do docker rm -f atlasmart-ch26-l2-r$n 2>/dev/null || true; docker volume rm atlasmart-ch26-l2-r$n-db 2>/dev/null || true; donedocker run -d --name atlasmart-ch26-l2-r1 --network atlasmart-ch26-l2-net -p 127.0.0.1:27203:27017 -v atlasmart-ch26-l2-r1-db:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet rs26l2 --bind_ip_alldocker run -d --name atlasmart-ch26-l2-r2 --network atlasmart-ch26-l2-net -p 127.0.0.1:27204:27017 -v atlasmart-ch26-l2-r2-db:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet rs26l2 --bind_ip_alldocker run -d --name atlasmart-ch26-l2-r3 --network atlasmart-ch26-l2-net -p 127.0.0.1:27205:27017 -v atlasmart-ch26-l2-r3-db:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet rs26l2 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27203/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27203/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"rs26l2",members:[{_id:0,host:"atlasmart-ch26-l2-r1:27017"},{_id:1,host:"atlasmart-ch26-l2-r2:27017"},{_id:2,host:"atlasmart-ch26-l2-r3:27017"}]})'until mongosh "mongodb://127.0.0.1:27203/admin?directConnection=true" --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; done
const a=db.getSiblingDB("admin");const rs=a.runCommand({replSetGetStatus:1});printjson(rs.members.map(m=>({name:m.name,stateStr:m.stateStr,optimeDate:m.optimeDate,lastHeartbeat:m.lastHeartbeatMessage})));printjson(a.runCommand({getParameter:1,featureCompatibilityVersion:1}));
3. Generate a bounded data/write rate and measure storage and oplog window
const d=db.getSiblingDB("atlasmart");d.events_ch26_l2.drop();d.events_ch26_l2.createIndex({tenantId:1,createdAt:-1},{name:"idx_tenant_created"});for (let b=0;b<50;b++) { const batch=[]; for (let i=0;i<1000;i++) { const n=b*1000+i; batch.push({tenantId:`tenant-${n%50}`,deviceId:`dev-${n%5000}`,kind:["view","cart","order","payment"][n%4],createdAt:new Date(Date.UTC(2026,8,1)+n*500),payload:"x".repeat(200+(n%800))}); } d.events_ch26_l2.insertMany(batch,{writeConcern:{w:"majority"}});}printjson(d.stats({scale:1024*1024}));printjson(d.events_ch26_l2.stats({scale:1024*1024}));
const local=db.getSiblingDB("local");const first=local.oplog.rs.find().sort({$natural:1}).limit(1).next();const last=local.oplog.rs.find().sort({$natural:-1}).limit(1).next();const windowSeconds=(last.wall-first.wall)/1000;const st=db.getSiblingDB("admin").runCommand({replSetGetStatus:1});printjson({ first:first.ts,last:last.ts,windowHours:windowSeconds/3600, members:st.members.map(m=>({name:m.name,state:m.stateStr,optimeDate:m.optimeDate}))});
The measured oplog window is specific to this write rate and configured oplog size. A larger oplog can still provide a dangerously short window if write volume spikes. Capacity planning should trend bytes/hour and window duration under peak-like workloads.
4. Correlate cache, disk, connections, and replication headroom
const a=db.getSiblingDB("admin");const d=db.getSiblingDB("atlasmart");const s=a.serverStatus();const dbs=d.stats({scale:1024*1024,freeStorage:0});printjson({ connections:s.connections, cache:{ maxBytes:s.wiredTiger.cache["maximum bytes configured"], currentBytes:s.wiredTiger.cache["bytes currently in the cache"], dirtyBytes:s.wiredTiger.cache["tracked dirty bytes in the cache"], pagesRead:s.wiredTiger.cache["pages read into cache"] }, queues:s.queues, dbMiB:{data:dbs.dataSize,storage:dbs.storageSize,indexes:dbs.indexSize,total:dbs.totalSize}});
docker stats --no-stream atlasmart-ch26-l2-r1 atlasmart-ch26-l2-r2 atlasmart-ch26-l2-r3docker system df -v | sed -n '1,160p'
Container metrics are supporting infrastructure evidence, not MongoDB semantics. On production hosts, collect disk latency/IOPS, filesystem free space, CPU steal, network saturation, and memory pressure from the infrastructure layer and align them to the same incident window.
5. Convert measurements into a capacity worksheet
capacity={ "disk_total_gib":1000, "disk_used_gib":620, "growth_gib_per_day":12, "oplog_window_hours":18, "planned_secondary_outage_hours":10, "current_connections":1800, "tested_connections":6000, "peak_connections":3200,}free=capacity["disk_total_gib"]-capacity["disk_used_gib"]days_to_full=free/capacity["growth_gib_per_day"]print({ "disk_free_gib":free, "days_to_full_at_observed_growth":round(days_to_full,1), "oplog_outage_margin_hours":capacity["oplog_window_hours"]-capacity["planned_secondary_outage_hours"], "connection_peak_margin":capacity["tested_connections"]-capacity["peak_connections"],})
The worksheet values are synthetic. Replace them with measured growth, peak workload, tested connection behavior, actual oplog window, restore/resync duration, and failure-domain assumptions. A percentage such as “always keep 30% disk free” is not a substitute for a time-to-exhaustion and recovery-headroom calculation.
6. Failure mode: zero headroom turns maintenance into an incident
Imagine a node that is already near its disk/connection/cache limits before an index build, resync, failover, or traffic burst. The maintenance operation is not the root cause; the absence of recovery headroom is. Repair the design by reducing data/index growth, extending the oplog window, right-sizing pools, adding disk/compute, rescheduling heavy work, or changing topology—whichever the measured bottleneck actually supports.
Check your understanding
- Why is current disk utilization insufficient for capacity planning?
- Why trend oplog window instead of oplog size alone?
- Why is a three-member lab on one laptop not a production HA model?
- What is connection headroom measured against?
- When is changing topology better than tuning a knob?
Review the answers
1. You also need growth rate, intervention time, burst behavior, and the extra work required during recovery or maintenance.
2. The same size represents very different amounts of history at different write rates.
3. The members still share host CPU, memory, storage, and failure domain.
4. A tested application/server pool behavior under expected peak and failover conditions, not just the current count.
5. When the evidence shows the workload, failure domain, or recovery requirement fundamentally exceeds the current node/topology capacity.
7. Production judgment
Capacity is a set of time-to-risk and recovery margins. Maintain headroom for ordinary peaks, one-member loss, resynchronization, index maintenance, backup/restore, elections, and traffic reconnection. Trend the rates that consume those margins. The next lesson turns the capacity model into a benchmark that preserves workload shape, skew, durability, and concurrency instead of generating an unrealistic maximum-throughput number.
Authoritative references
Operational fields, thresholds, upgrade paths, FCV behavior, Atlas metrics, and driver compatibility evolve. Re-check the current documentation for the exact server patch, deployment topology, driver, Atlas tier, and target upgrade/downgrade path before changing production systems.
- serverStatus command
- db.stats() / dbStats
- $collStats aggregation stage
- Database Profiler
- db.setProfilingLevel()
- $currentOp aggregation stage
- MongoDB Log Messages
- Explain Results
- Replication
- Replica Set Oplog
- Check Replica Set Replication Lag
- WiredTiger Storage Engine
- Atlas Monitoring and Alerts
- Atlas Monitoring and Alert Guidance
- Atlas Alert Basics
- Atlas Metrics
- PyMongo Release Notes
- PyMongo Upgrade Guidance
- MongoDB 8.3 Release Notes
- Upgrade 8.2 to 8.3
- Upgrade 8.2 Replica Set to 8.3
- Upgrade 8.2 Sharded Cluster to 8.3
- MongoDB 8.3 Compatibility Changes
- Downgrade 8.3 to 8.2
- MongoDB Versioning
- Backup Methods
- mongosh Release Notes