Capacity is the ability to survive the intended failure and maintenance state, not merely healthy steady state.
Capacity Models for Data Size, Replication, Indexes, Headroom, Growth, and Failure Scenarios
Build failure-aware capacity from data growth, compression, replication, indexes, storage-engine overhead, compaction/rebuild space, throughput headroom, and explicit node/zone failure scenarios.
Build storage capacity from raw data, compression, replication, indexes, metadata/tombstone overhead, and temporary compaction/rebuild space.
Combine growth and peak-throughput requirements with one-node/zone failure and maintenance/rebalance headroom.
Detect a healthy-state plan that fails during recovery even though normal disk utilization looks comfortable.
Create a capacity review that records assumptions, validates them with workload measurements, and bridges resource headroom to managed/cloud cost decisions.
1. Capacity planning is a failure-state worksheet
AtlasMart does not ask “does the dataset fit on the cluster today?” It asks whether the projected dataset, indexes, replicas, storage-engine overhead, temporary compaction/rebuild space, and peak workload still fit when a node or failure domain is unavailable and the system is doing recovery work. Headroom is reserved resource capacity for spikes, maintenance, failures, and estimation error—not unused waste.
Raw application bytes are only the first term. Compression changes stored bytes; replication multiplies them; secondary indexes and derived local structures add storage and write amplification; tombstones/obsolete versions and metadata add overhead; compaction may temporarily keep old and new SSTables simultaneously; repair/rebalance adds network, disk, and CPU work; backups may live outside the serving cluster but still consume storage/egress/restore bandwidth.
2. Write assumptions as equations, not folklore
For an illustrative LSM-style capacity model, AtlasMart projects one-year raw bytes, applies a measured compression ratio, multiplies by replication factor, adds index and engine overhead, then reserves temporary rewrite space. Separately it compares peak useful operations with measured safe per-node throughput for the same workload and guarantees. The target utilization is a planning policy, not a universal constant.
| Capacity term | Example question | Evidence source |
|---|---|---|
| Raw/growth | How many bytes next year? | retention, ingestion, cardinality forecasts |
| Compression | What ratio under representative values? | measured backup/SSTable/table footprint |
| Replication | How many durable copies and where? | topology/RF policy |
| Indexes/derived structures | What alternate access paths exist? | schema/index catalog and measured footprint |
| Engine overhead | Tombstones, metadata, obsolete versions? | production storage metrics and compaction history |
| Temporary work | Can compaction/rebuild coexist with current files? | strategy documentation + observed peaks |
| Throughput | Can survivors carry peak useful load? | representative benchmark with same guarantees |
| Failure headroom | Which node/zone loss must remain serviceable? | SLO/DR design and game-day evidence |
Some product metrics already include obsolete files or index bytes; others do not. Capacity worksheets must define every term. Validate with actual filesystem/object-store and engine metrics before purchasing or scaling.
3. Deliberately wrong approach: size for healthy steady-state utilization
The lab projects one-year data and calculates a steady serving footprint of about 4.01 TB plus ~1.20 TB temporary compaction/recovery space. A seven-node, 1-TB-per-node cluster looks comfortable in healthy steady state at ~57% disk utilization. After one node is unavailable, recovery-peak storage would require ~87% of surviving disk capacity, violating the lab's 75% policy. Healthy fit was not failure fit.
4. AtlasMart lab: find the minimum modeled node count
Python 3.13+ standard library only. Ratios and per-node throughput are explicit planning inputs, not recommendations or product limits.
import math
raw_tb = 1.20
growth = 1.25
compression = 0.55
replication = 3
index_ratio = 0.35
overhead_ratio = 0.20
compaction_temp_ratio = 0.30
node_disk_tb = 1.0
safe_disk_util = 0.75
peak_ops = 60000
safe_ops_per_node = 15000
future_raw = raw_tb * growth
compressed = future_raw * compression
replicated = compressed * replication
indexes = replicated * index_ratio
overhead = (replicated + indexes) * overhead_ratio
steady = replicated + indexes + overhead
compaction_temp = steady * compaction_temp_ratio
recovery_peak = steady + compaction_temp
print("ONE-YEAR FAILURE-AWARE CAPACITY WORKSHEET")
for k,v in {
"future_raw_tb":future_raw,
"compressed_tb":compressed,
"replicated_tb":replicated,
"indexes_tb":indexes,
"tombstone_metadata_overhead_tb":overhead,
"steady_state_tb":steady,
"compaction_temp_tb":compaction_temp,
"recovery_peak_tb":recovery_peak,
}.items(): print(k, round(v,3))
print("\nNODE COUNT CHECK (survive one node unavailable)")
minimum=None
for nodes in range(5,13):
survivors=nodes-1
disk_util = recovery_peak/(survivors*node_disk_tb)
throughput_util = peak_ops/(survivors*safe_ops_per_node)
ok = disk_util <= safe_disk_util and throughput_util <= .80
print(nodes, "nodes -> survivors", survivors,
"disk_util=",f"{disk_util:.1%}","throughput_util=",f"{throughput_util:.1%}","PASS" if ok else "FAIL")
if ok and minimum is None: minimum=nodes
print("minimum modeled node count:", minimum)
print("\nBROKEN HEALTHY-STATE ONLY PLAN")
healthy_nodes=7
print("7-node healthy disk util:", f"{steady/(healthy_nodes*node_disk_tb):.1%}")
print("7-node one-failure recovery-peak disk util:", f"{recovery_peak/((healthy_nodes-1)*node_disk_tb):.1%}")
print("healthy fit does not prove failure/rebalance fit")
Under the stated one-year growth, RF=3, index/overhead, temporary-space, 60k ops/s, and one-node-failure assumptions, 5–7 nodes fail the lab policy and 8 is the first modeled count that passes both disk and throughput checks. Different workloads or products require different measured inputs.
5. Production judgment: capacity is an iterated model with confidence bounds
Track forecast error and revisit the worksheet when retention, value size, replication, indexes, compression, compaction strategy, workload skew, or failure objectives change. Model zones, not only nodes: losing a whole zone can remove more capacity and quorum placement than a single server. Include repair/rebuild duration because remaining replicas stay exposed longer when recovery is slow. Include network and cloud egress limits where relevant.
Use observed high-percentile node utilization and hot-shard distributions rather than assuming perfect balance. Keep separate thresholds for alerting, autoscaling, and hard safety rejection. Test the intended degraded state—drain a disposable node or simulate capacity removal in a safe environment—and compare client SLOs, rebuild time, queues, and disk headroom. Chapter 24 will build on this worksheet by asking who owns these responsibilities in managed NoSQL systems and how provisioned/serverless pricing changes the cost model.
Check your understanding
- Why is raw data size insufficient for database capacity planning?
- Why model throughput and storage separately?
- What is failure headroom?
- Why can seven healthy nodes look safe yet fail the recovery model?
- What makes a capacity ratio trustworthy?
Review the answers
1. Serving storage also includes compression effects, replicas, indexes, engine metadata/obsolete versions, temporary compaction/rebuild space, and growth.
2. A cluster can have free disk but insufficient CPU/network/replica service rate, or enough throughput but insufficient storage/rewrite headroom.
3. Capacity intentionally reserved so the service can meet its objectives during a defined node/zone failure, maintenance, repair, or traffic spike.
4. Steady-state data is spread over all seven; after one node is unavailable and temporary recovery/compaction data must coexist, survivor utilization can exceed the safety policy.
5. Its inputs are measured or explicitly forecast for the real workload, every term is defined, balance/skew and failure domains are modeled, and the degraded state is validated in tests/game days.
References
Foundational claims use primary research or current official documentation where practical. Product references are implementation anchors only; the mandatory labs are vendor-neutral.
- Apache Cassandra 5.0 — Hardware Choices — Current official discussion of CPU, memory, disks, workload-dependent benchmarking, and storage layout considerations.
- Apache Cassandra 5.0 — Size Tiered Compaction Strategy — Official warning that compaction can require old and new SSTables simultaneously, creating space amplification.
- Apache Cassandra — Auto Repair — Current official example of repair backpressure and disk-headroom rejection semantics.
- YCSB — Benchmark framework used as a reminder that workload and guarantees must be explicit before throughput values feed a capacity model.