Chapter 12 · Compaction Strategies and Storage Amplification
Compaction Throughput, Concurrent Compactors, Pending Tasks, Disk Headroom, and Tuning by Evidence
Tune compaction throughput and concurrency from pending-work, disk-headroom and tail-latency evidence—not copied values.
Learning outcomes
AtlasMart's schema choice is reasonable, but pending compactions climb for hours during peak writes. An operator copies a blog recommendation, doubles concurrent compactors and removes the throughput cap. Queue depth falls briefly, then p99 reads and disk latency collapse. This lesson replaces knob-copying with a measurable control loop.
Interpret compaction throughput, concurrent compactors, pending tasks and history as a coupled resource system.
Measure compaction debt and disk headroom together with application tail latency.
Run a reversible local throughput experiment while capturing the original setting first.
Explain why concurrency is usually a second lever after throughput and storage capability.
Create production acceptance, stop and rollback criteria for compaction tuning or strategy migration.
The mandatory labs continue the disposable AtlasMart cluster
used by Chapters 01–11: pinned cassandra:5.0.9,
Docker network atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and Java 17 inside the image. Chapter 12 uses keyspace
atlasmart_compaction with
NetworkTopologyStrategy, replication factor (RF)
3, and normally LOCAL_QUORUM. Authentication,
client TLS, internode TLS, and remote JMX are disabled only
inside this isolated learning network. The Apache Cassandra
Java Driver 4.19.3 is optional; mandatory evidence uses
cqlsh, nodetool, Docker/Linux
filesystem tools, and Cassandra metrics. Unless a lesson
explicitly creates an STCS/LCS/TWCS comparison table, new
tables use UnifiedCompactionStrategy (UCS), reflecting current
Cassandra 5.0 guidance. Keep gc_grace_seconds at
its default unless a disposable purge experiment explicitly
changes it and performs repair first. Record the actual
compaction throughput, concurrent compactors, disk free space,
compression ratio, SSTable count, workload concurrency,
partition sizes, and latency distribution before changing
anything.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Compaction terms before changing a table
An SSTable (Sorted String Table) is immutable on-disk data. A memtable flush creates a new SSTable; updates and deletes therefore create newer versions instead of modifying older files in place. Compaction selects SSTables, reads their sorted contents, reconciles newer cells and tombstones, writes replacement SSTables, and retires input files after active readers no longer need them. Read amplification is the extra storage work a logical read performs across files/versions. Write amplification is the number of bytes rewritten in the background relative to application bytes written. Space amplification is temporary or steady-state disk usage above live logical data, including overlapping SSTables, replacement outputs, snapshots, streaming, repair and backup staging. A tombstone is a timestamped deletion/expiration marker; compaction can purge it only when Cassandra can do so safely. Compaction debt means eligible/pending work is accumulating faster than it is completed. UCS means UnifiedCompactionStrategy; STCS, LCS and TWCS mean SizeTiered, Leveled and TimeWindow compaction strategies respectively.
Current Cassandra 5.0 documentation recommends UCS for most
new workloads. At the same time, an unspecified
default_compaction still resolves to STCS in
cassandra.yaml. A course or legacy schema that
says “STCS is default” is therefore not proof that STCS is the
best choice for a new table.
1. The control loop
compaction_throughput caps aggregate compaction
throughput on a node; current Cassandra 5.0 configuration
documents 64 MiB/s as the default.
concurrent_compactors controls how many compactions
can run simultaneously. More concurrency can keep small SSTables
from accumulating during one long compaction, but it also
increases parallel CPU/disk pressure. Current guidance
explicitly suggests looking at throughput before concurrency
when compaction is too slow or too fast.
nodetool compactionstats shows active/pending work;
compactionhistory shows completed rewrites and
bytes. A queue is not automatically bad: the question is whether
debt trends upward without bound, whether SSTables/read and disk
occupancy worsen, and whether application SLOs are threatened.
| Signal | Healthy interpretation needs | Bad shortcut |
|---|---|---|
| Pending tasks | trend vs write rate and completion rate | any nonzero value is an incident |
| Throughput | device capacity + latency SLO + repair/streaming overlap | copy a blog MiB/s value |
| Concurrent compactors | CPU/disks + queue shape + task sizes | set equal to core count blindly |
| Disk free | rewrite/snapshot/streaming worst case | dataset size plus 10% |
| p99 latency | representative load and warmup | single cqlsh timing |
2. Baseline before touching knobs
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"
CREATE KEYSPACE IF NOT EXISTS atlasmart_compactionWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;
for i in 1 2 3 4 5 6; do date -Is docker exec atlasmart-cass-1 nodetool compactionstats docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra | tail -1" docker stats --no-stream atlasmart-cass-1 | tail -1 sleep 10donedocker exec atlasmart-cass-1 nodetool compactionhistory
In a real test, pair this with client p95/p99 read/write latency and timeout/error metrics. The mandatory local lab cannot create production-quality throughput numbers; its purpose is to prove that you can record settings, alter one variable, and restore it.
3. Reversible throughput experiment
First capture the current numeric throughput. The exact text format can vary, so verify the extracted number before executing the setter. The example changes only node 1 in the disposable cluster and immediately restores the original value after observing a small compaction. Do not run this pattern blindly in production.
RAW=$(docker exec atlasmart-cass-1 nodetool getcompactionthroughput)echo "$RAW"ORIGINAL=$(printf '%s\n' "$RAW" | grep -Eo '[0-9]+([.][0-9]+)?' | head -1)test -n "$ORIGINAL" || { echo "Could not parse baseline; stop."; exit 1; }echo "Parsed baseline MiB/s: $ORIGINAL"# Lab-only lower cap to make the change observable; not a recommendation.docker exec atlasmart-cass-1 nodetool setcompactionthroughput 16docker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool compact atlasmart_compaction compaction_ucsdocker exec atlasmart-cass-1 nodetool compactionstats# Restore exactly the recorded baseline.docker exec atlasmart-cass-1 nodetool setcompactionthroughput "$ORIGINAL"docker exec atlasmart-cass-1 nodetool getcompactionthroughput
If parsing does not match your version's output, stop and restore manually using the value you recorded. “Automation” is not a reason to risk an unknown production setting.
4. Concurrency and headroom
docker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_ucs
Increasing concurrent compactors can shorten a queue when the device has unused parallelism, but on a saturated disk it can worsen queueing and foreground latency. Headroom must cover more than the live dataset: include current SSTables, compaction outputs while inputs still exist, snapshots, streaming/repair/rebuild, commitlog, indexes such as SAI, backup/restore staging and failure scenarios. There is no safe universal free-space percentage.
Pending tasks are a symptom. First ask whether input write rate is temporarily high, SSTable geometry/strategy is wrong, disk is saturated, repair/streaming overlaps, or headroom is already constrained. Remove the bottleneck or adjust one control at a time with SLO stop conditions.
5. Strategy change is also a capacity event
ALTER TABLE ... WITH compaction=... looks like
metadata, but current Cassandra documentation warns that
changing compaction strategy on a populated table rewrites
existing SSTables using the new strategy. Treat an STCS/LCS/TWCS
→ UCS migration as an I/O and free-space project. Test on a
representative table or isolated node/JMX experiment where
appropriate, roll out deliberately, and avoid colliding with
repair, bootstrap, snapshot retention or backup windows.
| Before change | During | After |
|---|---|---|
| capture old schema/settings + backlog + headroom | watch compaction bytes, pending work, p99, disk | verify strategy, backlog convergence, SSTables/read, SLOs |
| estimate rewrite/temporary bytes | stop if free-space/latency threshold breached | retain rollback/migration notes |
| schedule around repair/streaming/backup | do not pile on major compaction | compare steady-state amplification, not first-minute speed |
6. Final evidence checklist
docker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_ucsdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_compaction compaction_ucsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 nodetool compactionhistorydocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra && find /var/lib/cassandra/data/atlasmart_compaction -type f -name '*Data.db' -printf '%s %p\n' | sort -n | tail -20"
Check your understanding
- Which knob should usually be examined before increasing concurrent compactors when compaction is too slow?
- Why is a falling pending-task count not enough to declare tuning successful?
- What makes changing strategy on a populated table risky?
- Why is 0/unlimited compaction throughput dangerous as a copied setting?
- What closes the loop after tuning?
Review the answers
1. Compaction throughput, alongside actual disk/CPU/SLO evidence; Cassandra documentation specifically directs operators to look there first.
2. Foreground p99/error rate, disk/CPU pressure, SSTables/read, headroom and long-term debt must also remain acceptable.
3. Existing SSTables are rewritten using the new strategy, creating potentially large I/O and temporary-space demand.
4. It can let compaction consume storage bandwidth needed by foreground reads/writes, repair and streaming.
5. Restore/confirm intended settings, verify backlog convergence and tail SLOs under representative load, document observed amplification/headroom, and retain rollback steps.
Production judgment
Compaction tuning is capacity engineering, not a table-property beauty contest. Evaluate application write rate, overwrite/delete/TTL rate, partition and clustering distribution, read/write p50/p95/p99 latency, SSTables per read, compression ratio, pending compactions, bytes compacted, compaction throughput, disk queue/throughput, CPU, JVM/GC, network, repair and streaming load, snapshots, SAI/vector index amplification, and free-space trajectory. RF and consistency level change how many replicas incur the physical work; a locally fast compaction configuration can still violate a cluster SLO under repair, node replacement, backup, or failure.
Every production change needs an acceptance window and rollback plan. Record the old schema/options, node-by-node rollout order, expected rewrite volume, required free space, compaction backlog limit, latency/error SLOs, and stop conditions. Managed Cassandra services may hide throughput/concurrency controls or select strategies for you; map the same concepts to the provider's exposed metrics rather than assuming identical knobs. Chapter 13 now focuses on the deletion side of this story: tombstone types, TTL, gc_grace, repair timing, zombie prevention and safe purge.
Summary and next bridge
Compaction tuning is the operational balance between debt and interference. Throughput, concurrency and strategy all consume the same finite disk/CPU/headroom budget, so evidence and rollback matter more than folklore values. Chapter 13 follows the tombstones that compaction is eventually asked to purge.
Authoritative references
Use these as the version-sensitive source of truth when regenerating the lesson; compaction recommendations and options can evolve between Cassandra releases.