Chapter 12 · Compaction Strategies and Storage Amplification

Compaction Throughput, Concurrent Compactors, Pending Tasks, Disk Headroom, and Tuning by Evidence

Tune compaction throughput and concurrency from pending-work, disk-headroom and tail-latency evidence—not copied values.

Intermediate120–160 minutesThroughput + debt tuning labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · UCS current guidance · STCS/LCS/TWCS comparisonsLast reviewed: September 2026

Learning outcomes

AtlasMart's schema choice is reasonable, but pending compactions climb for hours during peak writes. An operator copies a blog recommendation, doubles concurrent compactors and removes the throughput cap. Queue depth falls briefly, then p99 reads and disk latency collapse. This lesson replaces knob-copying with a measurable control loop.

01

Interpret compaction throughput, concurrent compactors, pending tasks and history as a coupled resource system.

02

Measure compaction debt and disk headroom together with application tail latency.

03

Run a reversible local throughput experiment while capturing the original setting first.

04

Explain why concurrency is usually a second lever after throughput and storage capability.

05

Create production acceptance, stop and rollback criteria for compaction tuning or strategy migration.

Chapter 12 lab baseline

The mandatory labs continue the disposable AtlasMart cluster used by Chapters 01–11: pinned cassandra:5.0.9, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes per node, and Java 17 inside the image. Chapter 12 uses keyspace atlasmart_compaction with NetworkTopologyStrategy, replication factor (RF) 3, and normally LOCAL_QUORUM. Authentication, client TLS, internode TLS, and remote JMX are disabled only inside this isolated learning network. The Apache Cassandra Java Driver 4.19.3 is optional; mandatory evidence uses cqlsh, nodetool, Docker/Linux filesystem tools, and Cassandra metrics. Unless a lesson explicitly creates an STCS/LCS/TWCS comparison table, new tables use UnifiedCompactionStrategy (UCS), reflecting current Cassandra 5.0 guidance. Keep gc_grace_seconds at its default unless a disposable purge experiment explicitly changes it and performs repair first. Record the actual compaction throughput, concurrent compactors, disk free space, compression ratio, SSTable count, workload concurrency, partition sizes, and latency distribution before changing anything.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Compaction terms before changing a table

An SSTable (Sorted String Table) is immutable on-disk data. A memtable flush creates a new SSTable; updates and deletes therefore create newer versions instead of modifying older files in place. Compaction selects SSTables, reads their sorted contents, reconciles newer cells and tombstones, writes replacement SSTables, and retires input files after active readers no longer need them. Read amplification is the extra storage work a logical read performs across files/versions. Write amplification is the number of bytes rewritten in the background relative to application bytes written. Space amplification is temporary or steady-state disk usage above live logical data, including overlapping SSTables, replacement outputs, snapshots, streaming, repair and backup staging. A tombstone is a timestamped deletion/expiration marker; compaction can purge it only when Cassandra can do so safely. Compaction debt means eligible/pending work is accumulating faster than it is completed. UCS means UnifiedCompactionStrategy; STCS, LCS and TWCS mean SizeTiered, Leveled and TimeWindow compaction strategies respectively.

Recommendation is not the same thing as default.

Current Cassandra 5.0 documentation recommends UCS for most new workloads. At the same time, an unspecified default_compaction still resolves to STCS in cassandra.yaml. A course or legacy schema that says “STCS is default” is therefore not proof that STCS is the best choice for a new table.

1. The control loop

compaction_throughput caps aggregate compaction throughput on a node; current Cassandra 5.0 configuration documents 64 MiB/s as the default. concurrent_compactors controls how many compactions can run simultaneously. More concurrency can keep small SSTables from accumulating during one long compaction, but it also increases parallel CPU/disk pressure. Current guidance explicitly suggests looking at throughput before concurrency when compaction is too slow or too fast.

nodetool compactionstats shows active/pending work; compactionhistory shows completed rewrites and bytes. A queue is not automatically bad: the question is whether debt trends upward without bound, whether SSTables/read and disk occupancy worsen, and whether application SLOs are threatened.

Signal Healthy interpretation needs Bad shortcut
Pending tasks trend vs write rate and completion rate any nonzero value is an incident
Throughput device capacity + latency SLO + repair/streaming overlap copy a blog MiB/s value
Concurrent compactors CPU/disks + queue shape + task sizes set equal to core count blindly
Disk free rewrite/snapshot/streaming worst case dataset size plus 10%
p99 latency representative load and warmup single cqlsh timing

2. Baseline before touching knobs

bash · verify the reusable local cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"
CQL · create the Chapter 12 keyspace
CREATE KEYSPACE IF NOT EXISTS atlasmart_compactionWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;
bash · 60-second observation sample
for i in 1 2 3 4 5 6; do  date -Is  docker exec atlasmart-cass-1 nodetool compactionstats  docker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra | tail -1"  docker stats --no-stream atlasmart-cass-1 | tail -1  sleep 10donedocker exec atlasmart-cass-1 nodetool compactionhistory

In a real test, pair this with client p95/p99 read/write latency and timeout/error metrics. The mandatory local lab cannot create production-quality throughput numbers; its purpose is to prove that you can record settings, alter one variable, and restore it.

3. Reversible throughput experiment

First capture the current numeric throughput. The exact text format can vary, so verify the extracted number before executing the setter. The example changes only node 1 in the disposable cluster and immediately restores the original value after observing a small compaction. Do not run this pattern blindly in production.

bash · capture, change, observe, restore
RAW=$(docker exec atlasmart-cass-1 nodetool getcompactionthroughput)echo "$RAW"ORIGINAL=$(printf '%s\n' "$RAW" | grep -Eo '[0-9]+([.][0-9]+)?' | head -1)test -n "$ORIGINAL" || { echo "Could not parse baseline; stop."; exit 1; }echo "Parsed baseline MiB/s: $ORIGINAL"# Lab-only lower cap to make the change observable; not a recommendation.docker exec atlasmart-cass-1 nodetool setcompactionthroughput 16docker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool compact atlasmart_compaction compaction_ucsdocker exec atlasmart-cass-1 nodetool compactionstats# Restore exactly the recorded baseline.docker exec atlasmart-cass-1 nodetool setcompactionthroughput "$ORIGINAL"docker exec atlasmart-cass-1 nodetool getcompactionthroughput

If parsing does not match your version's output, stop and restore manually using the value you recorded. “Automation” is not a reason to risk an unknown production setting.

4. Concurrency and headroom

bash · observe compactors without changing them
docker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_ucs

Increasing concurrent compactors can shorten a queue when the device has unused parallelism, but on a saturated disk it can worsen queueing and foreground latency. Headroom must cover more than the live dataset: include current SSTables, compaction outputs while inputs still exist, snapshots, streaming/repair/rebuild, commitlog, indexes such as SAI, backup/restore staging and failure scenarios. There is no safe universal free-space percentage.

Unsafe tuning sequence: “pending tasks high → set throughput unlimited → double compactors.”

Pending tasks are a symptom. First ask whether input write rate is temporarily high, SSTable geometry/strategy is wrong, disk is saturated, repair/streaming overlaps, or headroom is already constrained. Remove the bottleneck or adjust one control at a time with SLO stop conditions.

5. Strategy change is also a capacity event

ALTER TABLE ... WITH compaction=... looks like metadata, but current Cassandra documentation warns that changing compaction strategy on a populated table rewrites existing SSTables using the new strategy. Treat an STCS/LCS/TWCS → UCS migration as an I/O and free-space project. Test on a representative table or isolated node/JMX experiment where appropriate, roll out deliberately, and avoid colliding with repair, bootstrap, snapshot retention or backup windows.

Before change During After
capture old schema/settings + backlog + headroom watch compaction bytes, pending work, p99, disk verify strategy, backlog convergence, SSTables/read, SLOs
estimate rewrite/temporary bytes stop if free-space/latency threshold breached retain rollback/migration notes
schedule around repair/streaming/backup do not pile on major compaction compare steady-state amplification, not first-minute speed

6. Final evidence checklist

bash · capture compaction and storage evidence
docker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_ucsdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_compaction compaction_ucsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 nodetool compactionhistorydocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra && find /var/lib/cassandra/data/atlasmart_compaction -type f -name '*Data.db' -printf '%s %p\n' | sort -n | tail -20"

Check your understanding

  1. Which knob should usually be examined before increasing concurrent compactors when compaction is too slow?
  2. Why is a falling pending-task count not enough to declare tuning successful?
  3. What makes changing strategy on a populated table risky?
  4. Why is 0/unlimited compaction throughput dangerous as a copied setting?
  5. What closes the loop after tuning?
Review the answers

1. Compaction throughput, alongside actual disk/CPU/SLO evidence; Cassandra documentation specifically directs operators to look there first.

2. Foreground p99/error rate, disk/CPU pressure, SSTables/read, headroom and long-term debt must also remain acceptable.

3. Existing SSTables are rewritten using the new strategy, creating potentially large I/O and temporary-space demand.

4. It can let compaction consume storage bandwidth needed by foreground reads/writes, repair and streaming.

5. Restore/confirm intended settings, verify backlog convergence and tail SLOs under representative load, document observed amplification/headroom, and retain rollback steps.

Production judgment

Compaction tuning is capacity engineering, not a table-property beauty contest. Evaluate application write rate, overwrite/delete/TTL rate, partition and clustering distribution, read/write p50/p95/p99 latency, SSTables per read, compression ratio, pending compactions, bytes compacted, compaction throughput, disk queue/throughput, CPU, JVM/GC, network, repair and streaming load, snapshots, SAI/vector index amplification, and free-space trajectory. RF and consistency level change how many replicas incur the physical work; a locally fast compaction configuration can still violate a cluster SLO under repair, node replacement, backup, or failure.

Every production change needs an acceptance window and rollback plan. Record the old schema/options, node-by-node rollout order, expected rewrite volume, required free space, compaction backlog limit, latency/error SLOs, and stop conditions. Managed Cassandra services may hide throughput/concurrency controls or select strategies for you; map the same concepts to the provider's exposed metrics rather than assuming identical knobs. Chapter 13 now focuses on the deletion side of this story: tombstone types, TTL, gc_grace, repair timing, zombie prevention and safe purge.

Summary and next bridge

Compaction tuning is the operational balance between debt and interference. Throughput, concurrency and strategy all consume the same finite disk/CPU/headroom budget, so evidence and rollback matter more than folklore values. Chapter 13 follows the tombstones that compaction is eventually asked to purge.

Authoritative references

Use these as the version-sensitive source of truth when regenerating the lesson; compaction recommendations and options can evolve between Cassandra releases.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.