Chapter 12 · Compaction Strategies and Storage Amplification
Why Compaction Exists: Merge Immutable Files, Reclaim Tombstones, and Bound Read Amplification
Understand compaction as the merge stage that bounds immutable-file growth while trading read, write and space amplification.
Learning outcomes
AtlasMart sees its orders_by_customer table grow
from four SSTables to dozens after a burst of updates and TTL
expirations. Reads become less predictable and disk free space
falls even though the live row count barely changes. The correct
question is not “which compaction strategy is fastest?” but
“what immutable versions now exist, what work must be merged,
and which amplification dimension is hurting the workload?”
Trace why immutable SSTables require background compaction after updates, deletes, TTL and flushes.
Measure read, write and space amplification instead of using SSTable count alone.
Explain when a tombstone can be purged and why compaction cannot simply delete every expired marker.
Compare UCS with the legacy/specialized STCS/LCS/TWCS choices without treating any as timeless defaults.
Use compactionstats/history, tablestats, tablehistograms and disk evidence before/after a disposable compaction.
The mandatory labs continue the disposable AtlasMart cluster
used by Chapters 01–11: pinned cassandra:5.0.9,
Docker network atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and Java 17 inside the image. Chapter 12 uses keyspace
atlasmart_compaction with
NetworkTopologyStrategy, replication factor (RF)
3, and normally LOCAL_QUORUM. Authentication,
client TLS, internode TLS, and remote JMX are disabled only
inside this isolated learning network. The Apache Cassandra
Java Driver 4.19.3 is optional; mandatory evidence uses
cqlsh, nodetool, Docker/Linux
filesystem tools, and Cassandra metrics. Unless a lesson
explicitly creates an STCS/LCS/TWCS comparison table, new
tables use UnifiedCompactionStrategy (UCS), reflecting current
Cassandra 5.0 guidance. Keep gc_grace_seconds at
its default unless a disposable purge experiment explicitly
changes it and performs repair first. Record the actual
compaction throughput, concurrent compactors, disk free space,
compression ratio, SSTable count, workload concurrency,
partition sizes, and latency distribution before changing
anything.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Compaction terms before changing a table
An SSTable (Sorted String Table) is immutable on-disk data. A memtable flush creates a new SSTable; updates and deletes therefore create newer versions instead of modifying older files in place. Compaction selects SSTables, reads their sorted contents, reconciles newer cells and tombstones, writes replacement SSTables, and retires input files after active readers no longer need them. Read amplification is the extra storage work a logical read performs across files/versions. Write amplification is the number of bytes rewritten in the background relative to application bytes written. Space amplification is temporary or steady-state disk usage above live logical data, including overlapping SSTables, replacement outputs, snapshots, streaming, repair and backup staging. A tombstone is a timestamped deletion/expiration marker; compaction can purge it only when Cassandra can do so safely. Compaction debt means eligible/pending work is accumulating faster than it is completed. UCS means UnifiedCompactionStrategy; STCS, LCS and TWCS mean SizeTiered, Leveled and TimeWindow compaction strategies respectively.
Current Cassandra 5.0 documentation recommends UCS for most
new workloads. At the same time, an unspecified
default_compaction still resolves to STCS in
cassandra.yaml. A course or legacy schema that
says “STCS is default” is therefore not proof that STCS is the
best choice for a new table.
1. What compaction actually changes
Cassandra's write path favors append-style immutable storage: application writes become memtable state and later new SSTables. If a cell is updated five times, those physical versions can coexist until compaction reconciles them. If a row is deleted, a tombstone must suppress older replicas/files until it is safe to purge. Compaction performs a sorted merge, writes replacement SSTables, then retires inputs. It can reduce read amplification and reclaim obsolete bytes, but it creates write amplification and requires temporary disk headroom while old and new files overlap.
| Amplification | Observable question | Typical evidence |
|---|---|---|
| Read | How many SSTables/chunks/versions does one logical read touch? | tablehistograms, tracing, SSTables/read, tombstones scanned |
| Write | How many background bytes are rewritten for each foreground byte? | compactionhistory, bytes compacted, disk write throughput |
| Space | How much storage exists beyond live logical data? | df/du, snapshots, SSTable bytes, compaction/streaming temp state |
Tombstone removal is conditional, not automatic. Cassandra must
account for gc_grace_seconds, repaired/unrepaired
state, overlapping older data and replica convergence. A
compaction may rewrite an SSTable and still retain tombstones
because dropping them would risk resurrecting older data.
2. Current 5.0 strategy context
| Strategy | Current role | Core tradeoff |
|---|---|---|
| UCS | recommended starting point for most new workloads | tunable tiered/leveled behavior with sharding; evaluate measured amplification |
| STCS | supported default/fallback and common legacy write-oriented choice | lower write cost but potentially more overlapping SSTables/read amplification |
| LCS | supported read-heavy / predictable-read specialization | bounds overlap above L0 but rewrites data through levels |
| TWCS | supported mostly immutable TTL/time-series specialization | excellent expiration locality when timestamps/TTL are disciplined; poor fit for mutable/out-of-order data |
This chapter preserves the published STCS/LCS/TWCS titles because operators still inherit those tables. But every lesson asks whether UCS is the better new-table starting point on Cassandra 5.0.
3. Lab: create deliberate immutable generations
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"
CREATE KEYSPACE IF NOT EXISTS atlasmart_compactionWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;
CREATE TABLE IF NOT EXISTS atlasmart_compaction.compaction_ucs ( account_id text, day date, event_time timestamp, event_id uuid, status text, note text, PRIMARY KEY ((account_id,day),event_time,event_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'};INSERT INTO atlasmart_compaction.compaction_ucs VALUES ('acct-42','2026-09-07','2026-09-07T18:00:00Z',00000000-0000-0000-0000-000000000001,'CREATED','v1');
docker exec atlasmart-cass-1 nodetool flush atlasmart_compaction compaction_ucsdocker exec atlasmart-cass-2 nodetool flush atlasmart_compaction compaction_ucsdocker exec atlasmart-cass-3 nodetool flush atlasmart_compaction compaction_ucs
UPDATE atlasmart_compaction.compaction_ucs SET status='PAID', note='v2'WHERE account_id='acct-42' AND day='2026-09-07' AND event_time='2026-09-07T18:00:00Z'AND event_id=00000000-0000-0000-0000-000000000001;DELETE note FROM atlasmart_compaction.compaction_ucsWHERE account_id='acct-42' AND day='2026-09-07' AND event_time='2026-09-07T18:00:00Z'AND event_id=00000000-0000-0000-0000-000000000001;
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_compaction compaction_ucs; donedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_ucsdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_compaction compaction_ucsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 nodetool compactionhistorydocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"
4. Controlled before/after compaction
For this tiny disposable fixture, a manual compaction is useful because it makes the immutable-file rewrite visible. It is not a production performance prescription. Capture SSTable counts, disk bytes, read latency samples and compaction history before and after. Exact file counts can change before you intervene because background compaction is asynchronous.
docker exec atlasmart-cass-1 nodetool compact atlasmart_compaction compaction_ucsdocker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_ucsdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_compaction compaction_ucsdocker exec atlasmart-cass-1 nodetool compactionhistorydocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"
A manual major/forced compaction can generate a large rewrite, consume free space and I/O, interfere with normal compaction and create a new large SSTable that later has expensive overlap. Diagnose the backlog/model first; production compaction is usually background policy, not incident muscle memory.
5. What success proves—and does not prove
- Fewer SSTables after compaction proves physical consolidation, not that the strategy is optimal under production load.
- A lower SSTables/read sample can support a read-amplification hypothesis, but cache warmth and tiny data can dominate laptop timing.
- Disk free space may not fall or rise exactly as expected because compression, snapshots, active readers and asynchronous file deletion affect physical bytes.
- A tombstone remaining after compaction is not a bug if purge-safety conditions are not met.
Check your understanding
- Why is compaction unavoidable in an immutable SSTable engine?
- What is the central cost of compaction?
- Does compaction always purge tombstones?
- Why is UCS emphasized despite STCS being a default when no strategy is specified?
- What must be captured before a strategy experiment?
Review the answers
1. Updates/deletes create newer physical versions; compaction eventually merges versions, bounds file/read growth and reclaims obsolete bytes when safe.
2. Background rereads and rewrites create write amplification and temporary space/I/O demand.
3. No. Purge requires safety conditions such as grace/overlap/repair state; a compaction can legitimately retain them.
4. Cassandra 5.0 documentation recommends UCS for most new workloads; implementation defaults and current recommendations are different concepts.
5. Schema/strategy, SSTable count/size, pending work, throughput/concurrency, disk headroom, latency distributions, workload/partition shape, repair/snapshot/streaming context.
Production judgment
Compaction tuning is capacity engineering, not a table-property beauty contest. Evaluate application write rate, overwrite/delete/TTL rate, partition and clustering distribution, read/write p50/p95/p99 latency, SSTables per read, compression ratio, pending compactions, bytes compacted, compaction throughput, disk queue/throughput, CPU, JVM/GC, network, repair and streaming load, snapshots, SAI/vector index amplification, and free-space trajectory. RF and consistency level change how many replicas incur the physical work; a locally fast compaction configuration can still violate a cluster SLO under repair, node replacement, backup, or failure.
Every production change needs an acceptance window and rollback plan. Record the old schema/options, node-by-node rollout order, expected rewrite volume, required free space, compaction backlog limit, latency/error SLOs, and stop conditions. Managed Cassandra services may hide throughput/concurrency controls or select strategies for you; map the same concepts to the provider's exposed metrics rather than assuming identical knobs. Lesson 2 examines STCS as an inherited or explicitly chosen strategy and compares its size-tiered behavior to a UCS starting point.
Summary and next bridge
Compaction is the merge stage of Cassandra's log-structured storage model. It trades background I/O and space for bounded read cost and safe reclamation. Next, examine STCS on its own terms—then judge it against current UCS guidance rather than historical habit.
Authoritative references
Use these as the version-sensitive source of truth when regenerating the lesson; compaction recommendations and options can evolve between Cassandra releases.