Chapter 12 · Compaction Strategies and Storage Amplification
LeveledCompactionStrategy for Read Predictability and Its Write / Space Amplification Tradeoffs
Explain LCS level invariants and predictable reads alongside their repeated rewrite and space/I/O costs.
Learning outcomes
AtlasMart's product lookup table has a strict p99 read target. An older design selected LeveledCompactionStrategy (LCS) to bound overlapping SSTables, but storage writes are much higher than application writes. The team needs to understand that trade instead of assuming “read-optimized” means universally efficient.
Explain L0 and higher LCS levels, overlap guarantees and fanout behavior.
Connect predictable SSTables/read to write amplification through repeated promotion/rewrites.
Create a small LCS fixture and inspect compaction history/table metrics rather than inferring level state from latency alone.
Explain LCS space characteristics and very-large-partition edge cases.
Compare LCS with UCS configured toward read-heavy behavior and current 5.0 recommendations.
The mandatory labs continue the disposable AtlasMart cluster
used by Chapters 01–11: pinned cassandra:5.0.9,
Docker network atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and Java 17 inside the image. Chapter 12 uses keyspace
atlasmart_compaction with
NetworkTopologyStrategy, replication factor (RF)
3, and normally LOCAL_QUORUM. Authentication,
client TLS, internode TLS, and remote JMX are disabled only
inside this isolated learning network. The Apache Cassandra
Java Driver 4.19.3 is optional; mandatory evidence uses
cqlsh, nodetool, Docker/Linux
filesystem tools, and Cassandra metrics. Unless a lesson
explicitly creates an STCS/LCS/TWCS comparison table, new
tables use UnifiedCompactionStrategy (UCS), reflecting current
Cassandra 5.0 guidance. Keep gc_grace_seconds at
its default unless a disposable purge experiment explicitly
changes it and performs repair first. Record the actual
compaction throughput, concurrent compactors, disk free space,
compression ratio, SSTable count, workload concurrency,
partition sizes, and latency distribution before changing
anything.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Compaction terms before changing a table
An SSTable (Sorted String Table) is immutable on-disk data. A memtable flush creates a new SSTable; updates and deletes therefore create newer versions instead of modifying older files in place. Compaction selects SSTables, reads their sorted contents, reconciles newer cells and tombstones, writes replacement SSTables, and retires input files after active readers no longer need them. Read amplification is the extra storage work a logical read performs across files/versions. Write amplification is the number of bytes rewritten in the background relative to application bytes written. Space amplification is temporary or steady-state disk usage above live logical data, including overlapping SSTables, replacement outputs, snapshots, streaming, repair and backup staging. A tombstone is a timestamped deletion/expiration marker; compaction can purge it only when Cassandra can do so safely. Compaction debt means eligible/pending work is accumulating faster than it is completed. UCS means UnifiedCompactionStrategy; STCS, LCS and TWCS mean SizeTiered, Leveled and TimeWindow compaction strategies respectively.
Current Cassandra 5.0 documentation recommends UCS for most
new workloads. At the same time, an unspecified
default_compaction still resolves to STCS in
cassandra.yaml. A course or legacy schema that
says “STCS is default” is therefore not proof that STCS is the
best choice for a new table.
1. LCS mechanism: pay writes to bound overlap
Newly flushed SSTables enter level 0 (L0), where overlap is allowed. Compaction moves data into higher levels. Within a level above L0, SSTables are non-overlapping by token range, so a read generally needs at most one SSTable per level for a partition key, plus possible L0 candidates. Level capacity grows by a fanout factor. Maintaining this invariant means data can be rewritten repeatedly as it moves through levels: lower read amplification is purchased with higher write amplification and sustained compaction I/O.
A very large partition is an important boundary: Cassandra does not split one partition across arbitrary SSTables just to honor the configured LCS target size, so simplistic “every SSTable is exactly 160 MB” diagrams are wrong.
2. Create an LCS fixture
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"
CREATE KEYSPACE IF NOT EXISTS atlasmart_compactionWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;
CREATE TABLE IF NOT EXISTS atlasmart_compaction.compaction_lcs ( sku text, version_time timestamp, version_id uuid, price decimal, payload text, PRIMARY KEY (sku,version_time,version_id)) WITH compaction = {'class':'LeveledCompactionStrategy','sstable_size_in_mb':'160'};
for i in 1 2 3 4 5; do ID=$(printf "%012d" $i) docker exec atlasmart-cass-1 cqlsh -e "INSERT INTO atlasmart_compaction.compaction_lcs (sku,version_time,version_id,price,payload) VALUES ('sku-42','2026-09-07T18:0${i}:00Z',00000000-0000-0000-0000-${ID},${i}.0,'g${i}');" for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_compaction compaction_lcs; donedonefor i in 1 2 3 4 5; do docker exec atlasmart-cass-1 nodetool compactionstats; sleep 2; donedocker exec atlasmart-cass-1 nodetool compactionhistorydocker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_lcsdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_compaction compaction_lcs
3. Read predictability versus write/space amplification
| Effect | Why LCS does it | Measure |
|---|---|---|
| Fewer overlapping files at higher levels | non-overlap invariant per level | SSTables/read histogram, trace |
| More background rewrites | promotion merges with next-level overlapping ranges | compactionhistory bytes, disk writes |
| More uniform high-level I/O | fixed target sizes and level capacity | SSTable sizes, compaction duration |
| Temporary rewrite space | inputs coexist with outputs during compaction | df/du trajectory and headroom |
Do not benchmark LCS versus UCS with only five tiny rows. The lab proves mechanics and observability. A meaningful comparison needs representative partition sizes, update rate, read distribution, compression, concurrency and enough duration for the compaction steady state to emerge.
4. Current UCS comparison
UCS can express more tiered or more leveled behavior using scaling parameters and is recommended for most new Cassandra 5.0 workloads. That makes it the first comparison point for a new read-sensitive table. Existing LCS remains valid when its measured behavior, migration rewrite cost and operational history are favorable. Do not translate “UCS recommended” into “ALTER every table now”: current documentation warns that changing the compaction strategy of an existing table can rewrite existing SSTables and take hours.
The recommendation is for selection, not a waiver of migration physics. Estimate rewrite bytes, disk headroom, compaction debt, repair/snapshot overlap and p99 risk. Test a representative table/node first and define stop/rollback criteria.
5. Verification
docker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_lcsdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_compaction compaction_lcsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 nodetool compactionhistorydocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra && find /var/lib/cassandra/data/atlasmart_compaction -type f -name '*Data.db' -printf '%s %p\n' | sort -n | tail -20"
Check your understanding
- Why does LCS reduce SSTable overlap above L0?
- What is the price of that read predictability?
- Can one huge partition violate the target SSTable size intuition?
- Why is UCS relevant to an LCS lesson?
- What evidence should justify keeping LCS?
Review the answers
1. It organizes data into levels where SSTables in a level are non-overlapping by token range.
2. Repeated compaction and promotion cause additional write amplification and background I/O.
3. Yes. Cassandra does not arbitrarily split a single partition solely to meet the LCS SSTable target size.
4. Cassandra 5.0 recommends UCS for most new workloads and it can be tuned toward leveled behavior, so it is the modern comparison baseline.
5. Representative read/write tail latency, amplification, compaction debt, disk headroom and operational stability—not the strategy name alone.
Production judgment
Compaction tuning is capacity engineering, not a table-property beauty contest. Evaluate application write rate, overwrite/delete/TTL rate, partition and clustering distribution, read/write p50/p95/p99 latency, SSTables per read, compression ratio, pending compactions, bytes compacted, compaction throughput, disk queue/throughput, CPU, JVM/GC, network, repair and streaming load, snapshots, SAI/vector index amplification, and free-space trajectory. RF and consistency level change how many replicas incur the physical work; a locally fast compaction configuration can still violate a cluster SLO under repair, node replacement, backup, or failure.
Every production change needs an acceptance window and rollback plan. Record the old schema/options, node-by-node rollout order, expected rewrite volume, required free space, compaction backlog limit, latency/error SLOs, and stop conditions. Managed Cassandra services may hide throughput/concurrency controls or select strategies for you; map the same concepts to the provider's exposed metrics rather than assuming identical knobs. Lesson 4 shifts to TWCS, where time windows and TTL can make whole expired SSTables disposable—but only if data is mostly immutable and timestamp/retention discipline is strong.
Summary and next bridge
LCS is a deliberate exchange: more background rewrites for constrained file overlap and more predictable reads. It remains an important inherited/specialized strategy, but UCS is the current new-table comparison. Next, examine the very different retention-driven logic of TWCS.
Authoritative references
Use these as the version-sensitive source of truth when regenerating the lesson; compaction recommendations and options can evolve between Cassandra releases.
- Apache Cassandra releases
- Compaction overview and strategies
- Compaction CQL subproperties including UCS
- ALTER TABLE and compaction-strategy rewrite warning
- cassandra.yaml compaction throughput/concurrency/defaults
- Storage engine / LSM write amplification
- Repair and replica convergence
- LeveledCompactionStrategy