Chapter 12 · Compaction Strategies and Storage Amplification
SizeTieredCompactionStrategy for General Write-Oriented Workloads
Trace STCS size buckets and thresholds, then judge inherited write-oriented behavior against current UCS guidance.
Learning outcomes
AtlasMart inherits a large table created years ago with SizeTieredCompactionStrategy (STCS). Writes are smooth, but the team sees bursts of disk use and a growing number of similarly sized SSTables during peak ingestion. Before migrating anything, they need to understand exactly how STCS groups files and why its read/space behavior emerges.
Explain STCS size buckets, min/max thresholds and write-oriented behavior.
Create comparable-size SSTables with auto-compaction paused, then observe STCS selection after re-enable.
Measure pending tasks, SSTable counts, bytes compacted and read amplification before/after.
Explain why STCS can require substantial temporary headroom and why “default” is not “recommended for every new table.”
Compare inherited STCS behavior with a UCS new-table baseline without forcing a production migration.
The mandatory labs continue the disposable AtlasMart cluster
used by Chapters 01–11: pinned cassandra:5.0.9,
Docker network atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and Java 17 inside the image. Chapter 12 uses keyspace
atlasmart_compaction with
NetworkTopologyStrategy, replication factor (RF)
3, and normally LOCAL_QUORUM. Authentication,
client TLS, internode TLS, and remote JMX are disabled only
inside this isolated learning network. The Apache Cassandra
Java Driver 4.19.3 is optional; mandatory evidence uses
cqlsh, nodetool, Docker/Linux
filesystem tools, and Cassandra metrics. Unless a lesson
explicitly creates an STCS/LCS/TWCS comparison table, new
tables use UnifiedCompactionStrategy (UCS), reflecting current
Cassandra 5.0 guidance. Keep gc_grace_seconds at
its default unless a disposable purge experiment explicitly
changes it and performs repair first. Record the actual
compaction throughput, concurrent compactors, disk free space,
compression ratio, SSTable count, workload concurrency,
partition sizes, and latency distribution before changing
anything.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Compaction terms before changing a table
An SSTable (Sorted String Table) is immutable on-disk data. A memtable flush creates a new SSTable; updates and deletes therefore create newer versions instead of modifying older files in place. Compaction selects SSTables, reads their sorted contents, reconciles newer cells and tombstones, writes replacement SSTables, and retires input files after active readers no longer need them. Read amplification is the extra storage work a logical read performs across files/versions. Write amplification is the number of bytes rewritten in the background relative to application bytes written. Space amplification is temporary or steady-state disk usage above live logical data, including overlapping SSTables, replacement outputs, snapshots, streaming, repair and backup staging. A tombstone is a timestamped deletion/expiration marker; compaction can purge it only when Cassandra can do so safely. Compaction debt means eligible/pending work is accumulating faster than it is completed. UCS means UnifiedCompactionStrategy; STCS, LCS and TWCS mean SizeTiered, Leveled and TimeWindow compaction strategies respectively.
Current Cassandra 5.0 documentation recommends UCS for most
new workloads. At the same time, an unspecified
default_compaction still resolves to STCS in
cassandra.yaml. A course or legacy schema that
says “STCS is default” is therefore not proof that STCS is the
best choice for a new table.
1. STCS mechanism
STCS groups SSTables into size buckets. When enough similarly
sized SSTables are available—commonly governed by
min_threshold and max_threshold—it
merges a subset into a larger SSTable. The strategy tends to
write a generation a relatively small number of times compared
with leveled approaches, which is attractive for heavy writes,
but reads can face several overlapping files. Temporary disk
demand can also be high because input SSTables remain until the
output is safely installed.
| Property | Meaning | Do not infer |
|---|---|---|
| min_threshold | minimum candidate SSTables before normal minor compaction | that lowering it always reduces latency safely |
| max_threshold | caps how many candidates can join a minor compaction | that larger always means better throughput |
| bucket_low / bucket_high | relative size range used for grouping | that defaults fit every compression/data distribution |
| min_sstable_size | small-file bucketing floor | that tiny lab SSTables model production accurately |
2. Build a deterministic STCS queue
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"
CREATE KEYSPACE IF NOT EXISTS atlasmart_compactionWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;
CREATE TABLE IF NOT EXISTS atlasmart_compaction.compaction_stcs ( bucket text, event_time timestamp, event_id uuid, payload text, PRIMARY KEY (bucket,event_time,event_id)) WITH compaction = {'class':'SizeTieredCompactionStrategy','min_threshold':'4','max_threshold':'32'};
docker exec atlasmart-cass-1 nodetool disableautocompaction atlasmart_compaction compaction_stcsdocker exec atlasmart-cass-2 nodetool disableautocompaction atlasmart_compaction compaction_stcsdocker exec atlasmart-cass-3 nodetool disableautocompaction atlasmart_compaction compaction_stcs
INSERT INTO atlasmart_compaction.compaction_stcs VALUES ('b','2026-09-07T18:00:00Z',00000000-0000-0000-0000-000000000001,'generation-1');
for i in 1 2 3 4; do ID=$(printf "%012d" $i) docker exec atlasmart-cass-1 cqlsh -e "INSERT INTO atlasmart_compaction.compaction_stcs (bucket,event_time,event_id,payload) VALUES ('b','2026-09-07T18:0${i}:00Z',00000000-0000-0000-0000-${ID},'g${i}');" for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_compaction compaction_stcs; donedonedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_stcsdocker exec atlasmart-cass-1 nodetool compactionstats
3. Re-enable and observe selection
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool enableautocompaction atlasmart_compaction compaction_stcs; donefor i in 1 2 3 4 5; do docker exec atlasmart-cass-1 nodetool compactionstats sleep 2donedocker exec atlasmart-cass-1 nodetool compactionhistorydocker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_stcsdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_compaction compaction_stcs
On a tiny cluster, background compaction may finish before the first poll. That is not a failed lab: compare history and SSTable counts. On production-scale files, selection and rewrite can be long-running, and the critical signal is whether pending work grows faster than throughput clears it.
4. STCS tradeoffs versus UCS
| Question | STCS | UCS-oriented decision |
|---|---|---|
| Write amplification | often attractive for write-heavy general workloads | tunable toward tiered behavior while retaining a unified framework |
| Read amplification | can be higher because multiple tiers/files overlap | can tune scaling behavior based on measured read/write balance |
| Space during compaction | can need substantial headroom for large tier merges | still needs headroom; sharding/target sizing changes rewrite geometry |
| New Cassandra 5.0 table | supported but contextual choice | recommended starting point for most workloads |
The first clause can be true when no table/default_compaction override is set, while the conclusion is false. Current guidance recommends UCS for new workloads. Keep STCS where measured behavior and migration risk justify it; do not turn a legacy default into policy.
5. Verification and cleanup
- Auto-compaction is re-enabled on all three nodes.
- The table schema still explicitly says STCS.
- Capture before/after SSTable counts, compaction history and disk free space.
- Do not change global thresholds/throughput as part of this lesson.
Check your understanding
- What does STCS select primarily by?
- Why can STCS be write-friendly yet read-expensive?
- Why pause autocompaction in this lab?
- Is STCS being the unspecified default the same as Cassandra recommending it for new tables?
- What would block an STCS-to-UCS production migration?
Review the answers
1. Groups of SSTables with similar sizes, subject to strategy thresholds/options.
2. It avoids continuous leveling rewrites but can leave several overlapping SSTables that a read may consult.
3. Only to build observable candidate SSTables deterministically; it must be re-enabled because disabling compaction long term creates debt and space/read risk.
4. No. Cassandra 5.0 recommends UCS for most new workloads.
5. Insufficient disk/I/O headroom, rewrite volume, latency SLO risk, repair/backup/topology overlap, or lack of rollback/measurement plan.
Production judgment
Compaction tuning is capacity engineering, not a table-property beauty contest. Evaluate application write rate, overwrite/delete/TTL rate, partition and clustering distribution, read/write p50/p95/p99 latency, SSTables per read, compression ratio, pending compactions, bytes compacted, compaction throughput, disk queue/throughput, CPU, JVM/GC, network, repair and streaming load, snapshots, SAI/vector index amplification, and free-space trajectory. RF and consistency level change how many replicas incur the physical work; a locally fast compaction configuration can still violate a cluster SLO under repair, node replacement, backup, or failure.
Every production change needs an acceptance window and rollback plan. Record the old schema/options, node-by-node rollout order, expected rewrite volume, required free space, compaction backlog limit, latency/error SLOs, and stop conditions. Managed Cassandra services may hide throughput/concurrency controls or select strategies for you; map the same concepts to the provider's exposed metrics rather than assuming identical knobs. Lesson 3 examines LCS, where predictable overlap above L0 is purchased with repeated leveling rewrites and different space/I/O characteristics.
Summary and next bridge
STCS explains many inherited Cassandra fleets: bucket similar sizes, merge opportunistically, accept more overlap in exchange for a write-oriented profile. The right response is measurement, not automatic migration or automatic retention. Next, contrast this with LCS.
Authoritative references
Use these as the version-sensitive source of truth when regenerating the lesson; compaction recommendations and options can evolve between Cassandra releases.