Chapter 12 · Compaction Strategies and Storage Amplification
TimeWindowCompactionStrategy for TTL / Time-Series Workloads and Expiring Data
Use TWCS only when TTL, timestamp locality and mostly immutable time-series behavior make whole-window expiry safe and useful.
Learning outcomes
AtlasMart keeps clickstream events for seven days. The table is append-mostly, every event expires, and historical rows are rarely updated. That is the shape TWCS was designed for. But a separate team copied TWCS onto a mutable order table with no TTL and later discovered old/new timestamps mixed across SSTables, defeating efficient expiration.
Explain TWCS time windows, active-window STCS behavior and whole-expired-SSTable reclamation.
Connect TTL uniformity, mostly immutable writes and timestamp locality to TWCS success.
Explain out-of-order/mutable-data contamination and the zero-TTL guardrail.
Build a short-window disposable lab and inspect schema/SSTable/compaction evidence.
Compare TWCS with UCS for modern Cassandra 5.0 time-series workloads instead of assuming TWCS automatically wins.
The mandatory labs continue the disposable AtlasMart cluster
used by Chapters 01–11: pinned cassandra:5.0.9,
Docker network atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and Java 17 inside the image. Chapter 12 uses keyspace
atlasmart_compaction with
NetworkTopologyStrategy, replication factor (RF)
3, and normally LOCAL_QUORUM. Authentication,
client TLS, internode TLS, and remote JMX are disabled only
inside this isolated learning network. The Apache Cassandra
Java Driver 4.19.3 is optional; mandatory evidence uses
cqlsh, nodetool, Docker/Linux
filesystem tools, and Cassandra metrics. Unless a lesson
explicitly creates an STCS/LCS/TWCS comparison table, new
tables use UnifiedCompactionStrategy (UCS), reflecting current
Cassandra 5.0 guidance. Keep gc_grace_seconds at
its default unless a disposable purge experiment explicitly
changes it and performs repair first. Record the actual
compaction throughput, concurrent compactors, disk free space,
compression ratio, SSTable count, workload concurrency,
partition sizes, and latency distribution before changing
anything.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Compaction terms before changing a table
An SSTable (Sorted String Table) is immutable on-disk data. A memtable flush creates a new SSTable; updates and deletes therefore create newer versions instead of modifying older files in place. Compaction selects SSTables, reads their sorted contents, reconciles newer cells and tombstones, writes replacement SSTables, and retires input files after active readers no longer need them. Read amplification is the extra storage work a logical read performs across files/versions. Write amplification is the number of bytes rewritten in the background relative to application bytes written. Space amplification is temporary or steady-state disk usage above live logical data, including overlapping SSTables, replacement outputs, snapshots, streaming, repair and backup staging. A tombstone is a timestamped deletion/expiration marker; compaction can purge it only when Cassandra can do so safely. Compaction debt means eligible/pending work is accumulating faster than it is completed. UCS means UnifiedCompactionStrategy; STCS, LCS and TWCS mean SizeTiered, Leveled and TimeWindow compaction strategies respectively.
Current Cassandra 5.0 documentation recommends UCS for most
new workloads. At the same time, an unspecified
default_compaction still resolves to STCS in
cassandra.yaml. A course or legacy schema that
says “STCS is default” is therefore not proof that STCS is the
best choice for a new table.
1. TWCS mechanism
TWCS groups SSTables by the maximum timestamp of data they contain. Within the current time window, it uses size-tiered-style compaction to consolidate files. Once a window becomes historical, its files stop being mixed with newer windows. If an old SSTable becomes fully expired and no older live data requires its tombstones to shadow it, Cassandra can drop the whole file efficiently. This is powerful for TTL'd, mostly immutable time-series data.
It is fragile when application behavior violates the model. Late
writes with old timestamps can contaminate a newer SSTable;
mutable updates can place versions in different windows; mixed
TTL and non-TTL data can keep files from becoming fully expired.
Cassandra 5.0 includes a guardrail that warns about TWCS with
default_time_to_live=0 because a time window does
not itself expire data.
2. Short-window lab table
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra"
CREATE KEYSPACE IF NOT EXISTS atlasmart_compactionWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;
CREATE TABLE IF NOT EXISTS atlasmart_compaction.compaction_twcs ( sensor_id text, minute_bucket timestamp, event_time timestamp, event_id uuid, reading double, PRIMARY KEY ((sensor_id,minute_bucket),event_time,event_id)) WITH compaction = { 'class':'TimeWindowCompactionStrategy', 'compaction_window_unit':'MINUTES', 'compaction_window_size':'1'} AND default_time_to_live = 300;DESCRIBE TABLE atlasmart_compaction.compaction_twcs;
A one-minute window and five-minute TTL are intentionally tiny so the mechanism is observable. They are not recommendations. Production windows must be chosen from write rate, retention, SSTable size, repair cadence and query patterns.
for i in 1 2 3; do docker exec atlasmart-cass-1 cqlsh -e "INSERT INTO atlasmart_compaction.compaction_twcs (sensor_id,minute_bucket,event_time,event_id,reading) VALUES ('sensor-42','2026-09-07T19:00:00Z',dateOf(now()),now(),42.${i});" for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_compaction compaction_twcs; done sleep 2donedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_twcsdocker exec atlasmart-cass-1 nodetool compactionstats
3. The out-of-order boundary
TWCS decisions use cell timestamps, not your business
event_time column by itself. Clients that set
arbitrary USING TIMESTAMP values can therefore
place old logical versions into SSTables created now. Repairs
and streaming can also introduce historical data. The safe
operating rule is to understand timestamp sources and avoid
mixing old and new versions in the same SSTable when depending
on whole-file expiration.
-- This explicit timestamp is intentionally old and exists only to demonstrate TWCS risk.INSERT INTO atlasmart_compaction.compaction_twcs(sensor_id,minute_bucket,event_time,event_id,reading)USING TIMESTAMP 1757268000000000 AND TTL 300VALUES ('sensor-42','2026-09-07T18:00:00Z','2026-09-07T18:00:00Z',00000000-0000-0000-0000-000000000099,99.0);
Windows organize compaction; TTL expires data. A TWCS table with zero/default-no TTL does not magically enforce retention. Current Cassandra has a guardrail warning for this suspicious combination. Keep retention semantics explicit.
4. Reclamation and tombstones
After TTL expires, cells become tombstones and remain subject to
gc_grace_seconds and purge safety. Do not enable
unsafe_aggressive_sstable_expiration in this
course: its name is accurate—it relaxes shadowing checks and can
lose data if your assumptions are wrong. For a safe lab, inspect
sstableexpiredblockers or normal compaction
evidence and wait for the standard safety conditions. Chapter 13
will go much deeper into tombstone/grace/repair timing.
docker exec atlasmart-cass-1 nodetool tablestats atlasmart_compaction.compaction_twcsdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_compaction compaction_twcsdocker exec atlasmart-cass-1 nodetool compactionstatsdocker exec atlasmart-cass-1 nodetool compactionhistorydocker exec atlasmart-cass-1 nodetool getcompactionthroughputdocker exec atlasmart-cass-1 nodetool getconcurrentcompactorsdocker exec atlasmart-cass-1 sh -lc "df -h /var/lib/cassandra && find /var/lib/cassandra/data/atlasmart_compaction -type f -name '*Data.db' -printf '%s %p\n' | sort -n | tail -20"
5. TWCS versus UCS decision
| Workload property | TWCS implication | UCS comparison |
|---|---|---|
| Mostly immutable + uniform TTL | strong TWCS fit | UCS is still a current general-purpose starting point worth benchmarking |
| Frequent updates to old data | risks window contamination | UCS often simpler/safer |
| Mixed TTL/non-TTL | can block whole-file expiry | model/retention issue regardless of strategy |
| Late/out-of-order timestamps | can mix historical/new data | requires timestamp discipline either way |
| SAI/vector indexes | SSTable count also affects index work | include index amplification in benchmark |
Check your understanding
- What makes TWCS efficient for expiring time-series data?
- Does the compaction window expire data?
- Why are mutable old rows dangerous for TWCS?
- Why avoid unsafe_aggressive_sstable_expiration in the lab?
- Why still compare TWCS with UCS in Cassandra 5.0?
Review the answers
1. It isolates data by time windows so fully expired historical SSTables can often be dropped as units instead of repeatedly merging old/new data.
2. No. TTL/retention semantics expire data; the window only organizes SSTables.
3. New versions with old timestamps can contaminate windows/SSTables and prevent clean whole-file expiry.
4. It removes safety checks around shadowed data and can cause data loss when assumptions are wrong.
5. UCS is the current recommended strategy for most new workloads and can serve time-series patterns; benchmark representative retention/read/write amplification before specializing.
Production judgment
Compaction tuning is capacity engineering, not a table-property beauty contest. Evaluate application write rate, overwrite/delete/TTL rate, partition and clustering distribution, read/write p50/p95/p99 latency, SSTables per read, compression ratio, pending compactions, bytes compacted, compaction throughput, disk queue/throughput, CPU, JVM/GC, network, repair and streaming load, snapshots, SAI/vector index amplification, and free-space trajectory. RF and consistency level change how many replicas incur the physical work; a locally fast compaction configuration can still violate a cluster SLO under repair, node replacement, backup, or failure.
Every production change needs an acceptance window and rollback plan. Record the old schema/options, node-by-node rollout order, expected rewrite volume, required free space, compaction backlog limit, latency/error SLOs, and stop conditions. Managed Cassandra services may hide throughput/concurrency controls or select strategies for you; map the same concepts to the provider's exposed metrics rather than assuming identical knobs. Lesson 5 stops comparing strategy names and builds the operational control loop: throughput, compactors, pending tasks, headroom and SLO-driven tuning.
Summary and next bridge
TWCS can turn retention into an elegant physical layout when data is truly TTL'd, mostly immutable and timestamp-disciplined. It is not a generic “time-series switch.” The final lesson treats compaction as a measured node-level resource scheduler rather than a schema-only choice.
Authoritative references
Use these as the version-sensitive source of truth when regenerating the lesson; compaction recommendations and options can evolve between Cassandra releases.