Chapter 13 · Tombstones, TTL, gc_grace, Deletes, and Expiration
Redesign High-Delete / High-TTL Workloads with Bucketing, TWCS, and Retention-Aware Modeling
Redesign expiring workloads with bounded buckets, TWCS where appropriate, and retention-aware repair/grace policy.
Learning outcomes
AtlasMart stores short-lived clickstream/session events. The original table keeps years of event history in a tenant partition and relies on TTL for cleanup. That combines high write rate, high expiration rate, giant partitions, and mixed-age SSTables—the exact pattern that turns automatic expiration into read/compaction debt. The redesign aligns partition buckets, TTL horizon, and a specialized compaction strategy.
Redesign an unbounded high-TTL partition into bounded time buckets that match read and retention requirements.
Explain when TWCS is a reasonable specialized choice even though UCS is recommended for most new Cassandra 5.0 workloads.
Show how uniform TTL and time-windowed SSTables can make fully expired file reclamation easier.
Identify TWCS hazards: out-of-order writes, mixed TTLs, mutable old data, read repair of old data, and unsafe aggressive expiration.
Compare before/after partition shape, TTL countdown, SSTable counts, read scans, and disk behavior without claiming a tiny lab is a production benchmark.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the image,
cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. The chapter keyspace is
atlasmart_tombstones with
NetworkTopologyStrategy and replication factor
(RF) 3; reads and writes normally use
LOCAL_QUORUM. Authentication, client TLS,
internode TLS, and remote JMX remain disabled only inside this
isolated local learning network. UnifiedCompactionStrategy
(UCS) remains the default choice for ordinary chapter tables;
Lesson 5 intentionally uses TimeWindowCompactionStrategy
(TWCS) for a narrow expiring-data design. No application
driver is required for mandatory work; cqlsh and nodetool
provide the evidence. Exact metrics, tokens, timings, warning
text, and SSTable filenames are learner-captured rather than
invented.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms to define before reasoning about deletion
A coordinator is the Cassandra node handling one client request. A replica stores a copy of a partition according to the keyspace replication strategy. A partition is the unit identified by the partition key; its hash maps to a token and therefore a replica set. A consistency level (CL) specifies how many appropriately scoped replica responses the coordinator needs. A tombstone is a timestamped deletion marker written through Cassandra's normal write path. A Time To Live (TTL) is a duration after which a value expires and becomes logically deleted. gc_grace_seconds is the table-level grace period during which Cassandra normally retains deletion markers so replicas have time to converge before the markers become eligible for garbage collection. repair is anti-entropy synchronization between replicas. A hint is a best-effort mutation retained temporarily for an unavailable replica; it is not a replacement for repair. An SSTable (Sorted String Table) is immutable on-disk storage. Compaction rewrites SSTables, reconciles versions, and may purge tombstones only when Cassandra's safety conditions allow it.
The word delete therefore has two different meanings in Cassandra: a row can disappear from query results immediately while its physical deletion marker remains on disk for much longer. Conflating logical invisibility with physical reclamation is the root of many tombstone mistakes.
1. The redesign target: align the unit of query, retention, and expiry
A table keyed only by tenant_id can grow without
bound. TTL prevents old values from remaining logically live
forever, but it does not prevent the current partition from
spanning many SSTables or force expired rows to disappear from
disk immediately. Time bucketing places a predictable amount of
data in each partition. TWCS then groups SSTables by write
timestamp windows so similarly aged files can expire together.
Cassandra 5.0 recommends UCS for most new workloads, so TWCS should not be cargo-culted as “the TTL default.” Official documentation still recommends TWCS specifically for time-series and expiring TTL workloads where data is mostly append-only and timestamps/windows are well behaved. The decision is therefore workload-specific, not a contradiction.
| Design dimension | Unbounded TTL table | Bucketed TWCS table |
|---|---|---|
| Partition key | tenant only | tenant + time bucket |
| Read fan-out | one ever-growing logical partition, but much dead history | bounded number of explicit buckets |
| Expiry locality | mixed ages can coexist across overlapping SSTables | SSTables grouped by time windows |
| Mutation pattern | old rows may be rewritten/deleted repeatedly | prefer append-only/current-window writes |
| Compaction fit | UCS good general starting point | TWCS specialized for time-series/TTL expiry |
| Failure mode | wide dead scans and delayed reclamation | late writes can contaminate old/new windows |
2. Build a deliberately poor TTL table and a bounded TWCS alternative
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool status# Continue only after all three nodes are UN in dc1.# If the shared course cluster does not exist, recreate it with the same# Chapter 01 conventions before running this chapter. Do not expose CQL/JMX# to untrusted networks merely to make the lab convenient.
CREATE KEYSPACE IF NOT EXISTS atlasmart_tombstonesWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name='atlasmart_tombstones';
CREATE TABLE IF NOT EXISTS atlasmart_tombstones.ttl_events_unbounded ( tenant_id text, event_time timestamp, event_id uuid, payload text, PRIMARY KEY (tenant_id,event_time,event_id)) WITH CLUSTERING ORDER BY (event_time DESC,event_id ASC) AND default_time_to_live = 15 AND gc_grace_seconds = 5 AND compaction = {'class':'UnifiedCompactionStrategy'};CREATE TABLE IF NOT EXISTS atlasmart_tombstones.ttl_events_by_bucket ( tenant_id text, bucket_minute text, event_time timestamp, event_id uuid, payload text, PRIMARY KEY ((tenant_id,bucket_minute),event_time,event_id)) WITH CLUSTERING ORDER BY (event_time DESC,event_id ASC) AND default_time_to_live = 15 AND gc_grace_seconds = 5 AND compaction = { 'class':'TimeWindowCompactionStrategy', 'compaction_window_unit':'MINUTES', 'compaction_window_size':'1', 'unsafe_aggressive_sstable_expiration':'false' };
Only because these tables are disposable TTL-only teaching fixtures and the exercise needs visible expiration within one session. Production grace and TTL must be derived from retention, repair, outage, late-write, and recovery requirements. Do not copy these numbers.
3. Load the same events into both shapes and watch TTL count down
For a small reproducible lab, write three events into the unbounded partition and place those same events into three explicit minute buckets. Real bucket width comes from measured events/bytes per partition and query fan-out, not from a universal “one minute” rule.
CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_tombstones.ttl_events_unbounded(tenant_id,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:40:10Z',00000000-0000-0000-0000-000000000101,'e1');INSERT INTO atlasmart_tombstones.ttl_events_unbounded(tenant_id,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:41:10Z',00000000-0000-0000-0000-000000000102,'e2');INSERT INTO atlasmart_tombstones.ttl_events_unbounded(tenant_id,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:42:10Z',00000000-0000-0000-0000-000000000103,'e3');INSERT INTO atlasmart_tombstones.ttl_events_by_bucket(tenant_id,bucket_minute,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:40Z','2026-09-08T05:40:10Z',00000000-0000-0000-0000-000000000101,'e1');INSERT INTO atlasmart_tombstones.ttl_events_by_bucket(tenant_id,bucket_minute,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:41Z','2026-09-08T05:41:10Z',00000000-0000-0000-0000-000000000102,'e2');INSERT INTO atlasmart_tombstones.ttl_events_by_bucket(tenant_id,bucket_minute,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:42Z','2026-09-08T05:42:10Z',00000000-0000-0000-0000-000000000103,'e3');SELECT event_time,payload,TTL(payload)FROM atlasmart_tombstones.ttl_events_unboundedWHERE tenant_id='tenant-a';SELECT event_time,payload,TTL(payload)FROM atlasmart_tombstones.ttl_events_by_bucketWHERE tenant_id='tenant-a' AND bucket_minute='2026-09-08T05:41Z';
Run the TTL queries immediately and again after roughly 15 seconds. Values should disappear logically after expiry. The five-second lab grace means they become purge-eligible later, not necessarily instantly. Flush before expiry if you want each short-lived generation materialized as SSTables for inspection.
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones ttl_events_unbounded docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones ttl_events_by_bucketdonedocker exec atlasmart-cass-1 nodetool tablestats \ atlasmart_tombstones.ttl_events_unboundeddocker exec atlasmart-cass-1 nodetool tablestats \ atlasmart_tombstones.ttl_events_by_bucketdocker exec atlasmart-cass-1 nodetool tablehistograms \ atlasmart_tombstones ttl_events_unboundeddocker exec atlasmart-cass-1 nodetool tablehistograms \ atlasmart_tombstones ttl_events_by_bucket
4. Why TWCS can reclaim expiring windows efficiently—and how to break it
TWCS groups SSTables by time window. Once a closed window
contains only data that has expired and is safe to drop,
Cassandra can reclaim an entire SSTable rather than repeatedly
merging mixed-age data. That property depends on temporal
hygiene. Old and new data written through the same memtable can
be flushed together. Read repair of old data can also place old
timestamps in a current SSTable. Explicit
USING TIMESTAMP with old timestamps is another
contamination path. Mixed TTL horizons can keep a file alive
because some cells expire much later than others.
Current documentation warns that
unsafe_aggressive_sstable_expiration can cause data
loss or deleted data to reappear; it is false by default and
even requires a JVM-level opt-in. Leave it false unless you have
proved the shadowing/repair invariants and explicitly accept the
correctness risk.
# For this disposable table only, wait >20 seconds from insertion.sleep 22docker exec atlasmart-cass-1 nodetool compact \ atlasmart_tombstones ttl_events_unboundeddocker exec atlasmart-cass-1 nodetool compact \ atlasmart_tombstones ttl_events_by_bucketdocker exec atlasmart-cass-1 nodetool tablestats \ atlasmart_tombstones.ttl_events_unboundeddocker exec atlasmart-cass-1 nodetool tablestats \ atlasmart_tombstones.ttl_events_by_bucketdocker exec atlasmart-cass-1 nodetool compactionhistory
The tiny lab may not show a meaningful latency or disk-space win because there are only a few rows. The evidence goal is structural: bounded partitions, explicit bucket fan-out, TWCS table configuration, TTL disappearance, and compaction/SSTable state. Production validation requires representative retention horizons and many windows.
5. Before/after acceptance criteria for a real redesign
| Measure | Before | After target |
|---|---|---|
| Partition rows/bytes p95/max | growing with tenant age | bounded by bucket capacity policy |
| Buckets per request | implicit one giant partition | explicit finite fan-out known by application |
| Tombstones/read p95/p99 | large for historical scans | bounded because each query touches relevant buckets |
| SSTables/read | may grow with mixed history | controlled by time-window/compaction state |
| Expired disk lag | unpredictable mixed-age files | old windows become fully expired and droppable predictably |
| Late/mutable writes | common | restricted or handled with explicit policy |
| Repair cadence | not linked to retention | documented relative to grace/outage requirements |
TWCS assumes useful time locality. Frequent updates to historical rows, late writes, explicit old timestamps, or read-repair contamination can mix old and new data in the same SSTable and delay whole-file expiration. If the workload is genuinely mutable and general-purpose, UCS may be the better Cassandra 5.0 starting point.
DROP TABLE IF EXISTS atlasmart_tombstones.ttl_events_unbounded;DROP TABLE IF EXISTS atlasmart_tombstones.ttl_events_by_bucket;DROP TABLE IF EXISTS atlasmart_tombstones.tombstone_types;DROP TABLE IF EXISTS atlasmart_tombstones.delete_probe;
Check your understanding
- Why does time bucketing help a high-TTL workload even before considering compaction strategy?
- Why is TWCS still taught if UCS is recommended for most new Cassandra 5.0 workloads?
- What can contaminate a TWCS window with old data?
- Why leave unsafe_aggressive_sstable_expiration=false?
- What must a real migration benchmark compare?
Review the answers
1. It bounds partition rows/bytes and makes query fan-out explicit, so reads need not traverse an ever-growing tenant history.
2. TWCS remains a specialized recommended strategy for time-series/expiring TTL data whose temporal and mutation patterns fit its window model.
3. Out-of-order/old-timestamp writes, mutations of historical rows, and read repair that moves old data into a current memtable/SSTable.
4. Skipping shadowing checks can create correctness problems including data loss or reappearance of deleted data.
5. Representative partition distributions, tombstones/read, SSTables/read, read/write p95/p99, compaction debt, disk headroom/reclamation lag, repair behavior, and application bucket fan-out.
Production judgment
Deletion policy is part of the data model, repair plan, capacity
model, and failure model. Before changing TTL,
gc_grace_seconds, tombstone thresholds, compaction
options, or retention windows, record RF/CL, maximum tolerated
replica outage, hint window, repair cadence and success
evidence, partition rows/bytes, delete and TTL rate, clustering
scan width, SSTables-per-read, compaction debt, free disk, cache
state, JVM/heap pressure, p95/p99 read/write latency, and the
exact Cassandra patch/platform. Security and tenant boundaries
also matter: an abusive tenant that creates very wide partitions
or delete storms can create shared-node latency and heap risk.
Do not treat a local threshold or a blog value as a target. A large tombstone count can be harmless when queries touch narrow slices, while a smaller count can be damaging when every request scans a broad dead region. Hints reduce short-outage exposure but are not comprehensive anti-entropy. Lowering grace time transfers safety requirements to repair/outage discipline. SAI or vector indexes do not erase the storage semantics: indexed data still lives in SSTables and must coexist with deletes, compaction, repair, and retention. Driver timeouts/retries must be tested under real slow-read/failure conditions and only retried when operation semantics are safe. Chapter 14 now builds on this deletion/replication foundation to quantify how read and write consistency levels translate RF into client-visible availability, latency, and staleness guarantees.
Summary and next bridge
High-TTL workloads become manageable when partition boundaries, retention horizon, and storage organization agree. Bucketing bounds the query unit; TWCS can group similarly aged SSTables for expiring time-series data; repair/grace still define deletion safety. Chapter 14 turns to consistency-level math and the client guarantees those replica states provide.
Authoritative references
These official references are the source of truth for version-sensitive behavior. Re-check them when the course is regenerated because defaults, guardrails, tooling, and repair/compaction behavior can evolve.
- Apache Cassandra downloads / 5.0 release baseline
- Compaction overview — tombstones, TTL, gc_grace and purging
- CQL table options including gc_grace_seconds and default TTL
- ALTER TABLE — gc_grace_seconds and TTL considerations
- cassandra.yaml — tombstone warn/failure thresholds and guardrails
- Repair operations
- TimeWindowCompactionStrategy
- nodetool tablestats
- nodetool tablehistograms