Chapter 13 · Tombstones, TTL, gc_grace, Deletes, and Expiration

Redesign High-Delete / High-TTL Workloads with Bucketing, TWCS, and Retention-Aware Modeling

Redesign expiring workloads with bounded buckets, TWCS where appropriate, and retention-aware repair/grace policy.

Intermediate120–165 minutesBucketing + TWCS retention redesign labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · UCS baseline · TWCS specialized in Lesson 5Last reviewed: September 2026

Learning outcomes

AtlasMart stores short-lived clickstream/session events. The original table keeps years of event history in a tenant partition and relies on TTL for cleanup. That combines high write rate, high expiration rate, giant partitions, and mixed-age SSTables—the exact pattern that turns automatic expiration into read/compaction debt. The redesign aligns partition buckets, TTL horizon, and a specialized compaction strategy.

01

Redesign an unbounded high-TTL partition into bounded time buckets that match read and retention requirements.

02

Explain when TWCS is a reasonable specialized choice even though UCS is recommended for most new Cassandra 5.0 workloads.

03

Show how uniform TTL and time-windowed SSTables can make fully expired file reclamation easier.

04

Identify TWCS hazards: out-of-order writes, mixed TTLs, mutable old data, read repair of old data, and unsafe aggressive expiration.

05

Compare before/after partition shape, TTL countdown, SSTable counts, read scans, and disk behavior without claiming a tiny lab is a production benchmark.

Chapter 13 lab baseline

The mandatory labs continue the disposable AtlasMart course cluster: Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. The chapter keyspace is atlasmart_tombstones with NetworkTopologyStrategy and replication factor (RF) 3; reads and writes normally use LOCAL_QUORUM. Authentication, client TLS, internode TLS, and remote JMX remain disabled only inside this isolated local learning network. UnifiedCompactionStrategy (UCS) remains the default choice for ordinary chapter tables; Lesson 5 intentionally uses TimeWindowCompactionStrategy (TWCS) for a narrow expiring-data design. No application driver is required for mandatory work; cqlsh and nodetool provide the evidence. Exact metrics, tokens, timings, warning text, and SSTable filenames are learner-captured rather than invented.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms to define before reasoning about deletion

A coordinator is the Cassandra node handling one client request. A replica stores a copy of a partition according to the keyspace replication strategy. A partition is the unit identified by the partition key; its hash maps to a token and therefore a replica set. A consistency level (CL) specifies how many appropriately scoped replica responses the coordinator needs. A tombstone is a timestamped deletion marker written through Cassandra's normal write path. A Time To Live (TTL) is a duration after which a value expires and becomes logically deleted. gc_grace_seconds is the table-level grace period during which Cassandra normally retains deletion markers so replicas have time to converge before the markers become eligible for garbage collection. repair is anti-entropy synchronization between replicas. A hint is a best-effort mutation retained temporarily for an unavailable replica; it is not a replacement for repair. An SSTable (Sorted String Table) is immutable on-disk storage. Compaction rewrites SSTables, reconciles versions, and may purge tombstones only when Cassandra's safety conditions allow it.

The word delete therefore has two different meanings in Cassandra: a row can disappear from query results immediately while its physical deletion marker remains on disk for much longer. Conflating logical invisibility with physical reclamation is the root of many tombstone mistakes.

1. The redesign target: align the unit of query, retention, and expiry

A table keyed only by tenant_id can grow without bound. TTL prevents old values from remaining logically live forever, but it does not prevent the current partition from spanning many SSTables or force expired rows to disappear from disk immediately. Time bucketing places a predictable amount of data in each partition. TWCS then groups SSTables by write timestamp windows so similarly aged files can expire together.

Cassandra 5.0 recommends UCS for most new workloads, so TWCS should not be cargo-culted as “the TTL default.” Official documentation still recommends TWCS specifically for time-series and expiring TTL workloads where data is mostly append-only and timestamps/windows are well behaved. The decision is therefore workload-specific, not a contradiction.

Design dimension Unbounded TTL table Bucketed TWCS table
Partition key tenant only tenant + time bucket
Read fan-out one ever-growing logical partition, but much dead history bounded number of explicit buckets
Expiry locality mixed ages can coexist across overlapping SSTables SSTables grouped by time windows
Mutation pattern old rows may be rewritten/deleted repeatedly prefer append-only/current-window writes
Compaction fit UCS good general starting point TWCS specialized for time-series/TTL expiry
Failure mode wide dead scans and delayed reclamation late writes can contaminate old/new windows

2. Build a deliberately poor TTL table and a bounded TWCS alternative

Docker · verify the existing disposable cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool status# Continue only after all three nodes are UN in dc1.# If the shared course cluster does not exist, recreate it with the same# Chapter 01 conventions before running this chapter. Do not expose CQL/JMX# to untrusted networks merely to make the lab convenient.
CQL · create the chapter keyspace
CREATE KEYSPACE IF NOT EXISTS atlasmart_tombstonesWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name='atlasmart_tombstones';
CQL · bad unbounded table versus bounded time-window table
CREATE TABLE IF NOT EXISTS atlasmart_tombstones.ttl_events_unbounded (    tenant_id text,    event_time timestamp,    event_id uuid,    payload text,    PRIMARY KEY (tenant_id,event_time,event_id)) WITH CLUSTERING ORDER BY (event_time DESC,event_id ASC)  AND default_time_to_live = 15  AND gc_grace_seconds = 5  AND compaction = {'class':'UnifiedCompactionStrategy'};CREATE TABLE IF NOT EXISTS atlasmart_tombstones.ttl_events_by_bucket (    tenant_id text,    bucket_minute text,    event_time timestamp,    event_id uuid,    payload text,    PRIMARY KEY ((tenant_id,bucket_minute),event_time,event_id)) WITH CLUSTERING ORDER BY (event_time DESC,event_id ASC)  AND default_time_to_live = 15  AND gc_grace_seconds = 5  AND compaction = {      'class':'TimeWindowCompactionStrategy',      'compaction_window_unit':'MINUTES',      'compaction_window_size':'1',      'unsafe_aggressive_sstable_expiration':'false'  };
Why are TTL=15 seconds and grace=5 seconds acceptable here?

Only because these tables are disposable TTL-only teaching fixtures and the exercise needs visible expiration within one session. Production grace and TTL must be derived from retention, repair, outage, late-write, and recovery requirements. Do not copy these numbers.

3. Load the same events into both shapes and watch TTL count down

For a small reproducible lab, write three events into the unbounded partition and place those same events into three explicit minute buckets. Real bucket width comes from measured events/bytes per partition and query fan-out, not from a universal “one minute” rule.

CQL · identical logical events, different partition shapes
CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_tombstones.ttl_events_unbounded(tenant_id,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:40:10Z',00000000-0000-0000-0000-000000000101,'e1');INSERT INTO atlasmart_tombstones.ttl_events_unbounded(tenant_id,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:41:10Z',00000000-0000-0000-0000-000000000102,'e2');INSERT INTO atlasmart_tombstones.ttl_events_unbounded(tenant_id,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:42:10Z',00000000-0000-0000-0000-000000000103,'e3');INSERT INTO atlasmart_tombstones.ttl_events_by_bucket(tenant_id,bucket_minute,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:40Z','2026-09-08T05:40:10Z',00000000-0000-0000-0000-000000000101,'e1');INSERT INTO atlasmart_tombstones.ttl_events_by_bucket(tenant_id,bucket_minute,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:41Z','2026-09-08T05:41:10Z',00000000-0000-0000-0000-000000000102,'e2');INSERT INTO atlasmart_tombstones.ttl_events_by_bucket(tenant_id,bucket_minute,event_time,event_id,payload)VALUES ('tenant-a','2026-09-08T05:42Z','2026-09-08T05:42:10Z',00000000-0000-0000-0000-000000000103,'e3');SELECT event_time,payload,TTL(payload)FROM atlasmart_tombstones.ttl_events_unboundedWHERE tenant_id='tenant-a';SELECT event_time,payload,TTL(payload)FROM atlasmart_tombstones.ttl_events_by_bucketWHERE tenant_id='tenant-a' AND bucket_minute='2026-09-08T05:41Z';

Run the TTL queries immediately and again after roughly 15 seconds. Values should disappear logically after expiry. The five-second lab grace means they become purge-eligible later, not necessarily instantly. Flush before expiry if you want each short-lived generation materialized as SSTables for inspection.

nodetool · record physical state before and after expiry
for n in 1 2 3; do  docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones ttl_events_unbounded  docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones ttl_events_by_bucketdonedocker exec atlasmart-cass-1 nodetool tablestats \  atlasmart_tombstones.ttl_events_unboundeddocker exec atlasmart-cass-1 nodetool tablestats \  atlasmart_tombstones.ttl_events_by_bucketdocker exec atlasmart-cass-1 nodetool tablehistograms \  atlasmart_tombstones ttl_events_unboundeddocker exec atlasmart-cass-1 nodetool tablehistograms \  atlasmart_tombstones ttl_events_by_bucket

4. Why TWCS can reclaim expiring windows efficiently—and how to break it

TWCS groups SSTables by time window. Once a closed window contains only data that has expired and is safe to drop, Cassandra can reclaim an entire SSTable rather than repeatedly merging mixed-age data. That property depends on temporal hygiene. Old and new data written through the same memtable can be flushed together. Read repair of old data can also place old timestamps in a current SSTable. Explicit USING TIMESTAMP with old timestamps is another contamination path. Mixed TTL horizons can keep a file alive because some cells expire much later than others.

Current documentation warns that unsafe_aggressive_sstable_expiration can cause data loss or deleted data to reappear; it is false by default and even requires a JVM-level opt-in. Leave it false unless you have proved the shadowing/repair invariants and explicitly accept the correctness risk.

Docker/nodetool · after TTL + grace, compare compaction/reclamation evidence
# For this disposable table only, wait >20 seconds from insertion.sleep 22docker exec atlasmart-cass-1 nodetool compact \  atlasmart_tombstones ttl_events_unboundeddocker exec atlasmart-cass-1 nodetool compact \  atlasmart_tombstones ttl_events_by_bucketdocker exec atlasmart-cass-1 nodetool tablestats \  atlasmart_tombstones.ttl_events_unboundeddocker exec atlasmart-cass-1 nodetool tablestats \  atlasmart_tombstones.ttl_events_by_bucketdocker exec atlasmart-cass-1 nodetool compactionhistory

The tiny lab may not show a meaningful latency or disk-space win because there are only a few rows. The evidence goal is structural: bounded partitions, explicit bucket fan-out, TWCS table configuration, TTL disappearance, and compaction/SSTable state. Production validation requires representative retention horizons and many windows.

5. Before/after acceptance criteria for a real redesign

Measure Before After target
Partition rows/bytes p95/max growing with tenant age bounded by bucket capacity policy
Buckets per request implicit one giant partition explicit finite fan-out known by application
Tombstones/read p95/p99 large for historical scans bounded because each query touches relevant buckets
SSTables/read may grow with mixed history controlled by time-window/compaction state
Expired disk lag unpredictable mixed-age files old windows become fully expired and droppable predictably
Late/mutable writes common restricted or handled with explicit policy
Repair cadence not linked to retention documented relative to grace/outage requirements
Wrong redesign: “Use TWCS and keep rewriting old rows.”

TWCS assumes useful time locality. Frequent updates to historical rows, late writes, explicit old timestamps, or read-repair contamination can mix old and new data in the same SSTable and delay whole-file expiration. If the workload is genuinely mutable and general-purpose, UCS may be the better Cassandra 5.0 starting point.

CQL · cleanup chapter-specific TTL fixtures
DROP TABLE IF EXISTS atlasmart_tombstones.ttl_events_unbounded;DROP TABLE IF EXISTS atlasmart_tombstones.ttl_events_by_bucket;DROP TABLE IF EXISTS atlasmart_tombstones.tombstone_types;DROP TABLE IF EXISTS atlasmart_tombstones.delete_probe;

Check your understanding

  1. Why does time bucketing help a high-TTL workload even before considering compaction strategy?
  2. Why is TWCS still taught if UCS is recommended for most new Cassandra 5.0 workloads?
  3. What can contaminate a TWCS window with old data?
  4. Why leave unsafe_aggressive_sstable_expiration=false?
  5. What must a real migration benchmark compare?
Review the answers

1. It bounds partition rows/bytes and makes query fan-out explicit, so reads need not traverse an ever-growing tenant history.

2. TWCS remains a specialized recommended strategy for time-series/expiring TTL data whose temporal and mutation patterns fit its window model.

3. Out-of-order/old-timestamp writes, mutations of historical rows, and read repair that moves old data into a current memtable/SSTable.

4. Skipping shadowing checks can create correctness problems including data loss or reappearance of deleted data.

5. Representative partition distributions, tombstones/read, SSTables/read, read/write p95/p99, compaction debt, disk headroom/reclamation lag, repair behavior, and application bucket fan-out.

Production judgment

Deletion policy is part of the data model, repair plan, capacity model, and failure model. Before changing TTL, gc_grace_seconds, tombstone thresholds, compaction options, or retention windows, record RF/CL, maximum tolerated replica outage, hint window, repair cadence and success evidence, partition rows/bytes, delete and TTL rate, clustering scan width, SSTables-per-read, compaction debt, free disk, cache state, JVM/heap pressure, p95/p99 read/write latency, and the exact Cassandra patch/platform. Security and tenant boundaries also matter: an abusive tenant that creates very wide partitions or delete storms can create shared-node latency and heap risk.

Do not treat a local threshold or a blog value as a target. A large tombstone count can be harmless when queries touch narrow slices, while a smaller count can be damaging when every request scans a broad dead region. Hints reduce short-outage exposure but are not comprehensive anti-entropy. Lowering grace time transfers safety requirements to repair/outage discipline. SAI or vector indexes do not erase the storage semantics: indexed data still lives in SSTables and must coexist with deletes, compaction, repair, and retention. Driver timeouts/retries must be tested under real slow-read/failure conditions and only retried when operation semantics are safe. Chapter 14 now builds on this deletion/replication foundation to quantify how read and write consistency levels translate RF into client-visible availability, latency, and staleness guarantees.

Summary and next bridge

High-TTL workloads become manageable when partition boundaries, retention horizon, and storage organization agree. Bucketing bounds the query unit; TWCS can group similarly aged SSTables for expiring time-series data; repair/grace still define deletion safety. Chapter 14 turns to consistency-level math and the client guarantees those replica states provide.

Authoritative references

These official references are the source of truth for version-sensitive behavior. Re-check them when the course is regenerated because defaults, guardrails, tooling, and repair/compaction behavior can evolve.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.