Chapter 13 · Tombstones, TTL, gc_grace, Deletes, and Expiration
Row / Cell / Range Tombstones, Collection Tombstones, TTL Expiration, and Query Cost
Distinguish tombstone shapes and TTL expiration, then inspect their read and SSTable consequences safely.
Learning outcomes
AtlasMart's event timeline has several kinds of “deletion”: a user clears one note, a row is removed, a contiguous time slice is erased, one map key is removed, and temporary events expire automatically. All can produce deletion state, but the read cost and SSTable representation are not identical.
Distinguish cell, row, range, collection-element/collection, partition, and TTL-generated deletion markers.
Explain why range and collection operations can change tombstone shape even when the API intent sounds similar.
Observe TTL countdown and expiration without assuming expired data is physically gone.
Use tracing, tablestats/tablehistograms, and read-only sstabledump/sstablemetadata against a Cassandra snapshot to inspect disposable tombstone evidence safely.
Connect tombstone shape and clustering scan width to read amplification and heap/latency risk.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the image,
cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. The chapter keyspace is
atlasmart_tombstones with
NetworkTopologyStrategy and replication factor
(RF) 3; reads and writes normally use
LOCAL_QUORUM. Authentication, client TLS,
internode TLS, and remote JMX remain disabled only inside this
isolated local learning network. UnifiedCompactionStrategy
(UCS) remains the default choice for ordinary chapter tables;
Lesson 5 intentionally uses TimeWindowCompactionStrategy
(TWCS) for a narrow expiring-data design. No application
driver is required for mandatory work; cqlsh and nodetool
provide the evidence. Exact metrics, tokens, timings, warning
text, and SSTable filenames are learner-captured rather than
invented.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms to define before reasoning about deletion
A coordinator is the Cassandra node handling one client request. A replica stores a copy of a partition according to the keyspace replication strategy. A partition is the unit identified by the partition key; its hash maps to a token and therefore a replica set. A consistency level (CL) specifies how many appropriately scoped replica responses the coordinator needs. A tombstone is a timestamped deletion marker written through Cassandra's normal write path. A Time To Live (TTL) is a duration after which a value expires and becomes logically deleted. gc_grace_seconds is the table-level grace period during which Cassandra normally retains deletion markers so replicas have time to converge before the markers become eligible for garbage collection. repair is anti-entropy synchronization between replicas. A hint is a best-effort mutation retained temporarily for an unavailable replica; it is not a replacement for repair. An SSTable (Sorted String Table) is immutable on-disk storage. Compaction rewrites SSTables, reconciles versions, and may purge tombstones only when Cassandra's safety conditions allow it.
The word delete therefore has two different meanings in Cassandra: a row can disappear from query results immediately while its physical deletion marker remains on disk for much longer. Conflating logical invisibility with physical reclamation is the root of many tombstone mistakes.
1. Tombstone taxonomy is about what interval or cell is suppressed
| Kind | Typical CQL action | What it suppresses | Read-cost concern |
|---|---|---|---|
| Cell tombstone | DELETE note ... | one regular cell | many sparse deletions can create merge work |
| Row tombstone | DELETE FROM ... with full primary key | all non-PK cells for one clustering row | many dead rows in a scan accumulate tombstones |
| Range tombstone | DELETE ... with clustering slice | a contiguous clustering interval | large dead intervals can still intersect broad reads |
| Collection element tombstone | remove map/set element | one element in a non-frozen collection | repeated churn can accumulate element-level dead cells |
| Collection tombstone | replace/clear a non-frozen collection | previous collection contents at an older timestamp | large collections magnify mutation/read cost |
| Partition tombstone | delete a whole partition key | all rows in that partition at older timestamps | wide partition history still matters until purge |
| TTL expiration | INSERT/UPDATE USING TTL or table default TTL | the expiring value after its deadline | high TTL churn can create large dead regions |
These are implementation/storage semantics, not application API types. The safest design question is still query-first: how many dead cells or ranges must the read path traverse to return the live result? A few thousand tombstones spread across tiny independent partitions can behave very differently from the same count concentrated before the first live row in one huge partition.
2. Create one table that produces several deletion shapes
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool status# Continue only after all three nodes are UN in dc1.# If the shared course cluster does not exist, recreate it with the same# Chapter 01 conventions before running this chapter. Do not expose CQL/JMX# to untrusted networks merely to make the lab convenient.
CREATE KEYSPACE IF NOT EXISTS atlasmart_tombstonesWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name='atlasmart_tombstones';
CREATE TABLE IF NOT EXISTS atlasmart_tombstones.tombstone_types ( tenant_id text, event_day date, event_time timestamp, event_id uuid, status text, note text, attributes map<text,text>, tags set<text>, PRIMARY KEY ((tenant_id,event_day),event_time,event_id)) WITH CLUSTERING ORDER BY (event_time ASC,event_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'} AND gc_grace_seconds = 864000;CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_tombstones.tombstone_types(tenant_id,event_day,event_time,event_id,status,note,attributes,tags)VALUES ('tenant-a','2026-09-08','2026-09-08T05:00:00Z',00000000-0000-0000-0000-000000000001,'OPEN','cell-delete',{'legacy':'yes','channel':'web'},{'vip','new'});INSERT INTO atlasmart_tombstones.tombstone_types(tenant_id,event_day,event_time,event_id,status,note,attributes,tags)VALUES ('tenant-a','2026-09-08','2026-09-08T05:01:00Z',00000000-0000-0000-0000-000000000002,'OPEN','row-delete',{'channel':'app'},{'mobile'});INSERT INTO atlasmart_tombstones.tombstone_types(tenant_id,event_day,event_time,event_id,status,note)VALUES ('tenant-a','2026-09-08','2026-09-08T05:02:00Z',00000000-0000-0000-0000-000000000003,'OPEN','range-delete');INSERT INTO atlasmart_tombstones.tombstone_types(tenant_id,event_day,event_time,event_id,status,note)USING TTL 20VALUES ('tenant-a','2026-09-08','2026-09-08T05:03:00Z',00000000-0000-0000-0000-000000000004,'TEMP','ttl-expiration');SELECT event_time,status,note,TTL(status)FROM atlasmart_tombstones.tombstone_typesWHERE tenant_id='tenant-a' AND event_day='2026-09-08';
3. Produce cell, row, collection-element, range, and TTL deletion state
-- Cell tombstone.DELETE noteFROM atlasmart_tombstones.tombstone_typesWHERE tenant_id='tenant-a' AND event_day='2026-09-08' AND event_time='2026-09-08T05:00:00Z' AND event_id=00000000-0000-0000-0000-000000000001;-- Element-level deletion in a non-frozen map.UPDATE atlasmart_tombstones.tombstone_typesSET attributes = attributes - {'legacy'}WHERE tenant_id='tenant-a' AND event_day='2026-09-08' AND event_time='2026-09-08T05:00:00Z' AND event_id=00000000-0000-0000-0000-000000000001;-- Row tombstone.DELETE FROM atlasmart_tombstones.tombstone_typesWHERE tenant_id='tenant-a' AND event_day='2026-09-08' AND event_time='2026-09-08T05:01:00Z' AND event_id=00000000-0000-0000-0000-000000000002;-- Range tombstone across a clustering-time slice.DELETE FROM atlasmart_tombstones.tombstone_typesWHERE tenant_id='tenant-a' AND event_day='2026-09-08' AND event_time >= '2026-09-08T05:02:00Z' AND event_time < '2026-09-08T05:03:00Z';TRACING ON;SELECT event_time,event_id,status,note,attributes,tagsFROM atlasmart_tombstones.tombstone_typesWHERE tenant_id='tenant-a' AND event_day='2026-09-08';TRACING OFF;-- Run this immediately and again after about 20 seconds.SELECT event_time,status,TTL(status)FROM atlasmart_tombstones.tombstone_typesWHERE tenant_id='tenant-a' AND event_day='2026-09-08';
The TTL query should show a positive countdown for the temporary value before expiration and no live value after expiry. The physical SSTable bytes do not vanish at the TTL boundary; expired state participates in tombstone/compaction rules. Exact trace messages and tombstone counters are version- and state-dependent, so capture them rather than expecting a fixed transcript.
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones tombstone_typesdonedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_tombstones.tombstone_typesdocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_tombstones tombstone_types
4. Inspect a snapshot, never the live mutable path, with offline SSTable tools
sstabledump and sstablemetadata are
low-level offline tools. Treat their JSON/component details as
implementation evidence, not a stable application interface. The
following lab first takes a Cassandra snapshot, then mounts the
named volume read-only into a utility container
that does not start Cassandra. The tools read only snapshot
SSTables, not the live table path.
docker exec atlasmart-cass-1 nodetool snapshot \ -t ch13-l2 -cf tombstone_types atlasmart_tombstones# The utility container mounts the node-1 data volume read-only.docker run --rm \ -v atlasmart-cass-1-data:/var/lib/cassandra:ro \ cassandra:5.0.9 sh -lc ' f=$(find /var/lib/cassandra/data/atlasmart_tombstones \ -path "*/snapshots/ch13-l2/*Data.db" | head -1) echo "Snapshot Data component: $f" sstablemetadata "$f" | head -120 echo "--- selected dump ---" sstabledump "$f" | head -180 '# Remove only the named snapshot through nodetool.docker exec atlasmart-cass-1 nodetool clearsnapshot \ -t ch13-l2 atlasmart_tombstones
SSTable format and dump structure can vary. Look for deletion/expiration metadata and clustering ranges as evidence, but do not write scripts that depend on a particular field layout unless you own that compatibility contract. Do not edit the dump and copy bytes back into Cassandra.
TTL changes logical visibility at expiration. Physical reclamation waits for grace and safe compaction conditions. Mixed TTLs, live data, and older overlapping SSTables can delay space recovery substantially.
5. Verification checklist
- Observe a positive TTL countdown before the temporary row expires.
- Verify the expired value disappears logically without claiming instant disk reclamation.
- Distinguish the cell delete, row delete, collection-element removal, and clustering-range delete in CQL.
-
Record
tablestats/tablehistogramsbefore compaction. - Run SSTable tools only against the snapshot mounted read-only.
- Clear the named snapshot through nodetool.
Check your understanding
- Why can two DELETE statements create different tombstone shapes?
- What happens at the TTL deadline?
- Why inspect a snapshot instead of the live data directory with offline tools?
- Does a range tombstone necessarily mean thousands of independent row tombstone objects?
- Why can collection churn become expensive?
Review the answers
1. The primary-key/clustering restriction determines whether Cassandra is suppressing a single cell/row or a clustering interval; collection operations can also target individual elements.
2. The value becomes logically expired/deleted and is handled as tombstoned state; physical bytes are reclaimed later under compaction safety rules.
3. A snapshot gives a stable immutable view and avoids racing live Cassandra mutations; mounting it read-only further prevents accidental changes.
4. No. It represents a clustering interval, although reads crossing that dead interval still incur reconciliation/scan work.
5. Non-frozen collection elements are independent cells; repeated add/remove operations can create many versions and element tombstones that reads and compaction must reconcile.
Production judgment
Deletion policy is part of the data model, repair plan, capacity
model, and failure model. Before changing TTL,
gc_grace_seconds, tombstone thresholds, compaction
options, or retention windows, record RF/CL, maximum tolerated
replica outage, hint window, repair cadence and success
evidence, partition rows/bytes, delete and TTL rate, clustering
scan width, SSTables-per-read, compaction debt, free disk, cache
state, JVM/heap pressure, p95/p99 read/write latency, and the
exact Cassandra patch/platform. Security and tenant boundaries
also matter: an abusive tenant that creates very wide partitions
or delete storms can create shared-node latency and heap risk.
Do not treat a local threshold or a blog value as a target. A large tombstone count can be harmless when queries touch narrow slices, while a smaller count can be damaging when every request scans a broad dead region. Hints reduce short-outage exposure but are not comprehensive anti-entropy. Lowering grace time transfers safety requirements to repair/outage discipline. SAI or vector indexes do not erase the storage semantics: indexed data still lives in SSTables and must coexist with deletes, compaction, repair, and retention. Driver timeouts/retries must be tested under real slow-read/failure conditions and only retried when operation semantics are safe. Lesson 3 uses a deliberately shortened grace period in an isolated disposable table to show exactly how missed deletes plus premature purging can create zombie data.
Summary and next bridge
“Tombstone” is a family of deletion representations whose cost depends on what the query scans. TTL is another source of deletion state, not a disk-delete timer. Next, shorten grace deliberately in a disposable table and observe why repair timing is part of correctness.
Authoritative references
These official references are the source of truth for version-sensitive behavior. Re-check them when the course is regenerated because defaults, guardrails, tooling, and repair/compaction behavior can evolve.
- Apache Cassandra downloads / 5.0 release baseline
- Compaction overview — tombstones, TTL, gc_grace and purging
- CQL table options including gc_grace_seconds and default TTL
- ALTER TABLE — gc_grace_seconds and TTL considerations
- cassandra.yaml — tombstone warn/failure thresholds and guardrails
- Repair operations
- TimeWindowCompactionStrategy
- nodetool tablestats
- nodetool tablehistograms
- sstabledump tool
- sstablemetadata tool