Chapter 13 · Tombstones, TTL, gc_grace, Deletes, and Expiration
Tombstone Warning / Failure Thresholds, Wide Partitions, and Range-Scan Hazards
Trigger a warning-scale tombstone scan and diagnose why thresholds protect the server rather than define a design target.
Learning outcomes
AtlasMart's queue-like activity table returns only 50 live rows, yet the request scans a long prefix of deleted history and becomes slow. The result size looks tiny, but the storage work is not. Cassandra therefore has tombstone warning/failure thresholds and additional read/partition guardrails to expose or stop dangerous scans.
Explain why tombstones consume read-path memory/work even when the final result set is small.
Inspect Cassandra 5.0 tombstone warning/failure thresholds instead of memorizing undocumented values.
Create a bounded warning-scale lab with more than one thousand row tombstones but avoid a destructive hundred-thousand-tombstone failure test.
Use tracing, cqlsh warnings, logs, tablestats, tablehistograms, and partition shape to diagnose the scan.
Reject threshold-raising as a substitute for fixing partition/query/retention design.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the image,
cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. The chapter keyspace is
atlasmart_tombstones with
NetworkTopologyStrategy and replication factor
(RF) 3; reads and writes normally use
LOCAL_QUORUM. Authentication, client TLS,
internode TLS, and remote JMX remain disabled only inside this
isolated local learning network. UnifiedCompactionStrategy
(UCS) remains the default choice for ordinary chapter tables;
Lesson 5 intentionally uses TimeWindowCompactionStrategy
(TWCS) for a narrow expiring-data design. No application
driver is required for mandatory work; cqlsh and nodetool
provide the evidence. Exact metrics, tokens, timings, warning
text, and SSTable filenames are learner-captured rather than
invented.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms to define before reasoning about deletion
A coordinator is the Cassandra node handling one client request. A replica stores a copy of a partition according to the keyspace replication strategy. A partition is the unit identified by the partition key; its hash maps to a token and therefore a replica set. A consistency level (CL) specifies how many appropriately scoped replica responses the coordinator needs. A tombstone is a timestamped deletion marker written through Cassandra's normal write path. A Time To Live (TTL) is a duration after which a value expires and becomes logically deleted. gc_grace_seconds is the table-level grace period during which Cassandra normally retains deletion markers so replicas have time to converge before the markers become eligible for garbage collection. repair is anti-entropy synchronization between replicas. A hint is a best-effort mutation retained temporarily for an unavailable replica; it is not a replacement for repair. An SSTable (Sorted String Table) is immutable on-disk storage. Compaction rewrites SSTables, reconciles versions, and may purge tombstones only when Cassandra's safety conditions allow it.
The word delete therefore has two different meanings in Cassandra: a row can disappear from query results immediately while its physical deletion marker remains on disk for much longer. Conflating logical invisibility with physical reclamation is the root of many tombstone mistakes.
1. Thresholds are safety rails, not sizing targets
Current Cassandra 5.0 documentation lists
tombstone_warn_threshold at 1000 and
tombstone_failure_threshold at 100000 by default.
The warning threshold surfaces expensive scans; the failure
threshold protects a replica from processing an extreme number
of tombstones in one query. Those defaults are version-sensitive
configuration, so the lab reads the actual node's
cassandra.yaml first. Cassandra 5.0 also includes
separate guardrails for partition size and partition tombstone
count at SSTable-write time; those are disabled unless
configured in the current defaults.
Do not design “just under 1000.” Tombstone count is only one dimension. A query can be expensive because of wide partitions, many SSTables, range width, cache misses, replica filtering, payload size, or concurrent reads. Conversely, a narrow exact-key lookup may avoid most dead cells in a table that contains many tombstones elsewhere.
docker exec atlasmart-cass-1 sh -lc \ "grep -n -E 'tombstone_(warn|failure)_threshold|partition_tombstones_(warn|fail)_threshold|partition_size_(warn|fail)_threshold' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 nodetool version
2. Build a warning-scale wide partition without approaching the failure threshold
This is large enough to cross the default warning threshold after 1,200 row deletes but remains tiny compared with the default failure threshold. The mandatory lab intentionally does not generate 100,000+ tombstones because doing so adds resource risk without adding conceptual value.
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool status# Continue only after all three nodes are UN in dc1.# If the shared course cluster does not exist, recreate it with the same# Chapter 01 conventions before running this chapter. Do not expose CQL/JMX# to untrusted networks merely to make the lab convenient.
CREATE KEYSPACE IF NOT EXISTS atlasmart_tombstonesWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name='atlasmart_tombstones';
CREATE TABLE IF NOT EXISTS atlasmart_tombstones.wide_delete_probe ( bucket text, seq int, payload text, PRIMARY KEY (bucket,seq)) WITH compaction = {'class':'UnifiedCompactionStrategy'} AND gc_grace_seconds = 864000;
Generate the statements inside the Cassandra container so Windows users do not need a host Bash environment. One cqlsh process receives the entire stream.
docker exec atlasmart-cass-1 sh -lc '{ echo "CONSISTENCY LOCAL_QUORUM;" for i in $(seq 0 1499); do printf "INSERT INTO atlasmart_tombstones.wide_delete_probe (bucket,seq,payload) VALUES ('\''wide'\'',%s,'\''live-%s'\'');\n" "$i" "$i" done} | cqlsh'for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones wide_delete_probedone
Now delete the first 1,200 rows individually. Individual row deletes make the warning mechanism easy to reason about. A range delete would encode the dead clustering interval differently and is therefore not an equivalent tombstone-count experiment.
docker exec atlasmart-cass-1 sh -lc '{ echo "CONSISTENCY LOCAL_QUORUM;" for i in $(seq 0 1199); do printf "DELETE FROM atlasmart_tombstones.wide_delete_probe WHERE bucket='\''wide'\'' AND seq=%s;\n" "$i" done} | cqlsh'for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones wide_delete_probedone
3. Query past the dead prefix and capture the warning evidence
The first remaining live row is at clustering key 1200. A query
for the partition with LIMIT 50 may need to pass
the 1,200 deleted rows before it can return 50 live rows. That
is why final row count is a poor proxy for storage work.
CONSISTENCY LOCAL_QUORUM;TRACING ON;SELECT seq,payloadFROM atlasmart_tombstones.wide_delete_probeWHERE bucket='wide'LIMIT 50;TRACING OFF;
docker exec atlasmart-cass-1 nodetool tablestats \ atlasmart_tombstones.wide_delete_probedocker exec atlasmart-cass-1 nodetool tablehistograms \ atlasmart_tombstones wide_delete_probe# Review recent server logs. Search the output for "tombstone",# "Scanned over", warning, or read-failure text; exact wording varies.docker logs --since 10m atlasmart-cass-1
Depending on which replica/coordinator processes the read, cqlsh may also print a server warning. Capture the observed tombstone count and latency. If background compaction changed the physical layout, your exact numbers may differ; explain that state rather than forcing the expected transcript.
4. Why raising the threshold is usually the wrong first fix
Raising the warning threshold suppresses evidence. Raising the failure threshold asks a replica to hold/process even more deletion state before protecting itself. Neither changes the partition, delete rate, SSTables, clustering slice, or retention model. A correct remediation might be a bounded time bucket, a query that starts near the live clustering region, a separate “active items” table maintained at write time, lower delete churn, or a retention model that lets whole time windows expire together.
Likewise, deleting a huge clustering range can be logically efficient for the application but still leaves a large dead interval until safe compaction. The safest model often avoids repeatedly mutating a long-lived queue-like partition in the first place.
The threshold is reporting a data-model/query-path problem. First prove why the read crosses so much dead state. Change the model or access path, then validate latency and tombstone scans under representative load. Change the threshold only when you can explain the memory/latency budget and have tested failure behavior.
5. Cleanup and checks
DROP TABLE IF EXISTS atlasmart_tombstones.wide_delete_probe;
Check your understanding
- Why can a LIMIT 50 query scan more than a thousand tombstones?
- Why does the lab use individual row deletes instead of one range delete?
- Why not reproduce the default 100000 failure threshold?
- What does increasing tombstone_failure_threshold fix in the data model?
- Name two model-level remedies for queue-like dead prefixes.
Review the answers
1. LIMIT constrains returned live rows, not necessarily the number of dead cells/rows the storage engine must traverse to find them.
2. A range delete has a different tombstone representation; individual deletes make the warning-scale tombstone count observable and easier to correlate.
3. Generating that much dead state adds resource/heap/time risk with little educational benefit; inspecting the actual configured threshold and understanding its purpose is sufficient.
4. Nothing. It only changes when the server refuses the scan.
5. Time/bucket partitions, separate active-state query tables, bounded retention windows, or query shapes that do not repeatedly scan deleted history.
Production judgment
Deletion policy is part of the data model, repair plan, capacity
model, and failure model. Before changing TTL,
gc_grace_seconds, tombstone thresholds, compaction
options, or retention windows, record RF/CL, maximum tolerated
replica outage, hint window, repair cadence and success
evidence, partition rows/bytes, delete and TTL rate, clustering
scan width, SSTables-per-read, compaction debt, free disk, cache
state, JVM/heap pressure, p95/p99 read/write latency, and the
exact Cassandra patch/platform. Security and tenant boundaries
also matter: an abusive tenant that creates very wide partitions
or delete storms can create shared-node latency and heap risk.
Do not treat a local threshold or a blog value as a target. A large tombstone count can be harmless when queries touch narrow slices, while a smaller count can be damaging when every request scans a broad dead region. Hints reduce short-outage exposure but are not comprehensive anti-entropy. Lowering grace time transfers safety requirements to repair/outage discipline. SAI or vector indexes do not erase the storage semantics: indexed data still lives in SSTables and must coexist with deletes, compaction, repair, and retention. Driver timeouts/retries must be tested under real slow-read/failure conditions and only retried when operation semantics are safe. Lesson 5 redesigns the workload so expiration happens in bounded time buckets and TWCS can group similar-age SSTables, reducing the need to scan/delete giant mixed-age partitions.
Summary and next bridge
Tombstone thresholds protect the server from pathological scans; they are not recommended partition sizes. A small result can require large storage work when the read crosses dead history. The final lesson changes the model so retention, partition boundaries, and compaction windows cooperate.
Authoritative references
These official references are the source of truth for version-sensitive behavior. Re-check them when the course is regenerated because defaults, guardrails, tooling, and repair/compaction behavior can evolve.
- Apache Cassandra downloads / 5.0 release baseline
- Compaction overview — tombstones, TTL, gc_grace and purging
- CQL table options including gc_grace_seconds and default TTL
- ALTER TABLE — gc_grace_seconds and TTL considerations
- cassandra.yaml — tombstone warn/failure thresholds and guardrails
- Repair operations
- TimeWindowCompactionStrategy
- nodetool tablestats
- nodetool tablehistograms