Chapter 13 · Tombstones, TTL, gc_grace, Deletes, and Expiration

Tombstone Warning / Failure Thresholds, Wide Partitions, and Range-Scan Hazards

Trigger a warning-scale tombstone scan and diagnose why thresholds protect the server rather than define a design target.

Intermediate120–160 minutesWide-partition tombstone warning labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · UCS baseline · TWCS specialized in Lesson 5Last reviewed: September 2026

Learning outcomes

AtlasMart's queue-like activity table returns only 50 live rows, yet the request scans a long prefix of deleted history and becomes slow. The result size looks tiny, but the storage work is not. Cassandra therefore has tombstone warning/failure thresholds and additional read/partition guardrails to expose or stop dangerous scans.

01

Explain why tombstones consume read-path memory/work even when the final result set is small.

02

Inspect Cassandra 5.0 tombstone warning/failure thresholds instead of memorizing undocumented values.

03

Create a bounded warning-scale lab with more than one thousand row tombstones but avoid a destructive hundred-thousand-tombstone failure test.

04

Use tracing, cqlsh warnings, logs, tablestats, tablehistograms, and partition shape to diagnose the scan.

05

Reject threshold-raising as a substitute for fixing partition/query/retention design.

Chapter 13 lab baseline

The mandatory labs continue the disposable AtlasMart course cluster: Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. The chapter keyspace is atlasmart_tombstones with NetworkTopologyStrategy and replication factor (RF) 3; reads and writes normally use LOCAL_QUORUM. Authentication, client TLS, internode TLS, and remote JMX remain disabled only inside this isolated local learning network. UnifiedCompactionStrategy (UCS) remains the default choice for ordinary chapter tables; Lesson 5 intentionally uses TimeWindowCompactionStrategy (TWCS) for a narrow expiring-data design. No application driver is required for mandatory work; cqlsh and nodetool provide the evidence. Exact metrics, tokens, timings, warning text, and SSTable filenames are learner-captured rather than invented.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms to define before reasoning about deletion

A coordinator is the Cassandra node handling one client request. A replica stores a copy of a partition according to the keyspace replication strategy. A partition is the unit identified by the partition key; its hash maps to a token and therefore a replica set. A consistency level (CL) specifies how many appropriately scoped replica responses the coordinator needs. A tombstone is a timestamped deletion marker written through Cassandra's normal write path. A Time To Live (TTL) is a duration after which a value expires and becomes logically deleted. gc_grace_seconds is the table-level grace period during which Cassandra normally retains deletion markers so replicas have time to converge before the markers become eligible for garbage collection. repair is anti-entropy synchronization between replicas. A hint is a best-effort mutation retained temporarily for an unavailable replica; it is not a replacement for repair. An SSTable (Sorted String Table) is immutable on-disk storage. Compaction rewrites SSTables, reconciles versions, and may purge tombstones only when Cassandra's safety conditions allow it.

The word delete therefore has two different meanings in Cassandra: a row can disappear from query results immediately while its physical deletion marker remains on disk for much longer. Conflating logical invisibility with physical reclamation is the root of many tombstone mistakes.

1. Thresholds are safety rails, not sizing targets

Current Cassandra 5.0 documentation lists tombstone_warn_threshold at 1000 and tombstone_failure_threshold at 100000 by default. The warning threshold surfaces expensive scans; the failure threshold protects a replica from processing an extreme number of tombstones in one query. Those defaults are version-sensitive configuration, so the lab reads the actual node's cassandra.yaml first. Cassandra 5.0 also includes separate guardrails for partition size and partition tombstone count at SSTable-write time; those are disabled unless configured in the current defaults.

Do not design “just under 1000.” Tombstone count is only one dimension. A query can be expensive because of wide partitions, many SSTables, range width, cache misses, replica filtering, payload size, or concurrent reads. Conversely, a narrow exact-key lookup may avoid most dead cells in a table that contains many tombstones elsewhere.

Docker · inspect the actual Cassandra 5.0 threshold configuration
docker exec atlasmart-cass-1 sh -lc \  "grep -n -E 'tombstone_(warn|failure)_threshold|partition_tombstones_(warn|fail)_threshold|partition_size_(warn|fail)_threshold' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 nodetool version

2. Build a warning-scale wide partition without approaching the failure threshold

Why 1,500 rows?

This is large enough to cross the default warning threshold after 1,200 row deletes but remains tiny compared with the default failure threshold. The mandatory lab intentionally does not generate 100,000+ tombstones because doing so adds resource risk without adding conceptual value.

Docker · verify the existing disposable cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool status# Continue only after all three nodes are UN in dc1.# If the shared course cluster does not exist, recreate it with the same# Chapter 01 conventions before running this chapter. Do not expose CQL/JMX# to untrusted networks merely to make the lab convenient.
CQL · create the chapter keyspace
CREATE KEYSPACE IF NOT EXISTS atlasmart_tombstonesWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name='atlasmart_tombstones';
CQL · create one deliberately wide lab partition
CREATE TABLE IF NOT EXISTS atlasmart_tombstones.wide_delete_probe (    bucket text,    seq int,    payload text,    PRIMARY KEY (bucket,seq)) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND gc_grace_seconds = 864000;

Generate the statements inside the Cassandra container so Windows users do not need a host Bash environment. One cqlsh process receives the entire stream.

Docker · insert 1,500 clustering rows in one partition
docker exec atlasmart-cass-1 sh -lc '{  echo "CONSISTENCY LOCAL_QUORUM;"  for i in $(seq 0 1499); do    printf "INSERT INTO atlasmart_tombstones.wide_delete_probe (bucket,seq,payload) VALUES ('\''wide'\'',%s,'\''live-%s'\'');\n" "$i" "$i"  done} | cqlsh'for n in 1 2 3; do  docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones wide_delete_probedone

Now delete the first 1,200 rows individually. Individual row deletes make the warning mechanism easy to reason about. A range delete would encode the dead clustering interval differently and is therefore not an equivalent tombstone-count experiment.

Docker · create 1,200 row tombstones
docker exec atlasmart-cass-1 sh -lc '{  echo "CONSISTENCY LOCAL_QUORUM;"  for i in $(seq 0 1199); do    printf "DELETE FROM atlasmart_tombstones.wide_delete_probe WHERE bucket='\''wide'\'' AND seq=%s;\n" "$i"  done} | cqlsh'for n in 1 2 3; do  docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones wide_delete_probedone

3. Query past the dead prefix and capture the warning evidence

The first remaining live row is at clustering key 1200. A query for the partition with LIMIT 50 may need to pass the 1,200 deleted rows before it can return 50 live rows. That is why final row count is a poor proxy for storage work.

CQL · traced broad scan over a dead prefix
CONSISTENCY LOCAL_QUORUM;TRACING ON;SELECT seq,payloadFROM atlasmart_tombstones.wide_delete_probeWHERE bucket='wide'LIMIT 50;TRACING OFF;
nodetool/Docker · correlate warnings with table evidence
docker exec atlasmart-cass-1 nodetool tablestats \  atlasmart_tombstones.wide_delete_probedocker exec atlasmart-cass-1 nodetool tablehistograms \  atlasmart_tombstones wide_delete_probe# Review recent server logs. Search the output for "tombstone",# "Scanned over", warning, or read-failure text; exact wording varies.docker logs --since 10m atlasmart-cass-1

Depending on which replica/coordinator processes the read, cqlsh may also print a server warning. Capture the observed tombstone count and latency. If background compaction changed the physical layout, your exact numbers may differ; explain that state rather than forcing the expected transcript.

4. Why raising the threshold is usually the wrong first fix

Raising the warning threshold suppresses evidence. Raising the failure threshold asks a replica to hold/process even more deletion state before protecting itself. Neither changes the partition, delete rate, SSTables, clustering slice, or retention model. A correct remediation might be a bounded time bucket, a query that starts near the live clustering region, a separate “active items” table maintained at write time, lower delete churn, or a retention model that lets whole time windows expire together.

Likewise, deleting a huge clustering range can be logically efficient for the application but still leaves a large dead interval until safe compaction. The safest model often avoids repeatedly mutating a long-lived queue-like partition in the first place.

Wrong approach: “The query failed at the tombstone threshold, so double the threshold.”

The threshold is reporting a data-model/query-path problem. First prove why the read crosses so much dead state. Change the model or access path, then validate latency and tombstone scans under representative load. Change the threshold only when you can explain the memory/latency budget and have tested failure behavior.

5. Cleanup and checks

CQL · remove only the disposable wide-partition fixture
DROP TABLE IF EXISTS atlasmart_tombstones.wide_delete_probe;

Check your understanding

  1. Why can a LIMIT 50 query scan more than a thousand tombstones?
  2. Why does the lab use individual row deletes instead of one range delete?
  3. Why not reproduce the default 100000 failure threshold?
  4. What does increasing tombstone_failure_threshold fix in the data model?
  5. Name two model-level remedies for queue-like dead prefixes.
Review the answers

1. LIMIT constrains returned live rows, not necessarily the number of dead cells/rows the storage engine must traverse to find them.

2. A range delete has a different tombstone representation; individual deletes make the warning-scale tombstone count observable and easier to correlate.

3. Generating that much dead state adds resource/heap/time risk with little educational benefit; inspecting the actual configured threshold and understanding its purpose is sufficient.

4. Nothing. It only changes when the server refuses the scan.

5. Time/bucket partitions, separate active-state query tables, bounded retention windows, or query shapes that do not repeatedly scan deleted history.

Production judgment

Deletion policy is part of the data model, repair plan, capacity model, and failure model. Before changing TTL, gc_grace_seconds, tombstone thresholds, compaction options, or retention windows, record RF/CL, maximum tolerated replica outage, hint window, repair cadence and success evidence, partition rows/bytes, delete and TTL rate, clustering scan width, SSTables-per-read, compaction debt, free disk, cache state, JVM/heap pressure, p95/p99 read/write latency, and the exact Cassandra patch/platform. Security and tenant boundaries also matter: an abusive tenant that creates very wide partitions or delete storms can create shared-node latency and heap risk.

Do not treat a local threshold or a blog value as a target. A large tombstone count can be harmless when queries touch narrow slices, while a smaller count can be damaging when every request scans a broad dead region. Hints reduce short-outage exposure but are not comprehensive anti-entropy. Lowering grace time transfers safety requirements to repair/outage discipline. SAI or vector indexes do not erase the storage semantics: indexed data still lives in SSTables and must coexist with deletes, compaction, repair, and retention. Driver timeouts/retries must be tested under real slow-read/failure conditions and only retried when operation semantics are safe. Lesson 5 redesigns the workload so expiration happens in bounded time buckets and TWCS can group similar-age SSTables, reducing the need to scan/delete giant mixed-age partitions.

Summary and next bridge

Tombstone thresholds protect the server from pathological scans; they are not recommended partition sizes. A small result can require large storage work when the read crosses dead history. The final lesson changes the model so retention, partition boundaries, and compaction windows cooperate.

Authoritative references

These official references are the source of truth for version-sensitive behavior. Re-check them when the course is regenerated because defaults, guardrails, tooling, and repair/compaction behavior can evolve.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.