Chapter 13 · Tombstones, TTL, gc_grace, Deletes, and Expiration
gc_grace_seconds, Repair Timing, Zombie Data, and Safe Tombstone Purging
Connect gc_grace_seconds to repair timing and deliberately reproduce the mechanism that can create zombie data.
Learning outcomes
AtlasMart wants faster disk reclamation and proposes reducing
gc_grace_seconds. The dangerous question is not
“does compaction reclaim space sooner?” but “can every replica
learn every delete before another replica is allowed to forget
it?” This lesson makes that correctness dependency concrete.
Explain gc_grace_seconds as a repair/outage safety budget and derive zombie-data risk from missed tombstones.
Distinguish hints, read repair, and anti-entropy repair in the deletion-convergence story.
Create a reversible RF=3 missed-delete experiment with hints disabled only in the disposable lab.
Observe a short-grace tombstone before/after compaction and understand when zombie resurrection can occur.
Use only_purge_repaired_tombstone and Cassandra 5.0 data-resurrection startup checks as optional safety mechanisms rather than excuses to skip repair.
The mandatory labs continue the disposable AtlasMart course
cluster: Apache Cassandra 5.0.9 in the pinned
cassandra:5.0.9 image, Java 17 inside the image,
cluster atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, and 16 virtual nodes per
node. The chapter keyspace is
atlasmart_tombstones with
NetworkTopologyStrategy and replication factor
(RF) 3; reads and writes normally use
LOCAL_QUORUM. Authentication, client TLS,
internode TLS, and remote JMX remain disabled only inside this
isolated local learning network. UnifiedCompactionStrategy
(UCS) remains the default choice for ordinary chapter tables;
Lesson 5 intentionally uses TimeWindowCompactionStrategy
(TWCS) for a narrow expiring-data design. No application
driver is required for mandatory work; cqlsh and nodetool
provide the evidence. Exact metrics, tokens, timings, warning
text, and SSTable filenames are learner-captured rather than
invented.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms to define before reasoning about deletion
A coordinator is the Cassandra node handling one client request. A replica stores a copy of a partition according to the keyspace replication strategy. A partition is the unit identified by the partition key; its hash maps to a token and therefore a replica set. A consistency level (CL) specifies how many appropriately scoped replica responses the coordinator needs. A tombstone is a timestamped deletion marker written through Cassandra's normal write path. A Time To Live (TTL) is a duration after which a value expires and becomes logically deleted. gc_grace_seconds is the table-level grace period during which Cassandra normally retains deletion markers so replicas have time to converge before the markers become eligible for garbage collection. repair is anti-entropy synchronization between replicas. A hint is a best-effort mutation retained temporarily for an unavailable replica; it is not a replacement for repair. An SSTable (Sorted String Table) is immutable on-disk storage. Compaction rewrites SSTables, reconciles versions, and may purge tombstones only when Cassandra's safety conditions allow it.
The word delete therefore has two different meanings in Cassandra: a row can disappear from query results immediately while its physical deletion marker remains on disk for much longer. Conflating logical invisibility with physical reclamation is the root of many tombstone mistakes.
1. Zombie data is old live state with no surviving deletion state to defeat it
With RF=3, suppose all replicas store v='alive'.
Node 3 becomes unavailable. A delete succeeds on nodes 1 and 2.
If nodes 1 and 2 keep the tombstone until node 3 receives it
through hints or repair, the old value cannot win. If nodes 1
and 2 purge the tombstone first, node 3's old live value may be
the only surviving state. Anti-entropy repair can then copy that
old value back to nodes 1 and 2. This reappearance is commonly
called a zombie.
The safety inequality is operational rather than purely mathematical: your grace period must cover realistic time to detect a failed/missed replica, restore it or replace it, and successfully repair the affected data—plus margin for failed repair runs and maintenance delay. Hints usually cover a much shorter horizon and are not comprehensive; a healthy repair cadence is still required.
| Mechanism | What it helps with | Why it is insufficient alone |
|---|---|---|
| Hinted handoff | short replica outages during writes/deletes | bounded retention; hints can be lost/expire and do not compare all existing data |
| Read repair | stale replicas involved in qualifying reads | request-scoped; unread partitions may remain divergent |
| Anti-entropy repair | systematic replica comparison across token ranges | must be scheduled, completed, monitored and repeated |
| gc_grace_seconds | keeps deletion knowledge around long enough | does nothing if repair discipline exceeds the grace budget |
| only_purge_repaired_tombstone=true | adds repaired-state gate before purge | can retain tombstones indefinitely when repair is unhealthy |
2. Build a deliberately unsafe short-grace fixture — isolated lab only
Use only the disposable zombie_probe table. Do
not copy its 15-second grace period into production. Do not
run these commands against a shared or valuable cluster.
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool status# Continue only after all three nodes are UN in dc1.# If the shared course cluster does not exist, recreate it with the same# Chapter 01 conventions before running this chapter. Do not expose CQL/JMX# to untrusted networks merely to make the lab convenient.
CREATE KEYSPACE IF NOT EXISTS atlasmart_tombstonesWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name='atlasmart_tombstones';
CREATE TABLE IF NOT EXISTS atlasmart_tombstones.zombie_probe ( k text PRIMARY KEY, v text) WITH compaction = { 'class':'UnifiedCompactionStrategy', 'only_purge_repaired_tombstone':'false' } AND gc_grace_seconds = 15;CONSISTENCY ALL;INSERT INTO atlasmart_tombstones.zombie_probe (k,v)VALUES ('zombie-42','old-live-value');SELECT * FROM atlasmart_tombstones.zombie_probe WHERE k='zombie-42';
for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones zombie_probedonedocker exec atlasmart-cass-1 nodetool getendpoints \ atlasmart_tombstones zombie_probe zombie-42docker exec atlasmart-cass-1 nodetool status
3. Make node 3 miss the delete, then let the live replicas forget it
Connect cqlsh to node 1 so node 1 is the request coordinator.
Disable hinted handoff temporarily on the two live nodes, pause
node 3, and issue a LOCAL_QUORUM delete. The delete
can succeed with nodes 1 and 2. Re-enable hints immediately
after the delete so the rest of the cluster returns to normal
behavior; because hints were disabled during the missed
mutation, node 3 should not receive this deletion through that
path.
docker exec atlasmart-cass-1 nodetool disablehandoffdocker exec atlasmart-cass-2 nodetool disablehandoffdocker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e \ "CONSISTENCY LOCAL_QUORUM; DELETE FROM atlasmart_tombstones.zombie_probe WHERE k='zombie-42';"docker exec atlasmart-cass-1 nodetool enablehandoffdocker exec atlasmart-cass-2 nodetool enablehandoff# Persist the delete on the two available replicas.docker exec atlasmart-cass-1 nodetool flush atlasmart_tombstones zombie_probedocker exec atlasmart-cass-2 nodetool flush atlasmart_tombstones zombie_probe
Wait at least 20 seconds—longer than this table's deliberately
tiny grace—then compact only nodes 1 and 2. Because
only_purge_repaired_tombstone is false and the
relevant SSTables are compacted together, the deletion marker
can become purgeable. Exact purge timing remains
implementation/state dependent; verify rather than assume.
# Wait >15 seconds before continuing.sleep 20docker exec atlasmart-cass-1 nodetool compact atlasmart_tombstones zombie_probedocker exec atlasmart-cass-2 nodetool compact atlasmart_tombstones zombie_probedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_tombstones.zombie_probedocker exec atlasmart-cass-2 nodetool tablestats atlasmart_tombstones.zombie_probedocker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
4. Verify old state, then show why repair after purge can be dangerous
Node 3 was paused before the delete and should still have its old SSTable. The cleanest evidence is a read-only snapshot of node 3 rather than relying on a coordinator's replica choice. Take a snapshot, inspect it from a read-only utility container, and then clear the snapshot.
docker exec atlasmart-cass-3 nodetool snapshot \ -t ch13-zombie -cf zombie_probe atlasmart_tombstonesdocker run --rm \ -v atlasmart-cass-3-data:/var/lib/cassandra:ro \ cassandra:5.0.9 sh -lc ' f=$(find /var/lib/cassandra/data/atlasmart_tombstones \ -path "*/snapshots/ch13-zombie/*Data.db" | head -1) echo "$f" sstabledump "$f" | grep -n -A8 -B3 "old-live-value" || true 'docker exec atlasmart-cass-3 nodetool clearsnapshot \ -t ch13-zombie atlasmart_tombstones
If the tombstone has actually been purged from nodes 1 and 2, a subsequent full repair has no deletion marker to send to node 3. The old value can therefore be streamed back and become visible again. If your run does not resurrect the value, that does not invalidate the model: it means some safety/physical state (for example a surviving tombstone or compaction detail) prevented the exact sequence. Capture SSTable evidence and explain the state you actually observed.
docker exec atlasmart-cass-1 nodetool repair --full \ atlasmart_tombstones zombie_probefor n in 1 2 3; do docker exec atlasmart-cass-$n cqlsh -e \ "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_tombstones.zombie_probe WHERE k='zombie-42';"done
Repair is exposing a deletion-safety failure created earlier by premature purge. The correct response is to keep deletion markers long enough and run repair reliably enough—not to avoid repair. On real clusters, repair is the mechanism that prevents replicas from drifting indefinitely.
5. Safer production patterns
Keep the normal default unless you can justify a shorter grace
with measured repair/outage behavior. For tables where repair
discipline is uncertain,
only_purge_repaired_tombstone=true can add a safety
gate, at the cost of retaining tombstones when repaired state is
stale. Cassandra 5.0's optional startup
check_data_resurrection can also prevent a node
that appears to have been offline longer than grace from
starting, reducing one class of resurrection risk. Both features
shift operational consequences; neither replaces a repair
program.
After the exercise, reset the hazardous table and handoff state:
docker exec atlasmart-cass-1 nodetool enablehandoffdocker exec atlasmart-cass-2 nodetool enablehandoffdocker exec atlasmart-cass-3 nodetool enablehandoffdocker exec atlasmart-cass-1 cqlsh -e \ "DROP TABLE IF EXISTS atlasmart_tombstones.zombie_probe;"docker exec atlasmart-cass-1 nodetool status
Check your understanding
- What creates zombie data in this experiment?
- Why were hints disabled only briefly?
- Why is gc_grace_seconds a repair policy parameter?
- What does only_purge_repaired_tombstone=true change?
- Does a failed zombie reproduction mean the risk is imaginary?
Review the answers
1. A replica retains old live data while other replicas purge the newer deletion marker before convergence; later anti-entropy has no tombstone to suppress the old value.
2. To make node 3 miss the delete intentionally. They were re-enabled immediately because the goal is a controlled educational divergence, not long-lived cluster damage.
3. Its safety depends on whether unavailable replicas can be repaired or otherwise converged before tombstones become purgeable.
4. It prevents tombstone purge on unrepaired SSTables, adding safety when repair is delayed but potentially retaining tombstones longer.
5. No. Purge/compaction/repaired-state details can preserve a tombstone in one run. The correctness model remains: if all deletion state disappears before a stale replica converges, old data can return.
Production judgment
Deletion policy is part of the data model, repair plan, capacity
model, and failure model. Before changing TTL,
gc_grace_seconds, tombstone thresholds, compaction
options, or retention windows, record RF/CL, maximum tolerated
replica outage, hint window, repair cadence and success
evidence, partition rows/bytes, delete and TTL rate, clustering
scan width, SSTables-per-read, compaction debt, free disk, cache
state, JVM/heap pressure, p95/p99 read/write latency, and the
exact Cassandra patch/platform. Security and tenant boundaries
also matter: an abusive tenant that creates very wide partitions
or delete storms can create shared-node latency and heap risk.
Do not treat a local threshold or a blog value as a target. A large tombstone count can be harmless when queries touch narrow slices, while a smaller count can be damaging when every request scans a broad dead region. Hints reduce short-outage exposure but are not comprehensive anti-entropy. Lowering grace time transfers safety requirements to repair/outage discipline. SAI or vector indexes do not erase the storage semantics: indexed data still lives in SSTables and must coexist with deletes, compaction, repair, and retention. Driver timeouts/retries must be tested under real slow-read/failure conditions and only retried when operation semantics are safe. Lesson 4 turns from correctness risk to availability/performance risk: broad reads that must process thousands of tombstones can warn, fail, or exhaust useful latency/heap budget.
Summary and next bridge
Grace is a convergence budget. Shortening it without repair/outage proof can turn a successful delete into future zombie data. Next, keep correctness intact but create a deliberately tombstone-heavy partition to see how scan thresholds protect the server from pathological reads.
Authoritative references
These official references are the source of truth for version-sensitive behavior. Re-check them when the course is regenerated because defaults, guardrails, tooling, and repair/compaction behavior can evolve.
- Apache Cassandra downloads / 5.0 release baseline
- Compaction overview — tombstones, TTL, gc_grace and purging
- CQL table options including gc_grace_seconds and default TTL
- ALTER TABLE — gc_grace_seconds and TTL considerations
- cassandra.yaml — tombstone warn/failure thresholds and guardrails
- Repair operations
- TimeWindowCompactionStrategy
- nodetool tablestats
- nodetool tablehistograms
- cassandra.yaml startup check_data_resurrection
- Compaction common option only_purge_repaired_tombstone