Chapter 13 · Tombstones, TTL, gc_grace, Deletes, and Expiration

gc_grace_seconds, Repair Timing, Zombie Data, and Safe Tombstone Purging

Connect gc_grace_seconds to repair timing and deliberately reproduce the mechanism that can create zombie data.

Intermediate120–165 minutesZombie-data + repair timing labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · UCS baseline · TWCS specialized in Lesson 5Last reviewed: September 2026

Learning outcomes

AtlasMart wants faster disk reclamation and proposes reducing gc_grace_seconds. The dangerous question is not “does compaction reclaim space sooner?” but “can every replica learn every delete before another replica is allowed to forget it?” This lesson makes that correctness dependency concrete.

01

Explain gc_grace_seconds as a repair/outage safety budget and derive zombie-data risk from missed tombstones.

02

Distinguish hints, read repair, and anti-entropy repair in the deletion-convergence story.

03

Create a reversible RF=3 missed-delete experiment with hints disabled only in the disposable lab.

04

Observe a short-grace tombstone before/after compaction and understand when zombie resurrection can occur.

05

Use only_purge_repaired_tombstone and Cassandra 5.0 data-resurrection startup checks as optional safety mechanisms rather than excuses to skip repair.

Chapter 13 lab baseline

The mandatory labs continue the disposable AtlasMart course cluster: Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. The chapter keyspace is atlasmart_tombstones with NetworkTopologyStrategy and replication factor (RF) 3; reads and writes normally use LOCAL_QUORUM. Authentication, client TLS, internode TLS, and remote JMX remain disabled only inside this isolated local learning network. UnifiedCompactionStrategy (UCS) remains the default choice for ordinary chapter tables; Lesson 5 intentionally uses TimeWindowCompactionStrategy (TWCS) for a narrow expiring-data design. No application driver is required for mandatory work; cqlsh and nodetool provide the evidence. Exact metrics, tokens, timings, warning text, and SSTable filenames are learner-captured rather than invented.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms to define before reasoning about deletion

A coordinator is the Cassandra node handling one client request. A replica stores a copy of a partition according to the keyspace replication strategy. A partition is the unit identified by the partition key; its hash maps to a token and therefore a replica set. A consistency level (CL) specifies how many appropriately scoped replica responses the coordinator needs. A tombstone is a timestamped deletion marker written through Cassandra's normal write path. A Time To Live (TTL) is a duration after which a value expires and becomes logically deleted. gc_grace_seconds is the table-level grace period during which Cassandra normally retains deletion markers so replicas have time to converge before the markers become eligible for garbage collection. repair is anti-entropy synchronization between replicas. A hint is a best-effort mutation retained temporarily for an unavailable replica; it is not a replacement for repair. An SSTable (Sorted String Table) is immutable on-disk storage. Compaction rewrites SSTables, reconciles versions, and may purge tombstones only when Cassandra's safety conditions allow it.

The word delete therefore has two different meanings in Cassandra: a row can disappear from query results immediately while its physical deletion marker remains on disk for much longer. Conflating logical invisibility with physical reclamation is the root of many tombstone mistakes.

1. Zombie data is old live state with no surviving deletion state to defeat it

With RF=3, suppose all replicas store v='alive'. Node 3 becomes unavailable. A delete succeeds on nodes 1 and 2. If nodes 1 and 2 keep the tombstone until node 3 receives it through hints or repair, the old value cannot win. If nodes 1 and 2 purge the tombstone first, node 3's old live value may be the only surviving state. Anti-entropy repair can then copy that old value back to nodes 1 and 2. This reappearance is commonly called a zombie.

The safety inequality is operational rather than purely mathematical: your grace period must cover realistic time to detect a failed/missed replica, restore it or replace it, and successfully repair the affected data—plus margin for failed repair runs and maintenance delay. Hints usually cover a much shorter horizon and are not comprehensive; a healthy repair cadence is still required.

Mechanism What it helps with Why it is insufficient alone
Hinted handoff short replica outages during writes/deletes bounded retention; hints can be lost/expire and do not compare all existing data
Read repair stale replicas involved in qualifying reads request-scoped; unread partitions may remain divergent
Anti-entropy repair systematic replica comparison across token ranges must be scheduled, completed, monitored and repeated
gc_grace_seconds keeps deletion knowledge around long enough does nothing if repair discipline exceeds the grace budget
only_purge_repaired_tombstone=true adds repaired-state gate before purge can retain tombstones indefinitely when repair is unhealthy

2. Build a deliberately unsafe short-grace fixture — isolated lab only

This experiment intentionally creates zombie risk.

Use only the disposable zombie_probe table. Do not copy its 15-second grace period into production. Do not run these commands against a shared or valuable cluster.

Docker · verify the existing disposable cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool status# Continue only after all three nodes are UN in dc1.# If the shared course cluster does not exist, recreate it with the same# Chapter 01 conventions before running this chapter. Do not expose CQL/JMX# to untrusted networks merely to make the lab convenient.
CQL · create the chapter keyspace
CREATE KEYSPACE IF NOT EXISTS atlasmart_tombstonesWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name='atlasmart_tombstones';
CQL · short-grace zombie probe
CREATE TABLE IF NOT EXISTS atlasmart_tombstones.zombie_probe (    k text PRIMARY KEY,    v text) WITH compaction = {      'class':'UnifiedCompactionStrategy',      'only_purge_repaired_tombstone':'false'    }  AND gc_grace_seconds = 15;CONSISTENCY ALL;INSERT INTO atlasmart_tombstones.zombie_probe (k,v)VALUES ('zombie-42','old-live-value');SELECT * FROM atlasmart_tombstones.zombie_probe WHERE k='zombie-42';
Docker/nodetool · persist baseline and verify all replicas own the key
for n in 1 2 3; do  docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones zombie_probedonedocker exec atlasmart-cass-1 nodetool getendpoints \  atlasmart_tombstones zombie_probe zombie-42docker exec atlasmart-cass-1 nodetool status

3. Make node 3 miss the delete, then let the live replicas forget it

Connect cqlsh to node 1 so node 1 is the request coordinator. Disable hinted handoff temporarily on the two live nodes, pause node 3, and issue a LOCAL_QUORUM delete. The delete can succeed with nodes 1 and 2. Re-enable hints immediately after the delete so the rest of the cluster returns to normal behavior; because hints were disabled during the missed mutation, node 3 should not receive this deletion through that path.

Docker/nodetool · controlled divergence
docker exec atlasmart-cass-1 nodetool disablehandoffdocker exec atlasmart-cass-2 nodetool disablehandoffdocker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e \  "CONSISTENCY LOCAL_QUORUM; DELETE FROM atlasmart_tombstones.zombie_probe WHERE k='zombie-42';"docker exec atlasmart-cass-1 nodetool enablehandoffdocker exec atlasmart-cass-2 nodetool enablehandoff# Persist the delete on the two available replicas.docker exec atlasmart-cass-1 nodetool flush atlasmart_tombstones zombie_probedocker exec atlasmart-cass-2 nodetool flush atlasmart_tombstones zombie_probe

Wait at least 20 seconds—longer than this table's deliberately tiny grace—then compact only nodes 1 and 2. Because only_purge_repaired_tombstone is false and the relevant SSTables are compacted together, the deletion marker can become purgeable. Exact purge timing remains implementation/state dependent; verify rather than assume.

Docker/nodetool · cross the unsafe grace window and compact live replicas
# Wait >15 seconds before continuing.sleep 20docker exec atlasmart-cass-1 nodetool compact atlasmart_tombstones zombie_probedocker exec atlasmart-cass-2 nodetool compact atlasmart_tombstones zombie_probedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_tombstones.zombie_probedocker exec atlasmart-cass-2 nodetool tablestats atlasmart_tombstones.zombie_probedocker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status

4. Verify old state, then show why repair after purge can be dangerous

Node 3 was paused before the delete and should still have its old SSTable. The cleanest evidence is a read-only snapshot of node 3 rather than relying on a coordinator's replica choice. Take a snapshot, inspect it from a read-only utility container, and then clear the snapshot.

Docker · inspect node-3 snapshot for the old value
docker exec atlasmart-cass-3 nodetool snapshot \  -t ch13-zombie -cf zombie_probe atlasmart_tombstonesdocker run --rm \  -v atlasmart-cass-3-data:/var/lib/cassandra:ro \  cassandra:5.0.9 sh -lc '    f=$(find /var/lib/cassandra/data/atlasmart_tombstones \      -path "*/snapshots/ch13-zombie/*Data.db" | head -1)    echo "$f"    sstabledump "$f" | grep -n -A8 -B3 "old-live-value" || true  'docker exec atlasmart-cass-3 nodetool clearsnapshot \  -t ch13-zombie atlasmart_tombstones

If the tombstone has actually been purged from nodes 1 and 2, a subsequent full repair has no deletion marker to send to node 3. The old value can therefore be streamed back and become visible again. If your run does not resurrect the value, that does not invalidate the model: it means some safety/physical state (for example a surviving tombstone or compaction detail) prevented the exact sequence. Capture SSTable evidence and explain the state you actually observed.

nodetool · deliberately run repair, then verify all nodes
docker exec atlasmart-cass-1 nodetool repair --full \  atlasmart_tombstones zombie_probefor n in 1 2 3; do  docker exec atlasmart-cass-$n cqlsh -e \    "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_tombstones.zombie_probe WHERE k='zombie-42';"done
Wrong operational inference: “The zombie lab means repair is dangerous.”

Repair is exposing a deletion-safety failure created earlier by premature purge. The correct response is to keep deletion markers long enough and run repair reliably enough—not to avoid repair. On real clusters, repair is the mechanism that prevents replicas from drifting indefinitely.

5. Safer production patterns

Keep the normal default unless you can justify a shorter grace with measured repair/outage behavior. For tables where repair discipline is uncertain, only_purge_repaired_tombstone=true can add a safety gate, at the cost of retaining tombstones when repaired state is stale. Cassandra 5.0's optional startup check_data_resurrection can also prevent a node that appears to have been offline longer than grace from starting, reducing one class of resurrection risk. Both features shift operational consequences; neither replaces a repair program.

After the exercise, reset the hazardous table and handoff state:

CQL/Docker · cleanup the deliberately unsafe fixture
docker exec atlasmart-cass-1 nodetool enablehandoffdocker exec atlasmart-cass-2 nodetool enablehandoffdocker exec atlasmart-cass-3 nodetool enablehandoffdocker exec atlasmart-cass-1 cqlsh -e \  "DROP TABLE IF EXISTS atlasmart_tombstones.zombie_probe;"docker exec atlasmart-cass-1 nodetool status

Check your understanding

  1. What creates zombie data in this experiment?
  2. Why were hints disabled only briefly?
  3. Why is gc_grace_seconds a repair policy parameter?
  4. What does only_purge_repaired_tombstone=true change?
  5. Does a failed zombie reproduction mean the risk is imaginary?
Review the answers

1. A replica retains old live data while other replicas purge the newer deletion marker before convergence; later anti-entropy has no tombstone to suppress the old value.

2. To make node 3 miss the delete intentionally. They were re-enabled immediately because the goal is a controlled educational divergence, not long-lived cluster damage.

3. Its safety depends on whether unavailable replicas can be repaired or otherwise converged before tombstones become purgeable.

4. It prevents tombstone purge on unrepaired SSTables, adding safety when repair is delayed but potentially retaining tombstones longer.

5. No. Purge/compaction/repaired-state details can preserve a tombstone in one run. The correctness model remains: if all deletion state disappears before a stale replica converges, old data can return.

Production judgment

Deletion policy is part of the data model, repair plan, capacity model, and failure model. Before changing TTL, gc_grace_seconds, tombstone thresholds, compaction options, or retention windows, record RF/CL, maximum tolerated replica outage, hint window, repair cadence and success evidence, partition rows/bytes, delete and TTL rate, clustering scan width, SSTables-per-read, compaction debt, free disk, cache state, JVM/heap pressure, p95/p99 read/write latency, and the exact Cassandra patch/platform. Security and tenant boundaries also matter: an abusive tenant that creates very wide partitions or delete storms can create shared-node latency and heap risk.

Do not treat a local threshold or a blog value as a target. A large tombstone count can be harmless when queries touch narrow slices, while a smaller count can be damaging when every request scans a broad dead region. Hints reduce short-outage exposure but are not comprehensive anti-entropy. Lowering grace time transfers safety requirements to repair/outage discipline. SAI or vector indexes do not erase the storage semantics: indexed data still lives in SSTables and must coexist with deletes, compaction, repair, and retention. Driver timeouts/retries must be tested under real slow-read/failure conditions and only retried when operation semantics are safe. Lesson 4 turns from correctness risk to availability/performance risk: broad reads that must process thousands of tombstones can warn, fail, or exhaust useful latency/heap budget.

Summary and next bridge

Grace is a convergence budget. Shortening it without repair/outage proof can turn a successful delete into future zombie data. Next, keep correctness intact but create a deliberately tombstone-heavy partition to see how scan thresholds protect the server from pathological reads.

Authoritative references

These official references are the source of truth for version-sensitive behavior. Re-check them when the course is regenerated because defaults, guardrails, tooling, and repair/compaction behavior can evolve.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.