Chapter 13 · Tombstones, TTL, gc_grace, Deletes, and Expiration

Why Deletes Create Tombstones Instead of Immediate Physical Removal

Treat a delete as newer distributed state that must outlive replica divergence before physical reclamation is safe.

Intermediate105–145 minutesDelete-path + tombstone evidence labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · UCS baseline · TWCS specialized in Lesson 5Last reviewed: September 2026

Learning outcomes

AtlasMart receives a privacy deletion request for an order note. The API must stop returning the note immediately, yet one replica might be temporarily unavailable and immutable SSTables may still contain the older value. Cassandra cannot safely solve this by erasing one local file fragment. It must distribute a newer deletion state that can win during reads and repair.

01

Explain why Cassandra represents deletes as timestamped replicated mutations rather than immediate in-place physical removal.

02

Trace a cell or row delete through coordinator, replicas, memtables, commit log, SSTables, compaction, and later repair.

03

Explain the relationship between tombstone timestamps, last-write-wins reconciliation, gc_grace_seconds, and zombie-data risk.

04

Use cqlsh tracing, schema inspection, nodetool tablestats/tablehistograms, and SSTable counts to observe logical deletion versus physical persistence.

05

Reject shortcuts such as deleting SSTable files, setting gc_grace_seconds to zero casually, or assuming hints replace repair.

Chapter 13 lab baseline

The mandatory labs continue the disposable AtlasMart course cluster: Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, and 16 virtual nodes per node. The chapter keyspace is atlasmart_tombstones with NetworkTopologyStrategy and replication factor (RF) 3; reads and writes normally use LOCAL_QUORUM. Authentication, client TLS, internode TLS, and remote JMX remain disabled only inside this isolated local learning network. UnifiedCompactionStrategy (UCS) remains the default choice for ordinary chapter tables; Lesson 5 intentionally uses TimeWindowCompactionStrategy (TWCS) for a narrow expiring-data design. No application driver is required for mandatory work; cqlsh and nodetool provide the evidence. Exact metrics, tokens, timings, warning text, and SSTable filenames are learner-captured rather than invented.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms to define before reasoning about deletion

A coordinator is the Cassandra node handling one client request. A replica stores a copy of a partition according to the keyspace replication strategy. A partition is the unit identified by the partition key; its hash maps to a token and therefore a replica set. A consistency level (CL) specifies how many appropriately scoped replica responses the coordinator needs. A tombstone is a timestamped deletion marker written through Cassandra's normal write path. A Time To Live (TTL) is a duration after which a value expires and becomes logically deleted. gc_grace_seconds is the table-level grace period during which Cassandra normally retains deletion markers so replicas have time to converge before the markers become eligible for garbage collection. repair is anti-entropy synchronization between replicas. A hint is a best-effort mutation retained temporarily for an unavailable replica; it is not a replacement for repair. An SSTable (Sorted String Table) is immutable on-disk storage. Compaction rewrites SSTables, reconciles versions, and may purge tombstones only when Cassandra's safety conditions allow it.

The word delete therefore has two different meanings in Cassandra: a row can disappear from query results immediately while its physical deletion marker remains on disk for much longer. Conflating logical invisibility with physical reclamation is the root of many tombstone mistakes.

1. Cassandra writes a delete because replicas can disagree in time

Imagine RF=3 with the value note='gift' stored on replicas A, B, and C. Replica C becomes unavailable. A client deletes the note at LOCAL_QUORUM, so A and B accept a newer mutation that says the cell is deleted. If Cassandra merely removed bytes from A and B, C would still have the old value. A later repair would see C's live value with no competing deletion state and could copy it back. The deletion marker solves that ambiguity: its timestamp participates in reconciliation, so an older live value loses to the newer tombstone.

A tombstone follows the same durability path as other mutations: the coordinator sends it to replicas; each successful replica appends the mutation to its commit log and updates its memtable; flushes serialize it into immutable SSTables. Reads merge live cells and tombstones by timestamp. Compaction may eventually remove both the tombstone and older shadowed data, but only after the grace and overlap/repaired-state rules allow safe purging.

Moment Client-visible state Replica/storage state Key implication
Delete acknowledged value no longer returned at the requested CL tombstone exists in memtable/commit log and later SSTables logical deletion is immediate; reclamation is not
Replica outage available replicas carry newer tombstone down replica may still hold old live value repair/hints must deliver deletion state before purge
Grace elapsed still deleted logically tombstone becomes eligible, not guaranteed, for purge elapsed time alone does not remove bytes
Compaction safely covers overlap still deleted tombstone and shadowed older data may be removed physical reclamation is a compaction outcome
Unsafe purge before convergence may appear deleted temporarily old value can survive on isolated replica later repair can create zombie data

2. Build the smallest observable delete fixture

Docker · verify the existing disposable cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool status# Continue only after all three nodes are UN in dc1.# If the shared course cluster does not exist, recreate it with the same# Chapter 01 conventions before running this chapter. Do not expose CQL/JMX# to untrusted networks merely to make the lab convenient.
CQL · create the chapter keyspace
CREATE KEYSPACE IF NOT EXISTS atlasmart_tombstonesWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CONSISTENCY LOCAL_QUORUM;SELECT keyspace_name, replicationFROM system_schema.keyspacesWHERE keyspace_name='atlasmart_tombstones';
CQL · create a normal RF=3 tombstone table
CREATE TABLE IF NOT EXISTS atlasmart_tombstones.delete_probe (    customer_id text,    order_id text,    status text,    note text,    updated_at timestamp,    PRIMARY KEY (customer_id, order_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND gc_grace_seconds = 864000;CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_tombstones.delete_probe(customer_id,order_id,status,note,updated_at)VALUES ('cust-42','order-9001','PAID','gift wrap',toTimestamp(now()));SELECT * FROM atlasmart_tombstones.delete_probeWHERE customer_id='cust-42' AND order_id='order-9001';TRACING ON;DELETE noteFROM atlasmart_tombstones.delete_probeWHERE customer_id='cust-42' AND order_id='order-9001';SELECT * FROM atlasmart_tombstones.delete_probeWHERE customer_id='cust-42' AND order_id='order-9001';TRACING OFF;

The expected query result after the delete still contains the row but not the deleted note. Trace event wording and endpoints vary with coordinator choice and version. The trace proves a distributed request occurred; it does not by itself prove that every RF replica is already identical.

Docker/nodetool · record schema and physical evidence
docker exec atlasmart-cass-1 cqlsh -e "DESCRIBE TABLE atlasmart_tombstones.delete_probe"docker exec atlasmart-cass-1 nodetool tablestats atlasmart_tombstones.delete_probedocker exec atlasmart-cass-1 nodetool tablehistograms atlasmart_tombstones delete_probe# Force a flush only to make the disposable lab's physical state visible.for n in 1 2 3; do  docker exec atlasmart-cass-$n nodetool flush atlasmart_tombstones delete_probedonedocker exec atlasmart-cass-1 nodetool tablestats atlasmart_tombstones.delete_probe

After flush, the older live cell and/or the deletion marker may occupy immutable SSTables depending on timing and background compaction. The exact SSTable count is not a contract. Record it together with compaction state instead of expecting a magic number.

3. Why gc_grace_seconds is a convergence budget, not a cleanup timer

The default table value is 864000 seconds (ten days). That does not mean every tombstone consumes disk for exactly ten days, nor does it guarantee safety merely because the number is large. A tombstone must first age beyond the grace period and then be included in a compaction that can prove it is not still shadowing older data. If only_purge_repaired_tombstone is enabled, repaired-state requirements can additionally delay purging.

The operational meaning is more important than the default: your repair schedule and maximum outage must keep every replica from missing a delete until after other replicas have forgotten the deletion marker. Hinted handoff helps with shorter outages but hints are bounded and best effort. A node that is disconnected for too long can retain old data. Cassandra 5.0 also exposes an optional check_data_resurrection startup check that can prevent a node from starting when it appears to have been down longer than grace; whether to enable it is an operational policy decision, not a substitute for repair.

Wrong approach: “Deletes are expensive, so set gc_grace_seconds=0 everywhere.”

On a distributed table this removes the time safety margin that keeps deletion knowledge alive while replicas are unavailable. Zero can be appropriate in limited cases such as a single-node table, and official documentation discusses reduced values for carefully designed TTL-only tables, but a replicated delete-heavy table needs an explicit repair/outage proof before shortening grace.

4. Evidence checklist and reset

  • Confirm delete_probe uses RF=3 in dc1 and normal LOCAL_QUORUM.
  • Confirm the deleted cell no longer appears in the query result.
  • Record the table's actual gc_grace_seconds from schema, not memory.
  • Record SSTable count and tombstone/read metrics before and after the forced lab flush.
  • Do not delete any file beneath /var/lib/cassandra/data.
  • Leave all three nodes UN before moving to Lesson 2.
CQL/nodetool · final verification and local reset
docker exec atlasmart-cass-1 cqlsh -e \  "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_tombstones.delete_probe WHERE customer_id='cust-42' AND order_id='order-9001';"docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool tablestats \  atlasmart_tombstones.delete_probe# Keep the table for Lesson 2 context, but confirm no node is paused/down.

Check your understanding

  1. Why is an in-place byte deletion insufficient in a replicated Cassandra cluster?
  2. Does gc_grace_seconds mean the tombstone is physically deleted at that exact age?
  3. What is the ordering relation between an older live cell and a newer tombstone?
  4. Do hints replace repair?
  5. What did the forced flush prove?
Review the answers

1. A disconnected replica can retain the old live value. Without a newer deletion marker to win reconciliation, repair can copy that old value back.

2. No. It becomes eligible after the grace period; compaction and overlap/repaired-state safety determine when it can actually be purged.

3. The newer timestamped tombstone shadows the older live cell during reconciliation.

4. No. Hints are a bounded best-effort catch-up mechanism for unavailable replicas; repair is the anti-entropy mechanism that compares replica data.

5. It made memtable state eligible to appear in immutable SSTables for observation. It did not prove a production compaction schedule or a fixed SSTable count.

Production judgment

Deletion policy is part of the data model, repair plan, capacity model, and failure model. Before changing TTL, gc_grace_seconds, tombstone thresholds, compaction options, or retention windows, record RF/CL, maximum tolerated replica outage, hint window, repair cadence and success evidence, partition rows/bytes, delete and TTL rate, clustering scan width, SSTables-per-read, compaction debt, free disk, cache state, JVM/heap pressure, p95/p99 read/write latency, and the exact Cassandra patch/platform. Security and tenant boundaries also matter: an abusive tenant that creates very wide partitions or delete storms can create shared-node latency and heap risk.

Do not treat a local threshold or a blog value as a target. A large tombstone count can be harmless when queries touch narrow slices, while a smaller count can be damaging when every request scans a broad dead region. Hints reduce short-outage exposure but are not comprehensive anti-entropy. Lowering grace time transfers safety requirements to repair/outage discipline. SAI or vector indexes do not erase the storage semantics: indexed data still lives in SSTables and must coexist with deletes, compaction, repair, and retention. Driver timeouts/retries must be tested under real slow-read/failure conditions and only retried when operation semantics are safe. Lesson 2 now distinguishes the tombstone shapes created by row, cell, range, collection, and TTL operations and measures how those shapes change query cost.

Summary and next bridge

A Cassandra delete is distributed state. Tombstones let replicas remember that a value was deleted until convergence is sufficiently likely; compaction later reclaims physical space when safe. Next, identify exactly which tombstone shape a CQL operation produces and why wide scans across dead data are costly.

Authoritative references

These official references are the source of truth for version-sensitive behavior. Re-check them when the course is regenerated because defaults, guardrails, tooling, and repair/compaction behavior can evolve.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.