Trace deletes and TTL through tombstones, repair, and compaction; reproduce zombie resurrection safely and connect reclamation to operational evidence.
Tombstones, TTL, Compaction, Repair, and Why Deletes Can Be Operationally Expensive
Model AtlasMart documents with nested objects, arrays, dynamic fields, validation, and compatibility-safe schema evolution while distinguishing missing from explicit null.
Learning outcomes
Wide-column write paths often use immutable sorted files. An update or delete therefore does not erase every older byte immediately. A tombstone is deletion evidence that suppresses older values until the system can safely reclaim them. Time to live (TTL) is an expiry policy that eventually creates deletion state. Compaction rewrites immutable files to merge versions and reclaim obsolete data. Repair reconciles replicas that missed writes. The hard part is the interaction: deleting too aggressively can resurrect data, while retaining too much deletion state can make reads and compaction expensive.
Trace a delete/TTL through replicas, tombstones, reads, repair, and compaction.
Explain why safe tombstone reclamation depends on replica reconciliation.
Recognize tombstone-heavy queries and compaction debt as operational signals.
Inject a safe simulated failure that produces—and then prevents—zombie resurrection.
1. Delete is a distributed write, not an immediate absence
Suppose cart item cart-123 exists on replicas A, B,
and C. C is disconnected while A and B accept a delete. If A and
B merely erase their local copies, C still holds the old value.
A later repair sees C's value and cannot infer that “missing” on
A/B means “deliberately deleted” rather than “never received.” A
versioned tombstone carries that intent through the same
replication/repair machinery as a write.
The tombstone must remain long enough for the system's reconciliation assumptions. Cassandra exposes this through tombstone grace and repair behavior; exact settings and safe-purge rules are product/version/topology dependent. The vendor-neutral rule is stronger: do not discard deletion evidence until every stale copy that could legally reappear has been reconciled or otherwise made irrelevant.
2. The zombie resurrection failure, made observable
# Deterministic three-replica delete/repair model.
# Version tuples order writes: (logical_version, kind, value).
replicas = {
"A": (1, "value", "cart-123"),
"B": (1, "value", "cart-123"),
"C": (1, "value", "cart-123"),
}
# C is disconnected. Delete reaches A and B as tombstone version 2.
replicas["A"] = (2, "tombstone", None)
replicas["B"] = (2, "tombstone", None)
print("after delete while C is down:", replicas)
# Unsafe: A/B purge the tombstone before C has participated in repair.
unsafe = {"A": None, "B": None, "C": replicas["C"]}
print("unsafe post-purge state:", unsafe)
# Repair sees C's old value and no surviving deletion evidence, so the value can return.
for n in ("A", "B"):
unsafe[n] = unsafe["C"]
print("unsafe repair resurrects:", unsafe)
# Safe path: keep deletion evidence until disconnected replica has reconciled.
safe = {
"A": (2, "tombstone", None),
"B": (2, "tombstone", None),
"C": (1, "value", "cart-123"),
}
# Repair compares versions and propagates tombstone 2 to C.
safe["C"] = max(safe.values(), key=lambda x:x[0])
print("after repair convergence:", safe)
# Only after convergence and the operator's safe-reclamation policy may old state be compacted away.
safe = {k: None for k in safe}
print("after safe reclamation:", safe)
# Tombstone-heavy read-cost model: counts are illustrative, not a benchmark.
live = 120
tombstones = 880
print("scan candidates:", live+tombstones, "live rows:", live, "tombstones examined:", tombstones)
The unsafe history is concrete: C misses the tombstone, A/B purge it, and repair later copies C's older value back. In the safe history, version-2 deletion evidence reaches C before reclamation. The final counters also show a separate issue: a query that encounters 880 tombstones to return 120 live rows may be correct but operationally expensive. Those counts are an illustrative deterministic model, not a latency benchmark.
3. TTL, compaction, and repair form one lifecycle
TTL is not equivalent to cache eviction. TTL means data becomes logically expired according to the database's semantics; immutable storage may retain physical bytes and tombstones until maintenance can reclaim them. Compaction merges SSTables and can drop obsolete versions only when safety conditions hold. Repair reconciles replicas and therefore affects whether deletion evidence has reached the places that still contain old values. A topology change, extended outage, or missed repair can invalidate an operator's casual assumption that “the grace period passed, so purge is safe.”
| Signal | What it can indicate | What to investigate |
|---|---|---|
| tombstones/read rising | Delete/TTL-heavy access path or poor partition/bucket design | query slice, retention, bucket width, compaction state |
| compaction backlog | write/deletion rate exceeds maintenance capacity | disk bandwidth, strategy, SSTable count, headroom |
| repair age growing | replicas may retain divergent old state longer | repair schedule/failures/topology changes |
| disk usage remains high after TTL | expired data not yet reclaimable/compacted | SSTable overlap, compaction, repair, expiry layout |
4. Wrong approach: lower the deletion grace until disk graphs look better
Reducing a grace/reclamation interval can improve space metrics while silently weakening the assumption that every stale replica will see the deletion first. That is an availability/recovery trade, not housekeeping. The safe correction starts from maximum replica outage, repair frequency and success, topology changes, backup/restore behavior, and the product's purge rules. Test a disconnected replica in an isolated environment and prove deleted data stays deleted after repair and restart.
The lab simulates replica state in memory. Do not reproduce
the failure by stopping production nodes or changing
gc_grace_seconds without a product-specific
runbook, backups, repair evidence, and a rollback plan.
5. Production judgment
Delete/TTL-heavy workloads are viable only when compaction and repair are first-class capacity consumers. Measure tombstones scanned/returned row, SSTable count, compaction throughput/backlog, repair completion age, disk amplification, expired-but-not-reclaimed bytes, and read tail latency under realistic churn. Retention policy is also a governance control: legal holds and tenant deletion requirements may conflict with broad TTL defaults and backup retention. Security deletion is not proven until replicas, derived tables, backups, and restore procedures are considered.
Check your understanding
- Why does a distributed delete need explicit deletion evidence?
- What caused the zombie in the unsafe lab?
- Is TTL the same as immediate physical deletion?
- Why can tombstone-heavy reads be expensive?
- What evidence should precede aggressive tombstone reclamation?
Review the answers
1. A stale replica can otherwise reintroduce an older value because absence alone does not prove a deliberate newer delete.
2. The tombstone was purged before replica C reconciled, so C's old value became the only surviving version during repair.
3. No. Logical expiry can precede physical reclamation; tombstones/SSTables may remain until safe compaction.
4. The engine may examine many deletion markers/old versions to return a small number of live rows.
5. Successful repair/reconciliation, bounded outage assumptions, topology knowledge, backups/restore tests, and product-specific purge semantics.
Authoritative references
- Chang et al. — Bigtable: A Distributed Storage System for Structured Data (OSDI 2006) — primary research for the wide-column lineage, ordered row keys, column families, and sparse structured data.
- Apache Cassandra 5.0 — Data definition — current implementation reference for partition keys, clustering columns, composite primary keys, and clustering order.
- Apache Cassandra — Logical data modeling — query-first table design and primary-key reasoning.
- Apache Cassandra 5.0 — Tombstones — deletion markers, grace/reconciliation, repair interaction, and resurrection risk.
- Apache Cassandra 5.0 — Compaction overview — immutable SSTables, compaction, TTL/deletes, and read/space effects.
- Apache Cassandra releases — Cassandra 5.0.9 is the latest GA release as of the August 2026 review.