Chapter 24 · Repair, Anti-Entropy, Incremental/Full Strategies, and Data Convergence

Why Repair Is Required Even with Hinted Handoff and Read Repair

Prove why hints and request-scoped read repair cannot replace anti-entropy, then create and repair a controlled RF=3 replica divergence.

Advanced120–160 minutesControlled divergence + full preview/repair labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart has an uncomfortable incident report: a replica was down long enough to miss a write, hinted handoff later looked healthy, and common reads still succeeded. The operations team wants to mark the cluster “converged.” That conclusion is unsafe. This lesson proves why hints and read repair are useful but incomplete, then uses anti-entropy repair to compare actual replica ranges and stream the difference.

01

Explain exactly why hinted handoff and request-scoped read repair cannot prove that all replicated data has converged.

02

Create a reversible divergence in an RF=3 table without depending on wall-clock or host-firewall manipulation.

03

Use full repair preview as Merkle/range evidence before streaming and full repair as explicit convergence work.

04

Explain why one nodetool repair command repairs ranges for that node rather than automatically covering every node/range in a larger cluster.

05

Verify convergence with repair validation/replica-aware evidence instead of trusting only a zero exit code.

Chapter 24 anti-entropy lab baseline

Repair work is isolated from every earlier course cluster. Mandatory labs use Docker network atlasmart-cassandra-repair, cluster atlasmart-repair, nodes atlasmart-repair-1..3, pinned Docker Official Image cassandra:5.0.9, Java 17 inside the image, datacenter dc1, racks rack1..rack3, and 16 virtual nodes (vnodes) per node. Keyspace atlasmart_repair uses NetworkTopologyStrategy with replication factor (RF) 3 in dc1; ordinary test traffic uses LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no default time-to-live (TTL), gc_grace_seconds = 864000 unless a lesson explicitly creates a disposable shorter-grace comparison, and read_repair = 'NONE' on divergence fixtures so request-scoped read repair cannot hide the anti-entropy experiment. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only inside this isolated single-host learning network; no native/JMX port is published to the host. Recommended lab headroom is roughly 8 GiB of available host RAM plus at least 10 GiB free disk; resource-constrained learners can reduce seed rows while preserving the same mechanism. Exact tokens, Merkle-tree depth/hashes, repair session IDs, validation duration, stream bytes, SSTable counts, repaired percentages, disk/CPU/network utilization and p50/p95/p99 application latency are learner-captured evidence, not promised constants.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and anti-entropy mental model

Apache Cassandra stores a logical partition on multiple replicas chosen from its token ownership and keyspace replication strategy. A request-scoped coordinator is whichever node handles one client operation; it is not a permanent leader. A consistency level (CL) is the number/scope of replica responses required before that operation can succeed. A hint is a best-effort record of a mutation that a temporarily unavailable replica missed. Read repair is request-scoped reconciliation/write-back that may happen when replicas consulted by a read disagree; it does not scan unread data. Anti-entropy repair is the operator/scheduler-driven process that compares replicas for token ranges and streams differences so replicas converge even when no client happens to read the affected partition.

A token range is a portion of the partitioner hash space. During repair, replicas validate common ranges and summarize their contents with Merkle trees: hierarchical hashes that let Cassandra narrow mismatches without sending every row over the network. A mismatch causes streaming of data differences between replicas. Incremental repair, the current nodetool repair default, works on unrepaired/pending-repair data and—after a consistent session—separates repaired from unrepaired data through anticompaction or equivalent repaired-state handling. A full repair uses --full and compares all data in the selected ranges, including data already marked repaired.

An immutable SSTable (Sorted String Table) can be classified as repaired, unrepaired, or pending repair. repairedAt is persisted repair-state metadata associated with repaired SSTables/ranges. gc_grace_seconds is a table-level grace period after which old tombstones can become eligible for purge; repair cadence must leave enough margin that every replica receives deletes before tombstones disappear. Anticompaction can rewrite SSTables to separate data covered by a successful incremental repair from data that remains unrepaired, which is why incremental repair consumes temporary disk and I/O even when little network streaming is required.

1. Hints and read repair cover opportunities, not all data

Hinted handoff is bounded and best-effort. A coordinator may store a missed mutation for a replica that is unavailable, but hints can age out, be disabled, fail to replay, or never exist for writes coordinated elsewhere. Read repair only sees partitions and replicas that a read actually touches; a cold partition can remain inconsistent indefinitely if no anti-entropy process compares it. Therefore “no pending hints” and “our normal reads work” are useful observations, but neither establishes replica-wide equality.

Mechanism Trigger/scope Strength What it cannot prove
Hinted handoff missed write observed by a coordinator within hint policy fast best-effort catch-up after short outages that every missed mutation has a surviving hint or replayed successfully
Read reconciliation / read repair replicas consulted for one read returns/reconciles newest visible values for that request; may write back depending on table policy that unread partitions or uninvolved replicas are equal
Incremental anti-entropy repair unrepaired ranges selected by repair regularly converges new/unrepaired data and marks successful ranges repaired that previously repaired data is free from corruption/operator loss
Full anti-entropy repair all data in selected ranges rechecks both repaired and unrepaired data that unselected ranges/nodes/DCs have also been repaired
Misleading approach: “No pending hints means no repair is required.”

Hints are not a durable anti-entropy ledger. They optimize short outage recovery. The repair process is the mechanism that compares replica datasets for common token ranges and streams mismatches.

2. Build a known-consistent baseline, then force one missed mutation

Docker · create the isolated three-replica repair topology
docker network inspect atlasmart-cassandra-repair >/dev/null 2>&1 || docker network create atlasmart-cassandra-repairdocker volume create atlasmart-repair-1-datadocker volume create atlasmart-repair-2-datadocker volume create atlasmart-repair-3-datadocker run -d --name atlasmart-repair-1 --hostname atlasmart-repair-1 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-repair-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-repair-1 nodetool statusdocker run -d --name atlasmart-repair-2 --hostname atlasmart-repair-2 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-repair-3 --hostname atlasmart-repair-3 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-3-data:/var/lib/cassandra cassandra:5.0.9# Do not start a repair until at least two observers show all three nodes UN.docker exec atlasmart-repair-1 nodetool versiondocker exec atlasmart-repair-1 java -versiondocker exec atlasmart-repair-1 nodetool statusdocker exec atlasmart-repair-2 nodetool status
CQL · create the RF=3 repair fixture and seed consistent data
CREATE KEYSPACE IF NOT EXISTS atlasmart_repairWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_repair.repair_probe (    tenant_id text,    item_id int,    status text,    note text,    updated_at timestamp,    PRIMARY KEY ((tenant_id), item_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND gc_grace_seconds = 864000  AND read_repair = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_repair.repair_probe(tenant_id,item_id,status,note,updated_at)VALUES ('tenant-001',1,'BASELINE','present on all replicas','2026-09-08T10:00:00Z');DESCRIBE TABLE atlasmart_repair.repair_probe;SELECT * FROM atlasmart_repair.repair_probe WHERE tenant_id='tenant-001';
Docker/cqlsh · generate a small distributed fixture without giant CQL batches
docker exec atlasmart-repair-1 bash -lc 'rm -f /tmp/repair_seed.cqlprintf "CONSISTENCY ALL;\n" > /tmp/repair_seed.cqlfor t in $(seq -w 1 64); do  for i in $(seq 1 4); do    printf "INSERT INTO atlasmart_repair.repair_probe (tenant_id,item_id,status,note,updated_at) VALUES (\047tenant-%s\047,%s,\047BASELINE\047,\047seed\047,toTimestamp(now()));\n" "$t" "$i" >> /tmp/repair_seed.cql  donedonecqlsh -f /tmp/repair_seed.cql'# Force flushes only to make immutable repair state observable in this disposable lab.for n in 1 2 3; do docker exec atlasmart-repair-$n nodetool flush atlasmart_repair repair_probe; done
Failure injection · make node 3 miss a write without leaving a hint
# Hints are disabled only on the coordinator that will issue the mutation.docker exec atlasmart-repair-1 nodetool disablehandoffdocker pause atlasmart-repair-3docker exec atlasmart-repair-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_repair.repair_probe SET status='NEWER_ON_1_AND_2', note='node3 missed this mutation', updated_at=toTimestamp(now()) WHERE tenant_id='tenant-001' AND item_id=1;"# Re-enable future hint storage before the failed replica returns; no hint is created retroactively.docker exec atlasmart-repair-1 nodetool enablehandoffdocker unpause atlasmart-repair-3# Wait for all nodes to be UN again before repair/preview.docker exec atlasmart-repair-1 nodetool statusfor n in 1 2 3; do docker exec atlasmart-repair-$n nodetool flush atlasmart_repair repair_probe; done

The fixture disables hint storage only on the coordinator that issues the one divergent mutation, pauses only the disposable third container, restores future hint storage, then returns the node. The table also has read_repair='NONE' so an exploratory read does not silently write the newer value to the stale replica. Avoid claiming that a cqlsh connection to node 3 necessarily forces node 3 to serve the data at LOCAL_ONE; the contacted node is the coordinator and may choose another replica. Use tracing only to see which replica answered.

CQL · trace a low-CL read without repairing the stale replica
CONSISTENCY LOCAL_ONE;TRACING ON;SELECT tenant_id,item_id,status,note,updated_atFROM atlasmart_repair.repair_probeWHERE tenant_id='tenant-001' AND item_id=1;TRACING OFF;

3. Preview comparison, then repair explicitly

--preview builds and compares Merkle trees and estimates streaming without actually changing replica data. Add --full so the preview includes all data in the selected ranges rather than only unrepaired data. Tiny datasets can produce very small estimates; the important signal is whether the preview finds out-of-sync ranges/bytes.

nodetool · preview full anti-entropy work
docker exec atlasmart-repair-1 nodetool repair --preview --full atlasmart_repair repair_probe# Correlate the preview with logs and current resources.docker logs --since 10m atlasmart-repair-1 2>&1 | grep -Ei 'repair|validation|merkle|stream' || truedocker exec atlasmart-repair-1 nodetool netstats -Hdocker exec atlasmart-repair-1 nodetool proxyhistograms
nodetool · perform a full repair for node 1 replicated ranges
# Run foreground client traffic from another terminal if you want latency evidence.docker exec atlasmart-repair-1 nodetool repair --full atlasmart_repair repair_probedocker exec atlasmart-repair-1 nodetool repair_admin list --alldocker exec atlasmart-repair-1 nodetool netstats -Hdocker exec atlasmart-repair-1 nodetool tablestats atlasmart_repair.repair_probe

A successful command proves that the selected repair command completed from this coordinator. On this three-node RF=3 teaching topology, every range is replicated everywhere, so one-node full repair happens to compare a broad set. Do not generalize that shortcut: in a real larger cluster, one node owns/replicates only part of the ring. For a complete primary-range cycle, run the planned -pr sequence on every node in every relevant DC (or use a validated repair scheduler).

4. Verify convergence, not merely process completion

nodetool/CQL · verify after repair
# Validate repaired-data consistency where applicable; preview full data again for residual stream estimate.docker exec atlasmart-repair-1 nodetool repair --preview --full atlasmart_repair repair_probedocker exec atlasmart-repair-1 nodetool repair --validate atlasmart_repair repair_probe || true# Read at ALL so every replica must respond; tracing shows request participants.docker exec atlasmart-repair-1 cqlsh -e "CONSISTENCY ALL; TRACING ON; SELECT tenant_id,item_id,status,note FROM atlasmart_repair.repair_probe WHERE tenant_id='tenant-001' AND item_id=1; TRACING OFF;"docker exec atlasmart-repair-1 nodetool tablestats atlasmart_repair.repair_probedocker exec atlasmart-repair-1 nodetool repair_admin summarize-repaired -v

--validate checks repaired data without streaming; if the table has not yet accumulated incrementally repaired state, its usefulness is correspondingly limited. A post-repair full preview estimating no remaining differences plus a successful ALL read and stable repair/streaming evidence is stronger than any one signal alone. Production acceptance should also inspect application errors/latency and the whole scheduled token-space coverage.

Check your understanding

  1. Why can hints be empty while replicas still differ?
  2. Why can read repair miss stale data?
  3. What does repair preview do?
  4. Does one successful nodetool repair prove a whole large cluster is repaired?
  5. Why verify after the command exits?
Review the answers

1. Hints are bounded best-effort records, not a complete history of missed writes; a mutation may have no surviving hint or may have been coordinated where no useful hint remained.

2. It only operates on data/replicas involved in actual reads, so cold partitions can remain divergent.

3. It builds/compares Merkle trees for the selected repair scope and estimates streaming without changing replica data.

4. No. The command covers the ranges selected on the targeted repair coordinator; cluster-wide coverage requires a planned sequence/scheduler across nodes/DCs.

5. A successful process exit is necessary but not sufficient evidence for business-level convergence, range coverage and acceptable operational impact.

Docker · reset only the dedicated Chapter 24 repair lab
docker rm -f atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3 2>/dev/null || truedocker volume rm atlasmart-repair-1-data atlasmart-repair-2-data atlasmart-repair-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra-repair 2>/dev/null || true

Production judgment

Repair is a distributed maintenance workload that competes with foreground reads/writes for disk bandwidth, page cache, CPU, network, compaction capacity and JVM time. Plan it from the actual token/replica topology, RF and consistency levels, table sizes, SSTable overlap, tombstone/delete/TTL rate, shortest gc_grace_seconds, repair duration variance, inter-DC bandwidth, rack/zone maintenance, failure probability, SAI/vector indexes, snapshot/backup windows, compaction strategy, disk free space and business p95/p99 latency objectives. Incremental repair reduces repeated scope when run continuously but introduces repaired/unrepaired separation and anticompaction cost; full repair is broader and remains necessary for cases incremental repair intentionally skips, including periodically checking previously repaired data.

Do not use a single “repair every N days” value without measuring whether the entire required token-space/replica set actually completes within that interval. The safety condition is completion before tombstone grace can expire on unrepaired replicas, with margin for retries, outages and maintenance. Treat --force, aggressive parallelism, high -j, cross-DC sessions and throughput-cap changes as controlled operational choices with rollback and SLO gates. Managed Cassandra services may schedule repair internally or expose different controls; confirm who owns anti-entropy and what convergence evidence is available instead of assuming Apache nodetool semantics are exposed. Lesson 2 opens the repair mechanism itself: how Merkle trees summarize token ranges, how mismatched subranges become streaming sessions, and how repair session state is monitored/cancelled safely.

Summary and next bridge

Hints reduce outage-recovery cost and read repair can heal data that happens to be read, but neither is a complete anti-entropy strategy. Repair exists to compare replicated ranges deliberately. Next, trace the validation, Merkle-tree, streaming and repair-session mechanics that turn that comparison into convergence.

Authoritative references

Repair behavior and command options are version-sensitive. Re-check these sources before carrying a runbook to a newer Cassandra patch or managed service.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.