Chapter 24 · Repair, Anti-Entropy, Incremental/Full Strategies, and Data Convergence

Merkle Trees, Token Ranges, Streaming Differences, and Repair Session Coordination

Trace repair coordination from token-range validation and Merkle-tree comparison through streaming, pending repair state, and session administration.

Advanced120–170 minutesMerkle/range + repair-session labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart now accepts that repair is mandatory, but the runbook still says only “run nodetool repair.” That is not enough to operate safely. This lesson breaks one repair into validation/range comparison, Merkle-tree mismatch localization, streaming and session coordination so an operator can tell useful work from a stalled or over-broad repair.

01

Explain how token ranges and replica sets define the unit of anti-entropy comparison.

02

Explain what a Merkle tree summarizes and why a hash mismatch narrows work without identifying a business row directly.

03

Observe preview/repair logs, netstats, streaming bytes, repair_admin session state and table repair metrics together.

04

Use -pr to reason about non-duplicative cluster coverage rather than launching identical replicated-range repair everywhere.

05

Distinguish normal cancellation/recovery from force/skip behavior that can leave convergence incomplete.

Chapter 24 anti-entropy lab baseline

Repair work is isolated from every earlier course cluster. Mandatory labs use Docker network atlasmart-cassandra-repair, cluster atlasmart-repair, nodes atlasmart-repair-1..3, pinned Docker Official Image cassandra:5.0.9, Java 17 inside the image, datacenter dc1, racks rack1..rack3, and 16 virtual nodes (vnodes) per node. Keyspace atlasmart_repair uses NetworkTopologyStrategy with replication factor (RF) 3 in dc1; ordinary test traffic uses LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no default time-to-live (TTL), gc_grace_seconds = 864000 unless a lesson explicitly creates a disposable shorter-grace comparison, and read_repair = 'NONE' on divergence fixtures so request-scoped read repair cannot hide the anti-entropy experiment. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only inside this isolated single-host learning network; no native/JMX port is published to the host. Recommended lab headroom is roughly 8 GiB of available host RAM plus at least 10 GiB free disk; resource-constrained learners can reduce seed rows while preserving the same mechanism. Exact tokens, Merkle-tree depth/hashes, repair session IDs, validation duration, stream bytes, SSTable counts, repaired percentages, disk/CPU/network utilization and p50/p95/p99 application latency are learner-captured evidence, not promised constants.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and anti-entropy mental model

Apache Cassandra stores a logical partition on multiple replicas chosen from its token ownership and keyspace replication strategy. A request-scoped coordinator is whichever node handles one client operation; it is not a permanent leader. A consistency level (CL) is the number/scope of replica responses required before that operation can succeed. A hint is a best-effort record of a mutation that a temporarily unavailable replica missed. Read repair is request-scoped reconciliation/write-back that may happen when replicas consulted by a read disagree; it does not scan unread data. Anti-entropy repair is the operator/scheduler-driven process that compares replicas for token ranges and streams differences so replicas converge even when no client happens to read the affected partition.

A token range is a portion of the partitioner hash space. During repair, replicas validate common ranges and summarize their contents with Merkle trees: hierarchical hashes that let Cassandra narrow mismatches without sending every row over the network. A mismatch causes streaming of data differences between replicas. Incremental repair, the current nodetool repair default, works on unrepaired/pending-repair data and—after a consistent session—separates repaired from unrepaired data through anticompaction or equivalent repaired-state handling. A full repair uses --full and compares all data in the selected ranges, including data already marked repaired.

An immutable SSTable (Sorted String Table) can be classified as repaired, unrepaired, or pending repair. repairedAt is persisted repair-state metadata associated with repaired SSTables/ranges. gc_grace_seconds is a table-level grace period after which old tombstones can become eligible for purge; repair cadence must leave enough margin that every replica receives deletes before tombstones disappear. Anticompaction can rewrite SSTables to separate data covered by a successful incremental repair from data that remains unrepaired, which is why incremental repair consumes temporary disk and I/O even when little network streaming is required.

1. Repair is coordinated per ranges and replica sets

The repair coordinator determines token ranges in scope, identifies replicas that share those ranges, asks participants to validate their local data, compares Merkle trees and starts synchronization for mismatched sections. Merkle trees are summaries: a top-level mismatch recursively narrows which subranges differ, but it does not mean Cassandra knows “order 42 is wrong” from the root hash. Validation reads local SSTables and consumes disk/CPU/page-cache resources; streaming then transfers actual differences and competes for network/disk bandwidth.

Phase Observable evidence Primary resource cost Failure meaning
Prepare/session repair command/session ID, logs, repair_admin coordination/JMX/thread pools participant unavailable or session setup rejected
Validation ValidationTime, BytesValidated/PartitionsValidated, logs disk reads, CPU, page cache cannot build comparable range summaries
Merkle comparison preview/desync result, validation logs CPU/memory plus validation work hash mismatch means data differs for a subrange
Sync/streaming netstats, Streaming Incoming/OutgoingBytes, SyncTime network + disk read/write difference transfer incomplete/stalled
Finalize incremental PercentRepaired, BytesPendingRepair → repaired, anticompaction metrics disk rewrite/headroom, compaction pending repair data not promoted/finalized

2. Create enough immutable state for observable range work

Docker · create the isolated three-replica repair topology
docker network inspect atlasmart-cassandra-repair >/dev/null 2>&1 || docker network create atlasmart-cassandra-repairdocker volume create atlasmart-repair-1-datadocker volume create atlasmart-repair-2-datadocker volume create atlasmart-repair-3-datadocker run -d --name atlasmart-repair-1 --hostname atlasmart-repair-1 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-repair-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-repair-1 nodetool statusdocker run -d --name atlasmart-repair-2 --hostname atlasmart-repair-2 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-repair-3 --hostname atlasmart-repair-3 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-3-data:/var/lib/cassandra cassandra:5.0.9# Do not start a repair until at least two observers show all three nodes UN.docker exec atlasmart-repair-1 nodetool versiondocker exec atlasmart-repair-1 java -versiondocker exec atlasmart-repair-1 nodetool statusdocker exec atlasmart-repair-2 nodetool status
CQL · create the RF=3 repair fixture and seed consistent data
CREATE KEYSPACE IF NOT EXISTS atlasmart_repairWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_repair.repair_probe (    tenant_id text,    item_id int,    status text,    note text,    updated_at timestamp,    PRIMARY KEY ((tenant_id), item_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND gc_grace_seconds = 864000  AND read_repair = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_repair.repair_probe(tenant_id,item_id,status,note,updated_at)VALUES ('tenant-001',1,'BASELINE','present on all replicas','2026-09-08T10:00:00Z');DESCRIBE TABLE atlasmart_repair.repair_probe;SELECT * FROM atlasmart_repair.repair_probe WHERE tenant_id='tenant-001';
Docker/cqlsh · generate a small distributed fixture without giant CQL batches
docker exec atlasmart-repair-1 bash -lc 'rm -f /tmp/repair_seed.cqlprintf "CONSISTENCY ALL;\n" > /tmp/repair_seed.cqlfor t in $(seq -w 1 64); do  for i in $(seq 1 4); do    printf "INSERT INTO atlasmart_repair.repair_probe (tenant_id,item_id,status,note,updated_at) VALUES (\047tenant-%s\047,%s,\047BASELINE\047,\047seed\047,toTimestamp(now()));\n" "$t" "$i" >> /tmp/repair_seed.cql  donedonecqlsh -f /tmp/repair_seed.cql'# Force flushes only to make immutable repair state observable in this disposable lab.for n in 1 2 3; do docker exec atlasmart-repair-$n nodetool flush atlasmart_repair repair_probe; done
Failure injection · make node 3 miss a write without leaving a hint
# Hints are disabled only on the coordinator that will issue the mutation.docker exec atlasmart-repair-1 nodetool disablehandoffdocker pause atlasmart-repair-3docker exec atlasmart-repair-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_repair.repair_probe SET status='NEWER_ON_1_AND_2', note='node3 missed this mutation', updated_at=toTimestamp(now()) WHERE tenant_id='tenant-001' AND item_id=1;"# Re-enable future hint storage before the failed replica returns; no hint is created retroactively.docker exec atlasmart-repair-1 nodetool enablehandoffdocker unpause atlasmart-repair-3# Wait for all nodes to be UN again before repair/preview.docker exec atlasmart-repair-1 nodetool statusfor n in 1 2 3; do docker exec atlasmart-repair-$n nodetool flush atlasmart_repair repair_probe; done
nodetool · capture repair/resource evidence before changing state
docker exec atlasmart-repair-1 nodetool tablestats atlasmart_repair.repair_probedocker exec atlasmart-repair-1 nodetool repair_admin summarize-pending -vdocker exec atlasmart-repair-1 nodetool repair_admin summarize-repaired -vdocker exec atlasmart-repair-1 nodetool getstreamthroughput -mdocker exec atlasmart-repair-1 nodetool getcompactionthroughputdocker exec atlasmart-repair-1 nodetool compactionstats -Hdocker exec atlasmart-repair-1 nodetool proxyhistogramsdocker stats --no-stream atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3

Because the fixture is deliberately small, repair may finish before netstats catches a live session. That is not a failure of the mechanism. Capture the command output, logs, cumulative streaming counters and repair-admin history as durable evidence; increase the seed count only if your machine has measured headroom.

3. Compare full preview with incremental primary-range repair

Terminal A · preview full ranges without streaming
docker exec atlasmart-repair-1 nodetool repair --preview --full -pr atlasmart_repair repair_probe
Terminal A · run the default incremental repair on node 1 primary ranges
docker exec atlasmart-repair-1 nodetool repair -pr atlasmart_repair repair_probe
Terminal B · observe while repair is running
docker exec atlasmart-repair-1 nodetool netstats -Hdocker exec atlasmart-repair-2 nodetool netstats -Hdocker exec atlasmart-repair-1 nodetool repair_admin list --alldocker exec atlasmart-repair-1 nodetool repair_admin summarize-pending -vdocker exec atlasmart-repair-1 nodetool compactionstats -Hdocker stats --no-stream atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3

-pr means “primary ranges for this node,” not “all ranges in the cluster.” The normal cluster-wide pattern is to schedule primary-range repairs across all nodes so the ring's range space is covered without repairing the same replicated range repeatedly. In a multi-DC cluster, decide explicitly which DCs/replica sets each session should compare and how much cross-DC bandwidth is acceptable; do not copy a single-DC command blindly.

4. Repair session administration and safe failure handling

Incremental repair has explicit session state because participants may have data marked pending repair while the session is in progress. Cassandra 5.0's repair_admin can list sessions, summarize pending/repaired ranges, cancel an incremental session and clean up pending state. Cancellation is a recovery tool, not a way to make a failed repair “green.” After cancellation or participant recovery, rerun the intended scope and re-verify convergence.

nodetool · inspect/cancel only a known incremental session
docker exec atlasmart-repair-1 nodetool repair_admin list --alldocker exec atlasmart-repair-1 nodetool repair_admin summarize-pending -vdocker exec atlasmart-repair-1 nodetool repair_admin summarize-repaired -v# If (and only if) an active incremental session UUID is genuinely stuck:# docker exec atlasmart-repair-1 nodetool repair_admin cancel --session <session-uuid># Then inspect/clean pending state and rerun the intended repair scope.# docker exec atlasmart-repair-1 nodetool repair_admin cleanup
Wrong approach: run repair on every node simultaneously because “more parallelism finishes sooner.”

Replicas in the same repair can contend for validation reads, streaming, anticompaction and compaction. Uncoordinated parallel sessions also create duplicate range work. Schedule range coverage and replica concurrency deliberately; treat additional job threads/parallel sessions as measured capacity decisions.

5. Acceptance checklist

  • Repair scope names the keyspace/table and whether it is incremental/full/preview/primary-range/DC constrained.
  • Session IDs/logs and range coverage are recorded rather than only a terminal exit code.
  • netstats or cumulative streaming metrics are correlated with disk/network/latency.
  • Pending incremental repair state returns to an understood finalized/clean condition.
  • Post-repair preview/validation and representative CL reads support convergence.

Check your understanding

  1. Why use Merkle trees instead of sending every row for comparison?
  2. What does -pr change?
  3. Why might netstats show no active streams even though preview found differences?
  4. What is pending-repair state?
  5. Does cancelling a session mean convergence is complete?
Review the answers

1. Hierarchical hashes let replicas narrow mismatched token subranges and stream only differences rather than transmitting every row merely to compare equality.

2. It limits the repair coordinator to ranges for which that node is the primary/first replica, enabling a non-duplicative cluster-wide schedule when run across all nodes.

3. A small repair may finish before you sample it; use command output, logs and cumulative metrics as additional evidence.

4. Data isolated for an in-progress incremental repair session that has not yet been successfully finalized as repaired.

5. No. Cancellation recovers state/control; the intended range still needs a successful repair and verification.

Docker · reset only the dedicated Chapter 24 repair lab
docker rm -f atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3 2>/dev/null || truedocker volume rm atlasmart-repair-1-data atlasmart-repair-2-data atlasmart-repair-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra-repair 2>/dev/null || true

Production judgment

Repair is a distributed maintenance workload that competes with foreground reads/writes for disk bandwidth, page cache, CPU, network, compaction capacity and JVM time. Plan it from the actual token/replica topology, RF and consistency levels, table sizes, SSTable overlap, tombstone/delete/TTL rate, shortest gc_grace_seconds, repair duration variance, inter-DC bandwidth, rack/zone maintenance, failure probability, SAI/vector indexes, snapshot/backup windows, compaction strategy, disk free space and business p95/p99 latency objectives. Incremental repair reduces repeated scope when run continuously but introduces repaired/unrepaired separation and anticompaction cost; full repair is broader and remains necessary for cases incremental repair intentionally skips, including periodically checking previously repaired data.

Do not use a single “repair every N days” value without measuring whether the entire required token-space/replica set actually completes within that interval. The safety condition is completion before tombstone grace can expire on unrepaired replicas, with margin for retries, outages and maintenance. Treat --force, aggressive parallelism, high -j, cross-DC sessions and throughput-cap changes as controlled operational choices with rollback and SLO gates. Managed Cassandra services may schedule repair internally or expose different controls; confirm who owns anti-entropy and what convergence evidence is available instead of assuming Apache nodetool semantics are exposed. Lesson 3 focuses on the repaired/unrepaired boundary itself: why incremental repair is cheap only when sustained, why first adoption can be expensive, and why periodic full repair remains necessary.

Summary and next bridge

Repair is a distributed workflow with explicit preparation, validation, comparison, streaming and finalization state. Merkle trees reduce comparison traffic, but validation and anticompaction still consume real resources. Next, compare incremental and full repair using repaired/unrepaired SSTable evidence instead of treating them as command-line aliases.

Authoritative references

Repair behavior and command options are version-sensitive. Re-check these sources before carrying a runbook to a newer Cassandra patch or managed service.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.