Chapter 24 · Repair, Anti-Entropy, Incremental/Full Strategies, and Data Convergence

Build a Repair Runbook with Scheduling, Monitoring, Failure Recovery, and Verification

Turn Cassandra repair semantics into an auditable runbook with preflight gates, monitoring, failure recovery, verification and schedule ownership.

Advanced130–180 minutesFailure-recovery repair runbook labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart now needs an on-call runbook, not another conceptual diagram. The runbook must say what to check before repair, what scope to run, how to monitor it, how to respond when a replica fails mid-window, how to resume/retry safely, and which evidence proves convergence without hiding application impact.

01

Build a preflight checklist covering topology, schema agreement, disk headroom, compaction, backups, repair age and foreground SLOs.

02

Run a primary-range incremental cycle with explicit session/resource/latency evidence and periodic full/preview validation points.

03

Inject a reversible replica failure, classify the repair failure, restore the participant and rerun the intended scope instead of forcing success.

04

Use repair_admin, preview/validate, table metrics, streaming evidence and representative CL reads as post-repair acceptance gates.

05

Record an auditable schedule/history with clear ownership, escalation, rollback and managed-service boundaries.

Chapter 24 anti-entropy lab baseline

Repair work is isolated from every earlier course cluster. Mandatory labs use Docker network atlasmart-cassandra-repair, cluster atlasmart-repair, nodes atlasmart-repair-1..3, pinned Docker Official Image cassandra:5.0.9, Java 17 inside the image, datacenter dc1, racks rack1..rack3, and 16 virtual nodes (vnodes) per node. Keyspace atlasmart_repair uses NetworkTopologyStrategy with replication factor (RF) 3 in dc1; ordinary test traffic uses LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no default time-to-live (TTL), gc_grace_seconds = 864000 unless a lesson explicitly creates a disposable shorter-grace comparison, and read_repair = 'NONE' on divergence fixtures so request-scoped read repair cannot hide the anti-entropy experiment. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only inside this isolated single-host learning network; no native/JMX port is published to the host. Recommended lab headroom is roughly 8 GiB of available host RAM plus at least 10 GiB free disk; resource-constrained learners can reduce seed rows while preserving the same mechanism. Exact tokens, Merkle-tree depth/hashes, repair session IDs, validation duration, stream bytes, SSTable counts, repaired percentages, disk/CPU/network utilization and p50/p95/p99 application latency are learner-captured evidence, not promised constants.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and anti-entropy mental model

Apache Cassandra stores a logical partition on multiple replicas chosen from its token ownership and keyspace replication strategy. A request-scoped coordinator is whichever node handles one client operation; it is not a permanent leader. A consistency level (CL) is the number/scope of replica responses required before that operation can succeed. A hint is a best-effort record of a mutation that a temporarily unavailable replica missed. Read repair is request-scoped reconciliation/write-back that may happen when replicas consulted by a read disagree; it does not scan unread data. Anti-entropy repair is the operator/scheduler-driven process that compares replicas for token ranges and streams differences so replicas converge even when no client happens to read the affected partition.

A token range is a portion of the partitioner hash space. During repair, replicas validate common ranges and summarize their contents with Merkle trees: hierarchical hashes that let Cassandra narrow mismatches without sending every row over the network. A mismatch causes streaming of data differences between replicas. Incremental repair, the current nodetool repair default, works on unrepaired/pending-repair data and—after a consistent session—separates repaired from unrepaired data through anticompaction or equivalent repaired-state handling. A full repair uses --full and compares all data in the selected ranges, including data already marked repaired.

An immutable SSTable (Sorted String Table) can be classified as repaired, unrepaired, or pending repair. repairedAt is persisted repair-state metadata associated with repaired SSTables/ranges. gc_grace_seconds is a table-level grace period after which old tombstones can become eligible for purge; repair cadence must leave enough margin that every replica receives deletes before tombstones disappear. Anticompaction can rewrite SSTables to separate data covered by a successful incremental repair from data that remains unrepaired, which is why incremental repair consumes temporary disk and I/O even when little network streaming is required.

1. Runbook phase A: preflight and scope declaration

Every run should start by writing down the intended scope: incremental/full/preview, keyspace/tables, primary ranges or explicit token range, DC constraints, expected participant nodes and acceptable foreground impact. Check topology from more than one observer, schema agreement, disk free space, pending compactions, streaming caps, current repair state and application latency before starting. Do not begin merely because a cron timer fired.

Docker · create the isolated three-replica repair topology
docker network inspect atlasmart-cassandra-repair >/dev/null 2>&1 || docker network create atlasmart-cassandra-repairdocker volume create atlasmart-repair-1-datadocker volume create atlasmart-repair-2-datadocker volume create atlasmart-repair-3-datadocker run -d --name atlasmart-repair-1 --hostname atlasmart-repair-1 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-repair-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-repair-1 nodetool statusdocker run -d --name atlasmart-repair-2 --hostname atlasmart-repair-2 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-repair-3 --hostname atlasmart-repair-3 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-3-data:/var/lib/cassandra cassandra:5.0.9# Do not start a repair until at least two observers show all three nodes UN.docker exec atlasmart-repair-1 nodetool versiondocker exec atlasmart-repair-1 java -versiondocker exec atlasmart-repair-1 nodetool statusdocker exec atlasmart-repair-2 nodetool status
CQL · create the RF=3 repair fixture and seed consistent data
CREATE KEYSPACE IF NOT EXISTS atlasmart_repairWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_repair.repair_probe (    tenant_id text,    item_id int,    status text,    note text,    updated_at timestamp,    PRIMARY KEY ((tenant_id), item_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND gc_grace_seconds = 864000  AND read_repair = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_repair.repair_probe(tenant_id,item_id,status,note,updated_at)VALUES ('tenant-001',1,'BASELINE','present on all replicas','2026-09-08T10:00:00Z');DESCRIBE TABLE atlasmart_repair.repair_probe;SELECT * FROM atlasmart_repair.repair_probe WHERE tenant_id='tenant-001';
Docker/cqlsh · generate a small distributed fixture without giant CQL batches
docker exec atlasmart-repair-1 bash -lc 'rm -f /tmp/repair_seed.cqlprintf "CONSISTENCY ALL;\n" > /tmp/repair_seed.cqlfor t in $(seq -w 1 64); do  for i in $(seq 1 4); do    printf "INSERT INTO atlasmart_repair.repair_probe (tenant_id,item_id,status,note,updated_at) VALUES (\047tenant-%s\047,%s,\047BASELINE\047,\047seed\047,toTimestamp(now()));\n" "$t" "$i" >> /tmp/repair_seed.cql  donedonecqlsh -f /tmp/repair_seed.cql'# Force flushes only to make immutable repair state observable in this disposable lab.for n in 1 2 3; do docker exec atlasmart-repair-$n nodetool flush atlasmart_repair repair_probe; done
Preflight · collect the evidence bundle
for n in 1 2 3; do  echo "=== atlasmart-repair-$n ==="  docker exec atlasmart-repair-$n nodetool status atlasmart_repair  docker exec atlasmart-repair-$n nodetool describecluster  docker exec atlasmart-repair-$n nodetool compactionstats -H  docker exec atlasmart-repair-$n nodetool repair_admin list --all  docker exec atlasmart-repair-$n nodetool repair_admin summarize-pending -v  docker exec atlasmart-repair-$n nodetool getstreamthroughput -m  docker exec atlasmart-repair-$n df -h /var/lib/cassandradonedocker exec atlasmart-repair-1 nodetool proxyhistogramsdocker stats --no-stream atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3

Preflight rejection examples include a down replica needed by the intended range, unexplained pending repair state, insufficient disk for anticompaction/streaming, severe compaction backlog, active topology transition, a missed backup/recovery prerequisite, or already-breached p99 latency. Repair is important, but starting it blindly can turn maintenance into an outage.

2. Runbook phase B: preview, execute, monitor

Failure injection · make node 3 miss a write without leaving a hint
# Hints are disabled only on the coordinator that will issue the mutation.docker exec atlasmart-repair-1 nodetool disablehandoffdocker pause atlasmart-repair-3docker exec atlasmart-repair-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_repair.repair_probe SET status='NEWER_ON_1_AND_2', note='node3 missed this mutation', updated_at=toTimestamp(now()) WHERE tenant_id='tenant-001' AND item_id=1;"# Re-enable future hint storage before the failed replica returns; no hint is created retroactively.docker exec atlasmart-repair-1 nodetool enablehandoffdocker unpause atlasmart-repair-3# Wait for all nodes to be UN again before repair/preview.docker exec atlasmart-repair-1 nodetool statusfor n in 1 2 3; do docker exec atlasmart-repair-$n nodetool flush atlasmart_repair repair_probe; done
Preview + execute · one node primary range
# Preview broad/full comparison to estimate whether divergence remains.docker exec atlasmart-repair-1 nodetool repair --preview --full -pr atlasmart_repair repair_probe# Normal scheduled cycle: incremental is the current nodetool default.docker exec atlasmart-repair-1 nodetool repair -pr atlasmart_repair repair_probe
Monitor · sample session, streaming, repair and client-impact evidence
docker exec atlasmart-repair-1 nodetool repair_admin list --alldocker exec atlasmart-repair-1 nodetool repair_admin summarize-pending -vdocker exec atlasmart-repair-1 nodetool netstats -Hdocker exec atlasmart-repair-2 nodetool netstats -Hdocker exec atlasmart-repair-1 nodetool tablestats atlasmart_repair.repair_probedocker exec atlasmart-repair-1 nodetool compactionstats -Hdocker exec atlasmart-repair-1 nodetool proxyhistogramsdocker stats --no-stream atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3docker logs --since 15m atlasmart-repair-1 2>&1 | grep -Ei 'repair|validation|stream|anticompact|fail' || true

3. Runbook phase C: controlled failure and recovery

A repair that encounters a down participant should not be converted into apparent success with --force by default. Force can exclude down endpoints, which changes what “repair completed” means. The safer training path is to let the intended session fail, restore the node, inspect/clean any pending incremental state, and rerun the original scope. This preserves the convergence objective.

Failure injection · prove the runbook handles an unavailable participant
docker pause atlasmart-repair-3# Wait until failure detection from node 1 reports node 3 down.docker exec atlasmart-repair-1 nodetool status# This repair is expected to fail or be unable to complete the intended replica comparison.docker exec atlasmart-repair-1 nodetool repair -pr atlasmart_repair repair_probe || truedocker unpause atlasmart-repair-3# Wait for UN before recovery work.docker exec atlasmart-repair-1 nodetool statusdocker exec atlasmart-repair-1 nodetool repair_admin list --alldocker exec atlasmart-repair-1 nodetool repair_admin summarize-pending -v# If an active/stuck incremental session remains, cancel only that recorded UUID, then cleanup pending state.# docker exec atlasmart-repair-1 nodetool repair_admin cancel --session <session-uuid># docker exec atlasmart-repair-1 nodetool repair_admin cleanup# Rerun the intended scope after the participant is healthy.docker exec atlasmart-repair-1 nodetool repair -pr atlasmart_repair repair_probe
Wrong runbook: “If repair fails, rerun with --force until the command is green.”

A green command that intentionally excluded an unavailable replica can leave the very inconsistency anti-entropy was meant to resolve. Use force only for a documented failure scenario where the changed scope is understood, recorded and followed by later convergence work.

4. Runbook phase D: finish cluster coverage and verify

Cluster coverage · complete primary-range cycle and verify
# Node 1 was repaired above; finish primary-range coverage on nodes 2 and 3.for n in 2 3; do docker exec atlasmart-repair-$n nodetool repair -pr atlasmart_repair repair_probe; donefor n in 1 2 3; do  docker exec atlasmart-repair-$n nodetool repair_admin summarize-repaired -v  docker exec atlasmart-repair-$n nodetool repair_admin summarize-pending -v  docker exec atlasmart-repair-$n nodetool tablestats atlasmart_repair.repair_probe | grep -Ei 'Percent repaired|Bytes repaired|Bytes unrepaired|Bytes pending repair' || truedone# Verification layer 1: no expected full-repair differences remain.docker exec atlasmart-repair-1 nodetool repair --preview --full atlasmart_repair repair_probe# Verification layer 2: validate repaired data where incremental repaired state exists.docker exec atlasmart-repair-1 nodetool repair --validate atlasmart_repair repair_probe || true# Verification layer 3: require every replica response for a representative row.docker exec atlasmart-repair-1 cqlsh -e "CONSISTENCY ALL; TRACING ON; SELECT tenant_id,item_id,status,note FROM atlasmart_repair.repair_probe WHERE tenant_id='tenant-001' AND item_id=1; TRACING OFF;"# Verification layer 4: foreground/host state did not end degraded.docker exec atlasmart-repair-1 nodetool proxyhistogramsdocker exec atlasmart-repair-1 nodetool compactionstats -Hdocker stats --no-stream atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3

5. Evidence record, schedule ownership, and escalation

Runbook field Record before/during/after Escalate when
Scope repair type, -pr/ranges, keyspace/tables, DCs/hosts scope differs from approved plan or needs force/exclusion
Topology nodes/DC/racks/RF, state, schema agreement down/transitioning member or mismatch
Repair state session IDs, pending/repaired summaries, preview/validate stuck pending session or repaired-data desync
Resources disk free, compaction backlog, stream caps, CPU/network headroom/SLO gate breached
Application errors, p50/p95/p99 reads/writes, timeouts tail latency/error budget breached
Tombstone deadline table gc_grace, last completed coverage, retry buffer coverage cannot finish safely before grace margin
Outcome preview stream estimate, final CL read, remaining exceptions verification conflicts with command success

For Cassandra 5.0.9, Auto Repair is an optional scheduler available because CEP-37 was backported starting in 5.0.8. If a team adopts it, the same evidence/ownership principles remain: enablement is a migration, runtime scheduler settings are node-local unless consistently deployed, alerts need scheduler/history plus repair metrics, and periodic full/preview-repaired checks still matter. Managed services can hide nodetool entirely; the contract then becomes evidence from their documented repair/convergence APIs and support responsibility.

Check your understanding

  1. What should a repair runbook declare before execution?
  2. Why is --force not the default recovery action?
  3. What should happen after a failed incremental session?
  4. Why combine preview/validate, repair metrics and CL reads?
  5. What is the final scheduling invariant?
Review the answers

1. Exact scope/type/ranges/DCs/tables/participants plus preflight resource/topology/SLO gates and rollback/escalation conditions.

2. It can exclude down replicas and make a narrower repair succeed without achieving the intended convergence.

3. Restore participants, inspect session/pending state, cancel/cleanup only when needed, rerun the intended scope, then verify convergence.

4. Each proves a different layer; agreement across them is stronger than trusting a single command exit status.

5. Every required range/replica must complete convergence with enough margin before tombstone grace and operational failure windows can make missed deletes unsafe.

Docker · reset only the dedicated Chapter 24 repair lab
docker rm -f atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3 2>/dev/null || truedocker volume rm atlasmart-repair-1-data atlasmart-repair-2-data atlasmart-repair-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra-repair 2>/dev/null || true

Production judgment

Repair is a distributed maintenance workload that competes with foreground reads/writes for disk bandwidth, page cache, CPU, network, compaction capacity and JVM time. Plan it from the actual token/replica topology, RF and consistency levels, table sizes, SSTable overlap, tombstone/delete/TTL rate, shortest gc_grace_seconds, repair duration variance, inter-DC bandwidth, rack/zone maintenance, failure probability, SAI/vector indexes, snapshot/backup windows, compaction strategy, disk free space and business p95/p99 latency objectives. Incremental repair reduces repeated scope when run continuously but introduces repaired/unrepaired separation and anticompaction cost; full repair is broader and remains necessary for cases incremental repair intentionally skips, including periodically checking previously repaired data.

Do not use a single “repair every N days” value without measuring whether the entire required token-space/replica set actually completes within that interval. The safety condition is completion before tombstone grace can expire on unrepaired replicas, with margin for retries, outages and maintenance. Treat --force, aggressive parallelism, high -j, cross-DC sessions and throughput-cap changes as controlled operational choices with rollback and SLO gates. Managed Cassandra services may schedule repair internally or expose different controls; confirm who owns anti-entropy and what convergence evidence is available instead of assuming Apache nodetool semantics are exposed. Chapter 25 changes from convergence to recoverability: snapshots and incremental backups are local storage mechanisms until they are cataloged, protected off-host and proven through restore drills with measured RPO/RTO.

Summary and next bridge

A repair runbook is an operational control loop: declare scope, preflight, compare/stream, monitor, recover failures, complete token-space coverage, and verify both data convergence and application health. With anti-entropy ownership explicit, Chapter 25 can distinguish replica convergence from backup/disaster recovery—two responsibilities that solve different failure classes.

Authoritative references

Repair behavior and command options are version-sensitive. Re-check these sources before carrying a runbook to a newer Cassandra patch or managed service.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.