Chapter 24 · Repair, Anti-Entropy, Incremental/Full Strategies, and Data Convergence

Repair Cadence Relative to gc_grace, Token Ranges, Datacenters, and Resource Throttling

Derive a repair cadence from gc_grace, token-space coverage, DC topology, throughput, disk headroom and foreground latency rather than folklore.

Advanced110–150 minutesCadence + throttling planning labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart's runbook says “repair weekly.” One table has ten-day grace, another keeps heavy TTL/delete churn, one DC link is constrained, and the cluster sometimes needs two days to finish a repair cycle. A calendar label alone is not a safety model. This lesson derives cadence from the shortest relevant tombstone grace, actual full token-space completion time and resource margins.

01

Derive repair deadlines from gc_grace_seconds plus measured completion/retry margin instead of a universal calendar rule.

02

Plan -pr coverage across nodes/token ranges and reason about DC-scoped versus cross-DC repair traffic.

03

Measure stream/compaction throughput, pending compactions, disk headroom and foreground latency before changing concurrency.

04

Explain Cassandra 5.0 repair safety settings such as repair disk-headroom/compaction rejection and why defaults may be disabled for compatibility.

05

Distinguish optional Auto Repair scheduling in 5.0.8+ from the default/manual operational baseline and its irreversible feature enablement.

Chapter 24 anti-entropy lab baseline

Repair work is isolated from every earlier course cluster. Mandatory labs use Docker network atlasmart-cassandra-repair, cluster atlasmart-repair, nodes atlasmart-repair-1..3, pinned Docker Official Image cassandra:5.0.9, Java 17 inside the image, datacenter dc1, racks rack1..rack3, and 16 virtual nodes (vnodes) per node. Keyspace atlasmart_repair uses NetworkTopologyStrategy with replication factor (RF) 3 in dc1; ordinary test traffic uses LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no default time-to-live (TTL), gc_grace_seconds = 864000 unless a lesson explicitly creates a disposable shorter-grace comparison, and read_repair = 'NONE' on divergence fixtures so request-scoped read repair cannot hide the anti-entropy experiment. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only inside this isolated single-host learning network; no native/JMX port is published to the host. Recommended lab headroom is roughly 8 GiB of available host RAM plus at least 10 GiB free disk; resource-constrained learners can reduce seed rows while preserving the same mechanism. Exact tokens, Merkle-tree depth/hashes, repair session IDs, validation duration, stream bytes, SSTable counts, repaired percentages, disk/CPU/network utilization and p50/p95/p99 application latency are learner-captured evidence, not promised constants.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and anti-entropy mental model

Apache Cassandra stores a logical partition on multiple replicas chosen from its token ownership and keyspace replication strategy. A request-scoped coordinator is whichever node handles one client operation; it is not a permanent leader. A consistency level (CL) is the number/scope of replica responses required before that operation can succeed. A hint is a best-effort record of a mutation that a temporarily unavailable replica missed. Read repair is request-scoped reconciliation/write-back that may happen when replicas consulted by a read disagree; it does not scan unread data. Anti-entropy repair is the operator/scheduler-driven process that compares replicas for token ranges and streams differences so replicas converge even when no client happens to read the affected partition.

A token range is a portion of the partitioner hash space. During repair, replicas validate common ranges and summarize their contents with Merkle trees: hierarchical hashes that let Cassandra narrow mismatches without sending every row over the network. A mismatch causes streaming of data differences between replicas. Incremental repair, the current nodetool repair default, works on unrepaired/pending-repair data and—after a consistent session—separates repaired from unrepaired data through anticompaction or equivalent repaired-state handling. A full repair uses --full and compares all data in the selected ranges, including data already marked repaired.

An immutable SSTable (Sorted String Table) can be classified as repaired, unrepaired, or pending repair. repairedAt is persisted repair-state metadata associated with repaired SSTables/ranges. gc_grace_seconds is a table-level grace period after which old tombstones can become eligible for purge; repair cadence must leave enough margin that every replica receives deletes before tombstones disappear. Anticompaction can rewrite SSTables to separate data covered by a successful incremental repair from data that remains unrepaired, which is why incremental repair consumes temporary disk and I/O even when little network streaming is required.

1. Cadence is a coverage deadline, not a cron expression

The tombstone safety requirement is that replicas carrying older live data receive the delete before the tombstone can disappear elsewhere. The practical repair deadline therefore needs to be comfortably shorter than the relevant table's gc_grace_seconds, after subtracting worst-case repair duration, failed-session retries, node/rack maintenance and an operational buffer. If the shortest-grace table takes three days to cover across all required ranges and replicas, a repair that starts one day before grace expires is not safe even though “repair ran this week.”

Input Measure/record Why it changes cadence
gc_grace_seconds per table DESCRIBE schema / system schema defines tombstone grace; shortest delete-bearing table can dominate
cycle duration p95/p99 repair history/logs/scheduler history schedule must finish, not merely start, before the deadline
token-space coverage -pr plan per node/DC prevents gaps and duplicate work
delete/TTL rate table/app metrics higher churn raises consequence of missed delete convergence
stream/compaction capacity throughput, pending tasks, disk headroom controls how quickly repair can safely progress
maintenance/failure budget node/rack/DC calendar repair must survive delays/outages without crossing grace

2. Inspect table grace and current repair capacity

Docker · create the isolated three-replica repair topology
docker network inspect atlasmart-cassandra-repair >/dev/null 2>&1 || docker network create atlasmart-cassandra-repairdocker volume create atlasmart-repair-1-datadocker volume create atlasmart-repair-2-datadocker volume create atlasmart-repair-3-datadocker run -d --name atlasmart-repair-1 --hostname atlasmart-repair-1 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-repair-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-repair-1 nodetool statusdocker run -d --name atlasmart-repair-2 --hostname atlasmart-repair-2 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-repair-3 --hostname atlasmart-repair-3 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-3-data:/var/lib/cassandra cassandra:5.0.9# Do not start a repair until at least two observers show all three nodes UN.docker exec atlasmart-repair-1 nodetool versiondocker exec atlasmart-repair-1 java -versiondocker exec atlasmart-repair-1 nodetool statusdocker exec atlasmart-repair-2 nodetool status
CQL · create the RF=3 repair fixture and seed consistent data
CREATE KEYSPACE IF NOT EXISTS atlasmart_repairWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_repair.repair_probe (    tenant_id text,    item_id int,    status text,    note text,    updated_at timestamp,    PRIMARY KEY ((tenant_id), item_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND gc_grace_seconds = 864000  AND read_repair = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_repair.repair_probe(tenant_id,item_id,status,note,updated_at)VALUES ('tenant-001',1,'BASELINE','present on all replicas','2026-09-08T10:00:00Z');DESCRIBE TABLE atlasmart_repair.repair_probe;SELECT * FROM atlasmart_repair.repair_probe WHERE tenant_id='tenant-001';
Docker/cqlsh · generate a small distributed fixture without giant CQL batches
docker exec atlasmart-repair-1 bash -lc 'rm -f /tmp/repair_seed.cqlprintf "CONSISTENCY ALL;\n" > /tmp/repair_seed.cqlfor t in $(seq -w 1 64); do  for i in $(seq 1 4); do    printf "INSERT INTO atlasmart_repair.repair_probe (tenant_id,item_id,status,note,updated_at) VALUES (\047tenant-%s\047,%s,\047BASELINE\047,\047seed\047,toTimestamp(now()));\n" "$t" "$i" >> /tmp/repair_seed.cql  donedonecqlsh -f /tmp/repair_seed.cql'# Force flushes only to make immutable repair state observable in this disposable lab.for n in 1 2 3; do docker exec atlasmart-repair-$n nodetool flush atlasmart_repair repair_probe; done
CQL/nodetool · capture cadence inputs
docker exec atlasmart-repair-1 cqlsh -e "DESCRIBE TABLE atlasmart_repair.repair_probe;"for n in 1 2 3; do  echo "=== node $n ==="  docker exec atlasmart-repair-$n nodetool getstreamthroughput -m  docker exec atlasmart-repair-$n nodetool getinterdcstreamthroughput -m  docker exec atlasmart-repair-$n nodetool getcompactionthroughput  docker exec atlasmart-repair-$n nodetool compactionstats -H  docker exec atlasmart-repair-$n df -h /var/lib/cassandradonedocker stats --no-stream atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3

Cassandra 5.0 also exposes optional safety settings such as repair_disk_headroom_reject_ratio and reject_repair_compaction_threshold. The 5.0 defaults preserve backward compatibility and may leave the disk-headroom check disabled, so the operator still owns capacity gating unless deliberately configured otherwise.

CQL · inspect running repair-related settings where available
SELECT name, valueFROM system_views.settingsWHERE name IN ('repair_disk_headroom_reject_ratio','reject_repair_compaction_threshold');

3. Token ranges and datacenters: schedule coverage intentionally

For a single-DC ring, a common manual pattern is sequential/controlled -pr repair on each node so each primary token range is covered once. In a multi-DC keyspace, decide whether a repair session should compare across DCs or remain within a DC based on RF, failure model and WAN budget. -local/-dc constrain participants; that can reduce WAN cost but should not be described as proving global equality unless the total runbook explicitly covers the required replica relationships.

Runbook sketch · primary-range sequence for the three-node teaching ring
for n in 1 2 3; do  echo "repairing primary ranges from node $n"  docker exec atlasmart-repair-$n nodetool repair -pr atlasmart_repair repair_probe  docker exec atlasmart-repair-$n nodetool repair_admin summarize-repaired -vdone# Verify no active/pending session is left unexplained.for n in 1 2 3; do docker exec atlasmart-repair-$n nodetool repair_admin list --all; done
Multi-DC simulation instead of a mandatory six-node lab.

Map dc1=3 and dc2=3 on paper, assign each node's primary ranges, then estimate bytes/session from preview repairs and your inter-DC throughput cap. A full six-node WAN-like lab is optional because it requires substantially more RAM/disk. The learning objective is the scheduling decision and evidence, not merely creating six containers.

4. Throttle with measured state and restore what you change

Streaming throughput caps are node-local runtime controls. Lowering them can protect foreground latency but extend the repair cycle; raising them can shorten the cycle but saturate network/disk and worsen p99. Record the original value, change one disposable node only for a short experiment, compare proxyhistograms/docker stats/netstats, then restore. Never paste a throughput number from another cluster as a “best practice.”

Optional reversible lab · stream throughput cap experiment
# Record the current node-local value first.docker exec atlasmart-repair-1 nodetool getstreamthroughput -m# Example only: choose a conservative lab value that your machine can handle.docker exec atlasmart-repair-1 nodetool setstreamthroughput -m 25docker exec atlasmart-repair-1 nodetool getstreamthroughput -m# Run/preview repair and record proxyhistograms/docker stats/netstats.docker exec atlasmart-repair-1 nodetool repair --preview --full -pr atlasmart_repair repair_probedocker exec atlasmart-repair-1 nodetool proxyhistograms# Restore the exact value you recorded before the experiment, not a guessed default.# docker exec atlasmart-repair-1 nodetool setstreamthroughput -m <recorded-MiB-per-second>
Do not lower gc_grace_seconds simply to reduce tombstone storage unless repair can prove the new deadline.

A shorter grace shrinks the time available for missed deletes to converge. If repair, outages or maintenance can exceed that interval, stale data can survive after tombstones are purged and later resurrect.

5. Cassandra 5.0.8+ Auto Repair: current but opt-in

Auto Repair was introduced in Cassandra 6.0 and backported to 5.0.8, so it exists in the pinned 5.0.9 baseline. It is disabled by default and requires the JVM property -Dcassandra.autorepair.enable=true before startup. Official documentation states that this feature enablement is non-reversible because it creates required schema elements. For that reason the mandatory lab does not enable it. In production, evaluate it as a versioned migration with preflight full repair, scheduler capacity, history/alerting and rollback constraints—not as a harmless runtime toggle.

Check your understanding

  1. Why is “repair weekly” incomplete?
  2. What does -pr buy a manual cluster-wide schedule?
  3. Why can lowering stream throughput be unsafe even if p99 improves?
  4. Why is multi-DC -local not automatically global convergence proof?
  5. Why is Auto Repair not enabled in this disposable lesson?
Review the answers

1. Safety depends on whether all required token ranges/replicas actually complete within grace with retry/maintenance margin, not on a calendar label.

2. It lets each node cover its primary ranges so the ring can be covered without repeatedly repairing every replicated range from every node.

3. The repair cycle may become so long that it misses the convergence deadline relative to gc_grace and outage margin.

4. It constrains participants to a DC; the whole runbook must still cover every required replica relationship/failure domain.

5. On 5.0.8+ it requires a non-reversible feature JVM enablement/schema change, so it deserves explicit migration planning rather than a throwaway toggle.

Docker · reset only the dedicated Chapter 24 repair lab
docker rm -f atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3 2>/dev/null || truedocker volume rm atlasmart-repair-1-data atlasmart-repair-2-data atlasmart-repair-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra-repair 2>/dev/null || true

Production judgment

Repair is a distributed maintenance workload that competes with foreground reads/writes for disk bandwidth, page cache, CPU, network, compaction capacity and JVM time. Plan it from the actual token/replica topology, RF and consistency levels, table sizes, SSTable overlap, tombstone/delete/TTL rate, shortest gc_grace_seconds, repair duration variance, inter-DC bandwidth, rack/zone maintenance, failure probability, SAI/vector indexes, snapshot/backup windows, compaction strategy, disk free space and business p95/p99 latency objectives. Incremental repair reduces repeated scope when run continuously but introduces repaired/unrepaired separation and anticompaction cost; full repair is broader and remains necessary for cases incremental repair intentionally skips, including periodically checking previously repaired data.

Do not use a single “repair every N days” value without measuring whether the entire required token-space/replica set actually completes within that interval. The safety condition is completion before tombstone grace can expire on unrepaired replicas, with margin for retries, outages and maintenance. Treat --force, aggressive parallelism, high -j, cross-DC sessions and throughput-cap changes as controlled operational choices with rollback and SLO gates. Managed Cassandra services may schedule repair internally or expose different controls; confirm who owns anti-entropy and what convergence evidence is available instead of assuming Apache nodetool semantics are exposed. Lesson 5 turns these calculations into an executable runbook with preflight gates, session monitoring, safe failure injection, recovery, post-repair verification and evidence retention.

Summary and next bridge

Repair cadence is a measured race between data-convergence work and tombstone grace, constrained by token coverage and available resources. Cassandra provides throttles and newer scheduling capabilities, but none remove the need to measure completion. The final lesson converts those inputs into a runbook an operator can execute and audit.

Authoritative references

Repair behavior and command options are version-sensitive. Re-check these sources before carrying a runbook to a newer Cassandra patch or managed service.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.