Chapter 24 · Repair, Anti-Entropy, Incremental/Full Strategies, and Data Convergence

Incremental Repair vs Full Repair: Repaired / Unrepaired Data and Operational Tradeoffs

Make repaired, unrepaired and pending SSTable state observable and compare the persistent consequences of incremental versus full repair.

Advanced120–160 minutesIncremental/full repaired-state labApache Cassandra 5.0.9 · cqlsh/nodetool · RF=3 · dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart has two competing proposals: “incremental repair is cheaper, so run only incremental forever” and “full repair is safer, so run full repair every night.” Both ignore how Cassandra tracks repaired state. This lesson makes repaired/unrepaired bytes visible, then shows why incremental repair is a continuous operating model and full repair remains a different safety tool.

01

Explain repaired, unrepaired and pending-repair SSTable/range state and the role of anticompaction.

02

Measure PercentRepaired, BytesRepaired, BytesUnrepaired and BytesPendingRepair before/after incremental repair.

03

Explain why the first incremental repair on an old unrepaired dataset can be as broad/costly as a full-data pass.

04

Explain why incremental repair does not revisit already repaired data and therefore cannot replace occasional full repair.

05

Use offline SSTable metadata only on copied/snapshotted files, never against a running Cassandra data directory.

Chapter 24 anti-entropy lab baseline

Repair work is isolated from every earlier course cluster. Mandatory labs use Docker network atlasmart-cassandra-repair, cluster atlasmart-repair, nodes atlasmart-repair-1..3, pinned Docker Official Image cassandra:5.0.9, Java 17 inside the image, datacenter dc1, racks rack1..rack3, and 16 virtual nodes (vnodes) per node. Keyspace atlasmart_repair uses NetworkTopologyStrategy with replication factor (RF) 3 in dc1; ordinary test traffic uses LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no default time-to-live (TTL), gc_grace_seconds = 864000 unless a lesson explicitly creates a disposable shorter-grace comparison, and read_repair = 'NONE' on divergence fixtures so request-scoped read repair cannot hide the anti-entropy experiment. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only inside this isolated single-host learning network; no native/JMX port is published to the host. Recommended lab headroom is roughly 8 GiB of available host RAM plus at least 10 GiB free disk; resource-constrained learners can reduce seed rows while preserving the same mechanism. Exact tokens, Merkle-tree depth/hashes, repair session IDs, validation duration, stream bytes, SSTable counts, repaired percentages, disk/CPU/network utilization and p50/p95/p99 application latency are learner-captured evidence, not promised constants.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and anti-entropy mental model

Apache Cassandra stores a logical partition on multiple replicas chosen from its token ownership and keyspace replication strategy. A request-scoped coordinator is whichever node handles one client operation; it is not a permanent leader. A consistency level (CL) is the number/scope of replica responses required before that operation can succeed. A hint is a best-effort record of a mutation that a temporarily unavailable replica missed. Read repair is request-scoped reconciliation/write-back that may happen when replicas consulted by a read disagree; it does not scan unread data. Anti-entropy repair is the operator/scheduler-driven process that compares replicas for token ranges and streams differences so replicas converge even when no client happens to read the affected partition.

A token range is a portion of the partitioner hash space. During repair, replicas validate common ranges and summarize their contents with Merkle trees: hierarchical hashes that let Cassandra narrow mismatches without sending every row over the network. A mismatch causes streaming of data differences between replicas. Incremental repair, the current nodetool repair default, works on unrepaired/pending-repair data and—after a consistent session—separates repaired from unrepaired data through anticompaction or equivalent repaired-state handling. A full repair uses --full and compares all data in the selected ranges, including data already marked repaired.

An immutable SSTable (Sorted String Table) can be classified as repaired, unrepaired, or pending repair. repairedAt is persisted repair-state metadata associated with repaired SSTables/ranges. gc_grace_seconds is a table-level grace period after which old tombstones can become eligible for purge; repair cadence must leave enough margin that every replica receives deletes before tombstones disappear. Anticompaction can rewrite SSTables to separate data covered by a successful incremental repair from data that remains unrepaired, which is why incremental repair consumes temporary disk and I/O even when little network streaming is required.

1. Incremental repair changes persistent data classification

Incremental repair is “incremental” because successfully repaired data is marked repaired and skipped by later incremental cycles; newly written/unrepaired data becomes the next target. To preserve that boundary when SSTables contain both repaired-range and unrepaired-range data, Cassandra can anticompact/rewrite files. Therefore an incremental repair that streams almost no bytes can still consume significant disk I/O and temporary space.

Full repair deliberately ignores repaired-state filtering and compares all data in scope. This is why it remains useful for corruption, operator mistakes or bugs that affect data already marked repaired. A full repair's goal is convergence of selected data, not reclassifying all SSTables as newly incrementally repaired.

State Meaning Later incremental repair Operational consequence
Unrepaired not finalized by incremental repair eligible can compact with other unrepaired data
Pending repair isolated for an active consistent-repair session session-owned requires successful finalization or recovery/cleanup
Repaired successfully finalized by incremental repair normally skipped separated from unrepaired data; corruption needs full/validation attention

2. Observe repaired bytes over two incremental cycles

Docker · create the isolated three-replica repair topology
docker network inspect atlasmart-cassandra-repair >/dev/null 2>&1 || docker network create atlasmart-cassandra-repairdocker volume create atlasmart-repair-1-datadocker volume create atlasmart-repair-2-datadocker volume create atlasmart-repair-3-datadocker run -d --name atlasmart-repair-1 --hostname atlasmart-repair-1 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-repair-1-data:/var/lib/cassandra cassandra:5.0.9# Continue only after node 1 is UN.docker exec atlasmart-repair-1 nodetool statusdocker run -d --name atlasmart-repair-2 --hostname atlasmart-repair-2 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-repair-3 --hostname atlasmart-repair-3 --network atlasmart-cassandra-repair \  -e CASSANDRA_CLUSTER_NAME=atlasmart-repair -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-repair-1 -v atlasmart-repair-3-data:/var/lib/cassandra cassandra:5.0.9# Do not start a repair until at least two observers show all three nodes UN.docker exec atlasmart-repair-1 nodetool versiondocker exec atlasmart-repair-1 java -versiondocker exec atlasmart-repair-1 nodetool statusdocker exec atlasmart-repair-2 nodetool status
CQL · create the RF=3 repair fixture and seed consistent data
CREATE KEYSPACE IF NOT EXISTS atlasmart_repairWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_repair.repair_probe (    tenant_id text,    item_id int,    status text,    note text,    updated_at timestamp,    PRIMARY KEY ((tenant_id), item_id)) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND gc_grace_seconds = 864000  AND read_repair = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_repair.repair_probe(tenant_id,item_id,status,note,updated_at)VALUES ('tenant-001',1,'BASELINE','present on all replicas','2026-09-08T10:00:00Z');DESCRIBE TABLE atlasmart_repair.repair_probe;SELECT * FROM atlasmart_repair.repair_probe WHERE tenant_id='tenant-001';
Docker/cqlsh · generate a small distributed fixture without giant CQL batches
docker exec atlasmart-repair-1 bash -lc 'rm -f /tmp/repair_seed.cqlprintf "CONSISTENCY ALL;\n" > /tmp/repair_seed.cqlfor t in $(seq -w 1 64); do  for i in $(seq 1 4); do    printf "INSERT INTO atlasmart_repair.repair_probe (tenant_id,item_id,status,note,updated_at) VALUES (\047tenant-%s\047,%s,\047BASELINE\047,\047seed\047,toTimestamp(now()));\n" "$t" "$i" >> /tmp/repair_seed.cql  donedonecqlsh -f /tmp/repair_seed.cql'# Force flushes only to make immutable repair state observable in this disposable lab.for n in 1 2 3; do docker exec atlasmart-repair-$n nodetool flush atlasmart_repair repair_probe; done
nodetool · baseline repaired/unrepaired metrics
for n in 1 2 3; do  docker exec atlasmart-repair-$n nodetool tablestats atlasmart_repair.repair_probe | grep -Ei 'Percent repaired|Bytes repaired|Bytes unrepaired|Bytes pending repair' || truedone
nodetool · complete one primary-range incremental cycle on all three nodes
# Run sequentially in this small lab to avoid replica contention.for n in 1 2 3; do  docker exec atlasmart-repair-$n nodetool repair -pr atlasmart_repair repair_probedonefor n in 1 2 3; do  docker exec atlasmart-repair-$n nodetool repair_admin summarize-repaired -v  docker exec atlasmart-repair-$n nodetool tablestats atlasmart_repair.repair_probe | grep -Ei 'Percent repaired|Bytes repaired|Bytes unrepaired|Bytes pending repair' || truedone

After the cycle, the percentage is expected to move toward repaired for the ranges/files covered, but exact percentages and byte counts depend on vnodes, flush/compaction state and anticompaction output. Do not assert 100% without reading the actual metrics.

CQL/nodetool · create new unrepaired data after the first cycle
docker exec atlasmart-repair-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; INSERT INTO atlasmart_repair.repair_probe (tenant_id,item_id,status,note,updated_at) VALUES ('tenant-after-repair',1,'NEW','written after incremental cycle',toTimestamp(now()));"for n in 1 2 3; do docker exec atlasmart-repair-$n nodetool flush atlasmart_repair repair_probe; donefor n in 1 2 3; do docker exec atlasmart-repair-$n nodetool tablestats atlasmart_repair.repair_probe | grep -Ei 'Percent repaired|Bytes repaired|Bytes unrepaired|Bytes pending repair' || true; done

The new flushed mutation should contribute unrepaired bytes. A second incremental cycle targets that new state rather than revalidating every previously repaired byte.

3. Full repair revisits already repaired data

nodetool · compare preview scopes
# Incremental preview considers unrepaired scope.docker exec atlasmart-repair-1 nodetool repair --preview -pr atlasmart_repair repair_probe# Full preview considers all data in the same selected ranges.docker exec atlasmart-repair-1 nodetool repair --preview --full -pr atlasmart_repair repair_probe# Periodic full repair is a deliberate broader check, not a replacement for the incremental cadence.docker exec atlasmart-repair-1 nodetool repair --full -pr atlasmart_repair repair_probe
Wrong migration: enable incremental repair on a large old cluster and assume the first run is small.

The first incremental cycle has no repaired history to skip, so it can touch the entire unrepaired dataset and trigger substantial anticompaction. Cassandra 5.0's optional Auto Repair documentation explicitly warns about this adoption cost and recommends planning/preflight rather than switching on a scheduler casually.

4. Optional offline repairedAt evidence from a copied snapshot

sstablemetadata is an offline SSTable tool. Never aim it at files Cassandra is actively using. The safe course pattern is: flush, take a snapshot, copy the snapshot out, and inspect the copy from a utility container where no Cassandra daemon is running. The online tablestats/repair_admin metrics remain the mandatory cross-platform evidence.

Optional Bash/WSL · copy a snapshot and inspect repaired metadata read-only
docker exec atlasmart-repair-1 nodetool snapshot -t ch24-repaired-meta atlasmart_repair -cf repair_proberm -rf ./ch24-repaired-snapshot && mkdir ./ch24-repaired-snapshot# Copy only snapshot files out of the container; the exact table directory UUID is environment-specific.docker cp atlasmart-repair-1:/var/lib/cassandra/data/atlasmart_repair/. ./ch24-repaired-snapshot/docker run --rm -v "$PWD/ch24-repaired-snapshot:/inspect:ro" cassandra:5.0.9 bash -lc '  f=$(find /inspect -path "*snapshots/ch24-repaired-meta/*Data.db" | head -1)  test -n "$f" && sstablemetadata "$f" | grep -Ei "Repaired|Pending repair" || true'docker exec atlasmart-repair-1 nodetool clearsnapshot -t ch24-repaired-meta atlasmart_repair

5. Verification and reset

nodetool · capture repair/resource evidence before changing state
docker exec atlasmart-repair-1 nodetool tablestats atlasmart_repair.repair_probedocker exec atlasmart-repair-1 nodetool repair_admin summarize-pending -vdocker exec atlasmart-repair-1 nodetool repair_admin summarize-repaired -vdocker exec atlasmart-repair-1 nodetool getstreamthroughput -mdocker exec atlasmart-repair-1 nodetool getcompactionthroughputdocker exec atlasmart-repair-1 nodetool compactionstats -Hdocker exec atlasmart-repair-1 nodetool proxyhistogramsdocker stats --no-stream atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3

Check your understanding

  1. Why does incremental repair need repaired/unrepaired separation?
  2. Why can incremental repair consume disk even with little streaming?
  3. Why is the first incremental repair on an old cluster often large?
  4. What can full repair catch that later incremental repair may skip?
  5. Why inspect repairedAt only on copied/snapshotted files?
Review the answers

1. Later incremental cycles must skip data already successfully repaired while still targeting newly written/unrepaired data.

2. Anticompaction/finalization may rewrite SSTables to separate repaired ranges from unrepaired data.

3. All existing data starts unrepaired, so there is little or nothing to skip.

4. Divergence/corruption/operator loss in data already marked repaired because full repair compares all selected data.

5. Offline SSTable tools are not safe against files a running Cassandra process is actively managing.

Docker · reset only the dedicated Chapter 24 repair lab
docker rm -f atlasmart-repair-1 atlasmart-repair-2 atlasmart-repair-3 2>/dev/null || truedocker volume rm atlasmart-repair-1-data atlasmart-repair-2-data atlasmart-repair-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra-repair 2>/dev/null || true

Production judgment

Repair is a distributed maintenance workload that competes with foreground reads/writes for disk bandwidth, page cache, CPU, network, compaction capacity and JVM time. Plan it from the actual token/replica topology, RF and consistency levels, table sizes, SSTable overlap, tombstone/delete/TTL rate, shortest gc_grace_seconds, repair duration variance, inter-DC bandwidth, rack/zone maintenance, failure probability, SAI/vector indexes, snapshot/backup windows, compaction strategy, disk free space and business p95/p99 latency objectives. Incremental repair reduces repeated scope when run continuously but introduces repaired/unrepaired separation and anticompaction cost; full repair is broader and remains necessary for cases incremental repair intentionally skips, including periodically checking previously repaired data.

Do not use a single “repair every N days” value without measuring whether the entire required token-space/replica set actually completes within that interval. The safety condition is completion before tombstone grace can expire on unrepaired replicas, with margin for retries, outages and maintenance. Treat --force, aggressive parallelism, high -j, cross-DC sessions and throughput-cap changes as controlled operational choices with rollback and SLO gates. Managed Cassandra services may schedule repair internally or expose different controls; confirm who owns anti-entropy and what convergence evidence is available instead of assuming Apache nodetool semantics are exposed. Lesson 4 turns repaired-state mechanics into a schedule: derive cadence from gc_grace, token-space completion time, multi-DC bandwidth, repair headroom and foreground SLOs instead of copying folklore intervals.

Summary and next bridge

Incremental and full repair solve different operating problems. Incremental repair depends on continuously maintaining a repaired/unrepaired boundary; full repair deliberately revisits all selected data. The next lesson converts that distinction into a cadence that fits tombstone safety, DC topology and measurable capacity.

Authoritative references

Repair behavior and command options are version-sensitive. Re-check these sources before carrying a runbook to a newer Cassandra patch or managed service.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.