Chapter 15 · Failure Handling: Hints, Read Repair, Speculative Retry, Timeouts, and Unavailable Errors

Failure Drill: Stop Replicas, Vary Consistency Levels, Recover Nodes, and Verify Convergence

Run an end-to-end replica failure drill with CL changes, pending hints, stale replicas, read repair, explicit repair, recovery, and final convergence checks.

Intermediate125–175 minutesEnd-to-end failure drillApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart's release checklist now needs a failure drill that proves what happens—not just what architecture diagrams promise—when replicas stop, writes continue, reads use different CLs, hints accumulate, nodes recover, replicas diverge, and repair runs. This lesson turns Chapters 14–15 into a repeatable incident exercise with explicit acceptance criteria.

01

Execute a timed RF=3 failure drill with reversible replica pauses and explicit CL changes.

02

Observe successful LOCAL_QUORUM operations alongside deterministic ALL/LOCAL_QUORUM unavailability cases.

03

Prove hint-based convergence, then separately create divergence that requires read repair/repair.

04

Classify every client-visible error before deciding whether a retry would be safe.

05

Finish with post-repair equality and operational reset checks rather than declaring recovery when nodes merely return UN.

Chapter 15 lab baseline

The mandatory labs continue the established local AtlasMart cluster with Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 Docker image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes (vnodes) per node, NetworkTopologyStrategy, replication factor (RF) 3, and explicit per-request consistency levels (CLs). New Chapter 15 tables use UnifiedCompactionStrategy (UCS), no default Time To Live (TTL), and Cassandra's default gc_grace_seconds unless a lesson intentionally changes a setting. Authentication, client Transport Layer Security (TLS), internode TLS, and remote Java Management Extensions (JMX) are disabled only inside this isolated single-host learning network. The optional application examples use Apache Cassandra Java Driver 4.19.3. Windows learners should run the Linux containers through Docker Desktop/WSL rather than infer native Windows production support. No Storage-Attached Index (SAI) or vector index is required in this chapter, so failure-handling evidence is not confounded by index rebuild/query behavior.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Failure-handling vocabulary

A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores one copy of a partition according to the keyspace replication strategy. A partition is the set of rows sharing a partition key; Cassandra hashes that key to a token, and token ownership plus topology determines the replica set. RF (replication factor) is the number of replicas; CL (consistency level) is the response requirement for one operation. A hint is a durable record held by another node for a mutation a replica could not receive. Reconciliation chooses the newest visible cell/tombstone versions among replica responses. Read repair may write the reconciled result back to stale replicas involved in a read. Anti-entropy repair is the operator-run process that compares token-range data and streams differences; it is the comprehensive convergence mechanism. Speculation means starting redundant work before the original attempt has definitively failed. Idempotent means that repeating an operation produces the same intended database state as doing it once. A timeout means the operation did not complete within a deadline; it does not by itself prove that no replica applied a write. Unavailable means the coordinator already knows there are too few live replicas for the requested CL. Overloaded is a server response indicating the coordinator cannot currently process the request because its resources/backlog are exhausted.

1. Drill contract and safety envelope

Use only the disposable Docker cluster. Do not run these commands against production, managed Cassandra, or an unrelated local cluster. docker pause/unpause isolates the experiment without changing host firewall rules or clocks. Record timestamps for each step, nodetool status, requested CL, client result/error, pending hints, direct-replica values, and repair output. Exact failure-detection time, hint replay time, and error wording are runtime dependent.

bash · verify or recreate the three-node dc1 course cluster
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three replicas are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CQL · Chapter 15 RF=3 failure fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_failureWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_failure.order_state_by_id (    order_id text PRIMARY KEY,    status text,    version int,    note text,    updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND read_repair = 'BLOCKING'  AND speculative_retry = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_failure.order_state_by_id(order_id,status,version,note,updated_at)VALUES ('order-1501','CREATED',0,'baseline','2026-09-08T03:30:00Z');SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
bash · preflight: no node or handoff state is already abnormal
for n in 1 2 3; do  docker exec atlasmart-cass-$n nodetool statushandoff  docker exec atlasmart-cass-$n nodetool versiondonedocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool listpendinghints

2. Phase A — one replica down: quorum availability plus pending hints

bash · hold hint delivery on coordinator and pause node 3
docker exec atlasmart-cass-1 nodetool pausehandoffdocker pause atlasmart-cass-3# Wait for failure detector state before interpreting errors.docker exec atlasmart-cass-1 nodetool status
CQL · LOCAL_QUORUM succeeds; ALL fails while node 3 is unavailable
CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_failure.order_state_by_idSET status='PAID', version=10, note='phase-a', updated_at='2026-09-08T04:20:00Z'WHERE order_id='order-1501';SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';CONSISTENCY ALL;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';

Interpretation: the LOCAL_QUORUM write/read need two local responses and can succeed with one replica down. The ALL read requires all three and should become unavailable once failure detection knows node 3 is down. A longer timeout would not satisfy ALL. Meanwhile node 1 should hold a hint because its delivery process is paused.

bash · capture hint evidence and recover node 3 without replay yet
docker exec atlasmart-cass-1 nodetool listpendinghintsdocker unpause atlasmart-cass-3# Wait for node 3 to be UN, but node 1 is still not delivering hints.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"
bash · replay hints and verify convergence
docker exec atlasmart-cass-1 nodetool resumehandoff# Poll until the relevant hints drain.docker exec atlasmart-cass-1 nodetool listpendinghintsfor n in 1 2 3; do  docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"done

3. Phase B — create divergence without hints, then expose repair behavior

Phase A demonstrates the fast best-effort path. Phase B intentionally prevents hint creation on node 1 for one write so that the stale replica remains visible after recovery. This proves why node status alone is not data convergence.

bash · disable future hints on node 1 for this controlled mutation
docker exec atlasmart-cass-1 nodetool disablehandoffdocker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
CQL · update the two live replicas
CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_failure.order_state_by_idSET status='PACKED', version=11, note='phase-b-no-hint', updated_at='2026-09-08T04:25:00Z'WHERE order_id='order-1501';
bash · restore handoff configuration and bring the stale node back
docker exec atlasmart-cass-1 nodetool enablehandoffdocker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"

Node 3 can now be UN and still contain the older version. This is the central operational distinction between membership/health and data convergence. Use an ALL read to force the three replicas into reconciliation on this table, which is configured with read_repair='BLOCKING'.

CQL · force all three replicas into reconciliation
CONSISTENCY ALL;TRACING ON;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';TRACING OFF;
bash · verify, then run explicit anti-entropy repair anyway
for n in 1 2 3; do  docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"donedocker exec atlasmart-cass-1 nodetool repair --full atlasmart_failure order_state_by_idfor n in 1 2 3; do  docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"done

4. Phase C — two replicas down: classify before retrying

bash · make LOCAL_QUORUM structurally unavailable
docker pause atlasmart-cass-2docker pause atlasmart-cass-3# Wait until node 1 marks both peers down.docker exec atlasmart-cass-1 nodetool status
CQL · compare LOCAL_ONE and LOCAL_QUORUM under the same failure
CONSISTENCY LOCAL_ONE;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';CONSISTENCY LOCAL_QUORUM;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';

The same topology can satisfy LOCAL_ONE but not LOCAL_QUORUM. This does not mean LOCAL_ONE is the “better” setting: it is a different business contract with weaker freshness/overlap guarantees. An application that promised quorum semantics must surface/fail the request rather than silently downgrade it. Restore the nodes and let failure detection converge before the final acceptance check.

bash · final recovery and operational reset
docker unpause atlasmart-cass-2docker unpause atlasmart-cass-3for n in 1 2 3; do docker exec atlasmart-cass-$n nodetool enablehandoff; donefor n in 1 2 3; do docker exec atlasmart-cass-$n nodetool resumehandoff || true; done# Wait until all nodes are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool repair --full atlasmart_failure order_state_by_idfor n in 1 2 3; do  docker exec atlasmart-cass-$n cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"donedocker exec atlasmart-cass-1 nodetool listpendinghints

5. Acceptance record

Checkpoint Evidence to retain Pass condition
One replica down status + LOCAL_QUORUM result + ALL error quorum succeeds; ALL unavailable after detector convergence
Hint creation listpendinghints / read-only hints directory pending hint appears while delivery paused
Hint replay direct LOCAL_ONE values recovered replica converges after replay
No-hint divergence direct node 3 read node can be UN yet stale before reconciliation/repair
Read repair ALL trace + post-read direct values stale value is reconciled/updated when mismatch participates
Scheduled repair nodetool repair + three direct reads all replicas equal after repair
Two replicas down LOCAL_ONE vs LOCAL_QUORUM different CL availability matches RF arithmetic
Reset statushandoff/status/pending hints all nodes UN; handoff normal; no unexplained backlog
Do not turn the drill into a retry storm.

If the client receives UNAVAILABLE, retrying the same CL immediately while topology is unchanged is usually wasted work. If a mutation times out, treat the outcome as ambiguous and require idempotency/postcondition logic. If the system reports overload, reduce offered load and investigate capacity rather than launching simultaneous retries/speculations.

Check your understanding

  1. Why does node status UN not prove replica equality?
  2. What did Phase A prove that Phase B did not?
  3. Why does LOCAL_ONE succeed with one live RF=3 replica while LOCAL_QUORUM fails?
  4. Why run repair after a successful read-repair demonstration?
  5. What is the retry rule for an ambiguous write timeout?
Review the answers

1. Membership/failure detection says the node is reachable/alive; it does not compare every partition version across replicas.

2. Phase A isolated best-effort hinted handoff; Phase B deliberately removed hints so reconciliation/read repair and explicit repair could be observed separately.

3. LOCAL_ONE requires one local response; LOCAL_QUORUM requires a majority, which is two for RF=3.

4. Read repair is request-scoped and does not prove all other partitions/ranges are consistent; repair is the anti-entropy mechanism.

5. Do not infer failure. Retry only if the operation remains safe when the first attempt may already have applied, preferably with an idempotency/postcondition design.

Production judgment

Failure handling is part of the application contract, not a last-minute driver knob. Record the operation's business invariant, RF/CL, local/remote datacenter scope, partition size/cardinality, payload size, write type, TTL/delete rate, compaction/SSTable state, repair cadence, hint window and delivery backlog, driver timeout/retry/speculation policy, idempotency decision, concurrency, routing, p95/p99/p99.9 latency, server overload/failure counts, JVM/GC, disk/network saturation, and the exact failure injection. A successful retry can hide an incident; a failed timeout can hide a successful mutation. Neither result is enough without postcondition verification.

Hints consume disk and replay network/write capacity; read repair adds foreground read/write work; scheduled repair consumes disk/network/CPU; speculative execution duplicates requests; retries can amplify overload. SAI/vector queries can have different tail-latency and duplicate-work costs, managed services may hide or constrain nodetool/JMX/driver controls, and security/TLS/authentication failures must not be misclassified as ordinary replica availability. Do not lower consistency, enlarge timeouts, extend hint windows, raise overload limits, or enable aggressive speculation from folklore. Test rollback and recovery under representative load. Chapter 16 now focuses on the special case where an application truly needs a conditional invariant: Cassandra lightweight transactions use Paxos-backed compare-and-set rather than generic retries or ordinary quorum writes.

Summary and next bridge

A production failure drill should end with data equality, handoff state, repair evidence, and client error classification—not merely green node status. You can now distinguish hints, read repair, repair, speculation, retries, timeout, unavailable, and overload as separate layers. Chapter 16 builds on that discipline with Paxos-backed lightweight transactions for true conditional invariants.

Authoritative references

Version-sensitive claims in this lesson should be rechecked against these current Apache sources when the lesson is regenerated. Driver 4.19.3 is the pinned artifact; some manual pages remain published under the 4.19.0 documentation path but the relevant semantics are also represented in the current 4.19.x source/changelog.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.