Chapter 15 · Failure Handling: Hints, Read Repair, Speculative Retry, Timeouts, and Unavailable Errors

Timeout vs Unavailable vs Overloaded / Failure Responses and Correct Client Interpretation

Classify unavailable, server read/write timeout, overload/failure, and client deadline so retry behavior preserves correctness instead of amplifying incidents.

Intermediate110–150 minutesError classification + driver labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart's client currently catches every Cassandra exception and sleeps 200 ms before retrying three times. During a replica outage this turns UNAVAILABLE into more useless traffic; during a write timeout it can duplicate a mutation whose outcome is already unknown; during overload it can amplify the server queue. This lesson replaces “retry on exception” with evidence-based classification.

01

Distinguish UNAVAILABLE, READ_TIMEOUT, WRITE_TIMEOUT, overloaded, read/write failure, connection abort, and client request timeout.

02

Explain why a write timeout has an ambiguous commit outcome and why longer timeout does not solve unavailable.

03

Use controlled replica loss to make UNAVAILABLE deterministic and capture required/alive evidence.

04

Classify driver exceptions and gate automated retry on both error semantics and statement idempotency.

05

Design postcondition checks and application-level retry/backoff instead of infinite driver retries.

Chapter 15 lab baseline

The mandatory labs continue the established local AtlasMart cluster with Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 Docker image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes (vnodes) per node, NetworkTopologyStrategy, replication factor (RF) 3, and explicit per-request consistency levels (CLs). New Chapter 15 tables use UnifiedCompactionStrategy (UCS), no default Time To Live (TTL), and Cassandra's default gc_grace_seconds unless a lesson intentionally changes a setting. Authentication, client Transport Layer Security (TLS), internode TLS, and remote Java Management Extensions (JMX) are disabled only inside this isolated single-host learning network. The optional application examples use Apache Cassandra Java Driver 4.19.3. Windows learners should run the Linux containers through Docker Desktop/WSL rather than infer native Windows production support. No Storage-Attached Index (SAI) or vector index is required in this chapter, so failure-handling evidence is not confounded by index rebuild/query behavior.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Failure-handling vocabulary

A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores one copy of a partition according to the keyspace replication strategy. A partition is the set of rows sharing a partition key; Cassandra hashes that key to a token, and token ownership plus topology determines the replica set. RF (replication factor) is the number of replicas; CL (consistency level) is the response requirement for one operation. A hint is a durable record held by another node for a mutation a replica could not receive. Reconciliation chooses the newest visible cell/tombstone versions among replica responses. Read repair may write the reconciled result back to stale replicas involved in a read. Anti-entropy repair is the operator-run process that compares token-range data and streams differences; it is the comprehensive convergence mechanism. Speculation means starting redundant work before the original attempt has definitively failed. Idempotent means that repeating an operation produces the same intended database state as doing it once. A timeout means the operation did not complete within a deadline; it does not by itself prove that no replica applied a write. Unavailable means the coordinator already knows there are too few live replicas for the requested CL. Overloaded is a server response indicating the coordinator cannot currently process the request because its resources/backlog are exhausted.

1. Error classes are observations about different failure stages

The native protocol carries structured errors. UNAVAILABLE includes the requested CL plus the number of replicas required and known alive; it is decided before Cassandra can satisfy the operation at that CL. A write timeout means the coordinator did not receive enough acknowledgments before the server-side write deadline; some replicas may already have applied the mutation. A read timeout means the read failed to gather the required responses/data before its deadline. Overloaded means the coordinator is refusing/deferring work because resources/backlog are exhausted. ReadFailure/WriteFailure report failures from replicas rather than simply slow responses. Separately, a driver can hit its own request deadline or lose a connection without receiving a definitive coordinator response.

Signal What Cassandra/driver knows Unsafe interpretation Better first action
UNAVAILABLE alive replicas < required for CL “wait longer” restore availability or intentionally choose a different business contract
WRITE_TIMEOUT insufficient acks before server deadline “the write did not happen” treat outcome as ambiguous; verify/idempotency-aware retry
READ_TIMEOUT insufficient timely read responses/data “data is missing” inspect replica latency/load/tombstones/CL; retry only by policy
OVERLOADED coordinator cannot accept/process current work “retry immediately everywhere” reduce pressure/back off/fail over only if safe
READ/WRITE_FAILURE replica-side failures were reported “same as timeout” inspect server failure cause; default driver generally rethrows
Driver timeout/connection abort client did not receive a definitive result “server definitely rejected it” classify statement idempotency and verify outcome

2. Reproduce UNAVAILABLE instead of waiting for a timeout

bash · verify or recreate the three-node dc1 course cluster
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three replicas are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CQL · Chapter 15 RF=3 failure fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_failureWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_failure.order_state_by_id (    order_id text PRIMARY KEY,    status text,    version int,    note text,    updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND read_repair = 'BLOCKING'  AND speculative_retry = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_failure.order_state_by_id(order_id,status,version,note,updated_at)VALUES ('order-1501','CREATED',0,'baseline','2026-09-08T03:30:00Z');SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
bash · remove two replicas from the live set
docker pause atlasmart-cass-2docker pause atlasmart-cass-3# Wait until failure detection on node 1 marks them down; this is not instantaneous.docker exec atlasmart-cass-1 nodetool status
CQL · LOCAL_QUORUM cannot be satisfied with one live RF=3 replica
CONSISTENCY LOCAL_QUORUM;SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';UPDATE atlasmart_failure.order_state_by_idSET status='SHIPPED', version=3, updated_at='2026-09-08T04:10:00Z'WHERE order_id='order-1501';

Once the coordinator knows only one replica is alive, LOCAL_QUORUM needs two and returns unavailable without waiting for a longer read/write timeout. Increasing a client timeout cannot manufacture a second live replica. Restore both nodes before continuing.

bash · recover the disposable replicas
docker unpause atlasmart-cass-2docker unpause atlasmart-cass-3# Wait until all nodes return UN.docker exec atlasmart-cass-1 nodetool status

3. A timeout is ambiguous; a retry must be safe even if the first attempt succeeded

For writes, the core safety question is not “did the client see success?” but “what happens if this mutation is applied twice?” A server write timeout can arrive after one or more replicas persisted the mutation. A client-side deadline can expire while the coordinator is still finishing work. Therefore a blind second execution of a counter increment, list append, generated-ID insert, or external side effect can corrupt the business outcome.

Wrong retry loop

catch-all exception → retry erases the distinction among unavailable, ambiguous timeout, overload, validation errors, authentication errors, and non-idempotent mutations. It can also synchronize thousands of clients into a retry storm. Correctness requires operation classification, bounded attempts, jitter/backoff where appropriate, circuit/throttle behavior, and postcondition or idempotency-key design.

Java · classify coordinator/server and client failures
import com.datastax.oss.driver.api.core.DriverTimeoutException;import com.datastax.oss.driver.api.core.servererrors.*;try {    session.execute(statement);} catch (UnavailableException e) {    // Required replicas are not alive for this CL. Longer timeout is not the fix.    record("unavailable", e.getConsistencyLevel(), e.getRequired(), e.getAlive());    throw e;} catch (WriteTimeoutException e) {    // Outcome may be partially/applied. Retry only if the business operation is proven idempotent.    record("write_timeout", e.getConsistencyLevel(), e.getWriteType());    throw e;} catch (ReadTimeoutException e) {    record("read_timeout", e.getConsistencyLevel());    throw e;} catch (OverloadedException e) {    // Capacity/backpressure signal; do not create an immediate retry storm.    record("overloaded");    throw e;} catch (ReadFailureException | WriteFailureException e) {    record("replica_failure");    throw e;} catch (DriverTimeoutException e) {    // Client deadline expired; server outcome may be unknown for mutations.    record("driver_timeout");    throw e;}

The Apache Java Driver's default retry policy is intentionally conservative and retry-policy callbacks are only used for eligible/idempotent requests in ambiguous cases. The documentation warns that consistency-downgrading retry can break application invariants and datacenter locality; do not use it as an availability shortcut without a deliberate business contract.

4. Error response versus capacity response

Overload is often better handled by reducing offered load than by multiplying attempts: driver throttling, bounded queues, admission control, backoff/jitter, and capacity repair can protect the system. A retry on another node can sometimes succeed, but if the whole cluster is saturated, retries simply redistribute the same pressure. Likewise, increasing timeouts can reduce visible timeout counts while making queues and user latency worse.

Check your understanding

  1. Why does UNAVAILABLE not improve merely by increasing a timeout?
  2. Does WRITE_TIMEOUT prove the mutation was not applied?
  3. What does OVERLOADED communicate?
  4. Why must idempotency be checked before retrying an ambiguous write?
  5. Why is a consistency-downgrading retry policy dangerous?
Review the answers

1. The coordinator already knows fewer replicas are alive than the requested CL requires; time is not the missing resource.

2. No. Some replicas may already have persisted it, so the outcome can be ambiguous.

3. The coordinator cannot currently process the request because its resources/backlog are exhausted; it is a capacity/backpressure signal.

4. The first attempt may already have applied, so a repeated non-idempotent mutation can change state twice.

5. It can silently violate invariants and locality assumptions that were based on the originally requested consistency level.

Production judgment

Failure handling is part of the application contract, not a last-minute driver knob. Record the operation's business invariant, RF/CL, local/remote datacenter scope, partition size/cardinality, payload size, write type, TTL/delete rate, compaction/SSTable state, repair cadence, hint window and delivery backlog, driver timeout/retry/speculation policy, idempotency decision, concurrency, routing, p95/p99/p99.9 latency, server overload/failure counts, JVM/GC, disk/network saturation, and the exact failure injection. A successful retry can hide an incident; a failed timeout can hide a successful mutation. Neither result is enough without postcondition verification.

Hints consume disk and replay network/write capacity; read repair adds foreground read/write work; scheduled repair consumes disk/network/CPU; speculative execution duplicates requests; retries can amplify overload. SAI/vector queries can have different tail-latency and duplicate-work costs, managed services may hide or constrain nodetool/JMX/driver controls, and security/TLS/authentication failures must not be misclassified as ordinary replica availability. Do not lower consistency, enlarge timeouts, extend hint windows, raise overload limits, or enable aggressive speculation from folklore. Test rollback and recovery under representative load. Lesson 5 combines hints, CL changes, unavailable errors, stale replicas, read repair, driver-safe retry reasoning, recovery, and explicit anti-entropy repair into one timed failure drill.

Summary and next bridge

Unavailable, timeout, overload, replica failure, and client deadline are different evidence. Correct clients classify them, preserve the requested business invariant, and retry only when the operation remains safe under an ambiguous first outcome. The final lesson rehearses the full recovery timeline.

Authoritative references

Version-sensitive claims in this lesson should be rechecked against these current Apache sources when the lesson is regenerated. Driver 4.19.3 is the pinned artifact; some manual pages remain published under the 4.19.0 documentation path but the relevant semantics are also represented in the current 4.19.x source/changelog.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.