Chapter 15 · Failure Handling: Hints, Read Repair, Speculative Retry, Timeouts, and Unavailable Errors

Hinted Handoff: Best-Effort Delivery of Missed Writes and Retention Limits

Trace a missed RF=3 write into a durable hint, hold/replay the hint safely, and prove why hinted handoff remains best effort rather than backup or anti-entropy repair.

Intermediate105–145 minutesHint retention + replay labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · Java Driver 4.19.3 optional · RF=3 dc1 · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart deploys a pricing change while one Cassandra replica restarts. The checkout write succeeds at LOCAL_QUORUM, yet the restarted node initially lacks the mutation. The operating question is not simply “did Cassandra replicate it?” It is: which replica missed the write, which coordinator stored a hint, how long can new hints be generated, when are they replayed, and what mechanism guarantees convergence if hint delivery is incomplete?

01

Explain hinted handoff as best-effort delivery of missed mutations rather than a second copy of the database.

02

Distinguish hint creation, retention window, delivery pause/throttle, replay, and repair.

03

Observe pending hints and the hints directory without modifying hint files directly.

04

Create a reversible missed-write timeline and prove the stale replica converges after hint replay.

05

Explain why backups and scheduled repair remain necessary even when hinted handoff is healthy.

Chapter 15 lab baseline

The mandatory labs continue the established local AtlasMart cluster with Apache Cassandra 5.0.9 in the pinned cassandra:5.0.9 Docker image, Java 17 inside the image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes (vnodes) per node, NetworkTopologyStrategy, replication factor (RF) 3, and explicit per-request consistency levels (CLs). New Chapter 15 tables use UnifiedCompactionStrategy (UCS), no default Time To Live (TTL), and Cassandra's default gc_grace_seconds unless a lesson intentionally changes a setting. Authentication, client Transport Layer Security (TLS), internode TLS, and remote Java Management Extensions (JMX) are disabled only inside this isolated single-host learning network. The optional application examples use Apache Cassandra Java Driver 4.19.3. Windows learners should run the Linux containers through Docker Desktop/WSL rather than infer native Windows production support. No Storage-Attached Index (SAI) or vector index is required in this chapter, so failure-handling evidence is not confounded by index rebuild/query behavior.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Failure-handling vocabulary

A coordinator is the Cassandra node handling one client request; it is not a permanent leader. A replica stores one copy of a partition according to the keyspace replication strategy. A partition is the set of rows sharing a partition key; Cassandra hashes that key to a token, and token ownership plus topology determines the replica set. RF (replication factor) is the number of replicas; CL (consistency level) is the response requirement for one operation. A hint is a durable record held by another node for a mutation a replica could not receive. Reconciliation chooses the newest visible cell/tombstone versions among replica responses. Read repair may write the reconciled result back to stale replicas involved in a read. Anti-entropy repair is the operator-run process that compares token-range data and streams differences; it is the comprehensive convergence mechanism. Speculation means starting redundant work before the original attempt has definitively failed. Idempotent means that repeating an operation produces the same intended database state as doing it once. A timeout means the operation did not complete within a deadline; it does not by itself prove that no replica applied a write. Unavailable means the coordinator already knows there are too few live replicas for the requested CL. Overloaded is a server response indicating the coordinator cannot currently process the request because its resources/backlog are exhausted.

1. Why Cassandra records a hint

For a normal RF=3 write, the coordinator sends the mutation to all three natural replicas even if the requested CL needs fewer acknowledgments. With LOCAL_QUORUM in this one-DC lab, two local acknowledgments satisfy the client. If the third replica is unavailable, the coordinator can persist a hint on its own local filesystem that identifies the target endpoint and mutation. The client-visible write can therefore succeed while replica state is temporarily divergent.

Current Cassandra 5.0 defaults enable hinted handoff and generate new hints for a failed endpoint for up to the configured max_hint_window, documented as three hours by default. This is a creation window, not a promise that every outage shorter than three hours converges exclusively through hints. Hint storage can be lost, delivery can lag, an outage can exceed the window, and operational mistakes can remove pending hints. Cassandra's documentation therefore classifies hints and read repair as best effort and keeps operator-run repair as the mechanism that provides comprehensive replica synchronization.

Mechanism When it acts Scope Guarantee boundary
Hint write cannot reach a target replica one missed mutation / target endpoint best effort; bounded creation window
Hint replay target is reachable and delivery is enabled pending hints held by nodes reduces inconsistency duration; not a complete audit
Read repair a read observes differing replica versions replicas/data participating in that read request-scoped; unread data can remain stale
Repair operator schedules anti-entropy comparison selected token ranges/keyspaces/tables the comprehensive convergence mechanism
Backup independent recovery copy/process data recovery objective not replaced by hints or repair

2. Reproduce a missed write and hold delivery long enough to observe it

bash · verify or recreate the three-node dc1 course cluster
docker network inspect atlasmart-cassandra >/dev/null 2>&1 || docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker inspect atlasmart-cass-1 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 reports UN before starting peers.docker exec atlasmart-cass-1 nodetool statusdocker inspect atlasmart-cass-2 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker inspect atlasmart-cass-3 >/dev/null 2>&1 || docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -e CASSANDRA_SEEDS=atlasmart-cass-1 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only when all three replicas are UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 cqlsh -e "SHOW VERSION"
CQL · Chapter 15 RF=3 failure fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_failureWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_failure.order_state_by_id (    order_id text PRIMARY KEY,    status text,    version int,    note text,    updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'}  AND read_repair = 'BLOCKING'  AND speculative_retry = 'NONE';CONSISTENCY ALL;INSERT INTO atlasmart_failure.order_state_by_id(order_id,status,version,note,updated_at)VALUES ('order-1501','CREATED',0,'baseline','2026-09-08T03:30:00Z');SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
bash · inspect hint configuration before failure
docker exec atlasmart-cass-1 nodetool statushandoffdocker exec atlasmart-cass-1 nodetool getmaxhintwindowdocker exec atlasmart-cass-1 sh -lc "grep -n -E 'hinted_handoff_enabled|max_hint_window|hints_directory|hinted_handoff_throttle' /etc/cassandra/cassandra.yaml || true"docker exec atlasmart-cass-1 sh -lc "ls -lah /var/lib/cassandra/hints || true"

For evidence, pause delivery on node 1 while still allowing it to store hints. Then make node 3 unavailable and send the write through node 1. pausehandoff is different from disablehandoff: the former is useful here because it lets pending hints accumulate without immediately replaying them when the target returns.

bash · pause hint delivery and make replica 3 unavailable
docker exec atlasmart-cass-1 nodetool pausehandoffdocker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
CQL · LOCAL_QUORUM write through node 1
CONSISTENCY LOCAL_QUORUM;UPDATE atlasmart_failure.order_state_by_idSET status='PAID', version=1, note='missed-by-node3', updated_at='2026-09-08T03:40:00Z'WHERE order_id='order-1501';SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';
bash · observe pending hint evidence
docker exec atlasmart-cass-1 nodetool listpendinghintsdocker exec atlasmart-cass-1 sh -lc "ls -lah /var/lib/cassandra/hints || true"# Exact file names, bytes, endpoint IDs, and counts are runtime-dependent.

3. Recover the replica, prove it is stale, then replay hints

Bring node 3 back while node 1's hint delivery remains paused. A direct LOCAL_ONE query through node 3 can expose the old value if node 3 is a natural replica for this partition—which it is in this RF=3/three-node fixture. This is controlled evidence of temporary divergence, not a production recipe for forcing reads to a particular copy.

bash · recover node 3 but keep hint delivery paused
docker unpause atlasmart-cass-3# Wait until node 3 is UN.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"
bash · resume handoff, wait for replay, then verify
docker exec atlasmart-cass-1 nodetool resumehandoff# Poll rather than assuming replay is instantaneous.docker exec atlasmart-cass-1 nodetool listpendinghintsdocker exec atlasmart-cass-3 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_failure.order_state_by_id WHERE order_id='order-1501';"docker exec atlasmart-cass-1 nodetool status
Unsafe mental model: “Hints are my backup and repair system.”

Hints are stored on cluster nodes, share failure domains with the database, have bounded creation/retention behavior, and are delivered best effort. Extending max_hint_window is not free: it increases potential hint disk footprint and replay work. Never delete hint files by hand to free space. Diagnose why hints are accumulating, restore endpoint health/capacity, and maintain scheduled repair plus independent backup/recovery.

4. Verification and reset

  • All nodes return UN.
  • statushandoff reports enabled future handoff and delivery is resumed.
  • Pending hints for the exercise drain or are explained if still replaying.
  • Node 3 returns status='PAID' and version=1 after replay.
  • No hint files were edited/deleted manually.

Check your understanding

  1. Why can LOCAL_QUORUM succeed while one RF=3 replica is down?
  2. What does max_hint_window bound?
  3. Why use pausehandoff instead of disablehandoff in the observation lab?
  4. If all pending hints replay, can scheduled repair be abandoned?
  5. What new risk appears when the hint window is made much longer?
Review the answers

1. Two local replica acknowledgments satisfy LOCAL_QUORUM; Cassandra still attempts to deliver the mutation to all replicas.

2. How long Cassandra continues generating new hints for a failed endpoint; it is not a complete convergence guarantee or backup retention policy.

3. Pausing delivery lets hints still be stored so they can be observed; disabling handoff prevents future hint storage/delivery on that node.

4. No. Hints are best effort and only cover missed mutations that were actually hinted; repair is still needed for anti-entropy convergence.

5. More disk consumption and potentially larger replay bursts, which can add network/write load when the replica returns.

Production judgment

Failure handling is part of the application contract, not a last-minute driver knob. Record the operation's business invariant, RF/CL, local/remote datacenter scope, partition size/cardinality, payload size, write type, TTL/delete rate, compaction/SSTable state, repair cadence, hint window and delivery backlog, driver timeout/retry/speculation policy, idempotency decision, concurrency, routing, p95/p99/p99.9 latency, server overload/failure counts, JVM/GC, disk/network saturation, and the exact failure injection. A successful retry can hide an incident; a failed timeout can hide a successful mutation. Neither result is enough without postcondition verification.

Hints consume disk and replay network/write capacity; read repair adds foreground read/write work; scheduled repair consumes disk/network/CPU; speculative execution duplicates requests; retries can amplify overload. SAI/vector queries can have different tail-latency and duplicate-work costs, managed services may hide or constrain nodetool/JMX/driver controls, and security/TLS/authentication failures must not be misclassified as ordinary replica availability. Do not lower consistency, enlarge timeouts, extend hint windows, raise overload limits, or enable aggressive speculation from folklore. Test rollback and recovery under representative load. Lesson 2 deliberately removes hints from the equation so you can see reconciliation/read repair itself and why it still does not replace scheduled repair.

Summary and next bridge

Hinted handoff shortens the time a recovered replica stays stale, but it is bounded, best effort, and capacity-sensitive. The next lesson creates replica divergence with hints disabled so that reconciliation and read repair can be isolated from hint replay.

Authoritative references

Version-sensitive claims in this lesson should be rechecked against these current Apache sources when the lesson is regenerated. Driver 4.19.3 is the pinned artifact; some manual pages remain published under the 4.19.0 documentation path but the relevant semantics are also represented in the current 4.19.x source/changelog.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.