Chapter 21 · Drivers, Prepared Statements, Token Awareness, Paging, Retries, and Load Balancing

Retry Policies, Idempotency, Timeouts, Speculative Execution, and Error Classification

Classify Cassandra/client errors, prove idempotence before duplicate execution, and use retries/speculation conservatively with measured attempt and latency evidence.

Intermediate → Advanced120–165 minutesRetry/speculation/error labApache Cassandra 5.0.9 · Java 17 · Apache Java Driver 4.19.3 · RF=3 · LOCAL_QUORUM · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart's checkout API sees a timeout and blindly retries the mutation. Sometimes the retry is harmless; sometimes it duplicates a counter/list/event effect. A different team enables aggressive speculative execution for every query and doubles cluster work during overload. The driver needs an error/outcome model, not a blanket “retry on failure” rule.

01

Distinguish server unavailable/timeouts/failures/overload from client DriverTimeoutException and connectivity errors.

02

Classify statements as idempotent or non-idempotent before enabling retries or speculative execution.

03

Explain the Java Driver default retry policy and why it is deliberately conservative.

04

Enable a small speculative-execution experiment only for an idempotent read and inspect ExecutionInfo counts/coordinator attempts.

05

Build an application decision table for success, definite failure, ambiguous outcome, retry, reconcile and alert.

Chapter 21 lab baseline

The mandatory labs continue the established free/local AtlasMart cluster: Docker Official Image cassandra:5.0.9, Java 17 in that image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes per node, and named disposable data volumes. Keyspace atlasmart_driver uses NetworkTopologyStrategy with replication factor (RF) 3; ordinary reads/writes use LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no table default time-to-live (TTL), and Cassandra's normal gc_grace_seconds. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only on this isolated single-host learning network. Application examples use Apache Cassandra Java Driver 4.19.3 (org.apache.cassandra:java-driver-core:4.19.3) with Java 17 and one long-lived CqlSession. Driver documentation URLs still use the 4.19.0 documentation set, while the ASF release is 4.19.3. Exact coordinator choices, pool counts, traces, paging states, retry/speculation counts, and latency percentiles are learner-captured runtime evidence.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and request-path mental model

The native protocol is Cassandra's binary client/server protocol over TCP, normally on port 9042. A driver implements that protocol for an application language. A contact point is an initial address used to bootstrap discovery; it is not a permanent leader or a complete static node list. A session is the driver's long-lived view of a cluster and owns topology metadata, control-plane state, connection pools, policies, prepared-statement caches, and request execution. A connection pool is the driver's set of TCP connections to a node; unlike a JDBC-style blocking pool, each Cassandra connection multiplexes many in-flight requests using native-protocol stream identifiers.

A request is sent to a coordinator, the Cassandra node that handles that request. The driver chooses that coordinator using a load-balancing policy. Token-aware routing prefers replicas for the target partition when the statement supplies a keyspace and routing key/token. The partition key hashes to a token; replicas own token ranges according to the keyspace replication strategy. Prepared statements let the server parse CQL once and return metadata, including bind-variable types and partition-key variable positions, which enables the driver to compute routing keys for bound statements. Paging splits a large result into multiple protocol responses; the paging state is an opaque continuation token tied to the exact statement and values. Idempotent means repeating a request has the same final database effect as executing it once. Retry and speculative-execution policies use that property to avoid duplicating unsafe mutations.

1. Errors describe different failure stages

Unavailable means the coordinator knows too few replicas are alive to satisfy the requested consistency level; adding client wait time does not create replicas. Read/WriteTimeout means the coordinator could start the operation but did not receive enough required responses before the server deadline; a write timeout can be an ambiguous outcome because replicas may already have applied it. Read/WriteFailure means replicas reported failures. Overloaded is a server backpressure signal. DriverTimeoutException is the client's request deadline and does not prove the server did nothing. Connectivity/AllNodesFailed/NoNodeAvailable conditions have still different retry implications.

Class What is known Safe default question
Unavailable insufficient live replicas known fix availability/CL; do not just lengthen timeout
WriteTimeout some write work may have happened is mutation idempotent or can outcome be reconciled?
ReadTimeout read CL response not completed is retry safe/useful and cluster healthy?
Overloaded coordinator/backlog pressure back off; investigate load/capacity
Driver timeout client gave up outcome may still complete; reconcile unsafe writes
Validation/protocol error request itself invalid/incompatible do not retry unchanged

2. Idempotence is the gate for duplicate execution

Java · classify mutations explicitly
// Idempotent: repeating sets the same final value.BoundStatement setStatus = updateStatus.bind("PAID","tenant-a",orderId).setIdempotent(true);// Non-idempotent examples: counter increment, list prepend/append, generated-new-ID side effect.BoundStatement increment = incrementCounter.bind(1L,"tenant-a").setIdempotent(false);// Driver default idempotence is false unless configured/overridden.System.out.println(setStatus.isIdempotent());System.out.println(increment.isIdempotent());

The Apache Java Driver only applies its retry/speculative policies to requests considered idempotent in situations where duplicate execution could be possible. Marking a non-idempotent mutation true does not make it safe; it only lies to the policy.

3. Observe a definite unavailable error

bash / PowerShell-friendly Docker commands · verify the shared cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name,data_center,rack,release_version,native_protocol_version FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer,peer_port,data_center,rack,release_version FROM system.peers_v2;"
CQL · create the driver-focused AtlasMart query table
CREATE KEYSPACE IF NOT EXISTS atlasmart_driverWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_driver.orders_by_customer_month (  tenant_id text,  customer_id text,  order_month date,  order_time timestamp,  order_id uuid,  status text,  total decimal,  note text,  PRIMARY KEY ((tenant_id,customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:00:00Z',21000000-0000-0000-0000-000000000001,'PAID',129.90,'baseline');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:10:00Z',21000000-0000-0000-0000-000000000002,'PACKING',59.00,'fragile');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-99','2026-09-01','2026-09-08T08:20:00Z',21000000-0000-0000-0000-000000000003,'PAID',210.00,'priority');
bash · temporarily leave only one RF=3 replica available
docker pause atlasmart-cass-2docker pause atlasmart-cass-3# Wait until node 1's failure detector reports the paused peers down before testing UNAVAILABLE.docker exec atlasmart-cass-1 nodetool status# LOCAL_QUORUM requires two local replicas, so a read/write should be unavailable.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_driver.orders_by_customer_month WHERE tenant_id='tenant-a' AND customer_id='cust-42' AND order_month='2026-09-01';" || true# Always restore the isolated nodes.docker unpause atlasmart-cass-2docker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status

This deliberately demonstrates a known-insufficient-replica condition. It does not reproduce every timeout or overload class; those require timing/resource conditions that should not be faked by unsafe host manipulation.

4. Driver exception classification and conservative retry handling

XML · minimal pinned Maven project for the Apache Java Driver
<project xmlns="http://maven.apache.org/POM/4.0.0">  <modelVersion>4.0.0</modelVersion>  <groupId>academy.atlasmart</groupId><artifactId>cassandra-driver-lab</artifactId><version>1.0.0</version>  <properties><maven.compiler.release>17</maven.compiler.release></properties>  <dependencies>    <dependency>      <groupId>org.apache.cassandra</groupId>      <artifactId>java-driver-core</artifactId>      <version>4.19.3</version>    </dependency>    <dependency><groupId>org.slf4j</groupId><artifactId>slf4j-simple</artifactId><version>2.0.17</version></dependency>  </dependencies>  <build><plugins><plugin><groupId>org.codehaus.mojo</groupId><artifactId>exec-maven-plugin</artifactId><version>3.5.0</version></plugin></plugins></build></project>
Java · classify failures without blanket retry loops
try {  session.execute(statement);} catch (com.datastax.oss.driver.api.core.servererrors.UnavailableException e) {  // Definite lack of live replicas for CL: surface/route to availability handling.} catch (com.datastax.oss.driver.api.core.servererrors.WriteTimeoutException e) {  // Ambiguous write outcome: reconcile business state before replaying non-idempotent work.} catch (com.datastax.oss.driver.api.core.servererrors.ReadTimeoutException e) {  // Read didn't satisfy CL in time; application may retry only under bounded policy.} catch (com.datastax.oss.driver.api.core.servererrors.OverloadedException e) {  // Backpressure: shed/back off; don't create a retry storm.} catch (com.datastax.oss.driver.api.core.DriverTimeoutException e) {  // Client deadline expired; outcome is not proof of non-execution.} catch (com.datastax.oss.driver.api.core.AllNodesFailedException e) {  // Inspect per-node causes; this is not one homogeneous Cassandra error.}
bash / PowerShell · compile and run on the Cassandra Docker network
# From the Maven project directory. Docker Desktop/WSL or Linux/macOS shell:docker run --rm --network atlasmart-cassandra -v "$PWD:/work" -w /work maven:3.9.16-eclipse-temurin-17 \  mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab# PowerShell uses the same container and network; ${PWD} resolves to the current directory:# docker run --rm --network atlasmart-cassandra -v "${PWD}:/work" -w /work maven:3.9.16-eclipse-temurin-17 `#   mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab

The built-in default retry policy retries at most once in selected high-probability/safe cases. Do not replace it with “retry N times on any exception.” Also do not use consistency-downgrading retry unless the business explicitly accepts weaker guarantees.

5. Small speculative-execution experiment

HOCON · enable bounded speculation in an isolated execution profile
datastax-java-driver.advanced.speculative-execution-policy {  class = ConstantSpeculativeExecutionPolicy  max-executions = 2  delay = 100 milliseconds}# Only idempotent requests are eligible. Default policy is no speculative execution.
Java · inspect request execution evidence
BoundStatement read=preparedRead.bind("tenant-a","cust-42",LocalDate.parse("2026-09-01"))    .setIdempotent(true);ResultSet rs=session.execute(read);ExecutionInfo info=rs.getExecutionInfo();System.out.println("coordinator="+info.getCoordinator().getEndPoint());System.out.println("speculativeExecutions="+info.getSpeculativeExecutionCount());System.out.println("previousErrors="+info.getErrors());

On a healthy local lab, the speculative count may stay zero because the first coordinator answers before the delay. That is valid evidence. To test actual speculation safely, inject delay only inside a dedicated network simulator/container setup; do not alter host firewall rules. Measure extra attempts and cluster traffic alongside p95/p99 latency because speculation exchanges duplicate work for tail-latency protection.

Wrong approach: speculative writes + retries to hide overload.

When the cluster is already overloaded, duplicate attempts can intensify queueing. Backpressure, capacity/query modeling and bounded concurrency come before more parallel retries. Non-idempotent operations must never be made “eligible” merely to improve latency.

6. Application decision table and chapter reset

Outcome Idempotent request Non-idempotent request
success return success return success
unavailable before work bounded retry/failover only if policy/business allows usually fail/route availability; no blind replay
ambiguous timeout bounded retry may be safe if truly idempotent reconcile by request/business identifier before replay
overloaded backoff/shed + investigate backoff/shed; no retry storm
validation/protocol fix request/version fix request/version
CQL · reset the Chapter 21 fixture without touching earlier keyspaces
DROP KEYSPACE IF EXISTS atlasmart_driver;

Check your understanding

  1. Why does increasing a client timeout not fix UNAVAILABLE?
  2. Why can a timed-out write be dangerous to retry?
  3. What is the Java Driver default idempotence value?
  4. Does enabling speculation change retry policy?
  5. What should the application do with an ambiguous non-idempotent outcome?
Review the answers

1. The coordinator already knows there are too few live replicas for the requested CL; waiting longer does not change the replica count.

2. Some replicas may already have applied it, so replay can duplicate non-idempotent effects.

3. False unless configuration or the statement explicitly sets it.

4. No. Parallel executions can each experience their own retry behavior; together they can increase traffic.

5. Use request/business identifiers and reconciliation/read-back to determine state before deciding whether to compensate or replay.

Production judgment

Client correctness depends on more than successful TCP connectivity. Record Cassandra patch/native protocol negotiation, driver artifact/version, Java runtime, local datacenter, discovered topology, node distance, pool/in-flight metrics, prepared-statement/cache behavior, routing-key availability, page/fetch size, request/consistency/serial-consistency timeouts, retry policy, idempotence classification, speculative-execution policy, per-query execution profile, and p50/p95/p99 client latency. Correlate driver coordinator/attempt data with Cassandra tracing, server timeouts/failures, replica availability, RF/CL, partition size/cardinality, SSTable/compaction/tombstone state, disk/network/JVM pressure, repair state, and SAI/vector costs where those query paths are used.

Do not make a session per HTTP request, disable local-DC awareness casually, expose raw paging state as an authorization token, or enable broad retries/speculation to hide overload. Treat driver configuration as application production code: version it, test node/DC failures, test ambiguous timeouts with idempotent and non-idempotent mutations, measure extra traffic, and define rollback. Managed services may provide different endpoints, TLS/auth requirements, topology visibility, or restricted metrics while still using a Cassandra-compatible protocol; verify the service contract rather than assuming identical behavior. Chapter 22 moves from driver-side policy to server/operator configuration surfaces: cassandra.yaml, JVM/environment settings, seeds, addresses, nodetool/cqlsh, dynamic settings and configuration-drift control.

Summary and next bridge

A maintained driver is part of Cassandra correctness: session lifecycle, prepared routing metadata, local-DC policy, paging state, timeouts, retries and speculation shape every request. Configure these mechanisms from business semantics and measured failure behavior—not driver folklore. Chapter 22 turns to the server configuration and operational tools that define the other side of that contract.

Authoritative references

Driver behavior is language- and version-specific. Re-check the exact maintained driver and Cassandra patch before freezing production defaults or error-handling behavior.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.