Chapter 21 · Drivers, Prepared Statements, Token Awareness, Paging, Retries, and Load Balancing
Retry Policies, Idempotency, Timeouts, Speculative Execution, and Error Classification
Classify Cassandra/client errors, prove idempotence before duplicate execution, and use retries/speculation conservatively with measured attempt and latency evidence.
Learning outcomes
AtlasMart's checkout API sees a timeout and blindly retries the mutation. Sometimes the retry is harmless; sometimes it duplicates a counter/list/event effect. A different team enables aggressive speculative execution for every query and doubles cluster work during overload. The driver needs an error/outcome model, not a blanket “retry on failure” rule.
Distinguish server unavailable/timeouts/failures/overload from client DriverTimeoutException and connectivity errors.
Classify statements as idempotent or non-idempotent before enabling retries or speculative execution.
Explain the Java Driver default retry policy and why it is deliberately conservative.
Enable a small speculative-execution experiment only for an idempotent read and inspect ExecutionInfo counts/coordinator attempts.
Build an application decision table for success, definite failure, ambiguous outcome, retry, reconcile and alert.
The mandatory labs continue the established free/local
AtlasMart cluster: Docker Official Image
cassandra:5.0.9, Java 17 in that image, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and named disposable data volumes. Keyspace
atlasmart_driver uses
NetworkTopologyStrategy with replication factor
(RF) 3; ordinary reads/writes use LOCAL_QUORUM.
New tables explicitly use UnifiedCompactionStrategy (UCS), no
table default time-to-live (TTL), and Cassandra's normal
gc_grace_seconds. Authentication,
client/internode Transport Layer Security (TLS), and remote
Java Management Extensions (JMX) are disabled only on this
isolated single-host learning network. Application examples
use Apache Cassandra Java Driver
4.19.3
(org.apache.cassandra:java-driver-core:4.19.3)
with Java 17 and one long-lived CqlSession.
Driver documentation URLs still use the 4.19.0 documentation
set, while the ASF release is 4.19.3. Exact coordinator
choices, pool counts, traces, paging states, retry/speculation
counts, and latency percentiles are learner-captured runtime
evidence.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and request-path mental model
The native protocol is Cassandra's binary client/server protocol over TCP, normally on port 9042. A driver implements that protocol for an application language. A contact point is an initial address used to bootstrap discovery; it is not a permanent leader or a complete static node list. A session is the driver's long-lived view of a cluster and owns topology metadata, control-plane state, connection pools, policies, prepared-statement caches, and request execution. A connection pool is the driver's set of TCP connections to a node; unlike a JDBC-style blocking pool, each Cassandra connection multiplexes many in-flight requests using native-protocol stream identifiers.
A request is sent to a coordinator, the Cassandra node that handles that request. The driver chooses that coordinator using a load-balancing policy. Token-aware routing prefers replicas for the target partition when the statement supplies a keyspace and routing key/token. The partition key hashes to a token; replicas own token ranges according to the keyspace replication strategy. Prepared statements let the server parse CQL once and return metadata, including bind-variable types and partition-key variable positions, which enables the driver to compute routing keys for bound statements. Paging splits a large result into multiple protocol responses; the paging state is an opaque continuation token tied to the exact statement and values. Idempotent means repeating a request has the same final database effect as executing it once. Retry and speculative-execution policies use that property to avoid duplicating unsafe mutations.
1. Errors describe different failure stages
Unavailable means the coordinator knows too few replicas are alive to satisfy the requested consistency level; adding client wait time does not create replicas. Read/WriteTimeout means the coordinator could start the operation but did not receive enough required responses before the server deadline; a write timeout can be an ambiguous outcome because replicas may already have applied it. Read/WriteFailure means replicas reported failures. Overloaded is a server backpressure signal. DriverTimeoutException is the client's request deadline and does not prove the server did nothing. Connectivity/AllNodesFailed/NoNodeAvailable conditions have still different retry implications.
| Class | What is known | Safe default question |
|---|---|---|
| Unavailable | insufficient live replicas known | fix availability/CL; do not just lengthen timeout |
| WriteTimeout | some write work may have happened | is mutation idempotent or can outcome be reconciled? |
| ReadTimeout | read CL response not completed | is retry safe/useful and cluster healthy? |
| Overloaded | coordinator/backlog pressure | back off; investigate load/capacity |
| Driver timeout | client gave up | outcome may still complete; reconcile unsafe writes |
| Validation/protocol error | request itself invalid/incompatible | do not retry unchanged |
2. Idempotence is the gate for duplicate execution
// Idempotent: repeating sets the same final value.BoundStatement setStatus = updateStatus.bind("PAID","tenant-a",orderId).setIdempotent(true);// Non-idempotent examples: counter increment, list prepend/append, generated-new-ID side effect.BoundStatement increment = incrementCounter.bind(1L,"tenant-a").setIdempotent(false);// Driver default idempotence is false unless configured/overridden.System.out.println(setStatus.isIdempotent());System.out.println(increment.isIdempotent());
The Apache Java Driver only applies its retry/speculative
policies to requests considered idempotent in situations where
duplicate execution could be possible. Marking a non-idempotent
mutation true does not make it safe; it only lies
to the policy.
3. Observe a definite unavailable error
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name,data_center,rack,release_version,native_protocol_version FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer,peer_port,data_center,rack,release_version FROM system.peers_v2;"
CREATE KEYSPACE IF NOT EXISTS atlasmart_driverWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_driver.orders_by_customer_month ( tenant_id text, customer_id text, order_month date, order_time timestamp, order_id uuid, status text, total decimal, note text, PRIMARY KEY ((tenant_id,customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:00:00Z',21000000-0000-0000-0000-000000000001,'PAID',129.90,'baseline');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:10:00Z',21000000-0000-0000-0000-000000000002,'PACKING',59.00,'fragile');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-99','2026-09-01','2026-09-08T08:20:00Z',21000000-0000-0000-0000-000000000003,'PAID',210.00,'priority');
docker pause atlasmart-cass-2docker pause atlasmart-cass-3# Wait until node 1's failure detector reports the paused peers down before testing UNAVAILABLE.docker exec atlasmart-cass-1 nodetool status# LOCAL_QUORUM requires two local replicas, so a read/write should be unavailable.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_driver.orders_by_customer_month WHERE tenant_id='tenant-a' AND customer_id='cust-42' AND order_month='2026-09-01';" || true# Always restore the isolated nodes.docker unpause atlasmart-cass-2docker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
This deliberately demonstrates a known-insufficient-replica condition. It does not reproduce every timeout or overload class; those require timing/resource conditions that should not be faked by unsafe host manipulation.
4. Driver exception classification and conservative retry handling
<project xmlns="http://maven.apache.org/POM/4.0.0"> <modelVersion>4.0.0</modelVersion> <groupId>academy.atlasmart</groupId><artifactId>cassandra-driver-lab</artifactId><version>1.0.0</version> <properties><maven.compiler.release>17</maven.compiler.release></properties> <dependencies> <dependency> <groupId>org.apache.cassandra</groupId> <artifactId>java-driver-core</artifactId> <version>4.19.3</version> </dependency> <dependency><groupId>org.slf4j</groupId><artifactId>slf4j-simple</artifactId><version>2.0.17</version></dependency> </dependencies> <build><plugins><plugin><groupId>org.codehaus.mojo</groupId><artifactId>exec-maven-plugin</artifactId><version>3.5.0</version></plugin></plugins></build></project>
try { session.execute(statement);} catch (com.datastax.oss.driver.api.core.servererrors.UnavailableException e) { // Definite lack of live replicas for CL: surface/route to availability handling.} catch (com.datastax.oss.driver.api.core.servererrors.WriteTimeoutException e) { // Ambiguous write outcome: reconcile business state before replaying non-idempotent work.} catch (com.datastax.oss.driver.api.core.servererrors.ReadTimeoutException e) { // Read didn't satisfy CL in time; application may retry only under bounded policy.} catch (com.datastax.oss.driver.api.core.servererrors.OverloadedException e) { // Backpressure: shed/back off; don't create a retry storm.} catch (com.datastax.oss.driver.api.core.DriverTimeoutException e) { // Client deadline expired; outcome is not proof of non-execution.} catch (com.datastax.oss.driver.api.core.AllNodesFailedException e) { // Inspect per-node causes; this is not one homogeneous Cassandra error.}
# From the Maven project directory. Docker Desktop/WSL or Linux/macOS shell:docker run --rm --network atlasmart-cassandra -v "$PWD:/work" -w /work maven:3.9.16-eclipse-temurin-17 \ mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab# PowerShell uses the same container and network; ${PWD} resolves to the current directory:# docker run --rm --network atlasmart-cassandra -v "${PWD}:/work" -w /work maven:3.9.16-eclipse-temurin-17 `# mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab
The built-in default retry policy retries at most once in selected high-probability/safe cases. Do not replace it with “retry N times on any exception.” Also do not use consistency-downgrading retry unless the business explicitly accepts weaker guarantees.
5. Small speculative-execution experiment
datastax-java-driver.advanced.speculative-execution-policy { class = ConstantSpeculativeExecutionPolicy max-executions = 2 delay = 100 milliseconds}# Only idempotent requests are eligible. Default policy is no speculative execution.
BoundStatement read=preparedRead.bind("tenant-a","cust-42",LocalDate.parse("2026-09-01")) .setIdempotent(true);ResultSet rs=session.execute(read);ExecutionInfo info=rs.getExecutionInfo();System.out.println("coordinator="+info.getCoordinator().getEndPoint());System.out.println("speculativeExecutions="+info.getSpeculativeExecutionCount());System.out.println("previousErrors="+info.getErrors());
On a healthy local lab, the speculative count may stay zero because the first coordinator answers before the delay. That is valid evidence. To test actual speculation safely, inject delay only inside a dedicated network simulator/container setup; do not alter host firewall rules. Measure extra attempts and cluster traffic alongside p95/p99 latency because speculation exchanges duplicate work for tail-latency protection.
When the cluster is already overloaded, duplicate attempts can intensify queueing. Backpressure, capacity/query modeling and bounded concurrency come before more parallel retries. Non-idempotent operations must never be made “eligible” merely to improve latency.
6. Application decision table and chapter reset
| Outcome | Idempotent request | Non-idempotent request |
|---|---|---|
| success | return success | return success |
| unavailable before work | bounded retry/failover only if policy/business allows | usually fail/route availability; no blind replay |
| ambiguous timeout | bounded retry may be safe if truly idempotent | reconcile by request/business identifier before replay |
| overloaded | backoff/shed + investigate | backoff/shed; no retry storm |
| validation/protocol | fix request/version | fix request/version |
DROP KEYSPACE IF EXISTS atlasmart_driver;
Check your understanding
- Why does increasing a client timeout not fix UNAVAILABLE?
- Why can a timed-out write be dangerous to retry?
- What is the Java Driver default idempotence value?
- Does enabling speculation change retry policy?
- What should the application do with an ambiguous non-idempotent outcome?
Review the answers
1. The coordinator already knows there are too few live replicas for the requested CL; waiting longer does not change the replica count.
2. Some replicas may already have applied it, so replay can duplicate non-idempotent effects.
3. False unless configuration or the statement explicitly sets it.
4. No. Parallel executions can each experience their own retry behavior; together they can increase traffic.
5. Use request/business identifiers and reconciliation/read-back to determine state before deciding whether to compensate or replay.
Production judgment
Client correctness depends on more than successful TCP connectivity. Record Cassandra patch/native protocol negotiation, driver artifact/version, Java runtime, local datacenter, discovered topology, node distance, pool/in-flight metrics, prepared-statement/cache behavior, routing-key availability, page/fetch size, request/consistency/serial-consistency timeouts, retry policy, idempotence classification, speculative-execution policy, per-query execution profile, and p50/p95/p99 client latency. Correlate driver coordinator/attempt data with Cassandra tracing, server timeouts/failures, replica availability, RF/CL, partition size/cardinality, SSTable/compaction/tombstone state, disk/network/JVM pressure, repair state, and SAI/vector costs where those query paths are used.
Do not make a session per HTTP request, disable local-DC awareness casually, expose raw paging state as an authorization token, or enable broad retries/speculation to hide overload. Treat driver configuration as application production code: version it, test node/DC failures, test ambiguous timeouts with idempotent and non-idempotent mutations, measure extra traffic, and define rollback. Managed services may provide different endpoints, TLS/auth requirements, topology visibility, or restricted metrics while still using a Cassandra-compatible protocol; verify the service contract rather than assuming identical behavior. Chapter 22 moves from driver-side policy to server/operator configuration surfaces: cassandra.yaml, JVM/environment settings, seeds, addresses, nodetool/cqlsh, dynamic settings and configuration-drift control.
Summary and next bridge
A maintained driver is part of Cassandra correctness: session lifecycle, prepared routing metadata, local-DC policy, paging state, timeouts, retries and speculation shape every request. Configure these mechanisms from business semantics and measured failure behavior—not driver folklore. Chapter 22 turns to the server configuration and operational tools that define the other side of that contract.
Authoritative references
Driver behavior is language- and version-specific. Re-check the exact maintained driver and Cassandra patch before freezing production defaults or error-handling behavior.
- Apache Cassandra downloads / current 5.0 patch
- Docker Official Cassandra 5.0 image source
- Apache Cassandra Java Driver overview
- Java Driver pooling
- Java Driver prepared statements
- Java Driver load balancing / token awareness
- Java Driver paging
- Java Driver retries
- Java Driver idempotence
- Java Driver speculative execution
- Maven Central Java Driver 4.19.3