Chapter 21 · Drivers, Prepared Statements, Token Awareness, Paging, Retries, and Load Balancing
Load Balancing by Local Datacenter, Host Distance, Failure Detection, and Request Routing
Inspect local-DC load balancing, node distance/state, token-aware replica preference, and reversible node failure from both server and Java-driver viewpoints.
Learning outcomes
AtlasMart adds a second datacenter later, but the application must not send normal requests across the WAN by accident. Even in today's single-DC lab, the driver already classifies nodes by distance and uses node state plus token information to build each query plan. This lesson makes that policy visible.
Explain local datacenter, node distance, query plan, token awareness, coordinator choice and driver node-state monitoring.
Inspect node DC/rack/state/distance and correlate selected coordinators with prepared routing keys.
Demonstrate reversible node failure and distinguish server gossip/failure state from the driver connection-based view.
Explain why the built-in Java Driver policy avoids remote-DC traffic by default and why local-DC configuration is mandatory.
Reject naive round-robin or manual node pinning as substitutes for topology-aware routing.
The mandatory labs continue the established free/local
AtlasMart cluster: Docker Official Image
cassandra:5.0.9, Java 17 in that image, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and named disposable data volumes. Keyspace
atlasmart_driver uses
NetworkTopologyStrategy with replication factor
(RF) 3; ordinary reads/writes use LOCAL_QUORUM.
New tables explicitly use UnifiedCompactionStrategy (UCS), no
table default time-to-live (TTL), and Cassandra's normal
gc_grace_seconds. Authentication,
client/internode Transport Layer Security (TLS), and remote
Java Management Extensions (JMX) are disabled only on this
isolated single-host learning network. Application examples
use Apache Cassandra Java Driver
4.19.3
(org.apache.cassandra:java-driver-core:4.19.3)
with Java 17 and one long-lived CqlSession.
Driver documentation URLs still use the 4.19.0 documentation
set, while the ASF release is 4.19.3. Exact coordinator
choices, pool counts, traces, paging states, retry/speculation
counts, and latency percentiles are learner-captured runtime
evidence.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and request-path mental model
The native protocol is Cassandra's binary client/server protocol over TCP, normally on port 9042. A driver implements that protocol for an application language. A contact point is an initial address used to bootstrap discovery; it is not a permanent leader or a complete static node list. A session is the driver's long-lived view of a cluster and owns topology metadata, control-plane state, connection pools, policies, prepared-statement caches, and request execution. A connection pool is the driver's set of TCP connections to a node; unlike a JDBC-style blocking pool, each Cassandra connection multiplexes many in-flight requests using native-protocol stream identifiers.
A request is sent to a coordinator, the Cassandra node that handles that request. The driver chooses that coordinator using a load-balancing policy. Token-aware routing prefers replicas for the target partition when the statement supplies a keyspace and routing key/token. The partition key hashes to a token; replicas own token ranges according to the keyspace replication strategy. Prepared statements let the server parse CQL once and return metadata, including bind-variable types and partition-key variable positions, which enables the driver to compute routing keys for bound statements. Paging splits a large result into multiple protocol responses; the paging state is an opaque continuation token tied to the exact statement and values. Idempotent means repeating a request has the same final database effect as executing it once. Retry and speculative-execution policies use that property to avoid duplicating unsafe mutations.
1. Locality is a policy decision built from topology metadata
The Java driver's default load-balancing policy assigns node distance and builds a query plan for each request. Built-in policies prioritize the configured local datacenter and, by default, avoid opening normal request pools to remote-DC nodes. When routing information is present, token awareness moves replicas for the partition to the front of the local plan. The result is locality plus ownership awareness without turning any node into a leader.
Node state is also a client observation. Cassandra gossip may suspect a node while an application still has a working TCP connection, or the driver may be reconnecting after a connection failure. Diagnose server and driver views separately.
| Signal | Server side | Driver side |
|---|---|---|
| topology | system.local/peers_v2, nodetool status | Metadata nodes, getDatacenter/getRack |
| health | gossip/failure detector | NodeState, open connections, reconnect state |
| proximity | snitch/DC/rack labels | NodeDistance from load-balancing policy |
| ownership | tokens/replicas | token metadata + routing key |
| coordinator | trace events | ExecutionInfo.getCoordinator() |
2. Print the driver's routing view
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name,data_center,rack,release_version,native_protocol_version FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer,peer_port,data_center,rack,release_version FROM system.peers_v2;"
CREATE KEYSPACE IF NOT EXISTS atlasmart_driverWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_driver.orders_by_customer_month ( tenant_id text, customer_id text, order_month date, order_time timestamp, order_id uuid, status text, total decimal, note text, PRIMARY KEY ((tenant_id,customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:00:00Z',21000000-0000-0000-0000-000000000001,'PAID',129.90,'baseline');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:10:00Z',21000000-0000-0000-0000-000000000002,'PACKING',59.00,'fragile');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-99','2026-09-01','2026-09-08T08:20:00Z',21000000-0000-0000-0000-000000000003,'PAID',210.00,'priority');
<project xmlns="http://maven.apache.org/POM/4.0.0"> <modelVersion>4.0.0</modelVersion> <groupId>academy.atlasmart</groupId><artifactId>cassandra-driver-lab</artifactId><version>1.0.0</version> <properties><maven.compiler.release>17</maven.compiler.release></properties> <dependencies> <dependency> <groupId>org.apache.cassandra</groupId> <artifactId>java-driver-core</artifactId> <version>4.19.3</version> </dependency> <dependency><groupId>org.slf4j</groupId><artifactId>slf4j-simple</artifactId><version>2.0.17</version></dependency> </dependencies> <build><plugins><plugin><groupId>org.codehaus.mojo</groupId><artifactId>exec-maven-plugin</artifactId><version>3.5.0</version></plugin></plugins></build></project>
package academy.atlasmart;import com.datastax.oss.driver.api.core.*; import com.datastax.oss.driver.api.core.cql.*;import java.net.*; import java.time.*;public class DriverLab { public static void main(String[] a){ try(CqlSession s=CqlSession.builder().addContactPoint(new InetSocketAddress("atlasmart-cass-1",9042)).withLocalDatacenter("dc1").build()){ s.getMetadata().getNodes().values().forEach(n -> System.out.printf("%s state=%s dc=%s rack=%s distance=%s conns=%d%n",n.getEndPoint(),n.getState(),n.getDatacenter(),n.getRack(),n.getDistance(),n.getOpenConnections())); PreparedStatement ps=s.prepare("SELECT * FROM atlasmart_driver.orders_by_customer_month WHERE tenant_id=? AND customer_id=? AND order_month=?"); BoundStatement bs=ps.bind("tenant-a","cust-42",LocalDate.parse("2026-09-01")).setIdempotent(true); for(int i=0;i<12;i++){ ResultSet rs=s.execute(bs); System.out.println("coordinator="+rs.getExecutionInfo().getCoordinator().getEndPoint()); } } }}
# From the Maven project directory. Docker Desktop/WSL or Linux/macOS shell:docker run --rm --network atlasmart-cassandra -v "$PWD:/work" -w /work maven:3.9.16-eclipse-temurin-17 \ mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab# PowerShell uses the same container and network; ${PWD} resolves to the current directory:# docker run --rm --network atlasmart-cassandra -v "${PWD}:/work" -w /work maven:3.9.16-eclipse-temurin-17 `# mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab
3. Controlled node failure: compare server and driver observations
docker pause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status# Run the Java RoutingView repeatedly while node 3 is paused.# Then restore it:docker unpause atlasmart-cass-3docker exec atlasmart-cass-1 nodetool status
With RF=3 and LOCAL_QUORUM, reads can still succeed
with one replica unavailable if two local replicas respond.
Coordinator choices should exclude a node the driver considers
unavailable, but exact transition timing depends on TCP, driver
heartbeat/reconnection, gossip events and the pause duration.
Record timestamps rather than writing a fixed “down in N
seconds” claim.
4. Wrong approach: plain round-robin or no local datacenter
CqlSession session=CqlSession.builder() .addContactPoint(new InetSocketAddress("atlasmart-cass-1",9042)) .withLocalDatacenter("dc1") .build();// The default load-balancing policy is token-aware and local-DC oriented.// Do not replace it with a naive global round-robin simply to "spread load".
In a multi-DC cluster, indiscriminate coordinators add cross-DC latency and can interact badly with LOCAL consistency levels, failover assumptions and capacity. Local-DC routing is not authorization or disaster-recovery correctness either; application failover to another DC requires explicit traffic, RF/CL and recovery design.
5. Verification and topology-change acceptance
- Every discovered node prints DC/rack/state/distance and open connections.
- The prepared statement exposes routing information and coordinators are chosen by the policy, not pinned.
- During the paused-node drill, requests continue at LOCAL_QUORUM and node recovery is observed from both server and driver.
- No host firewall or system clock is modified.
- A future second DC would require explicit local-DC/failover design and capacity tests before traffic cutover.
Check your understanding
- What decides NodeDistance in the driver?
- Does LOCAL_QUORUM itself choose the coordinator?
- Why can driver and gossip state disagree briefly?
- What does token awareness add to local-DC routing?
- Why is cross-DC round-robin not a generic failover strategy?
Review the answers
1. The load-balancing policy using topology/proximity configuration such as the local datacenter.
2. No. The driver chooses a coordinator; the consistency level governs replica responses required by Cassandra.
3. They observe health through different mechanisms; active connections/reconnection logic can differ from peer failure suspicion.
4. It can prioritize local replicas that own the target partition when routing metadata is available.
5. It adds WAN coordination and bypasses explicit application/RF/CL/capacity/recovery design.
Production judgment
Client correctness depends on more than successful TCP connectivity. Record Cassandra patch/native protocol negotiation, driver artifact/version, Java runtime, local datacenter, discovered topology, node distance, pool/in-flight metrics, prepared-statement/cache behavior, routing-key availability, page/fetch size, request/consistency/serial-consistency timeouts, retry policy, idempotence classification, speculative-execution policy, per-query execution profile, and p50/p95/p99 client latency. Correlate driver coordinator/attempt data with Cassandra tracing, server timeouts/failures, replica availability, RF/CL, partition size/cardinality, SSTable/compaction/tombstone state, disk/network/JVM pressure, repair state, and SAI/vector costs where those query paths are used.
Do not make a session per HTTP request, disable local-DC awareness casually, expose raw paging state as an authorization token, or enable broad retries/speculation to hide overload. Treat driver configuration as application production code: version it, test node/DC failures, test ambiguous timeouts with idempotent and non-idempotent mutations, measure extra traffic, and define rollback. Managed services may provide different endpoints, TLS/auth requirements, topology visibility, or restricted metrics while still using a Cassandra-compatible protocol; verify the service contract rather than assuming identical behavior. Lesson 5 completes the client path by classifying errors and making retry/speculation decisions only after statement idempotence and outcome ambiguity are understood.
Summary and next bridge
Load balancing is a topology-aware query-plan mechanism, not “pick any node.” The correct local DC, node health/distance and token routing determine coordinator preference. The final lesson adds safe retry, timeout and speculative-execution behavior.
Authoritative references
Driver behavior is language- and version-specific. Re-check the exact maintained driver and Cassandra patch before freezing production defaults or error-handling behavior.
- Apache Cassandra downloads / current 5.0 patch
- Docker Official Cassandra 5.0 image source
- Apache Cassandra Java Driver overview
- Java Driver pooling
- Java Driver prepared statements
- Java Driver load balancing / token awareness
- Java Driver paging
- Java Driver retries
- Java Driver idempotence
- Java Driver speculative execution
- Maven Central Java Driver 4.19.3