Chapter 21 · Drivers, Prepared Statements, Token Awareness, Paging, Retries, and Load Balancing
Native Protocol, Driver Sessions, Contact Points, Topology Discovery, and Connection Pools
Trace Cassandra client startup from native protocol and contact points through topology discovery, a long-lived CqlSession, multiplexed per-node pools, and request coordinator selection.
Learning outcomes
AtlasMart's API team sees intermittent connection spikes because each web request opens a new Cassandra client. A second team hard-codes all three node addresses and assumes that is “load balancing.” The real fix starts by understanding what a driver session owns and how native-protocol discovery/pooling works.
Explain native-protocol multiplexing, contact points, topology discovery, control-plane metadata, sessions, and per-node connection pools.
Use one long-lived Apache Java Driver 4.19.3 CqlSession and inspect discovered nodes, DC/rack, driver state/distance, and open connections.
Distinguish seed/contact-point discovery from request leadership or a permanent coordinator.
Explain why Cassandra connections are multiplexed and why JDBC-style pool intuition can cause over-connection.
Observe server topology with cqlsh/nodetool and correlate it with driver metadata.
The mandatory labs continue the established free/local
AtlasMart cluster: Docker Official Image
cassandra:5.0.9, Java 17 in that image, cluster
atlasmart-course, Docker network
atlasmart-cassandra, nodes
atlasmart-cass-1..3, datacenter dc1,
racks rack1..rack3, 16 virtual nodes per node,
and named disposable data volumes. Keyspace
atlasmart_driver uses
NetworkTopologyStrategy with replication factor
(RF) 3; ordinary reads/writes use LOCAL_QUORUM.
New tables explicitly use UnifiedCompactionStrategy (UCS), no
table default time-to-live (TTL), and Cassandra's normal
gc_grace_seconds. Authentication,
client/internode Transport Layer Security (TLS), and remote
Java Management Extensions (JMX) are disabled only on this
isolated single-host learning network. Application examples
use Apache Cassandra Java Driver
4.19.3
(org.apache.cassandra:java-driver-core:4.19.3)
with Java 17 and one long-lived CqlSession.
Driver documentation URLs still use the 4.19.0 documentation
set, while the ASF release is 4.19.3. Exact coordinator
choices, pool counts, traces, paging states, retry/speculation
counts, and latency percentiles are learner-captured runtime
evidence.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and request-path mental model
The native protocol is Cassandra's binary client/server protocol over TCP, normally on port 9042. A driver implements that protocol for an application language. A contact point is an initial address used to bootstrap discovery; it is not a permanent leader or a complete static node list. A session is the driver's long-lived view of a cluster and owns topology metadata, control-plane state, connection pools, policies, prepared-statement caches, and request execution. A connection pool is the driver's set of TCP connections to a node; unlike a JDBC-style blocking pool, each Cassandra connection multiplexes many in-flight requests using native-protocol stream identifiers.
A request is sent to a coordinator, the Cassandra node that handles that request. The driver chooses that coordinator using a load-balancing policy. Token-aware routing prefers replicas for the target partition when the statement supplies a keyspace and routing key/token. The partition key hashes to a token; replicas own token ranges according to the keyspace replication strategy. Prepared statements let the server parse CQL once and return metadata, including bind-variable types and partition-key variable positions, which enables the driver to compute routing keys for bound statements. Paging splits a large result into multiple protocol responses; the paging state is an opaque continuation token tied to the exact statement and values. Idempotent means repeating a request has the same final database effect as executing it once. Retry and speculative-execution policies use that property to avoid duplicating unsafe mutations.
1. Contact points bootstrap; the session learns the cluster
At startup the Java driver opens a connection to one or more configured contact points. Those addresses only get the session into the cluster. The driver then learns node identity, endpoint, datacenter, rack, schema and token metadata and subscribes to topology/state changes. A contact point is not a leader, and a request need not return to the same node that helped the session start.
The driver keeps a control connection for cluster events/metadata and regular pools for application traffic. Each application request still has a Cassandra coordinator, but that coordinator is selected per request by the load-balancing plan. A healthy application normally creates a small number of long-lived sessions—often one per cluster/security identity—not one session per query or HTTP request.
| Concept | Correct role | Wrong inference |
|---|---|---|
| Contact point | bootstrap discovery | permanent coordinator/leader |
| CqlSession | long-lived cluster/policy/pool owner | cheap request-scoped object |
| Pool | connections to one connected node | one blocking connection per query |
| Native stream id | multiplex concurrent requests on one TCP connection | proof that more TCP sockets always improve throughput |
| Control connection | metadata/events | application data leader |
2. Verify the server view first
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name,data_center,rack,release_version,native_protocol_version FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer,peer_port,data_center,rack,release_version FROM system.peers_v2;"
CREATE KEYSPACE IF NOT EXISTS atlasmart_driverWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_driver.orders_by_customer_month ( tenant_id text, customer_id text, order_month date, order_time timestamp, order_id uuid, status text, total decimal, note text, PRIMARY KEY ((tenant_id,customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC) AND compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:00:00Z',21000000-0000-0000-0000-000000000001,'PAID',129.90,'baseline');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:10:00Z',21000000-0000-0000-0000-000000000002,'PACKING',59.00,'fragile');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-99','2026-09-01','2026-09-08T08:20:00Z',21000000-0000-0000-0000-000000000003,'PAID',210.00,'priority');
The server output should show three nodes in
dc1 across three racks and Cassandra 5.0.9. Exact
endpoint addresses depend on the Docker network and must be
captured locally.
3. Build one long-lived Java session and inspect topology/pools
<project xmlns="http://maven.apache.org/POM/4.0.0"> <modelVersion>4.0.0</modelVersion> <groupId>academy.atlasmart</groupId><artifactId>cassandra-driver-lab</artifactId><version>1.0.0</version> <properties><maven.compiler.release>17</maven.compiler.release></properties> <dependencies> <dependency> <groupId>org.apache.cassandra</groupId> <artifactId>java-driver-core</artifactId> <version>4.19.3</version> </dependency> <dependency><groupId>org.slf4j</groupId><artifactId>slf4j-simple</artifactId><version>2.0.17</version></dependency> </dependencies> <build><plugins><plugin><groupId>org.codehaus.mojo</groupId><artifactId>exec-maven-plugin</artifactId><version>3.5.0</version></plugin></plugins></build></project>
package academy.atlasmart;import com.datastax.oss.driver.api.core.CqlSession;import com.datastax.oss.driver.api.core.metadata.Node;import java.net.InetSocketAddress;public class DriverLab { public static void main(String[] args) { try (CqlSession session = CqlSession.builder() .addContactPoint(new InetSocketAddress("atlasmart-cass-1", 9042)) .withLocalDatacenter("dc1") .build()) { System.out.println("protocol=" + session.getContext().getProtocolVersion()); session.getMetadata().getNodes().values().forEach(n -> System.out.printf("endpoint=%s state=%s dc=%s rack=%s distance=%s openConnections=%d%n", n.getEndPoint(), n.getState(), n.getDatacenter(), n.getRack(), n.getDistance(), n.getOpenConnections())); var rs = session.execute("SELECT release_version FROM system.local"); System.out.println("coordinator=" + rs.getExecutionInfo().getCoordinator().getEndPoint()); } }}
# From the Maven project directory. Docker Desktop/WSL or Linux/macOS shell:docker run --rm --network atlasmart-cassandra -v "$PWD:/work" -w /work maven:3.9.16-eclipse-temurin-17 \ mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab# PowerShell uses the same container and network; ${PWD} resolves to the current directory:# docker run --rm --network atlasmart-cassandra -v "${PWD}:/work" -w /work maven:3.9.16-eclipse-temurin-17 `# mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab
On the default driver policy in a single-DC cluster, the known nodes should be local/active and typically have one connection per node unless configuration or startup timing differs. The driver can multiplex many requests over each connection, so “only one open connection” is not evidence of serial execution.
4. Deliberately wrong: one session per web request
// WRONG: handshake, discovery, pools and threads are rebuilt per HTTP request.void handleRequestWrong() { try (CqlSession s = CqlSession.builder() .addContactPoint(new InetSocketAddress("atlasmart-cass-1",9042)) .withLocalDatacenter("dc1").build()) { s.execute("SELECT now() FROM system.local"); }}// CORRECT: create once at service startup, inject/share it, close once at shutdown.final class CassandraGateway implements AutoCloseable { private final CqlSession session; CassandraGateway(CqlSession session) { this.session = session; } public void close() { session.close(); }}
Session construction creates control/pool connections, metadata state, event loops and background resources. Repeating that lifecycle multiplies connection churn and can overwhelm both application and Cassandra before useful query work begins. Measure pool/in-flight metrics and session creation rate; do not “fix” churn by merely raising server connection limits.
5. Verification and reset
-
Server and driver both see three nodes in
dc1and their racks. - The session reports its negotiated protocol instead of hard-coding a protocol number into the lesson.
-
Node#getOpenConnections()is interpreted as multiplexed TCP connections, not request capacity by itself. - Repeated queries reuse one session and their coordinator may vary.
- Closing the application session releases driver resources without stopping Cassandra.
Check your understanding
- Why are contact points not a complete long-term routing list?
- Why is one Cassandra TCP connection able to serve many requests?
- Does a discovered node necessarily have an active pool?
- Why is a session per HTTP request harmful?
- What proves which node coordinated one query?
Review the answers
1. They bootstrap discovery; the driver learns and continuously updates cluster topology/metadata after connecting.
2. The native protocol is asynchronous and multiplexes in-flight requests with stream identifiers.
3. No. Metadata can include down or ignored nodes; node distance/state and load-balancing policy determine connectivity.
4. It repeatedly builds discovery/control/pool/event-loop resources and creates connection churn.
5. Driver ExecutionInfo/request tracking or Cassandra tracing for that execution, not the original contact point.
Production judgment
Client correctness depends on more than successful TCP connectivity. Record Cassandra patch/native protocol negotiation, driver artifact/version, Java runtime, local datacenter, discovered topology, node distance, pool/in-flight metrics, prepared-statement/cache behavior, routing-key availability, page/fetch size, request/consistency/serial-consistency timeouts, retry policy, idempotence classification, speculative-execution policy, per-query execution profile, and p50/p95/p99 client latency. Correlate driver coordinator/attempt data with Cassandra tracing, server timeouts/failures, replica availability, RF/CL, partition size/cardinality, SSTable/compaction/tombstone state, disk/network/JVM pressure, repair state, and SAI/vector costs where those query paths are used.
Do not make a session per HTTP request, disable local-DC awareness casually, expose raw paging state as an authorization token, or enable broad retries/speculation to hide overload. Treat driver configuration as application production code: version it, test node/DC failures, test ambiguous timeouts with idempotent and non-idempotent mutations, measure extra traffic, and define rollback. Managed services may provide different endpoints, TLS/auth requirements, topology visibility, or restricted metrics while still using a Cassandra-compatible protocol; verify the service contract rather than assuming identical behavior. Lesson 2 uses prepared-statement metadata to supply routing keys so the default token-aware load-balancing policy can prefer replicas instead of choosing coordinators blindly.
Summary and next bridge
The session is the driver's long-lived cluster client, not a query object. It discovers topology and owns multiplexed per-node pools; contact points only bootstrap that process. Next, use prepared metadata to make routing partition-aware.
Authoritative references
Driver behavior is language- and version-specific. Re-check the exact maintained driver and Cassandra patch before freezing production defaults or error-handling behavior.
- Apache Cassandra downloads / current 5.0 patch
- Docker Official Cassandra 5.0 image source
- Apache Cassandra Java Driver overview
- Java Driver pooling
- Java Driver prepared statements
- Java Driver load balancing / token awareness
- Java Driver paging
- Java Driver retries
- Java Driver idempotence
- Java Driver speculative execution
- Maven Central Java Driver 4.19.3