Chapter 21 · Drivers, Prepared Statements, Token Awareness, Paging, Retries, and Load Balancing

Native Protocol, Driver Sessions, Contact Points, Topology Discovery, and Connection Pools

Trace Cassandra client startup from native protocol and contact points through topology discovery, a long-lived CqlSession, multiplexed per-node pools, and request coordinator selection.

Intermediate → Advanced110–150 minutesSession/topology/pool labApache Cassandra 5.0.9 · Java 17 · Apache Java Driver 4.19.3 · RF=3 · LOCAL_QUORUM · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart's API team sees intermittent connection spikes because each web request opens a new Cassandra client. A second team hard-codes all three node addresses and assumes that is “load balancing.” The real fix starts by understanding what a driver session owns and how native-protocol discovery/pooling works.

01

Explain native-protocol multiplexing, contact points, topology discovery, control-plane metadata, sessions, and per-node connection pools.

02

Use one long-lived Apache Java Driver 4.19.3 CqlSession and inspect discovered nodes, DC/rack, driver state/distance, and open connections.

03

Distinguish seed/contact-point discovery from request leadership or a permanent coordinator.

04

Explain why Cassandra connections are multiplexed and why JDBC-style pool intuition can cause over-connection.

05

Observe server topology with cqlsh/nodetool and correlate it with driver metadata.

Chapter 21 lab baseline

The mandatory labs continue the established free/local AtlasMart cluster: Docker Official Image cassandra:5.0.9, Java 17 in that image, cluster atlasmart-course, Docker network atlasmart-cassandra, nodes atlasmart-cass-1..3, datacenter dc1, racks rack1..rack3, 16 virtual nodes per node, and named disposable data volumes. Keyspace atlasmart_driver uses NetworkTopologyStrategy with replication factor (RF) 3; ordinary reads/writes use LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no table default time-to-live (TTL), and Cassandra's normal gc_grace_seconds. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only on this isolated single-host learning network. Application examples use Apache Cassandra Java Driver 4.19.3 (org.apache.cassandra:java-driver-core:4.19.3) with Java 17 and one long-lived CqlSession. Driver documentation URLs still use the 4.19.0 documentation set, while the ASF release is 4.19.3. Exact coordinator choices, pool counts, traces, paging states, retry/speculation counts, and latency percentiles are learner-captured runtime evidence.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and request-path mental model

The native protocol is Cassandra's binary client/server protocol over TCP, normally on port 9042. A driver implements that protocol for an application language. A contact point is an initial address used to bootstrap discovery; it is not a permanent leader or a complete static node list. A session is the driver's long-lived view of a cluster and owns topology metadata, control-plane state, connection pools, policies, prepared-statement caches, and request execution. A connection pool is the driver's set of TCP connections to a node; unlike a JDBC-style blocking pool, each Cassandra connection multiplexes many in-flight requests using native-protocol stream identifiers.

A request is sent to a coordinator, the Cassandra node that handles that request. The driver chooses that coordinator using a load-balancing policy. Token-aware routing prefers replicas for the target partition when the statement supplies a keyspace and routing key/token. The partition key hashes to a token; replicas own token ranges according to the keyspace replication strategy. Prepared statements let the server parse CQL once and return metadata, including bind-variable types and partition-key variable positions, which enables the driver to compute routing keys for bound statements. Paging splits a large result into multiple protocol responses; the paging state is an opaque continuation token tied to the exact statement and values. Idempotent means repeating a request has the same final database effect as executing it once. Retry and speculative-execution policies use that property to avoid duplicating unsafe mutations.

1. Contact points bootstrap; the session learns the cluster

At startup the Java driver opens a connection to one or more configured contact points. Those addresses only get the session into the cluster. The driver then learns node identity, endpoint, datacenter, rack, schema and token metadata and subscribes to topology/state changes. A contact point is not a leader, and a request need not return to the same node that helped the session start.

The driver keeps a control connection for cluster events/metadata and regular pools for application traffic. Each application request still has a Cassandra coordinator, but that coordinator is selected per request by the load-balancing plan. A healthy application normally creates a small number of long-lived sessions—often one per cluster/security identity—not one session per query or HTTP request.

Concept Correct role Wrong inference
Contact point bootstrap discovery permanent coordinator/leader
CqlSession long-lived cluster/policy/pool owner cheap request-scoped object
Pool connections to one connected node one blocking connection per query
Native stream id multiplex concurrent requests on one TCP connection proof that more TCP sockets always improve throughput
Control connection metadata/events application data leader

2. Verify the server view first

bash / PowerShell-friendly Docker commands · verify the shared cluster
docker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -versiondocker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name,data_center,rack,release_version,native_protocol_version FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer,peer_port,data_center,rack,release_version FROM system.peers_v2;"
CQL · create the driver-focused AtlasMart query table
CREATE KEYSPACE IF NOT EXISTS atlasmart_driverWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_driver.orders_by_customer_month (  tenant_id text,  customer_id text,  order_month date,  order_time timestamp,  order_id uuid,  status text,  total decimal,  note text,  PRIMARY KEY ((tenant_id,customer_id,order_month),order_time,order_id)) WITH CLUSTERING ORDER BY (order_time DESC,order_id ASC)  AND compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:00:00Z',21000000-0000-0000-0000-000000000001,'PAID',129.90,'baseline');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-42','2026-09-01','2026-09-08T08:10:00Z',21000000-0000-0000-0000-000000000002,'PACKING',59.00,'fragile');INSERT INTO atlasmart_driver.orders_by_customer_month(tenant_id,customer_id,order_month,order_time,order_id,status,total,note)VALUES ('tenant-a','cust-99','2026-09-01','2026-09-08T08:20:00Z',21000000-0000-0000-0000-000000000003,'PAID',210.00,'priority');

The server output should show three nodes in dc1 across three racks and Cassandra 5.0.9. Exact endpoint addresses depend on the Docker network and must be captured locally.

3. Build one long-lived Java session and inspect topology/pools

XML · minimal pinned Maven project for the Apache Java Driver
<project xmlns="http://maven.apache.org/POM/4.0.0">  <modelVersion>4.0.0</modelVersion>  <groupId>academy.atlasmart</groupId><artifactId>cassandra-driver-lab</artifactId><version>1.0.0</version>  <properties><maven.compiler.release>17</maven.compiler.release></properties>  <dependencies>    <dependency>      <groupId>org.apache.cassandra</groupId>      <artifactId>java-driver-core</artifactId>      <version>4.19.3</version>    </dependency>    <dependency><groupId>org.slf4j</groupId><artifactId>slf4j-simple</artifactId><version>2.0.17</version></dependency>  </dependencies>  <build><plugins><plugin><groupId>org.codehaus.mojo</groupId><artifactId>exec-maven-plugin</artifactId><version>3.5.0</version></plugin></plugins></build></project>
Java · SessionProbe.java
package academy.atlasmart;import com.datastax.oss.driver.api.core.CqlSession;import com.datastax.oss.driver.api.core.metadata.Node;import java.net.InetSocketAddress;public class DriverLab {  public static void main(String[] args) {    try (CqlSession session = CqlSession.builder()        .addContactPoint(new InetSocketAddress("atlasmart-cass-1", 9042))        .withLocalDatacenter("dc1")        .build()) {      System.out.println("protocol=" + session.getContext().getProtocolVersion());      session.getMetadata().getNodes().values().forEach(n ->        System.out.printf("endpoint=%s state=%s dc=%s rack=%s distance=%s openConnections=%d%n",          n.getEndPoint(), n.getState(), n.getDatacenter(), n.getRack(), n.getDistance(), n.getOpenConnections()));      var rs = session.execute("SELECT release_version FROM system.local");      System.out.println("coordinator=" + rs.getExecutionInfo().getCoordinator().getEndPoint());    }  }}
bash / PowerShell · compile and run on the Cassandra Docker network
# From the Maven project directory. Docker Desktop/WSL or Linux/macOS shell:docker run --rm --network atlasmart-cassandra -v "$PWD:/work" -w /work maven:3.9.16-eclipse-temurin-17 \  mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab# PowerShell uses the same container and network; ${PWD} resolves to the current directory:# docker run --rm --network atlasmart-cassandra -v "${PWD}:/work" -w /work maven:3.9.16-eclipse-temurin-17 `#   mvn -q -DskipTests compile exec:java -Dexec.mainClass=academy.atlasmart.DriverLab

On the default driver policy in a single-DC cluster, the known nodes should be local/active and typically have one connection per node unless configuration or startup timing differs. The driver can multiplex many requests over each connection, so “only one open connection” is not evidence of serial execution.

4. Deliberately wrong: one session per web request

Java · anti-pattern versus corrected lifecycle
// WRONG: handshake, discovery, pools and threads are rebuilt per HTTP request.void handleRequestWrong() {  try (CqlSession s = CqlSession.builder()      .addContactPoint(new InetSocketAddress("atlasmart-cass-1",9042))      .withLocalDatacenter("dc1").build()) {    s.execute("SELECT now() FROM system.local");  }}// CORRECT: create once at service startup, inject/share it, close once at shutdown.final class CassandraGateway implements AutoCloseable {  private final CqlSession session;  CassandraGateway(CqlSession session) { this.session = session; }  public void close() { session.close(); }}
Why this fails under load

Session construction creates control/pool connections, metadata state, event loops and background resources. Repeating that lifecycle multiplies connection churn and can overwhelm both application and Cassandra before useful query work begins. Measure pool/in-flight metrics and session creation rate; do not “fix” churn by merely raising server connection limits.

5. Verification and reset

  • Server and driver both see three nodes in dc1 and their racks.
  • The session reports its negotiated protocol instead of hard-coding a protocol number into the lesson.
  • Node#getOpenConnections() is interpreted as multiplexed TCP connections, not request capacity by itself.
  • Repeated queries reuse one session and their coordinator may vary.
  • Closing the application session releases driver resources without stopping Cassandra.

Check your understanding

  1. Why are contact points not a complete long-term routing list?
  2. Why is one Cassandra TCP connection able to serve many requests?
  3. Does a discovered node necessarily have an active pool?
  4. Why is a session per HTTP request harmful?
  5. What proves which node coordinated one query?
Review the answers

1. They bootstrap discovery; the driver learns and continuously updates cluster topology/metadata after connecting.

2. The native protocol is asynchronous and multiplexes in-flight requests with stream identifiers.

3. No. Metadata can include down or ignored nodes; node distance/state and load-balancing policy determine connectivity.

4. It repeatedly builds discovery/control/pool/event-loop resources and creates connection churn.

5. Driver ExecutionInfo/request tracking or Cassandra tracing for that execution, not the original contact point.

Production judgment

Client correctness depends on more than successful TCP connectivity. Record Cassandra patch/native protocol negotiation, driver artifact/version, Java runtime, local datacenter, discovered topology, node distance, pool/in-flight metrics, prepared-statement/cache behavior, routing-key availability, page/fetch size, request/consistency/serial-consistency timeouts, retry policy, idempotence classification, speculative-execution policy, per-query execution profile, and p50/p95/p99 client latency. Correlate driver coordinator/attempt data with Cassandra tracing, server timeouts/failures, replica availability, RF/CL, partition size/cardinality, SSTable/compaction/tombstone state, disk/network/JVM pressure, repair state, and SAI/vector costs where those query paths are used.

Do not make a session per HTTP request, disable local-DC awareness casually, expose raw paging state as an authorization token, or enable broad retries/speculation to hide overload. Treat driver configuration as application production code: version it, test node/DC failures, test ambiguous timeouts with idempotent and non-idempotent mutations, measure extra traffic, and define rollback. Managed services may provide different endpoints, TLS/auth requirements, topology visibility, or restricted metrics while still using a Cassandra-compatible protocol; verify the service contract rather than assuming identical behavior. Lesson 2 uses prepared-statement metadata to supply routing keys so the default token-aware load-balancing policy can prefer replicas instead of choosing coordinators blindly.

Summary and next bridge

The session is the driver's long-lived cluster client, not a query object. It discovers topology and owns multiplexed per-node pools; contact points only bootstrap that process. Next, use prepared metadata to make routing partition-aware.

Authoritative references

Driver behavior is language- and version-specific. Re-check the exact maintained driver and Cassandra patch before freezing production defaults or error-handling behavior.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.