Chapter 01 · Cassandra Foundations, Version 5.0, Architecture, and Lab Setup
Peer-to-Peer Architecture vs Primary / Replica Systems: Coordination Without a Permanent Leader
Separate four often-confused roles—peer, coordinator, replica, and seed—and prove why Cassandra can coordinate requests through different healthy nodes without assigning a permanent primary.
Learning outcomes
AtlasMart's architecture review contains a dangerous sentence: “send all Cassandra traffic to the seed server because it is the cluster leader.” Cassandra does use seeds, and every request does have a coordinator, but neither concept means permanent leader. This lesson replaces that shortcut with a request-by-request peer model and makes the distinction observable.
Contrast Cassandra peer-to-peer request handling with a primary/replica topology without claiming that every database operation is leaderless internally.
Explain coordinator, replica, seed, contact point, token owner, gossip, and failure detector as separate responsibilities.
Demonstrate that different healthy peers can coordinate client requests while replicas are selected from partition ownership and RF.
Explain what seeds do during discovery and why treating seeds as traffic leaders creates unnecessary concentration and fragile client configuration.
Recognize the limits of a single-host Docker cluster as evidence for real rack, zone, or datacenter resilience.
The current generally available Apache Cassandra line is 5.0
and the current patch verified for this chapter is
5.0.9. Reproducible container examples pin
cassandra:5.0.9 instead of using a moving
latest tag. Cassandra 5.0 binary releases support
documented Java 11/17 runtime paths; the lesson always asks
you to record the Java runtime actually present in your
installation or image rather than inferring it from the
Cassandra version.
The generation environment used to build this chapter does not
contain Docker, Cassandra, cqlsh, or
nodetool. Commands were checked against current
official documentation but were not executed here. Expected
output is therefore described by shape and invariant, never
presented as captured output. The lab uses an isolated Docker
network and course-specific containers/volumes. Do not point
any cleanup, failure, configuration, or schema command at
unrelated or production Cassandra data.
1. Peer-to-peer means symmetric server roles for ordinary client access
In a classic primary/replica database, a designated primary commonly accepts writes and secondaries follow a replicated log. Cassandra does not assign one permanent cluster-wide primary for ordinary CQL reads and writes. Healthy nodes participate as peers. A client driver can connect through several contact points, discover the topology, and send requests to appropriate nodes according to its routing policy.
The contacted node becomes the coordinator for that request. It is responsible for determining the partition token, locating replicas, forwarding work, waiting for enough responses for the requested CL, and returning a result or error. The coordinator may itself be a replica, but the two roles are conceptually independent. If the next request is routed to a different node, that peer becomes coordinator.
This model removes a permanent request leader, but it does not mean “no coordination.” Some Cassandra features—schema propagation, topology changes, lightweight transactions, repair, streaming—have their own protocols and constraints. Peer-to-peer is a topology and request-routing model, not a promise that every distributed operation is a coordination-free commutative write.
2. Seeds are discovery helpers, not leaders
A seed is an initial contact used by Cassandra nodes during startup/discovery. Seeds help a joining/restarting node learn cluster membership. They do not own special token ranges merely because they are seeds, they are not required coordinators for application traffic, and they do not become a primary. Overloading seeds with the word “master” leads to two practical mistakes: teams route all application connections through them, and operators believe losing a seed implies immediate loss of an already formed cluster.
Gossip is Cassandra's distributed mechanism for sharing membership and state information among peers. A distributed failure detector uses heartbeat/state timing to suspect endpoints. Chapter 02 covers those mechanisms precisely; for now, the key boundary is that seed configuration assists discovery, while request coordination and replica ownership are separate.
| Role | Lifetime | Selection basis | Common misconception |
|---|---|---|---|
| Seed | Configuration/discovery role | Operator-provided seed list | “The seed is the leader.” |
| Coordinator | One request | The node the client/driver sends that request to | “Only special coordinators can accept writes.” |
| Replica | Per partition/token placement | Partition token + keyspace replication strategy | “Replica means passive secondary.” |
| Contact point | Client bootstrap/discovery | Driver/application configuration | “The driver can only use those nodes forever.” |
3. Two requests can have different coordinators
Use the same AtlasMart naming convention but start three nodes so the peer relationship is visible. This is still one physical Docker host, so it is a topology simulation—not a zone-failure test. Start node 1 first and wait for it to become queryable before adding peers. Nodes 2 and 3 use node 1 as a seed to discover the cluster.
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until: docker exec atlasmart-cass-1 cqlsh -e "SELECT now() FROM system.local;" succeeds.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9
Wait for all three nodes to become Up/Normal. Then compare the local metadata from two different endpoints:
docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name, host_id, data_center, rack, listen_address FROM system.local;"docker exec atlasmart-cass-2 cqlsh atlasmart-cass-2 9042 -e "SELECT cluster_name, host_id, data_center, rack, listen_address FROM system.local;"docker exec atlasmart-cass-3 cqlsh atlasmart-cass-3 9042 -e "SELECT cluster_name, host_id, data_center, rack, listen_address FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer, data_center, rack, release_version FROM system.peers_v2;"
Each system.local query is answered through a
different contacted node and therefore reports a different
node-local host_id/address/rack while retaining the
same cluster name. system.peers_v2 should list the
other peers known to the queried node after membership
converges. Exact IPs and rows are environment-dependent. This
evidence demonstrates peer identities; it does not yet prove
that a business partition has RF=3.
4. Add replicated AtlasMart data and vary the contact point
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_peer_lab WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_peer_lab.order_status_by_order (order_id text, changed_at timestamp, status text, PRIMARY KEY ((order_id), changed_at)) WITH CLUSTERING ORDER BY (changed_at DESC);"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; INSERT INTO atlasmart_peer_lab.order_status_by_order (order_id,changed_at,status) VALUES ('o-9001','2026-09-07T09:00:00Z','PLACED');"docker exec atlasmart-cass-2 cqlsh atlasmart-cass-2 9042 -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_peer_lab.order_status_by_order WHERE order_id='o-9001';"docker exec atlasmart-cass-3 cqlsh atlasmart-cass-3 9042 -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_peer_lab.order_status_by_order WHERE order_id='o-9001';"
The important observation is not that “node 2 contains the row.” At RF=3 in a three-node single-DC lab, all three nodes are intended replicas for the partition, but each CQL session also contacts a particular coordinator. A driver in production should usually use topology-aware routing and multiple contact points rather than hard-code all traffic to one seed.
cqlsh shows the host it connected to; tracing or
driver request metrics can provide deeper coordinator/replica
evidence. Chapter 23 covers driver routing, token awareness,
local-DC policy, retries, idempotency, paging, and
observability. Do not infer production driver behavior from
manual shell sessions alone.
5. Wrong pattern: route everything to the seed
Suppose AtlasMart configures a load balancer that sends every
CQL request only to atlasmart-cass-1 because that
node appears in CASSANDRA_SEEDS. The cluster can
still replicate data, but the application has introduced a
coordinator hot spot and a single client ingress dependency that
Cassandra itself did not require. The seed setting solved node
discovery, not client traffic distribution.
The repair is to configure clients with appropriate contact points and a supported topology-aware driver policy, then measure per-node request distribution and failure behavior. Avoid turning seed nodes into permanent application gateways unless an intentional external architecture requires that—and then treat the gateway as its own availability/capacity component.
Safe failure observation
docker stop atlasmart-cass-1# From a different running peer, wait until membership/failure detection reflects node 1 as down.docker exec atlasmart-cass-2 nodetool statusdocker exec atlasmart-cass-2 cqlsh atlasmart-cass-2 9042 -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_peer_lab.order_status_by_order WHERE order_id='o-9001';"docker start atlasmart-cass-1
With RF=3 and two healthy replicas, a QUORUM read can normally still be satisfied after failure detection converges. The exact time to classify the node down and the exact error if you query too early are runtime-dependent. This is not proof of rack or datacenter resilience because all containers share one host.
Cleanup
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
Production judgment
Peer-to-peer removes a permanent request primary but increases the importance of correct client routing, replica placement, failure detection, repair, and observability. Capacity planning must consider both coordinator work and replica work. A driver timeout can also be ambiguous: an operation may have reached replicas even if the client did not receive a response. Later application lessons will tie retries to idempotency rather than assuming every timeout is safe to repeat.
The next lesson freezes the version/configuration baseline so that architecture experiments are reproducible instead of depending on whatever Cassandra, Java, or configuration a machine happens to contain.
Check your understanding
- What is the difference between a seed and a coordinator?
- If a node is a seed, does it own more token ranges automatically?
- Why can two cqlsh sessions use different coordinators?
- Why is a three-container cluster on one laptop not proof of availability-zone tolerance?
- What new bottleneck appears if all application traffic is forced through one seed?
Review the answers
1. A seed assists node discovery during startup. A coordinator is the peer handling one client request and routing it to the relevant replicas.
2. No. Seed status does not grant special partition ownership or leader status.
3. Each session connects to a node; the contacted node coordinates requests from that session. Different reachable peers can therefore coordinate different requests.
4. All nodes share the same physical host, power, storage subsystem, Docker daemon, and often network path. It simulates logical node failure, not independent failure domains.
5. That node becomes an avoidable coordinator/ingress concentration and the application may lose access when that one endpoint fails even though other Cassandra peers remain healthy.
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Apache Cassandra 5.0 Baseline, Java Requirements, Configuration Layout, and Release Awareness.
Authoritative references
- Apache Cassandra 5.0 documentation — Current official documentation entry point for the 5.0 line.
- Apache Cassandra downloads — Official release page used to verify the current 5.0 patch.
- Cassandra architecture overview — Official architecture and wide-column/distributed design framing.
- Cassandra quickstart — Official Docker-oriented learning workflow and isolated-network approach.
- Cassandra configuration reference — Official cassandra.yaml semantics, including native transport and security-related settings.
- CQL querying and cqlsh — Official CQL/cqlsh connection and system.local examples.
- Java support for Cassandra 5.0 — Official Java build/runtime compatibility notes for Cassandra 5.0.
- Docker Official Image: Cassandra — Container-image usage and tag information for the Docker Official Image.
- Dynamo-style architecture and gossip — Official background for distributed membership, gossip and Dynamo-influenced design.