Chapter 02 · Clusters, Datacenters, Racks, Snitches, Gossip, and Failure Detection
Cluster Topology: Nodes, Datacenters, Racks, Endpoints, and Failure Domains
Turn topology vocabulary into an evidence-backed map of where Cassandra thinks peers live—and what that map can and cannot guarantee.
Learning outcomes
AtlasMart is moving from a single-node proof of concept to a cluster that must survive host and zone failures without creating a mythical “leader node.” Before replication-factor math or consistency levels can mean anything, the team needs a topology vocabulary that is tied to observable metadata and to real failure domains.
Distinguish node, endpoint, cluster, datacenter, rack and physical failure domain, and explain which are labels versus infrastructure facts.
Explain why every request has a coordinator while no node is a permanent request leader.
Inspect local and peer DC/rack metadata with system tables and nodetool.
Explain how a bad rack map can correlate replicas even when the cluster has multiple nodes.
Design a reversible three-node topology lab that later lessons can use for gossip and failure detection.
Apache Cassandra 5.0.9 is the current GA 5.0 patch on the
official download page. These labs pin
cassandra:5.0.9. The chapter uses one logical
datacenter, dc1, and three logical racks,
rack1–rack3, on the isolated Docker
network atlasmart-cassandra. Topology labels are
learning metadata; they do not create real
host/rack/availability-zone isolation by themselves.
The generation environment does not contain Docker or Cassandra, so commands were documentation- and syntax-checked rather than executed here. Record your actual addresses, host IDs, tokens, gossip generations, phi values, startup times, and driver events. Never treat the example output shapes as captured measurements. All failure injection is confined to disposable course containers; no host firewall, clock manipulation, or production endpoint is required.
1. Topology is a map Cassandra uses; failure domains are reality
A Cassandra node is one server process. Its endpoint is the network identity other nodes use to communicate with it. A cluster is the peer group sharing cluster identity and schema. A datacenter (DC) is Cassandra's largest placement/locality label, while a rack is a smaller failure-domain label within a DC. Those labels are operational metadata. They become useful only when they accurately mirror physical or cloud failure boundaries such as availability zones, racks, power domains, or other correlated-failure groups.
If three containers all run on one laptop but are labeled rack1, rack2 and rack3, Cassandra can demonstrate rack-aware placement logic, but the laptop remains one physical failure domain. Conversely, if three production nodes really sit in three zones but are all labeled rack1, Cassandra cannot infer that separation. The application has resilient infrastructure but a misleading placement map.
| Concept | What Cassandra sees | What operators must guarantee |
|---|---|---|
| Datacenter | Case-sensitive topology label used by replication/routing policy | Meaningful locality/failure boundary for the deployment |
| Rack | Sub-DC label used to spread replicas | Maps to a correlated failure domain, often an AZ |
| Endpoint | Address/port identity for a peer | Reachable, correctly advertised network path |
| Seed | Discovery contact during startup | Available enough for bootstrap; not a leader |
2. Coordinator, replica and topology answer different questions
The coordinator is whichever node handles a specific client request. A replica is a node that owns a copy of the requested partition under the keyspace's replication strategy. A seed helps a joining/restarting node find the cluster. None of those roles is a permanent primary. A client can connect to node 2, which coordinates a read whose replicas are nodes 1 and 3; the next request can use node 1 as coordinator.
Topology metadata does not itself choose every replica. The
endpoint snitch tells Cassandra how nodes map to DCs/racks; the
keyspace replication strategy—normally
NetworkTopologyStrategy in production—uses that
topology plus token ownership to choose replicas. Driver
locality is another layer: a driver can prefer a local DC, but
that does not redefine server-side replica ownership.
“rack1,” “dc1,” seed membership, coordinator role, replica ownership, gossip state and driver host state are related but independent observations. Debugging becomes much faster when each piece of evidence is attached to the mechanism that owns it.
3. Build the observable three-rack course topology
For a smooth three-node lab, budget roughly four logical CPUs, several gigabytes of RAM available to Docker, and a few gigabytes of free disk; exact memory pressure varies by host/runtime. If your machine cannot sustain three Cassandra processes, read the commands and use the single-node Chapter 01 lab for metadata inspection, but do not claim the smaller simulation proves rack-failure tolerance.
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait for all three nodes to become Up/Normal, then verify from two peers.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool status
When all nodes are Up/Normal, verify the local and peer views:
docker exec atlasmart-cass-1 cqlsh -e "SELECT cluster_name, listen_address, data_center, rack, host_id, release_version FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer, peer_port, data_center, rack, host_id, release_version FROM system.peers_v2;"docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool status
Expected invariants, not fabricated values: one local row
identifies node 1 as dc1/rack1; two peer rows
identify the other hosts as rack2 and
rack3; all three nodes converge on the same
membership set. Exact IPs, host IDs, token counts and load
figures are environment-specific.
4. Deliberately wrong topology: three zones labeled as one rack
Imagine AtlasMart has hosts in zones A, B and C but
configuration says dc1/rack1 on all three. The node
count looks healthy and RF=3 can still create three copies, yet
the snitch has no logical evidence that those copies occupy
distinct failure domains. If a future cluster has many nodes per
zone, replica placement can repeatedly choose nodes from the
same zone because the mapping erased the distinction.
The safe repair is not “edit rack labels on a live populated cluster until nodetool looks nicer.” Cassandra's production guidance warns that rack/snitch changes after provisioning can be unsupported and can cause data-loss risks. Correct topology should be designed before loading data, or migrated with an explicit supported topology-change plan, streaming/repair validation and rollback.
| Wrong belief | Observable correction |
|---|---|
| “Three racks in cassandra-rackdc.properties means three real AZs.” | Prove host placement independently in the infrastructure layer. |
| “Seeds are leaders.” | Send requests through different peers; coordinator role moves per request. |
| “DOWN means dead forever.” | Later pause/unpause a node and observe state recover. |
5. Production judgment and verification
Topology is part of correctness because RF, CL, local-DC routing, repair, maintenance and disaster planning are interpreted through it. A rack map that lies about infrastructure can silently remove the correlated-failure protection operators thought they bought. Treat DC/rack names as a contract between Cassandra configuration, deployment automation and physical/cloud placement.
For AtlasMart, keep dc1 stable for this single-DC
learning chapter and use rack labels that model three
independent zones. Later multi-DC lessons can add a second DC
explicitly. Do not prematurely use multi-DC labels as a
substitute for understanding replication and consistency.
Verification checklist
- All three nodes are Up/Normal before failure tests.
-
system.localandsystem.peers_v2show the intended DC/rack map. - You can explain why rack labels do not create physical isolation.
- You can distinguish seed, coordinator and replica roles.
- No Cassandra or JMX port was broadly published to the host by the chapter commands.
Check your understanding
- Why is a Cassandra rack label not proof of an availability zone?
- Is a seed node the coordinator for every request?
- Which component supplies DC/rack topology and which component chooses replicas?
- Why should you avoid casually changing rack labels on a populated cluster?
- What does a three-container laptop lab prove?
Review the answers
1. Because it is metadata supplied to Cassandra; the infrastructure must independently place nodes in the intended failure domains.
2. No. Seeds help discovery. Any suitable contacted peer can coordinate a request.
3. The snitch supplies topology; the keyspace replication strategy uses topology and token ownership to choose replicas.
4. Changing topology can alter placement assumptions and may require supported migration, streaming and repair procedures; editing labels is not a cosmetic operation.
5. It proves Cassandra topology metadata and placement logic under a simulated map, not independent physical AZ failure tolerance.
Cleanup
docker unpause atlasmart-cass-1 2>/dev/null || truedocker unpause atlasmart-cass-2 2>/dev/null || truedocker unpause atlasmart-cass-3 2>/dev/null || truedocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
docker unpause atlasmart-cass-1 2>$nulldocker unpause atlasmart-cass-2 2>$nulldocker unpause atlasmart-cass-3 2>$nulldocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Endpoint Snitches and Replication Placement: Mapping Logical Topology to Physical Infrastructure.
Authoritative references
- Apache Cassandra 5.0 documentation — Official documentation entry point for the current 5.0 line.
- Apache Cassandra downloads — Official release page used to verify Cassandra 5.0.9 as the current GA patch.
- Snitch — Official explanation of topology/proximity and rack-aware replica placement.
- cassandra-rackdc.properties — Official DC/rack configuration for GossipingPropertyFileSnitch.
- Dynamo architecture and replication — Official NetworkTopologyStrategy and replica-placement semantics.
- nodetool gossipinfo — Official gossip-state inspection command.
- nodetool failuredetector — Official failure-detector inspection command.
- Java Driver 4.19 node metadata — Apache Java Driver documentation for driver-side node state and topology metadata.
- Production recommendations — Official guidance for NetworkTopologyStrategy and rack/snitch planning.