Chapter 02 · Clusters, Datacenters, Racks, Snitches, Gossip, and Failure Detection
Design Rack / Zone Placement That Survives Host, Rack, and Availability-Zone Failure
Convert topology mechanics into a fault-domain design: place replicas intentionally, fail one logical rack, verify what survives, and document what the laptop simulation cannot prove.
Learning outcomes
AtlasMart can now observe topology, placement, gossip and failure suspicion. The chapter closes by turning those mechanisms into a rack/zone placement design. The goal is not “three racks because Cassandra tutorials use three,” but an explicit statement of what fails together, where replicas land, which CLs remain available, and what the single-host lab cannot prove.
Design a DC/rack map from real host/rack/AZ failure domains instead of arbitrary labels.
Use NetworkTopologyStrategy and RF to reason about replica survival under host and rack failure.
Validate replica endpoints and client-visible QUORUM behavior during a simulated rack-node outage.
Distinguish logical rack survival in the lab from real availability-zone isolation in production.
Produce a topology acceptance checklist that can be carried into token/replication chapters.
Apache Cassandra 5.0.9 is the current GA 5.0 patch on the
official download page. These labs pin
cassandra:5.0.9. The chapter uses one logical
datacenter, dc1, and three logical racks,
rack1–rack3, on the isolated Docker
network atlasmart-cassandra. Topology labels are
learning metadata; they do not create real
host/rack/availability-zone isolation by themselves.
The generation environment does not contain Docker or Cassandra, so commands were documentation- and syntax-checked rather than executed here. Record your actual addresses, host IDs, tokens, gossip generations, phi values, startup times, and driver events. Never treat the example output shapes as captured measurements. All failure injection is confined to disposable course containers; no host firewall, clock manipulation, or production endpoint is required.
1. Start with correlated-failure questions
Before choosing rack names, ask what can fail together: one process, one VM, one hypervisor, one physical rack, one top-of-rack switch, one availability zone, one region, one shared storage plane, or one Kubernetes node pool. Cassandra's rack label should represent the smaller placement domain across which replicas should be spread inside a DC. In many cloud deployments, operators map one AZ to one Cassandra rack.
A useful topology design document contains both the Cassandra labels and the infrastructure proof behind them. “rack2” alone is weak. “rack2 = eu-west-zone-b; anti-affinity keeps one Cassandra node per host; independent power/network domain verified by platform inventory” is an operational statement.
| Failure | Desired mapping | What Cassandra can help with |
|---|---|---|
| Single host | Replicas on different hosts | Replica placement if host/rack mapping is correct |
| One AZ/rack | Replica copies spread to other racks | NTS rack-aware placement |
| Entire region/DC | Replicas in another DC + client failover plan | Multi-DC NTS, but application routing/repair still required |
2. Validate the three-rack RF=3 design
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait for all three nodes to become Up/Normal, then verify from two peers.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool status
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_resilience WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_resilience.shipment_by_id (shipment_id text PRIMARY KEY, order_id text, state text, updated_at timestamp);"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; INSERT INTO atlasmart_resilience.shipment_by_id (shipment_id,order_id,state,updated_at) VALUES ('s-3001','o-2001','IN_TRANSIT','2026-09-07T12:30:00Z');"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_resilience shipment_by_id s-3001docker exec atlasmart-cass-1 nodetool status
In this three-node topology, RF=3 necessarily places the partition on all three nodes. Because the nodes claim three distinct racks, the logical placement also spans all three racks. On a laptop, that proves rack-aware metadata/placement—not physical fault isolation.
3. Simulate one rack/node failure and verify service behavior
Each lab rack contains one node, so pausing node 2 models “the only node in rack2 is unavailable.” In production, a rack can contain many nodes; a rack failure would remove every node in that failure domain at once. Keep that scale distinction explicit.
docker pause atlasmart-cass-2# Wait for node 1 to mark the peer down.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool failuredetector# RF=3, QUORUM needs two replicas. rack1 + rack3 remain responsive.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_resilience.shipment_by_id WHERE shipment_id='s-3001';"docker unpause atlasmart-cass-2# Wait for recovery, then verify all peers are Up/Normal again.docker exec atlasmart-cass-1 nodetool status
If the QUORUM read succeeds after node 2 is down, the evidence demonstrates this particular RF/CL/topology combination can serve the tested read with one replica unavailable. It does not prove all writes, all tables, all workloads, or all failure modes behave identically.
4. Design anti-patterns and their consequences
Anti-pattern 1: every node in one rack. Rack-aware placement has no smaller failure-domain distinction to use. Anti-pattern 2: fake racks that do not match infrastructure. Cassandra spreads replicas across labels while the underlying systems can still fail together. Anti-pattern 3: one tiny new rack. Rack-aware placement can load that rack disproportionately. Anti-pattern 4: treat DOWN as permission to replace/remove. Failure suspicion is reversible and not a lifecycle decision.
Another trap is assuming “three AZs, RF=3” automatically makes the application highly available. CL, driver local-DC configuration, timeouts/retries, repair strategy, capacity headroom during failure, network dependencies, DNS/service discovery and deployment automation all contribute. A cluster can have perfect replica placement and still fail because every client uses one unreachable contact point or because surviving nodes have no performance headroom.
When one rack disappears, the remaining nodes inherit more coordinator/replica work. Availability testing must include latency and saturation, not only whether a query eventually returns.
5. AtlasMart topology acceptance checklist and chapter bridge
For a future production design, AtlasMart should version-control a topology inventory containing node identity, DC/rack, actual zone/host, advertised endpoints, seed set, keyspace RFs, application CLs, driver local DC, maintenance ownership and failure-drill evidence. Changes to those fields should be reviewed like data-placement changes.
Acceptance checklist
- DC/rack labels map to documented real failure domains.
- Production keyspaces use NetworkTopologyStrategy with explicit per-DC RF.
- Replica endpoints for representative partition keys span intended racks.
- One-rack failure tests verify both correctness and latency/capacity headroom.
- Driver contact points and local-DC policy do not create a hidden single point of failure.
- DOWN detection is separated from replace/remove/decommission decisions.
- Recovery includes return-to-service verification and later repair/convergence planning.
Chapter 03 now adds the missing placement dimension: tokens. You already know where Cassandra thinks nodes live; next you will see how a partition key hashes to a token, how vnodes divide the token space, and how token ownership combines with NTS topology to determine replicas.
Check your understanding
- Why is one node per rack in this lab only a simulation of AZ survival?
- With RF=3 and three racks, what does NTS try to achieve when enough racks exist?
- Why is a successful QUORUM read after one rack failure not a complete HA test?
- What should happen before changing rack labels in production?
- What new mechanism will Chapter 03 add to the placement model?
Review the answers
1. All containers still share one physical host; the rack labels exercise Cassandra placement logic but not independent infrastructure failure domains.
2. One replica per distinct rack for a token range, avoiding correlated rack failure where possible.
3. It does not measure every operation, latency under load, driver failover, surviving-node headroom, repair, or other shared dependencies.
4. A supported topology migration plan with placement/streaming/repair validation, backups, capacity checks and rollback—not an ad-hoc edit.
5. Partition-key hashing, tokens, token ranges and vnodes, which combine with replication topology to determine replicas.
Cleanup
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_resilience;"
docker unpause atlasmart-cass-1 2>/dev/null || truedocker unpause atlasmart-cass-2 2>/dev/null || truedocker unpause atlasmart-cass-3 2>/dev/null || truedocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
docker unpause atlasmart-cass-1 2>$nulldocker unpause atlasmart-cass-2 2>$nulldocker unpause atlasmart-cass-3 2>$nulldocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Partition Key Hashing and Murmur3Partitioner: From Key to Token.
Authoritative references
- Apache Cassandra 5.0 documentation — Official documentation entry point for the current 5.0 line.
- Apache Cassandra downloads — Official release page used to verify Cassandra 5.0.9 as the current GA patch.
- Snitch — Official explanation of topology/proximity and rack-aware replica placement.
- cassandra-rackdc.properties — Official DC/rack configuration for GossipingPropertyFileSnitch.
- Dynamo architecture and replication — Official NetworkTopologyStrategy and replica-placement semantics.
- nodetool gossipinfo — Official gossip-state inspection command.
- nodetool failuredetector — Official failure-detector inspection command.
- Java Driver 4.19 node metadata — Apache Java Driver documentation for driver-side node state and topology metadata.
- Production recommendations — Official guidance on NetworkTopologyStrategy and rack configuration before production use.