Chapter 02 · Clusters, Datacenters, Racks, Snitches, Gossip, and Failure Detection

Endpoint Snitches and Replication Placement: Mapping Logical Topology to Physical Infrastructure

Map logical topology to placement correctly: the snitch says where peers are; the replication strategy decides how many copies belong in each DC and rack-aware selection constrains where they land.

Intermediate110–130 minutesRack-aware replica placement labApache Cassandra 5.0.9 · 3-node dc1/rack1–3 labLast reviewed: September 2026

Learning outcomes

AtlasMart now has three visible peers. The next risk is subtle: engineers often say “the snitch replicates data across racks.” That sentence assigns responsibility to the wrong mechanism. This lesson separates the endpoint snitch's topology map from the replication strategy's replica-selection rules and then proves the distinction with a keyspace and real endpoint evidence.

01

Explain what GossipingPropertyFileSnitch reads locally and what it disseminates to peers.

02

Distinguish endpoint topology/proximity from NetworkTopologyStrategy replication factor and replica placement.

03

Inspect cassandra-rackdc.properties and peer topology without assuming container names are authoritative.

04

Use nodetool getendpoints to connect one partition key to concrete replica endpoints.

05

Reason about mis-labeled racks and why adding a tiny new rack can produce surprising ownership pressure.

Chapter baseline reviewed 7 September 2026

Apache Cassandra 5.0.9 is the current GA 5.0 patch on the official download page. These labs pin cassandra:5.0.9. The chapter uses one logical datacenter, dc1, and three logical racks, rack1–rack3, on the isolated Docker network atlasmart-cassandra. Topology labels are learning metadata; they do not create real host/rack/availability-zone isolation by themselves.

Execution and safety note

The generation environment does not contain Docker or Cassandra, so commands were documentation- and syntax-checked rather than executed here. Record your actual addresses, host IDs, tokens, gossip generations, phi values, startup times, and driver events. Never treat the example output shapes as captured measurements. All failure injection is confined to disposable course containers; no host firewall, clock manipulation, or production endpoint is required.

1. The snitch answers “where is this endpoint?”

Cassandra's endpoint snitch classifies peers by network topology. In the course baseline, GossipingPropertyFileSnitch reads the local node's dc and rack values from cassandra-rackdc.properties and propagates that information through gossip. The snitch can also contribute proximity information used for request behavior. It does not independently decide “store three copies.”

The replication strategy belongs to the keyspace. NetworkTopologyStrategy (NTS) accepts an RF per datacenter and uses token ownership plus snitch-supplied rack information to choose replicas. If a DC has at least as many racks as its RF, NTS can place each replica in a distinct rack. If there are fewer racks than RF, some racks necessarily hold multiple replicas.

Mechanism Owns Does not own
GossipingPropertyFileSnitch DC/rack identity and topology/proximity knowledge RF, consistency level, token ownership
NetworkTopologyStrategy Replica selection per DC using token+topology Physical cloud placement or driver connection pools
Driver local-DC policy Which coordinator hosts the client prefers Server-side replica ownership

2. Inspect the topology source, not just its effects

bash / PowerShell · reusable three-node topology lab
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait for all three nodes to become Up/Normal, then verify from two peers.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool status

After convergence, inspect the effective rack/DC file inside each container. The Docker Official Image maps the chapter's environment variables into Cassandra configuration, but the evidence should be the effective file and system metadata.

configuration provenance · local files plus gossiped peer metadata
docker exec atlasmart-cass-1 sh -lc "grep -E '^(dc|rack)=' /etc/cassandra/cassandra-rackdc.properties"docker exec atlasmart-cass-2 sh -lc "grep -E '^(dc|rack)=' /etc/cassandra/cassandra-rackdc.properties"docker exec atlasmart-cass-3 sh -lc "grep -E '^(dc|rack)=' /etc/cassandra/cassandra-rackdc.properties"docker exec atlasmart-cass-1 cqlsh -e "SELECT data_center, rack FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer, data_center, rack FROM system.peers_v2;"

The local file proves what each node was configured to claim; system.peers_v2 proves what node 1 currently knows about its peers. Neither command proves those rack labels map to three real power/network/AZ domains.

3. Prove replica placement for an AtlasMart partition

Create a small RF=3 keyspace and one query-shaped table. This lesson is not yet teaching the full token ring, but nodetool getendpoints lets you connect one partition key to the replica endpoints selected by the current topology.

replica evidence · one partition key, RF=3, three racks
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_topology WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_topology.orders_by_customer (customer_id text, order_id text, created_at timestamp, status text, PRIMARY KEY ((customer_id), created_at, order_id)) WITH CLUSTERING ORDER BY (created_at DESC);"docker exec atlasmart-cass-1 cqlsh -e "INSERT INTO atlasmart_topology.orders_by_customer (customer_id,order_id,created_at,status) VALUES ('c-1001','o-2001','2026-09-07T10:00:00Z','PAID');"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_topology orders_by_customer c-1001docker exec atlasmart-cass-1 nodetool status

With three Up/Normal nodes and RF=3 in one DC, this tiny cluster normally reports all three endpoints for the partition. The important proof is not the endpoint order but that replica selection is derived from the keyspace's RF and the topology/token map. Chapter 03 will explain the token calculation itself.

4. Why a new rack is not automatically “more balanced”

Official Cassandra documentation calls out a counter-intuitive case: if a new rack contains only one node, rack-aware replica placement can cause that node to become a replica for a very large fraction of token ranges. A rack label changes placement constraints; it is not a harmless grouping tag. Uneven node counts per rack can therefore create uneven data/load pressure.

The deliberately wrong approach is to invent a fresh rack name whenever a host moves, or to “fix” an imbalance by relabeling live nodes without a topology plan. The safer pattern is to decide the rack model from infrastructure failure domains, keep rack populations operationally sensible, and make topology changes through documented lifecycle/streaming procedures.

Do not run a relabel experiment on a populated production cluster.

This course demonstrates placement on disposable containers. Production rack/snitch changes must be treated as data-placement changes with backups, supported procedure, repair/streaming validation, and rollback—not as a metadata edit.

5. Production judgment, verification, and cleanup

Choose NTS for production keyspaces and make DC/rack labels match the failure domains that matter to the service. Replication factor must be chosen with consistency levels, failure tolerance, storage cost and repair cost—not by copying “RF=3” as folklore. A three-rack RF=3 pattern is useful because one replica per rack can tolerate a rack loss for some CLs, but the exact client-visible result depends on CL and remaining replica availability.

Verification checklist

  • Each node's effective rack/DC file matches the intended map.
  • system.peers_v2 shows the same topology from a peer's perspective.
  • The keyspace uses NetworkTopologyStrategy, not SimpleStrategy.
  • nodetool getendpoints returns concrete replica endpoints for the test partition.
  • You can explain why snitch, NTS and driver routing are separate layers.

Check your understanding

  1. Does GossipingPropertyFileSnitch set the replication factor?
  2. What happens if RF is greater than the number of racks in a DC?
  3. Why can a one-node new rack be operationally surprising?
  4. What does nodetool getendpoints prove?
  5. Does a driver preferring dc1 change which nodes own the data?
Review the answers

1. No. It supplies topology metadata; RF is configured by the keyspace replication strategy.

2. Some replicas must share racks; NTS cannot create failure domains that do not exist in the topology.

3. Rack-aware placement can make that single node a replica for a disproportionate set of ranges.

4. It shows which endpoints are replicas for a specific key under current token ownership and replication configuration.

5. No. It changes coordinator preference, not server-side replica ownership.

Cleanup

remove chapter-specific keyspace
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_topology;"
cleanup · Bash form; remove only course-owned resources
docker unpause atlasmart-cass-1 2>/dev/null || truedocker unpause atlasmart-cass-2 2>/dev/null || truedocker unpause atlasmart-cass-3 2>/dev/null || truedocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
cleanup · PowerShell form
docker unpause atlasmart-cass-1 2>$nulldocker unpause atlasmart-cass-2 2>$nulldocker unpause atlasmart-cass-3 2>$nulldocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Gossip for Membership and State Dissemination: What Nodes Learn About Each Other.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.