Chapter 04 · Keyspaces, Replication Strategies, Replication Factor, and Placement
NetworkTopologyStrategy: Replication Factor per Datacenter and Rack-Aware Placement
Make NetworkTopologyStrategy rack-aware placement observable and separate topology labels from real failure domains.
Learning outcomes
AtlasMart can pay for three replicas and still receive poor fault isolation if Cassandra topology labels do not match real failure domains. This lesson makes the division between topology metadata and placement policy explicit.
Explain snitch/locator versus replication-strategy responsibilities.
Show how NTS sets RF independently per DC.
Map replica endpoints back to rack labels for multiple keys.
Separate logical rack labels from physical rack/AZ isolation.
Design a multi-DC RF map without pretending the local Docker host is multi-DC.
Apache Cassandra 5.0.9 is the current GA 5.0 patch as of 7
September 2026. The pinned Docker Official Image
cassandra:5.0.9 uses Java 17. The course cluster
is atlasmart-course, one DC (dc1),
three rack labels, 16 vnodes per node, no TLS/auth in this
isolated chapter lab, and no host-published CQL/JMX ports.
Docker and Cassandra are unavailable in this generation environment. Commands and result shapes were checked against current official Cassandra documentation, but outputs are expected shapes rather than fabricated captured runs. Record exact values on your own disposable cluster.
Reusable Chapter 04 lab topology
Continue the Chapter 01–03 naming and failure-domain conventions. The rack labels let Cassandra exercise rack-aware placement, but all containers still share one physical host, so the lab demonstrates placement semantics rather than true rack/AZ isolation. Require all three nodes to be Up/Normal before interpreting replica evidence.
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
Capture
SELECT cluster_name, data_center, rack, release_version FROM
system.local;,
SELECT peer, data_center, rack, release_version FROM
system.peers_v2;, and nodetool status. If those disagree with the
intended topology, stop and fix the lab rather than explaining
placement from bad metadata.
1. Snitch says where; NTS says how many
The endpoint snitch/locator provides each node's datacenter and rack metadata. NTS consumes those labels plus token ownership and the keyspace RF map to choose natural replicas. A driver local-DC policy is another layer again: it influences which coordinator the client prefers, not where the server stores natural replicas.
| Layer | Responsibility |
|---|---|
| Snitch/locator | Node DC/rack identity and proximity metadata |
| NTS | Replica placement per keyspace/DC while preferring rack diversity |
| Driver policy | Coordinator routing/locality |
2. Use RF=3 and observe rack-aware placement
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_replication; CREATE KEYSPACE atlasmart_replication WITH replication = {'class':'NetworkTopologyStrategy','dc1':3}; CREATE TABLE atlasmart_replication.inventory_by_sku (sku text PRIMARY KEY, quantity int); INSERT INTO atlasmart_replication.inventory_by_sku (sku,quantity) VALUES ('sku-1001',42); INSERT INTO atlasmart_replication.inventory_by_sku (sku,quantity) VALUES ('sku-1002',17); INSERT INTO atlasmart_replication.inventory_by_sku (sku,quantity) VALUES ('sku-1003',9);"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_replication';"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_replication inventory_by_sku sku-1001docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_replication inventory_by_sku sku-1002docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_replication inventory_by_sku sku-1003docker exec atlasmart-cass-1 cqlsh -e "SELECT peer, data_center, rack FROM system.peers_v2;"
With three nodes in three rack labels and RF=3, each partition should resolve to all three nodes. The order can vary. Map endpoint→rack and record that the logical rack diversity is visible.
3. Per-DC RF is explicit
A production two-DC keyspace might use
{'class':'NetworkTopologyStrategy','dc1':3,'dc2':3}. That means three intended replicas in each named DC.
Cassandra does not infer RF from node count and cannot invent
physical independence when a DC lacks enough racks/nodes.
Example production intent:{'class':'NetworkTopologyStrategy','dc1':3,'dc2':3}Verify before rollout:- exact snitch-reported DC/rack names- enough independent racks/AZs per DC- local and failover CLs- driver local-DC configuration- repair, storage and backup capacity- application RPO/RTO under DC loss
4. Deliberately wrong: labels that lie
Three containers labeled rack1/rack2/rack3 on one laptop do not create three physical failure domains. Conversely, three real AZs mislabeled as one rack hide available diversity from NTS. Changing labels on an established cluster is not cosmetic; it can alter placement expectations and require topology/data movement work.
Pair Cassandra topology evidence with infrastructure evidence. Cassandra sees labels, not power feeds, hypervisors, switches, AZ boundaries, or correlated failure.
5. Production judgment
Use NTS for production keyspaces and choose RF from tolerated failures, CLs, latency SLOs, storage/network cost, and repair headroom. RF=3 is common, but only useful as fault tolerance when the three placements are genuinely independent enough for the stated failure model.
Verification checklist
- NTS RF=3 is visible in schema metadata.
- You mapped three sample keys to endpoints/racks.
- You can state which layer supplies topology and which chooses replicas.
- You did not claim the single Docker host survives a real rack outage.
- You can explain what a multi-DC RF map adds.
Check your understanding
- Does the snitch choose RF?
- What does NTS configure independently?
- Can rack labels create physical isolation?
- Does driver local-DC routing change natural replicas?
- What does Lesson 3 contrast?
Review the answers
1. No; it supplies topology metadata.
2. Replication factor for each named datacenter.
3. No; they only describe intended topology to Cassandra.
4. No; it changes coordinator preference.
5. SimpleStrategy, which has no production rack/DC-aware placement contract.
Cleanup/reset
Remove only the disposable course-owned containers, volumes, and network. Never copy these cleanup commands into production.
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>/dev/null || truedocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra 2>/dev/null || true
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>$nulldocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>$nulldocker network rm atlasmart-cassandra 2>$null
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Why SimpleStrategy Is for Limited/Legacy Scenarios and Not Production Multi-Rack Design.
Authoritative references
- Apache Cassandra downloads — current GA version evidence.
- CQL data definition — CREATE/ALTER KEYSPACE, NTS, SimpleStrategy, RF, durable_writes.
- Repair — incremental/full repair, preview, and RF increase implications.
- Cassandra FAQ — RF increase/decrease operational guidance.
- nodetool — status, getendpoints, repair, cleanup and related operator commands.
- Docker Official Image 5.0 Dockerfile — Cassandra 5.0.9 + Java 17 image baseline.