Chapter 04 · Keyspaces, Replication Strategies, Replication Factor, and Placement

Replication Factor, Failure Tolerance, Storage Cost, and Consistency-Level Math

Derive RF/consistency-level failure behavior from replica counts and test controlled replica loss safely.

Intermediate110–160 minutesMechanism-first replication labApache Cassandra 5.0.9 · Java 17 · NTS · 16 vnodes/nodeLast reviewed: September 2026

Learning outcomes

AtlasMart now has topology-aware replicas, but request availability still depends on how many replicas each operation requires. This lesson derives behavior from RF/CL math instead of relying on generic “high availability” language.

01

Calculate ONE, TWO, QUORUM, LOCAL_QUORUM, and ALL requirements for RF=3.

02

Separate intended replica count from per-request response/acknowledgment count.

03

Test one-node and two-node loss safely.

04

Explain quorum overlap without overclaiming linearizability.

05

Connect RF to storage, repair, streaming, and cost.

Version/topology baseline

Apache Cassandra 5.0.9 is the current GA 5.0 patch as of 7 September 2026. The pinned Docker Official Image cassandra:5.0.9 uses Java 17. The course cluster is atlasmart-course, one DC (dc1), three rack labels, 16 vnodes per node, no TLS/auth in this isolated chapter lab, and no host-published CQL/JMX ports.

Execution note

Docker and Cassandra are unavailable in this generation environment. Commands and result shapes were checked against current official Cassandra documentation, but outputs are expected shapes rather than fabricated captured runs. Record exact values on your own disposable cluster.

Reusable Chapter 04 lab topology

Continue the Chapter 01–03 naming and failure-domain conventions. The rack labels let Cassandra exercise rack-aware placement, but all containers still share one physical host, so the lab demonstrates placement semantics rather than true rack/AZ isolation. Require all three nodes to be Up/Normal before interpreting replica evidence.

setup · Bash; pinned three-node cluster
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
setup · PowerShell; same pinned cluster
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version

Capture SELECT cluster_name, data_center, rack, release_version FROM system.local;, SELECT peer, data_center, rack, release_version FROM system.peers_v2;, and nodetool status. If those disagree with the intended topology, stop and fix the lab rather than explaining placement from bad metadata.

1. Derive the counts first

For RF=N, QUORUM requires floor(N/2)+1. At RF=3, QUORUM and LOCAL_QUORUM in this one-DC lab each require two replicas. ALL requires three; ONE requires one. In multi-DC deployments the scope of LOCAL_QUORUM differs from global QUORUM even when a local count looks similar.

RF=3 CL Required 1 replica down 2 replicas down
ONE / LOCAL_ONE 1 Can succeed Can succeed if remaining replica is reachable
TWO 2 Can succeed Cannot meet CL
QUORUM / LOCAL_QUORUM 2 Can succeed Cannot meet CL
ALL 3 Cannot meet CL Cannot meet CL

2. Create an RF=3 fixture

CQL · consistency fixture
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_consistency; CREATE KEYSPACE atlasmart_consistency WITH replication = {'class':'NetworkTopologyStrategy','dc1':3}; CREATE TABLE atlasmart_consistency.stock_by_sku (sku text PRIMARY KEY, quantity int); INSERT INTO atlasmart_consistency.stock_by_sku (sku,quantity) VALUES ('sku-1001',42);"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_consistency stock_by_sku sku-1001docker exec atlasmart-cass-1 nodetool status atlasmart_consistency

Because there are three nodes and RF=3, all three are replicas for this partition. That makes the failure boundary easy to observe.

3. One replica down: LOCAL_QUORUM survives, ALL does not

failure injection · stop one replica
docker stop atlasmart-cass-3# Wait until a live node reflects the failure.docker exec atlasmart-cass-1 nodetool status atlasmart_consistencydocker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_consistency.stock_by_sku WHERE sku='sku-1001';"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; UPDATE atlasmart_consistency.stock_by_sku SET quantity=41 WHERE sku='sku-1001';"# Expected to fail because ALL needs all 3 replicas.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY ALL; SELECT * FROM atlasmart_consistency.stock_by_sku WHERE sku='sku-1001';"

Capture the actual exception rather than hard-coding one message. Failure detection timing can affect whether the client sees an unavailable versus timeout-style outcome.

4. Two replicas down: cross the quorum boundary

failure injection · stop second replica, then recover
docker stop atlasmart-cass-2docker exec atlasmart-cass-1 nodetool status atlasmart_consistency# Expected to fail: LOCAL_QUORUM needs two RF=3 replicas.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_consistency.stock_by_sku WHERE sku='sku-1001';"docker start atlasmart-cass-2 atlasmart-cass-3# Wait for Up/Normal before accepting recovery.docker exec atlasmart-cass-1 nodetool status atlasmart_consistencydocker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_consistency.stock_by_sku WHERE sku='sku-1001';"

The lesson is not that lower CL is “better.” Lower CL can preserve availability with fewer replicas, but it changes participation/freshness guarantees and can expose more ambiguous failure/retry behavior.

5. Quorum overlap and cost

The condition R + W > RF guarantees overlap between the read and successful write replica sets. At RF=3, QUORUM+QUORUM gives 2+2>3. This is useful reasoning, but it does not make every non-LWT Cassandra workload globally linearizable: concurrent timestamps, topology scope, retries, outages, and repair still matter.

Higher RF also means more stored data, write fan-out, streaming/repair work, backup volume and capacity headroom. Choose RF/CL from tolerated failures, latency SLOs, RPO/RTO and operating cost.

Verification checklist

  • You derived the counts before testing.
  • One node down preserved LOCAL_QUORUM but not ALL.
  • Two nodes down prevented LOCAL_QUORUM.
  • Both nodes were restored and returned Up/Normal.
  • You did not equate RF with backup or CL with universal correctness.

Check your understanding

  1. How many replicas does QUORUM require at RF=3?
  2. Why can ALL fail while other traffic works?
  3. Does RF=5 improve availability for ALL?
  4. What does R+W>RF prove?
  5. What later mechanism repairs divergence?
Review the answers

1. Two.

2. ALL is an operation-specific requirement for every replica.

3. Not necessarily; ALL would then depend on all five replicas.

4. Replica-set overlap, not every concurrency or linearizability property.

5. Anti-entropy repair, with hints/reconciliation serving separate roles.

Cleanup/reset

Remove only the disposable course-owned containers, volumes, and network. Never copy these cleanup commands into production.

cleanup · Bash
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>/dev/null || truedocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra 2>/dev/null || true
cleanup · PowerShell
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>$nulldocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>$nulldocker network rm atlasmart-cassandra 2>$null

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Alter Replication Safely: Repair Requirements, Data Movement, and Validation.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.