Chapter 04 · Keyspaces, Replication Strategies, Replication Factor, and Placement
Alter Replication Safely: Repair Requirements, Data Movement, and Validation
Change replication factor safely through schema change, repair or cleanup, convergence evidence, validation, and rollback planning.
Learning outcomes
AtlasMart needs to increase a keyspace from RF=1 to RF=3 without confusing an immediate schema change with physical convergence. The safe workflow captures before/after placement, performs a full repair, validates results, and understands cleanup after an RF decrease.
Explain why RF changes have schema and physical-data phases.
Capture natural replicas before and after an RF increase.
Use full repair for historical data after increasing RF.
Explain cleanup after decreasing RF.
Build acceptance and rollback criteria that include data movement.
Apache Cassandra 5.0.9 is the current GA 5.0 patch as of 7
September 2026. The pinned Docker Official Image
cassandra:5.0.9 uses Java 17. The course cluster
is atlasmart-course, one DC (dc1),
three rack labels, 16 vnodes per node, no TLS/auth in this
isolated chapter lab, and no host-published CQL/JMX ports.
Docker and Cassandra are unavailable in this generation environment. Commands and result shapes were checked against current official Cassandra documentation, but outputs are expected shapes rather than fabricated captured runs. Record exact values on your own disposable cluster.
Reusable Chapter 04 lab topology
Continue the Chapter 01–03 naming and failure-domain conventions. The rack labels let Cassandra exercise rack-aware placement, but all containers still share one physical host, so the lab demonstrates placement semantics rather than true rack/AZ isolation. Require all three nodes to be Up/Normal before interpreting replica evidence.
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
Capture
SELECT cluster_name, data_center, rack, release_version FROM
system.local;,
SELECT peer, data_center, rack, release_version FROM
system.peers_v2;, and nodetool status. If those disagree with the
intended topology, stop and fix the lab rather than explaining
placement from bad metadata.
1. Create historical data at RF=1
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_rfchange; CREATE KEYSPACE atlasmart_rfchange WITH replication = {'class':'NetworkTopologyStrategy','dc1':1}; CREATE TABLE atlasmart_rfchange.inventory_by_sku (sku text PRIMARY KEY, quantity int); INSERT INTO atlasmart_rfchange.inventory_by_sku (sku,quantity) VALUES ('sku-1001',42); INSERT INTO atlasmart_rfchange.inventory_by_sku (sku,quantity) VALUES ('sku-1002',17); INSERT INTO atlasmart_rfchange.inventory_by_sku (sku,quantity) VALUES ('sku-1003',9);"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_rfchange';"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_rfchange inventory_by_sku sku-1001docker exec atlasmart-cass-1 nodetool flush atlasmart_rfchange inventory_by_sku
Record the single natural endpoint for sku-1001.
The flush is only an observability aid; it is not what makes the
RF workflow correct.
2. ALTER RF=3 changes desired placement now
docker exec atlasmart-cass-1 cqlsh -e "ALTER KEYSPACE atlasmart_rfchange WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_rfchange';"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_rfchange inventory_by_sku sku-1001docker exec atlasmart-cass-1 nodetool status atlasmart_rfchange
Current Cassandra warns that increasing RF requires a full repair to distribute existing data. Seeing RF=3 and three natural endpoints proves the new placement policy, not historical convergence.
Treat “all nodes agree on RF=3” and “historical rows exist on the new replicas” as separate acceptance gates.
3. Preview and run a full repair
docker exec atlasmart-cass-1 nodetool repair --preview --full atlasmart_rfchange inventory_by_skudocker exec atlasmart-cass-1 nodetool repair --full atlasmart_rfchange inventory_by_sku# Re-run preview after repair; tiny datasets may report little/no streaming.docker exec atlasmart-cass-1 nodetool repair --preview --full atlasmart_rfchange inventory_by_skudocker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_rfchange.inventory_by_sku WHERE sku='sku-1001';"
Full repair is required for an RF increase because already-repaired SSTables must not be skipped. In production, repair is rolling and resource-aware: record disk headroom, network, compaction, GC, pending tasks and p95/p99 latency.
4. RF decrease requires cleanup, not the same workflow
After lowering RF, nodes that are no longer natural replicas can
still retain surplus SSTable data. Official guidance is to run
nodetool cleanup on relevant nodes in a controlled
rolling plan. Cleanup is per node and can be I/O intensive.
1. Verify current RF/topology/CLs/repair health/backups.2. ALTER KEYSPACE to the lower RF.3. Confirm schema agreement and new replica maps.4. Validate application reads/writes at intended CLs.5. Run nodetool cleanup on relevant nodes in a controlled rolling plan.6. Track reclaimed disk, compaction pressure, errors and tail latency.7. Revalidate and update capacity/DR documentation.
5. Rollback and production judgment
Rollback is not merely reverse CQL. If repair has already copied new replicas, reducing RF leaves surplus data until cleanup. If cleanup has already removed data and you restore the old RF, you must repopulate replicas again. Any production RF change therefore needs a capacity, repair/cleanup, validation and rollback plan.
Replication remains distinct from backup. RF can tolerate some node/rack failures; it does not protect against application-wide deletes, credential compromise, bad timestamps, or correlated loss without independent recovery copies.
Verification checklist
- You captured RF=1 and one endpoint before change.
- You captured RF=3 and three endpoints immediately after ALTER.
- You did not mistake metadata for convergence.
- You ran full repair and recorded real evidence.
- You can explain cleanup after RF decrease.
- You can explain why rollback can require new data movement.
Check your understanding
- Why is full repair required after increasing RF?
- What does ALTER KEYSPACE prove immediately?
- What follows an RF decrease?
- Why is rollback more than CQL?
- Does RF replace backup planning?
Review the answers
1. To populate newly selected replicas with existing data and avoid skipping already-repaired SSTables.
2. Desired replication metadata changed.
3. Controlled nodetool cleanup to remove no-longer-owned surplus data.
4. Repair/cleanup may already have changed physical replica contents.
5. No; replication and backup cover different failure classes.
Cleanup/reset
Remove only the disposable course-owned containers, volumes, and network. Never copy these cleanup commands into production.
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>/dev/null || truedocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra 2>/dev/null || true
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>$nulldocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>$nulldocker network rm atlasmart-cassandra 2>$null
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to CREATE/ALTER/DROP TABLE and Keyspace Objects with Schema Agreement Awareness.
Authoritative references
- Apache Cassandra downloads — current GA version evidence.
- CQL data definition — CREATE/ALTER KEYSPACE, NTS, SimpleStrategy, RF, durable_writes.
- Repair — incremental/full repair, preview, and RF increase implications.
- Cassandra FAQ — RF increase/decrease operational guidance.
- nodetool — status, getendpoints, repair, cleanup and related operator commands.
- Docker Official Image 5.0 Dockerfile — Cassandra 5.0.9 + Java 17 image baseline.