Chapter 04 · Keyspaces, Replication Strategies, Replication Factor, and Placement

Alter Replication Safely: Repair Requirements, Data Movement, and Validation

Change replication factor safely through schema change, repair or cleanup, convergence evidence, validation, and rollback planning.

Intermediate110–160 minutesMechanism-first replication labApache Cassandra 5.0.9 · Java 17 · NTS · 16 vnodes/nodeLast reviewed: September 2026

Learning outcomes

AtlasMart needs to increase a keyspace from RF=1 to RF=3 without confusing an immediate schema change with physical convergence. The safe workflow captures before/after placement, performs a full repair, validates results, and understands cleanup after an RF decrease.

01

Explain why RF changes have schema and physical-data phases.

02

Capture natural replicas before and after an RF increase.

03

Use full repair for historical data after increasing RF.

04

Explain cleanup after decreasing RF.

05

Build acceptance and rollback criteria that include data movement.

Version/topology baseline

Apache Cassandra 5.0.9 is the current GA 5.0 patch as of 7 September 2026. The pinned Docker Official Image cassandra:5.0.9 uses Java 17. The course cluster is atlasmart-course, one DC (dc1), three rack labels, 16 vnodes per node, no TLS/auth in this isolated chapter lab, and no host-published CQL/JMX ports.

Execution note

Docker and Cassandra are unavailable in this generation environment. Commands and result shapes were checked against current official Cassandra documentation, but outputs are expected shapes rather than fabricated captured runs. Record exact values on your own disposable cluster.

Reusable Chapter 04 lab topology

Continue the Chapter 01–03 naming and failure-domain conventions. The rack labels let Cassandra exercise rack-aware placement, but all containers still share one physical host, so the lab demonstrates placement semantics rather than true rack/AZ isolation. Require all three nodes to be Up/Normal before interpreting replica evidence.

setup · Bash; pinned three-node cluster
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version
setup · PowerShell; same pinned cluster
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool versiondocker exec atlasmart-cass-1 java -version

Capture SELECT cluster_name, data_center, rack, release_version FROM system.local;, SELECT peer, data_center, rack, release_version FROM system.peers_v2;, and nodetool status. If those disagree with the intended topology, stop and fix the lab rather than explaining placement from bad metadata.

1. Create historical data at RF=1

baseline · RF=1 data before change
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_rfchange; CREATE KEYSPACE atlasmart_rfchange WITH replication = {'class':'NetworkTopologyStrategy','dc1':1}; CREATE TABLE atlasmart_rfchange.inventory_by_sku (sku text PRIMARY KEY, quantity int); INSERT INTO atlasmart_rfchange.inventory_by_sku (sku,quantity) VALUES ('sku-1001',42); INSERT INTO atlasmart_rfchange.inventory_by_sku (sku,quantity) VALUES ('sku-1002',17); INSERT INTO atlasmart_rfchange.inventory_by_sku (sku,quantity) VALUES ('sku-1003',9);"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_rfchange';"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_rfchange inventory_by_sku sku-1001docker exec atlasmart-cass-1 nodetool flush atlasmart_rfchange inventory_by_sku

Record the single natural endpoint for sku-1001. The flush is only an observability aid; it is not what makes the RF workflow correct.

2. ALTER RF=3 changes desired placement now

schema change · new RF/new replica map
docker exec atlasmart-cass-1 cqlsh -e "ALTER KEYSPACE atlasmart_rfchange WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "SELECT keyspace_name, replication FROM system_schema.keyspaces WHERE keyspace_name='atlasmart_rfchange';"docker exec atlasmart-cass-1 nodetool getendpoints atlasmart_rfchange inventory_by_sku sku-1001docker exec atlasmart-cass-1 nodetool status atlasmart_rfchange

Current Cassandra warns that increasing RF requires a full repair to distribute existing data. Seeing RF=3 and three natural endpoints proves the new placement policy, not historical convergence.

Schema agreement ≠ data convergence

Treat “all nodes agree on RF=3” and “historical rows exist on the new replicas” as separate acceptance gates.

3. Preview and run a full repair

repair · preview then full convergence
docker exec atlasmart-cass-1 nodetool repair --preview --full atlasmart_rfchange inventory_by_skudocker exec atlasmart-cass-1 nodetool repair --full atlasmart_rfchange inventory_by_sku# Re-run preview after repair; tiny datasets may report little/no streaming.docker exec atlasmart-cass-1 nodetool repair --preview --full atlasmart_rfchange inventory_by_skudocker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_rfchange.inventory_by_sku WHERE sku='sku-1001';"

Full repair is required for an RF increase because already-repaired SSTables must not be skipped. In production, repair is rolling and resource-aware: record disk headroom, network, compaction, GC, pending tasks and p95/p99 latency.

4. RF decrease requires cleanup, not the same workflow

After lowering RF, nodes that are no longer natural replicas can still retain surplus SSTable data. Official guidance is to run nodetool cleanup on relevant nodes in a controlled rolling plan. Cleanup is per node and can be I/O intensive.

guarded RF-decrease workflow
1. Verify current RF/topology/CLs/repair health/backups.2. ALTER KEYSPACE to the lower RF.3. Confirm schema agreement and new replica maps.4. Validate application reads/writes at intended CLs.5. Run nodetool cleanup on relevant nodes in a controlled rolling plan.6. Track reclaimed disk, compaction pressure, errors and tail latency.7. Revalidate and update capacity/DR documentation.

5. Rollback and production judgment

Rollback is not merely reverse CQL. If repair has already copied new replicas, reducing RF leaves surplus data until cleanup. If cleanup has already removed data and you restore the old RF, you must repopulate replicas again. Any production RF change therefore needs a capacity, repair/cleanup, validation and rollback plan.

Replication remains distinct from backup. RF can tolerate some node/rack failures; it does not protect against application-wide deletes, credential compromise, bad timestamps, or correlated loss without independent recovery copies.

Verification checklist

  • You captured RF=1 and one endpoint before change.
  • You captured RF=3 and three endpoints immediately after ALTER.
  • You did not mistake metadata for convergence.
  • You ran full repair and recorded real evidence.
  • You can explain cleanup after RF decrease.
  • You can explain why rollback can require new data movement.

Check your understanding

  1. Why is full repair required after increasing RF?
  2. What does ALTER KEYSPACE prove immediately?
  3. What follows an RF decrease?
  4. Why is rollback more than CQL?
  5. Does RF replace backup planning?
Review the answers

1. To populate newly selected replicas with existing data and avoid skipping already-repaired SSTables.

2. Desired replication metadata changed.

3. Controlled nodetool cleanup to remove no-longer-owned surplus data.

4. Repair/cleanup may already have changed physical replica contents.

5. No; replication and backup cover different failure classes.

Cleanup/reset

Remove only the disposable course-owned containers, volumes, and network. Never copy these cleanup commands into production.

cleanup · Bash
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>/dev/null || truedocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra 2>/dev/null || true
cleanup · PowerShell
docker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>$nulldocker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>$nulldocker network rm atlasmart-cassandra 2>$null

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to CREATE/ALTER/DROP TABLE and Keyspace Objects with Schema Agreement Awareness.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.