Chapter 21 · Redis Cluster: Hash Slots, Routing, Resharding, and Cluster Availability

Add/Remove Nodes, Reshard Slots, Migrate Keys, Fail Over Primaries, and Maintain Headroom

Add/remove capacity, reshard slots, use current atomic migration tooling, and perform safe planned primary failover with headroom.

Advanced220–320 minutesadd-node, reshard, atomic migration, del-node, failoverRedis Open Source 8.10.1Docker + redis-cli + redis-py 8.1.06-node Cluster + 1 spare nodeDB 0 · AOF everysec · maxmemory 0/noevictionFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

By the end of this lesson, you should be able to:

01

Add an empty Cluster node, verify propagation, and move a bounded set of slots to it.

02

Explain Redis 8.4+ atomic slot migration and Redis 8.10 redis-cli reshard behavior.

03

Move slots off a node before deleting it and preserve replica/capacity headroom during maintenance.

04

Perform a coordinated CLUSTER FAILOVER on a selected replica and verify config-epoch/role changes.

05

Measure reshard/failover latency and slot hotness instead of assuming maintenance is free.

Reproducible Chapter 21 baseline

Redis Open Source 8.10.1 using redis:8.10.1; seven named containers are defined but the normal cluster starts six nodes (three primaries + three replicas) on private Docker network atlasmart-redis-ch21-net; the seventh node is started only for add/remove exercises; Redis Cluster uses logical database 0 only; AOF everysec; maxmemory 0/noeviction for this bounded lab; TLS is off only because all Cluster client/bus traffic is confined to one private single-host Docker network; default ACL user is disabled; named academy-admin and atlasmart-app users use disposable lab passwords. No Search/JSON/vector/time-series/probabilistic feature is required. All keys use atlasmart:ch21:*.

1. Practical problem: scale without a stop-the-world repartition

AtlasMart needs more primary capacity. Cluster can add an empty node and transfer slot ownership while clients continue operating. The hard part is not the command syntax; it is maintaining spare memory, network, CPU, replicas, and client routing headroom during the move.

2. Start the spare n7 and join it without slots

n7 is defined in Compose but was not part of the initial six-node cluster. Start it, use --cluster add-node, and verify it appears as a primary with zero slots before moving anything.

Shell · add n7 as an empty primary
docker compose -f ch21-compose.yaml up -d n7docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 \  redis-cli --user academy-admin --cluster add-node atlasmart-redis-ch21-n7:6379 atlasmart-redis-ch21-n1:6379sleep 2docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 redis-cli --user academy-admin CLUSTER SHARDS

3. Capture slot/resource baselines before reshard

Record slot ownership, hottest slots, memory, network counters, client p95/p99, and replica coverage. Do not begin a large move when a source/destination is already near memory or network saturation. The migration temporarily needs both control-plane and data-movement capacity.

4. Move 64 slots from n1 to n7

For a bounded lab, move only 64 slots. Redis 8.10 redis-cli reshard uses the Redis 8.4+ server-side atomic migration path. Current tooling can therefore have different transient details from classic key-by-key MIGRATING/IMPORTING examples; record the actual 8.10 evidence rather than expecting an old trace.

Shell · bounded reshard to n7
SRC=$(docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 redis-cli --user academy-admin --raw CLUSTER MYID)DST=$(docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n7 redis-cli --user academy-admin --raw CLUSTER MYID)docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 \  redis-cli --user academy-admin --cluster reshard atlasmart-redis-ch21-n1:6379 \  --cluster-from "$SRC" --cluster-to "$DST" --cluster-slots 64 --cluster-yesdocker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 redis-cli --user academy-admin CLUSTER SHARDS

5. Verify ownership and actual workload impact

Compare CLUSTER SHARDS and CLUSTER INFO before/after. If you run a bounded client loop during migration, report actual retries/redirections and p50/p95/p99 latency. A successful command exit alone does not prove the application met its SLO during the move.

6. Remove a node only after evacuating its slots

A primary serving slots cannot simply disappear from topology without availability consequences. Move n7’s 64 slots back to n1, verify n7 owns zero slots, then delete its node ID from the cluster and stop only that container.

Shell · evacuate then delete n7
SRC=$(docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n7 redis-cli --user academy-admin --raw CLUSTER MYID)DST=$(docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 redis-cli --user academy-admin --raw CLUSTER MYID)docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 \  redis-cli --user academy-admin --cluster reshard atlasmart-redis-ch21-n1:6379 \  --cluster-from "$SRC" --cluster-to "$DST" --cluster-slots 64 --cluster-yesdocker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 \  redis-cli --user academy-admin --cluster del-node atlasmart-redis-ch21-n1:6379 "$SRC"docker stop atlasmart-redis-ch21-n7

7. Planned primary maintenance: fail over to a replica first

For a healthy planned maintenance, select the actual replica of a primary from CLUSTER SHARDS or CLUSTER NODES, then run CLUSTER FAILOVER on that replica. Normal manual failover coordinates with the primary and waits for the replica to consume the primary replication stream before promotion. FORCE and especially TAKEOVER have weaker safety assumptions and are not used in the mandatory lab.

Shell · inspect first, then run normal manual failover
# Example assumes n4 is currently a replica of n1; VERIFY this first.docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 redis-cli --user academy-admin CLUSTER NODESdocker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n4 redis-cli --user academy-admin CLUSTER FAILOVERsleep 3docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 redis-cli --user academy-admin CLUSTER NODES

8. Why TAKEOVER is excluded from the normal playbook

CLUSTER FAILOVER TAKEOVER can promote without majority authorization and unilaterally create a new epoch. It exists for exceptional disaster-recovery scenarios and can violate normal last-failover-wins assumptions. Treat it as a specialist recovery command with explicit split-brain risk, not a shortcut for maintenance.

9. Headroom checklist before topology change

Require spare primary memory, replica availability, network capacity for slot movement/full sync, AOF/fork headroom, stable Cluster bus links, client retry budget, and rollback time. Resharding big keys can create latency because key migration itself has work proportional to what must be transferred.

Check your understanding

  1. Why must n7 be empty before del-node?
  2. What changed in Redis 8.10 redis-cli resharding?
  3. Why prefer normal CLUSTER FAILOVER for planned maintenance?
  4. Why is TAKEOVER dangerous?
Review the answers

Removing a node still responsible for slots would leave ownership/coverage unavailable or force unsafe recovery.

It uses the newer server-side atomic slot migration capability introduced in Redis 8.4.

It coordinates with the primary and waits for the replica to catch up before promotion.

It bypasses normal cluster consensus/epoch authorization and can create conflicting topology during partitions.

10. Production judgment and bridge

Automate repeatable resharding/failover with preflight health/capacity checks and post-change topology/client validation. Move bounded batches, watch tail latency and hot slots, and preserve redundancy throughout maintenance. Lesson 5 expands this to unplanned node/zone partitions and replica migration.

Summary and next step

Add/Remove Nodes, Reshard Slots, Migrate Keys, Fail Over Primaries, and Maintain Headroom is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Cluster Partitions, Replica Migration, Availability Rules, and Designing for Node/Zone Failure.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.