Chapter 21 · Redis Cluster: Hash Slots, Routing, Resharding, and Cluster Availability

Cluster Partitions, Replica Migration, Availability Rules, and Designing for Node/Zone Failure

Test node failure, availability rules, replica migration, majority behavior, and real failure-domain placement before calling the Cluster highly available.

Advanced220–320 minutespartitions, replica migration, coverage, failover game dayRedis Open Source 8.10.1Docker + redis-cli + redis-py 8.1.06-node Cluster + 1 spare nodeDB 0 · AOF everysec · maxmemory 0/noevictionFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

By the end of this lesson, you should be able to:

01

Explain Cluster PFAIL, FAIL, config epochs, and availability when primaries or slots become unreachable.

02

Distinguish automatic replica failover from replica migration and explain cluster-migration-barrier.

03

Reason about cluster-require-full-coverage and cluster-allow-reads-when-down as policy choices with correctness consequences.

04

Design primary/replica placement across real failure domains instead of counting containers.

05

Run a bounded primary-failure drill and measure cluster/application recovery without claiming zero acknowledged-write loss.

Reproducible Chapter 21 baseline

Redis Open Source 8.10.1 using redis:8.10.1; seven named containers are defined but the normal cluster starts six nodes (three primaries + three replicas) on private Docker network atlasmart-redis-ch21-net; the seventh node is started only for add/remove exercises; Redis Cluster uses logical database 0 only; AOF everysec; maxmemory 0/noeviction for this bounded lab; TLS is off only because all Cluster client/bus traffic is confined to one private single-host Docker network; default ACL user is disabled; named academy-admin and atlasmart-app users use disposable lab passwords. No Search/JSON/vector/time-series/probabilistic feature is required. All keys use atlasmart:ch21:*.

1. Practical problem: six nodes can still fail as one system

A Cluster with three primaries and three replicas sounds redundant, but availability depends on which nodes fail together. If a primary and its only replica share one host or zone, that failure can make the primary’s slots unavailable. Node count alone is not fault-domain safety.

2. PFAIL, FAIL, and cluster_state

A node first suspects another node as PFAIL when it is not reachable. Failure reports can promote that condition to FAIL. CLUSTER INFO exposes slots in PFAIL/FAIL and whether the local node considers the cluster ok or fail. Configuration epochs decide which primary claim wins after failover/topology change; they do not version individual writes.

3. Automatic replica failover needs reachable primary majority

When a primary fails and an eligible replica can reach enough primary voters, the replica can be promoted and claim its primary’s slots with a new configuration epoch. Asynchronous replication means promotion is an availability mechanism, not proof that every acknowledged write survived. Measure replica lag and post-failover application state.

4. Bounded failure drill: stop one verified primary

First verify n1’s current role because Lesson 4 may have changed roles. If n1 is a replica now, choose another verified primary and its container instead. Stop only one primary, keep all other nodes alive, and measure: first client failure, PFAIL/FAIL observation, replica promotion, first successful cluster-aware write, and topology convergence.

Shell · example single-primary failure drill
docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n1 redis-cli --user academy-admin CLUSTER NODES# Only if n1 is currently a primary:docker stop atlasmart-redis-ch21-n1sleep 6docker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n2 redis-cli --user academy-admin CLUSTER INFOdocker exec -e REDISCLI_AUTH=AtlasMart-Ch21-Admin-Lab-Only-2026 atlasmart-redis-ch21-n2 redis-cli --user academy-admin CLUSTER NODESdocker start atlasmart-redis-ch21-n1

5. Replica migration repairs future redundancy, not past data loss

Redis Cluster can move a replica from a well-covered primary to an orphaned primary. This is replica migration, separate from slot migration. cluster-migration-barrier controls how many good replicas the donor primary must keep before a replica can migrate away. Migration improves future survivability; it does not recover writes that were never replicated.

6. Full coverage policy changes blast radius

With default cluster-require-full-coverage yes, loss of slot coverage can make the cluster stop serving writes rather than continue a partially available keyspace. Setting it to no can keep unaffected slots available, but the application must tolerate some keys being unavailable. cluster-allow-reads-when-down is another separate policy; changing either without application-level design can turn infrastructure availability into semantic inconsistency.

7. Majority partition behavior

A primary partition that cannot communicate with a majority of primaries cannot safely authorize ordinary failover epochs. The majority side can usually continue by promoting eligible replicas for failed primaries it can replace; a minority side may stop serving as the cluster is marked down. This is why three primaries are the minimum useful teaching topology and why three primaries all in one zone are still poor HA design.

8. Zone design: place each shard across failure domains

For each shard, separate primary and replica across hosts/zones. Also distribute primary voters so one zone loss does not remove a majority. A production topology may need more than one replica per primary when the organization wants tolerance for sequential failures and automated replica migration. Redis Open Source Cluster itself does not magically know cloud zones in this single-host lab; placement is an operator/orchestrator responsibility.

9. Hot-slot and failover capacity interact

When one primary fails, its replica inherits that primary’s hottest slots. A replica sized only for average background replication may become CPU/memory/network saturated immediately after promotion. Headroom planning therefore combines per-slot statistics with failover placement, not just total dataset size.

10. Game-day evidence to capture

Record CLUSTER INFO, SHARDS, NODES, slot statistics, role changes, config epochs, client MOVED/connection errors, p50/p95/p99, successful/failed writes, recovery duration, and exact acknowledged business state before/after. Do not claim zero loss unless your own measured protocol/application invariant proves it under the tested failure.

11. Deliberately wrong design: one primary + replica pair per zone

If each primary’s only replica sits in the same zone as that primary, a zone outage can remove whole shards. Repair by cross-zone pairing and majority-aware primary distribution; verify by simulating the intended failure domain, not just killing one process on one host.

12. Cleanup and bridge to security

After collecting evidence, remove only Chapter 21 containers/volumes/network. Chapter 22 assumes the topology concepts are understood and focuses on ACL/TLS/network security boundaries. Cluster routing does not replace authentication, authorization, TLS, or tenant isolation.

Shell · remove only Chapter 21 resources
docker compose -f ch21-compose.yaml down -v --remove-orphansdocker network rm atlasmart-redis-ch21-net 2>/dev/null || true# Never use docker system prune for this lab.

Check your understanding

  1. What is replica migration?
  2. What does cluster-require-full-coverage yes do?
  3. Does automatic failover guarantee every acknowledged write survives?
  4. Why combine slot statistics with failover planning?
Review the answers

Automatic reassignment of a replica from a well-covered primary to an orphaned primary to improve future availability; it is not slot migration.

By default, missing slot coverage can make the cluster stop accepting writes rather than serve a partially covered keyspace.

No. Replication remains asynchronous unless additional application/durability mechanisms prove a tighter bound.

A promoted replica inherits the failed primary’s actual hot-slot load, not just an average share of keys.

Summary and next step

Cluster Partitions, Replica Migration, Availability Rules, and Designing for Node/Zone Failure is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Default User, ACL Users, Passwords, Command Categories, Key Patterns, and Channel Patterns.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.