Chapter 20 · Redis Sentinel: Monitoring, Automatic Failover, and Service Discovery

Leader Election, Failover State Machine, Replica Promotion, and Reconfiguration

Trace Sentinel election, replica selection, promotion, configuration epochs, and reconfiguration of replicas and the returning old primary.

Advanced180–260 minuteselection, config epoch, replica priority, promotionRedis Open Source 8.10.1Docker + redis-cli + redis-py 8.1.06-process isolated Sentinel topologyFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

By the end of this lesson, you should be able to:

01

Trace Sentinel failover from SDOWN/ODOWN through election, replica selection, promotion, and replica reconfiguration.

02

Explain configuration epochs and why majority authorization prevents simultaneous failovers in minority partitions.

03

Predict how replica freshness, priority, replication offset, and run ID influence promotion selection.

04

Inspect Sentinel event logs instead of treating failover as a black box.

05

Verify that the old primary is reconfigured as a replica when it returns.

Reproducible Chapter 20 baseline

Redis Open Source 8.10.1 using redis:8.10.1; three Redis data nodes and three Sentinel processes on one private Docker network; logical database 0; AOF everysec on data nodes; Redis Cluster is not enabled; TLS is off only because the mandatory lab is single-host and Docker-private with host ports bound to 127.0.0.1; Redis ACLs protect both data-node and Sentinel control connections; all passwords are disposable lab values; no Search/JSON/vector/time-series/probabilistic feature is required. Failure injection is confined to the named Chapter 20 topology; synthetic application fixtures use atlasmart:ch20:*.

1. Practical problem: “Sentinel promoted something” is not enough evidence

AtlasMart operators must know which replica was promoted, why it was eligible, how long writes were unavailable, whether other replicas were reconfigured, and whether the previous primary returned safely as a replica. The failover state machine exposes those decisions in Sentinel logs and API state.

2. Election and configuration epochs

After ODOWN, Sentinels try to authorize exactly one failover leader for a configuration epoch. A majority vote gives that leader a unique epoch for the new topology. If a majority cannot communicate, Sentinel deliberately avoids performing a failover in that minority partition.

This voting protects topology configuration, not application data linearizability. Asynchronous replication remains the source of any acknowledged-write loss window.

3. How Sentinel chooses a replica

Sentinel first rejects replicas considered too stale/unreliable. Among eligible replicas it sorts by replica-priority (lower is preferred, zero means never promote), then by replication offset (more processed history is preferred), then by run ID as a deterministic tie-breaker.

Criterion Preference Operational implication
Connectivity/freshness Eligible/recent enough A disconnected replica can be excluded
replica-priority Lower nonzero wins Operator expresses placement preference
Replication offset Higher wins if priority ties Fresher processed stream is preferred
Run ID Lexicographic tie-break Deterministic, not a business preference

4. Capture the pre-failure baseline

Write bounded AtlasMart state, wait for both replicas to catch up, then record the current primary and replica priorities. This is necessary to reason about the later promotion rather than guessing from container names.

Shell · baseline the candidate set
export REDISCLI_AUTH=AtlasMart-Ch20-Admin-Lab-Only-2026printf 'SET atlasmart:ch20:order:9001 state:paidINCR atlasmart:ch20:sequenceWAIT 2 5000' | redis-cli -h 127.0.0.1 -p 6411 --user academy-admin# SET, INCR, and WAIT above use one redis-cli process and therefore one connection.redis-cli -h 127.0.0.1 -p 6412 --user academy-admin INFO replicationredis-cli -h 127.0.0.1 -p 6413 --user academy-admin INFO replicationunset REDISCLI_AUTH

5. Trigger the controlled failover by stopping only the primary

Stopping one named primary container is a bounded process failure on the isolated lab. Do not stop the Docker daemon or sever host networking. Start log tails first so you can see +sdown, +odown, +try-failover, +elected-leader, replica selection, promotion, reconfiguration, and +switch-master events.

Shell · controlled primary failure
docker logs --since 1m -f atlasmart-redis-ch20-sentinel-1# In a second terminal, resolve the exact current primary through Sentinel:export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-client --raw SENTINEL get-master-addr-by-name atlasmart-primary > ch20-current-master.txtunset REDISCLI_AUTHFAILED_PRIMARY=$(head -n 1 ch20-current-master.txt)echo "stopping $FAILED_PRIMARY"docker stop "$FAILED_PRIMARY"# Watch until SENTINEL get-master-addr-by-name atlasmart-primary changes.

6. Observe—not assume—the new primary

Query at least two Sentinels and the promoted Redis node. Replica A is expected to win because priority 50 is lower than replica B’s 100, but the lesson deliberately says expected: an ineligible/stale replica must not be promoted merely because its priority is lower.

Shell · verify promotion and reconfiguration
export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026for p in 26401 26402 26403; do  redis-cli -h 127.0.0.1 -p "$p" --user sentinel-client SENTINEL get-master-addr-by-name atlasmart-primarydoneunset REDISCLI_AUTHexport REDISCLI_AUTH=AtlasMart-Ch20-Admin-Lab-Only-2026redis-cli -h 127.0.0.1 -p 6412 --user academy-admin ROLEredis-cli -h 127.0.0.1 -p 6413 --user academy-admin ROLEunset REDISCLI_AUTH

7. Restart the old primary and watch Sentinel impose current topology

A failed-over primary does not automatically remain an independent master forever. When the old primary becomes reachable, Sentinels try to impose the current configuration and reconfigure it as a replica of the new primary. Verify this with ROLE and Sentinel logs instead of assuming split-brain disappeared.

Shell · return the old node
FAILED_PRIMARY=$(head -n 1 ch20-current-master.txt)docker start "$FAILED_PRIMARY"# Allow Sentinel time to rediscover/reconfigure it.export REDISCLI_AUTH=AtlasMart-Ch20-Admin-Lab-Only-2026redis-cli -h 127.0.0.1 -p 6411 --user academy-admin ROLEunset REDISCLI_AUTH# Expected eventual role: replica of the new primary, not an independent writable primary.

8. Deliberate edge case: replica-priority 0

A replica with replica-priority 0 may replicate normally but is never selected for promotion. This can be useful for a reporting replica that must not become primary. It also means “I have two replicas” may actually mean “only one promotion candidate.” Count eligible failure domains, not just node count.

9. What the state machine does not solve

A successful +switch-master event proves a topology transition, not zero data loss, application recovery, DNS/TLS correctness, or write idempotency. Those are separate observable contracts. Lesson 4 focuses on client rediscovery; Lesson 5 combines them in a game day.

Check your understanding

  1. What happens if replica-priority is zero?
  2. Why can replica A lose despite having lower priority?
  3. What does a configuration epoch protect?
  4. Why restart the old primary after failover?
Review the answers

That replica remains usable as a replica but Sentinel will not promote it to primary.

It can be excluded as stale/unreachable; priority is considered only among suitable replicas.

It versions/authorizes a failover topology so Sentinel leaders coordinate configuration changes; it is not a per-write data-consistency epoch.

To verify Sentinel reconfigures it as a replica of the new primary and the topology converges rather than leaving an old writable primary.

10. Reset for the next lesson

If replica A was promoted, leave that topology in place. Sentinel clients should not care which original container is primary; that is exactly what Lesson 4 will prove. If the lab is inconsistent, destroy only the Chapter 20 topology and recreate it from Lesson 2.

Summary and next step

Leader Election, Failover State Machine, Replica Promotion, and Reconfiguration is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Client Discovery Through Sentinel, Reconnect Behavior, DNS/NAT/Container Considerations.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.