Chapter 20 · Redis Sentinel: Monitoring, Automatic Failover, and Service Discovery
Leader Election, Failover State Machine, Replica Promotion, and Reconfiguration
Trace Sentinel election, replica selection, promotion, configuration epochs, and reconfiguration of replicas and the returning old primary.
Learning outcomes
By the end of this lesson, you should be able to:
Trace Sentinel failover from SDOWN/ODOWN through election, replica selection, promotion, and replica reconfiguration.
Explain configuration epochs and why majority authorization prevents simultaneous failovers in minority partitions.
Predict how replica freshness, priority, replication offset, and run ID influence promotion selection.
Inspect Sentinel event logs instead of treating failover as a black box.
Verify that the old primary is reconfigured as a replica when it returns.
Redis Open Source 8.10.1 using
redis:8.10.1; three Redis data nodes and three
Sentinel processes on one private Docker network; logical
database 0; AOF everysec on data nodes; Redis
Cluster is not enabled; TLS is off only because the mandatory
lab is single-host and Docker-private with host ports bound to
127.0.0.1; Redis ACLs protect both data-node and
Sentinel control connections; all passwords are disposable lab
values; no Search/JSON/vector/time-series/probabilistic feature
is required. Failure injection is confined to the named Chapter
20 topology; synthetic application fixtures use
atlasmart:ch20:*.
1. Practical problem: “Sentinel promoted something” is not enough evidence
AtlasMart operators must know which replica was promoted, why it was eligible, how long writes were unavailable, whether other replicas were reconfigured, and whether the previous primary returned safely as a replica. The failover state machine exposes those decisions in Sentinel logs and API state.
2. Election and configuration epochs
After ODOWN, Sentinels try to authorize exactly one failover leader for a configuration epoch. A majority vote gives that leader a unique epoch for the new topology. If a majority cannot communicate, Sentinel deliberately avoids performing a failover in that minority partition.
This voting protects topology configuration, not application data linearizability. Asynchronous replication remains the source of any acknowledged-write loss window.
3. How Sentinel chooses a replica
Sentinel first rejects replicas considered too stale/unreliable.
Among eligible replicas it sorts by
replica-priority (lower is preferred, zero means
never promote), then by replication offset (more processed
history is preferred), then by run ID as a deterministic
tie-breaker.
| Criterion | Preference | Operational implication |
|---|---|---|
| Connectivity/freshness | Eligible/recent enough | A disconnected replica can be excluded |
| replica-priority | Lower nonzero wins | Operator expresses placement preference |
| Replication offset | Higher wins if priority ties | Fresher processed stream is preferred |
| Run ID | Lexicographic tie-break | Deterministic, not a business preference |
4. Capture the pre-failure baseline
Write bounded AtlasMart state, wait for both replicas to catch up, then record the current primary and replica priorities. This is necessary to reason about the later promotion rather than guessing from container names.
export REDISCLI_AUTH=AtlasMart-Ch20-Admin-Lab-Only-2026printf 'SET atlasmart:ch20:order:9001 state:paidINCR atlasmart:ch20:sequenceWAIT 2 5000' | redis-cli -h 127.0.0.1 -p 6411 --user academy-admin# SET, INCR, and WAIT above use one redis-cli process and therefore one connection.redis-cli -h 127.0.0.1 -p 6412 --user academy-admin INFO replicationredis-cli -h 127.0.0.1 -p 6413 --user academy-admin INFO replicationunset REDISCLI_AUTH
5. Trigger the controlled failover by stopping only the primary
Stopping one named primary container is a bounded process
failure on the isolated lab. Do not stop the Docker daemon or
sever host networking. Start log tails first so you can see
+sdown, +odown,
+try-failover, +elected-leader,
replica selection, promotion, reconfiguration, and
+switch-master events.
docker logs --since 1m -f atlasmart-redis-ch20-sentinel-1# In a second terminal, resolve the exact current primary through Sentinel:export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-client --raw SENTINEL get-master-addr-by-name atlasmart-primary > ch20-current-master.txtunset REDISCLI_AUTHFAILED_PRIMARY=$(head -n 1 ch20-current-master.txt)echo "stopping $FAILED_PRIMARY"docker stop "$FAILED_PRIMARY"# Watch until SENTINEL get-master-addr-by-name atlasmart-primary changes.
6. Observe—not assume—the new primary
Query at least two Sentinels and the promoted Redis node. Replica A is expected to win because priority 50 is lower than replica B’s 100, but the lesson deliberately says expected: an ineligible/stale replica must not be promoted merely because its priority is lower.
export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026for p in 26401 26402 26403; do redis-cli -h 127.0.0.1 -p "$p" --user sentinel-client SENTINEL get-master-addr-by-name atlasmart-primarydoneunset REDISCLI_AUTHexport REDISCLI_AUTH=AtlasMart-Ch20-Admin-Lab-Only-2026redis-cli -h 127.0.0.1 -p 6412 --user academy-admin ROLEredis-cli -h 127.0.0.1 -p 6413 --user academy-admin ROLEunset REDISCLI_AUTH
7. Restart the old primary and watch Sentinel impose current topology
A failed-over primary does not automatically remain an
independent master forever. When the old primary becomes
reachable, Sentinels try to impose the current configuration and
reconfigure it as a replica of the new primary. Verify this with
ROLE and Sentinel logs instead of assuming
split-brain disappeared.
FAILED_PRIMARY=$(head -n 1 ch20-current-master.txt)docker start "$FAILED_PRIMARY"# Allow Sentinel time to rediscover/reconfigure it.export REDISCLI_AUTH=AtlasMart-Ch20-Admin-Lab-Only-2026redis-cli -h 127.0.0.1 -p 6411 --user academy-admin ROLEunset REDISCLI_AUTH# Expected eventual role: replica of the new primary, not an independent writable primary.
8. Deliberate edge case: replica-priority 0
A replica with replica-priority 0 may replicate
normally but is never selected for promotion. This can be useful
for a reporting replica that must not become primary. It also
means “I have two replicas” may actually mean “only one
promotion candidate.” Count eligible failure domains, not just
node count.
9. What the state machine does not solve
A successful +switch-master event proves a topology
transition, not zero data loss, application recovery, DNS/TLS
correctness, or write idempotency. Those are separate observable
contracts. Lesson 4 focuses on client rediscovery; Lesson 5
combines them in a game day.
Check your understanding
- What happens if replica-priority is zero?
- Why can replica A lose despite having lower priority?
- What does a configuration epoch protect?
- Why restart the old primary after failover?
Review the answers
That replica remains usable as a replica but Sentinel will not promote it to primary.
It can be excluded as stale/unreachable; priority is considered only among suitable replicas.
It versions/authorizes a failover topology so Sentinel leaders coordinate configuration changes; it is not a per-write data-consistency epoch.
To verify Sentinel reconfigures it as a replica of the new primary and the topology converges rather than leaving an old writable primary.
10. Reset for the next lesson
If replica A was promoted, leave that topology in place. Sentinel clients should not care which original container is primary; that is exactly what Lesson 4 will prove. If the lab is inconsistent, destroy only the Chapter 20 topology and recreate it from Lesson 2.
Summary and next step
Leader Election, Failover State Machine, Replica Promotion, and Reconfiguration is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Client Discovery Through Sentinel, Reconnect Behavior, DNS/NAT/Container Considerations.