Chapter 20 · Redis Sentinel: Monitoring, Automatic Failover, and Service Discovery
Sentinel Architecture, Quorum vs Majority, Objective Down, and Failure Detection
Understand Sentinel monitoring, SDOWN/ODOWN, quorum, majority authorization, and why a robust deployment needs multiple independent Sentinels.
Learning outcomes
By the end of this lesson, you should be able to:
Describe Sentinel as monitoring, notification, automatic failover, and service discovery for non-Cluster Redis.
Distinguish SDOWN (one Sentinel's local judgment) from ODOWN (quorum agreement about a master).
Explain why configured quorum and failover majority authorization are separate thresholds.
Inspect Sentinel health with SENTINEL MASTER,
REPLICAS, SENTINELS, and
CKQUORUM.
Reject the misleading claim that “one Sentinel makes Redis highly available.”
Redis Open Source 8.10.1 using
redis:8.10.1; three Redis data nodes and three
Sentinel processes on one private Docker network; logical
database 0; AOF everysec on data nodes; Redis
Cluster is not enabled; TLS is off only because the mandatory
lab is single-host and Docker-private with host ports bound to
127.0.0.1; Redis ACLs protect both data-node and
Sentinel control connections; all passwords are disposable lab
values; no Search/JSON/vector/time-series/probabilistic feature
is required. Failure injection is confined to the named Chapter
20 topology; synthetic application fixtures use
atlasmart:ch20:*.
1. Practical problem: monitoring a replica set is not yet automatic recovery
Chapter 19 gave AtlasMart one primary and two replicas. That topology can copy data, but an application hard-coded to the failed primary still stops. Sentinel adds a distributed control plane around non-Cluster Redis: it monitors instances, shares observations, coordinates promotion, reconfigures replicas, and answers clients asking which node is primary now.
Sentinel does not turn asynchronous replication into synchronous consensus. It decides which reachable replica should become primary; it cannot invent writes that never reached that replica. That distinction is the durability boundary for this chapter.
2. Mental model: three judgments, not one “down” bit
Each Sentinel independently sends health probes. If one Sentinel
cannot get an acceptable response for the configured
down-after-milliseconds, that Sentinel marks the
instance Subjectively Down (SDOWN). For a
master, SDOWN can escalate to
Objectively Down (ODOWN) only when enough
Sentinels report the master unreachable to satisfy the
configured quorum.
ODOWN is still not authorization to rewrite topology. A Sentinel must be elected for a configuration epoch, which requires votes from a majority of known Sentinels (or a larger configured quorum if quorum itself exceeds the majority). Quorum answers “is the master objectively down?”; majority answers “may one Sentinel perform this failover?”
| State / threshold | Scope | What it permits | What it does not prove |
|---|---|---|---|
| SDOWN | Local to one Sentinel | Records its own failure suspicion | Other Sentinels agree |
| ODOWN | Master-level, quorum-backed | Allows failover attempt to be triggered | Leader has majority authorization |
| Quorum | Configured per monitored master | Number needed for ODOWN | Always equals majority |
| Majority authorization | Sentinel voting/election | Authorizes one failover leader for an epoch | Zero data loss |
3. The six-process AtlasMart topology
The lab deliberately runs three Sentinels,
because a one-Sentinel demo cannot teach majority authorization
or survive a Sentinel-process failure. The topology is
atlasmart-redis-ch20-primary, replicas
...-replica-a/...-replica-b, and
Sentinels ...-sentinel-1 through
...-sentinel-3 on Docker network
atlasmart-redis-ch20-net. Host mappings are only
for human inspection: Redis data ports
6411–6413 and Sentinel ports
26401–26403. Sentinel-aware application code runs
on the same Docker network so the names announced by Sentinel
are routable without pretending host port-remapping is
transparent.
For a robust local learning topology, the three Sentinels share no process with each other even though they run on one physical computer. That teaches process-level majority semantics but is not a true failure-domain design: one laptop, Docker daemon, power source, and network remain a common failure domain.
4. Observe a healthy Sentinel view before injecting failure
Sentinel listens on port 26379 by default and speaks the Redis
protocol, so redis-cli can inspect it. The
restricted Sentinel client ACL user is used for read-only
discovery commands; administration uses the dedicated lab
Sentinel administrator.
export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026for p in 26401 26402 26403; do echo "=== Sentinel $p ===" redis-cli -h 127.0.0.1 -p "$p" --user sentinel-client SENTINEL MASTER atlasmart-primary redis-cli -h 127.0.0.1 -p "$p" --user sentinel-client SENTINEL REPLICAS atlasmart-primary redis-cli -h 127.0.0.1 -p "$p" --user sentinel-client SENTINEL SENTINELS atlasmart-primarydoneunset REDISCLI_AUTH
5. Evidence: flags, peer counts, and CKQUORUM
A healthy master entry should include the
master flag without s_down or
o_down, two discovered replicas, and two “other”
Sentinels from each Sentinel's perspective.
SENTINEL CKQUORUM atlasmart-primary checks both
whether the configured quorum is reachable and whether enough
Sentinels exist to authorize a failover. It is a
deployment-health check, not a data-durability check.
export REDISCLI_AUTH=AtlasMart-Ch20-SentinelAdmin-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-admin SENTINEL CKQUORUM atlasmart-primaryunset REDISCLI_AUTH# Expected when all three Sentinels are mutually visible: OK ... usable Sentinels.
6. Deliberate failure: quorum 2 is not the same as a majority of 3
Stop two Sentinel processes but leave all Redis data nodes running. The remaining Sentinel can still make a local SDOWN observation if a Redis node fails later, but it cannot obtain a majority of three Sentinel votes. This is the concrete reason “quorum = 2” does not mean one surviving Sentinel can perform a failover.
docker stop atlasmart-redis-ch20-sentinel-2 atlasmart-redis-ch20-sentinel-3export REDISCLI_AUTH=AtlasMart-Ch20-SentinelAdmin-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-admin SENTINEL CKQUORUM atlasmart-primaryunset REDISCLI_AUTH# Expected: an error explaining that quorum and/or majority authorization cannot be reached.docker start atlasmart-redis-ch20-sentinel-2 atlasmart-redis-ch20-sentinel-3
It proves that Sentinel availability depends on the Sentinel voting topology as well as Redis data nodes. It does not prove that failover will preserve every acknowledged write, that client discovery works, or that these six containers are independent failure domains. Those are separate tests later in the chapter.
7. Wrong approach: “run one Sentinel next to the primary”
A single Sentinel can monitor and report a problem, but it has no independent peer majority. Co-locating the only Sentinel with the primary also creates an obvious correlated failure. Redis documentation recommends at least three Sentinel instances placed in independent failure domains for robust deployments.
The repair is not merely “add two processes.” Place Sentinels
where failures are independent enough, verify
CKQUORUM, verify peer discovery, and test
application reconnection. Sentinel high availability is an
end-to-end property of Redis nodes, Sentinels, networking,
credentials, and clients.
8. Production judgment
Sentinel fits non-Cluster Redis when one writable primary plus replicas is the desired topology and the application can tolerate failover/reconnect behavior. It does not shard data. Redis Cluster, managed service control planes, or an external orchestrator solve different topology problems.
Treat down-after-milliseconds, quorum, and failover
timeouts as failure-detector parameters to measure against real
latency and failure domains—not folklore constants. Overly
aggressive thresholds can cause false suspicions; overly slow
thresholds increase unavailability.
Check your understanding
- What is the difference between SDOWN and ODOWN?
- Why can quorum 2 and majority 2 both matter in a three-Sentinel deployment?
- Does CKQUORUM prove that no acknowledged writes will be lost?
- Why is six containers on one laptop not six independent failure domains?
Review the answers
SDOWN is one Sentinel’s local failure judgment; ODOWN is a master failure state backed by enough Sentinel reports to satisfy the configured quorum.
Quorum 2 establishes ODOWN; a majority of the three known Sentinels must authorize one leader to execute failover.
No. It checks Sentinel failover governance reachability, not replication durability of a particular write.
The Docker host, daemon, power, storage, and network remain shared dependencies.
9. Lab verification and cleanup
Before moving on, restore all three Sentinels, verify each sees
two peers and two replicas, and confirm CKQUORUM is
healthy. Keep the topology running for Lessons 2–5.
Summary and next step
Sentinel Architecture, Quorum vs Majority, Objective Down, and Failure Detection is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Configure Multiple Sentinels, Monitored Primary, Replicas, Authentication, and Networking.