Chapter 20 · Redis Sentinel: Monitoring, Automatic Failover, and Service Discovery
Run a Sentinel Failover Game Day and Verify Application Recovery and Data-Loss Bounds
Run an end-to-end Sentinel failover game day and measure application recovery and acknowledged-write bounds without claiming zero-loss failover.
Learning outcomes
By the end of this lesson, you should be able to:
Run a bounded Sentinel failure game day with predeclared blast radius and reset path.
Measure detection, failover, client recovery, and replica reconfiguration as separate intervals.
Compare acknowledged client operations with state present after promotion and recovery.
Report actual availability/data-loss evidence without claiming asynchronous replication provides zero-loss failover.
Convert the exercise into production acceptance criteria and recurring operational tests.
Redis Open Source 8.10.1 using
redis:8.10.1; three Redis data nodes and three
Sentinel processes on one private Docker network; logical
database 0; AOF everysec on data nodes; Redis
Cluster is not enabled; TLS is off only because the mandatory
lab is single-host and Docker-private with host ports bound to
127.0.0.1; Redis ACLs protect both data-node and
Sentinel control connections; all passwords are disposable lab
values; no Search/JSON/vector/time-series/probabilistic feature
is required. Failure injection is confined to the named Chapter
20 topology; synthetic application fixtures use
atlasmart:ch20:*.
1. Game-day objective: prove the whole chain, not only Sentinel promotion
AtlasMart now tests one end-to-end claim: when the current primary process disappears, a three-Sentinel control plane should detect failure, authorize one failover, promote an eligible replica, publish the new primary, let a Sentinel-aware client reconnect, reconfigure remaining/returning replicas, and converge on one writable primary. The drill also records which client operations were acknowledged and which values are present after recovery.
2. Pre-flight acceptance criteria
Do not start failure injection until all checks are green.
- Three Redis data nodes: exactly one primary, two connected replicas.
- Three Sentinels: each knows two peer Sentinels and two replicas.
-
SENTINEL CKQUORUM atlasmart-primarysucceeds. -
Sentinel-aware client can write
atlasmart:ch20:gameday:ack-seq. - Redis and Sentinel ACL authentication is working; no default user is enabled.
- AOF everysec is enabled on data nodes; this is persistence, not a zero-loss failover guarantee.
- The operator knows the exact current primary container before issuing the stop command.
3. Instrument the client before failure
The client loop records per-operation success/failure, returned counter value, discovered primary, and latency. The output is evidence; this course does not provide fabricated p50/p95/p99 or loss counts because they depend on the learner's host, Docker scheduler, replication lag, and timing.
import json, timefrom redis.sentinel import Sentinelfrom redis.exceptions import RedisErrors = Sentinel( [("atlasmart-redis-ch20-sentinel-1",26379), ("atlasmart-redis-ch20-sentinel-2",26379), ("atlasmart-redis-ch20-sentinel-3",26379)], min_other_sentinels=1, sentinel_kwargs={"username":"sentinel-client","password":"AtlasMart-Ch20-SentinelClient-Lab-Only-2026","socket_timeout":0.5}, username="atlasmart-app", password="AtlasMart-Ch20-App-Lab-Only-2026", socket_timeout=0.5, socket_connect_timeout=0.5, decode_responses=True)r = s.master_for("atlasmart-primary")records=[]start=time.perf_counter()for i in range(240): t=time.perf_counter() try: value=r.incr("atlasmart:ch20:gameday:ack-seq") records.append({"i":i,"ok":True,"value":value,"master":s.discover_master("atlasmart-primary"),"ms":round((time.perf_counter()-t)*1000,3)}) except RedisError as exc: records.append({"i":i,"ok":False,"error":type(exc).__name__,"ms":round((time.perf_counter()-t)*1000,3)}) time.sleep(0.1)print(json.dumps({"elapsed_s":round(time.perf_counter()-start,3),"records":records}, indent=2))
4. Record Sentinel and Redis baselines
Capture SENTINEL MASTER, REPLICAS,
SENTINELS, data-node ROLE/INFO replication, and the pre-failure counter. Save logs with timestamps. These
samples let you explain the transition after the fact.
mkdir -p ch20-evidenceexport REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-client SENTINEL MASTER atlasmart-primary > ch20-evidence/sentinel-master-before.txtredis-cli -h 127.0.0.1 -p 26401 --user sentinel-client SENTINEL REPLICAS atlasmart-primary > ch20-evidence/sentinel-replicas-before.txtunset REDISCLI_AUTHdocker logs --since 2m atlasmart-redis-ch20-sentinel-1 > ch20-evidence/sentinel1-before.log 2>&1
5. Inject exactly one failure: stop the current primary
This is a process failure, not host chaos. Stop only the current primary container while the client writer is active. Do not stop replicas, Sentinels, Docker networking, or the host clock.
export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-client SENTINEL get-master-addr-by-name atlasmart-primaryunset REDISCLI_AUTHredis-cli -h 127.0.0.1 -p 26401 --user sentinel-client --raw SENTINEL get-master-addr-by-name atlasmart-primary > ch20-evidence/failed-primary.txtFAILED_PRIMARY=$(head -n 1 ch20-evidence/failed-primary.txt)echo "stopping $FAILED_PRIMARY"docker stop "$FAILED_PRIMARY"
6. Observe four clocks
Do not compress failover into one number.
| Interval | Start | End | Why it matters |
|---|---|---|---|
| Failure detection | Primary stops responding | Sentinel SDOWN/ODOWN | Failure detector sensitivity |
| Control-plane failover | ODOWN/election begins | +switch-master | Election + promotion time |
| Application recovery | First client failure | First later successful write | User-visible outage |
| Topology convergence | New primary selected | Other/returned nodes become replicas | Steady-state recovery |
7. Measure data-loss bounds with acknowledged application evidence
A counter sequence can reveal gaps or rollbacks after promotion, but interpret it carefully. The client only knows a write was acknowledged by the primary that served it; default asynchronous replication means an acknowledged write may fail to reach the promoted replica before the primary disappears. Sentinel does not guarantee zero acknowledged-write loss.
If your application requires tighter bounds, combine
design-level measures such as replica lag admission
(min-replicas-to-write), same-connection
WAIT/WAITAOF where appropriate,
persistence, idempotent request semantics, and business
reconciliation. None of those turns Sentinel into a universal
linearizable consensus database.
8. Bring the old primary back and verify convergence
Restart the failed node and wait for Sentinel to impose the current topology. Verify it becomes a replica and catches up. Then check that all Sentinels report the same primary and that the client continues using that service without hard-coded address changes.
FAILED_PRIMARY=$(head -n 1 ch20-evidence/failed-primary.txt)docker start "$FAILED_PRIMARY"sleep 5export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026for p in 26401 26402 26403; do redis-cli -h 127.0.0.1 -p "$p" --user sentinel-client SENTINEL get-master-addr-by-name atlasmart-primary; doneunset REDISCLI_AUTHexport REDISCLI_AUTH=AtlasMart-Ch20-Admin-Lab-Only-2026for p in 6411 6412 6413; do echo "=== $p ==="; redis-cli -h 127.0.0.1 -p "$p" --user academy-admin ROLE; doneunset REDISCLI_AUTH
9. Deliberately wrong conclusion: “failover worked, therefore HA is done”
A green Sentinel promotion does not cover independent failure
domains, TLS/certificate renewal, DNS/NAT reachability,
credential rotation, client retry semantics, replica freshness
policy, persistence/backup/restore, capacity during promotion,
or application reconciliation. Production HA is an operational
system, not a screenshot of +switch-master.
10. Production acceptance report
Record actual measured values and environment rather than copying targets from this lesson.
- Redis/redis-cli/redis-py exact versions and container image digest if available.
- Data-node and Sentinel placement/failure domains.
- Detection, promotion, application-recovery, and topology-convergence distributions across repeated drills.
- Acknowledged-operation comparison before/after failover and any observed gaps.
- Replica selection evidence and maximum observed lag before failure.
- Client reconnect/retry errors and whether any operation may have been duplicated.
- Sentinel quorum/majority health, auth/TLS state, DNS/NAT assumptions, and monitoring alerts.
- Restore/backup evidence from Chapter 17—because replica failover is still not backup.
11. Cleanup/reset
After capturing evidence, remove only the named Chapter 20 lab. If you want to repeat the game day, recreate it from Lesson 2 to avoid carrying mutated Sentinel configuration epochs into a supposedly fresh baseline.
docker compose -f ch20-compose.yaml down -v --remove-orphansdocker network rm atlasmart-redis-ch20-net 2>/dev/null || true# Do not run docker system prune; cleanup is deliberately name-scoped.
12. Bridge to Chapter 21
Sentinel gives automatic failover for a non-sharded primary/replica topology. Chapter 21 changes the data-placement model itself: Redis Cluster partitions keys across 16,384 hash slots, introduces MOVED/ASK routing and Cluster-aware clients, and has different availability/failover rules. Do not treat Sentinel and Cluster as interchangeable control planes.
Check your understanding
- Why record acknowledged client operations rather than only Redis key counts?
- Which interval should define user-visible availability?
- Does a successful switch-master prove backup readiness?
- Why rebuild the topology before a fresh experiment?
Review the answers
The business question is what the client was told succeeded; asynchronous replication can lose some acknowledged writes during failure.
The application recovery interval—from failed client operations until the first successful operation after rediscovery/reconnect.
No. Replication/failover and backup/restore are separate controls.
Sentinel rewrites configuration and epochs during failover; rebuilding gives a controlled baseline rather than hidden state from the previous game day.
Summary and next step
Run a Sentinel Failover Game Day and Verify Application Recovery and Data-Loss Bounds is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with 16384 Hash Slots, CRC16, Key Distribution, Nodes, Primaries, Replicas, and Cluster Bus.