Chapter 20 · Redis Sentinel: Monitoring, Automatic Failover, and Service Discovery

Run a Sentinel Failover Game Day and Verify Application Recovery and Data-Loss Bounds

Run an end-to-end Sentinel failover game day and measure application recovery and acknowledged-write bounds without claiming zero-loss failover.

Advanced220–320 minutesgame day, recovery timing, acknowledged-write bounds, reconciliationRedis Open Source 8.10.1Docker + redis-cli + redis-py 8.1.06-process isolated Sentinel topologyFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

By the end of this lesson, you should be able to:

01

Run a bounded Sentinel failure game day with predeclared blast radius and reset path.

02

Measure detection, failover, client recovery, and replica reconfiguration as separate intervals.

03

Compare acknowledged client operations with state present after promotion and recovery.

04

Report actual availability/data-loss evidence without claiming asynchronous replication provides zero-loss failover.

05

Convert the exercise into production acceptance criteria and recurring operational tests.

Reproducible Chapter 20 baseline

Redis Open Source 8.10.1 using redis:8.10.1; three Redis data nodes and three Sentinel processes on one private Docker network; logical database 0; AOF everysec on data nodes; Redis Cluster is not enabled; TLS is off only because the mandatory lab is single-host and Docker-private with host ports bound to 127.0.0.1; Redis ACLs protect both data-node and Sentinel control connections; all passwords are disposable lab values; no Search/JSON/vector/time-series/probabilistic feature is required. Failure injection is confined to the named Chapter 20 topology; synthetic application fixtures use atlasmart:ch20:*.

1. Game-day objective: prove the whole chain, not only Sentinel promotion

AtlasMart now tests one end-to-end claim: when the current primary process disappears, a three-Sentinel control plane should detect failure, authorize one failover, promote an eligible replica, publish the new primary, let a Sentinel-aware client reconnect, reconfigure remaining/returning replicas, and converge on one writable primary. The drill also records which client operations were acknowledged and which values are present after recovery.

2. Pre-flight acceptance criteria

Do not start failure injection until all checks are green.

  • Three Redis data nodes: exactly one primary, two connected replicas.
  • Three Sentinels: each knows two peer Sentinels and two replicas.
  • SENTINEL CKQUORUM atlasmart-primary succeeds.
  • Sentinel-aware client can write atlasmart:ch20:gameday:ack-seq.
  • Redis and Sentinel ACL authentication is working; no default user is enabled.
  • AOF everysec is enabled on data nodes; this is persistence, not a zero-loss failover guarantee.
  • The operator knows the exact current primary container before issuing the stop command.

3. Instrument the client before failure

The client loop records per-operation success/failure, returned counter value, discovered primary, and latency. The output is evidence; this course does not provide fabricated p50/p95/p99 or loss counts because they depend on the learner's host, Docker scheduler, replication lag, and timing.

Python · bounded game-day writer
import json, timefrom redis.sentinel import Sentinelfrom redis.exceptions import RedisErrors = Sentinel(    [("atlasmart-redis-ch20-sentinel-1",26379),     ("atlasmart-redis-ch20-sentinel-2",26379),     ("atlasmart-redis-ch20-sentinel-3",26379)],    min_other_sentinels=1,    sentinel_kwargs={"username":"sentinel-client","password":"AtlasMart-Ch20-SentinelClient-Lab-Only-2026","socket_timeout":0.5},    username="atlasmart-app", password="AtlasMart-Ch20-App-Lab-Only-2026",    socket_timeout=0.5, socket_connect_timeout=0.5, decode_responses=True)r = s.master_for("atlasmart-primary")records=[]start=time.perf_counter()for i in range(240):    t=time.perf_counter()    try:        value=r.incr("atlasmart:ch20:gameday:ack-seq")        records.append({"i":i,"ok":True,"value":value,"master":s.discover_master("atlasmart-primary"),"ms":round((time.perf_counter()-t)*1000,3)})    except RedisError as exc:        records.append({"i":i,"ok":False,"error":type(exc).__name__,"ms":round((time.perf_counter()-t)*1000,3)})    time.sleep(0.1)print(json.dumps({"elapsed_s":round(time.perf_counter()-start,3),"records":records}, indent=2))

4. Record Sentinel and Redis baselines

Capture SENTINEL MASTER, REPLICAS, SENTINELS, data-node ROLE/INFO replication, and the pre-failure counter. Save logs with timestamps. These samples let you explain the transition after the fact.

Shell · baseline snapshots
mkdir -p ch20-evidenceexport REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-client SENTINEL MASTER atlasmart-primary > ch20-evidence/sentinel-master-before.txtredis-cli -h 127.0.0.1 -p 26401 --user sentinel-client SENTINEL REPLICAS atlasmart-primary > ch20-evidence/sentinel-replicas-before.txtunset REDISCLI_AUTHdocker logs --since 2m atlasmart-redis-ch20-sentinel-1 > ch20-evidence/sentinel1-before.log 2>&1

5. Inject exactly one failure: stop the current primary

This is a process failure, not host chaos. Stop only the current primary container while the client writer is active. Do not stop replicas, Sentinels, Docker networking, or the host clock.

Shell · controlled failure
export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-client SENTINEL get-master-addr-by-name atlasmart-primaryunset REDISCLI_AUTHredis-cli -h 127.0.0.1 -p 26401 --user sentinel-client --raw SENTINEL get-master-addr-by-name atlasmart-primary > ch20-evidence/failed-primary.txtFAILED_PRIMARY=$(head -n 1 ch20-evidence/failed-primary.txt)echo "stopping $FAILED_PRIMARY"docker stop "$FAILED_PRIMARY"

6. Observe four clocks

Do not compress failover into one number.

Interval Start End Why it matters
Failure detection Primary stops responding Sentinel SDOWN/ODOWN Failure detector sensitivity
Control-plane failover ODOWN/election begins +switch-master Election + promotion time
Application recovery First client failure First later successful write User-visible outage
Topology convergence New primary selected Other/returned nodes become replicas Steady-state recovery

7. Measure data-loss bounds with acknowledged application evidence

A counter sequence can reveal gaps or rollbacks after promotion, but interpret it carefully. The client only knows a write was acknowledged by the primary that served it; default asynchronous replication means an acknowledged write may fail to reach the promoted replica before the primary disappears. Sentinel does not guarantee zero acknowledged-write loss.

If your application requires tighter bounds, combine design-level measures such as replica lag admission (min-replicas-to-write), same-connection WAIT/WAITAOF where appropriate, persistence, idempotent request semantics, and business reconciliation. None of those turns Sentinel into a universal linearizable consensus database.

8. Bring the old primary back and verify convergence

Restart the failed node and wait for Sentinel to impose the current topology. Verify it becomes a replica and catches up. Then check that all Sentinels report the same primary and that the client continues using that service without hard-coded address changes.

Shell · recovery proof
FAILED_PRIMARY=$(head -n 1 ch20-evidence/failed-primary.txt)docker start "$FAILED_PRIMARY"sleep 5export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026for p in 26401 26402 26403; do redis-cli -h 127.0.0.1 -p "$p" --user sentinel-client SENTINEL get-master-addr-by-name atlasmart-primary; doneunset REDISCLI_AUTHexport REDISCLI_AUTH=AtlasMart-Ch20-Admin-Lab-Only-2026for p in 6411 6412 6413; do echo "=== $p ==="; redis-cli -h 127.0.0.1 -p "$p" --user academy-admin ROLE; doneunset REDISCLI_AUTH

9. Deliberately wrong conclusion: “failover worked, therefore HA is done”

A green Sentinel promotion does not cover independent failure domains, TLS/certificate renewal, DNS/NAT reachability, credential rotation, client retry semantics, replica freshness policy, persistence/backup/restore, capacity during promotion, or application reconciliation. Production HA is an operational system, not a screenshot of +switch-master.

10. Production acceptance report

Record actual measured values and environment rather than copying targets from this lesson.

  • Redis/redis-cli/redis-py exact versions and container image digest if available.
  • Data-node and Sentinel placement/failure domains.
  • Detection, promotion, application-recovery, and topology-convergence distributions across repeated drills.
  • Acknowledged-operation comparison before/after failover and any observed gaps.
  • Replica selection evidence and maximum observed lag before failure.
  • Client reconnect/retry errors and whether any operation may have been duplicated.
  • Sentinel quorum/majority health, auth/TLS state, DNS/NAT assumptions, and monitoring alerts.
  • Restore/backup evidence from Chapter 17—because replica failover is still not backup.

11. Cleanup/reset

After capturing evidence, remove only the named Chapter 20 lab. If you want to repeat the game day, recreate it from Lesson 2 to avoid carrying mutated Sentinel configuration epochs into a supposedly fresh baseline.

Shell · remove only the Chapter 20 lab
docker compose -f ch20-compose.yaml down -v --remove-orphansdocker network rm atlasmart-redis-ch20-net 2>/dev/null || true# Do not run docker system prune; cleanup is deliberately name-scoped.

12. Bridge to Chapter 21

Sentinel gives automatic failover for a non-sharded primary/replica topology. Chapter 21 changes the data-placement model itself: Redis Cluster partitions keys across 16,384 hash slots, introduces MOVED/ASK routing and Cluster-aware clients, and has different availability/failover rules. Do not treat Sentinel and Cluster as interchangeable control planes.

Check your understanding

  1. Why record acknowledged client operations rather than only Redis key counts?
  2. Which interval should define user-visible availability?
  3. Does a successful switch-master prove backup readiness?
  4. Why rebuild the topology before a fresh experiment?
Review the answers

The business question is what the client was told succeeded; asynchronous replication can lose some acknowledged writes during failure.

The application recovery interval—from failed client operations until the first successful operation after rediscovery/reconnect.

No. Replication/failover and backup/restore are separate controls.

Sentinel rewrites configuration and epochs during failover; rebuilding gives a controlled baseline rather than hidden state from the previous game day.

Summary and next step

Run a Sentinel Failover Game Day and Verify Application Recovery and Data-Loss Bounds is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with 16384 Hash Slots, CRC16, Key Distribution, Nodes, Primaries, Replicas, and Cluster Bus.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.