Chapter 20 · Redis Sentinel: Monitoring, Automatic Failover, and Service Discovery

Client Discovery Through Sentinel, Reconnect Behavior, DNS/NAT/Container Considerations

Use a Sentinel-aware redis-py client to rediscover the promoted primary and measure reconnect behavior across Docker/DNS/NAT boundaries.

Advanced180–260 minutesredis-py Sentinel, service discovery, reconnect, DNS/NATRedis Open Source 8.10.1Docker + redis-cli + redis-py 8.1.06-process isolated Sentinel topologyFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

By the end of this lesson, you should be able to:

01

Explain why applications must resolve the current primary through Sentinel after reconnect/failure.

02

Use redis-py Sentinel discovery and master_for() rather than hard-coding a data-node address.

03

Separate Sentinel credentials from data-node credentials.

04

Reason about Docker/NAT/DNS/TLS name reachability as part of failover correctness.

05

Record actual application outage/recovery traces instead of equating server promotion with application availability.

Reproducible Chapter 20 baseline

Redis Open Source 8.10.1 using redis:8.10.1; three Redis data nodes and three Sentinel processes on one private Docker network; logical database 0; AOF everysec on data nodes; Redis Cluster is not enabled; TLS is off only because the mandatory lab is single-host and Docker-private with host ports bound to 127.0.0.1; Redis ACLs protect both data-node and Sentinel control connections; all passwords are disposable lab values; no Search/JSON/vector/time-series/probabilistic feature is required. Failure injection is confined to the named Chapter 20 topology; synthetic application fixtures use atlasmart:ch20:*.

1. Practical problem: Sentinel can succeed while the application still fails

Suppose Sentinel promotes replica A successfully. An application that reconnects to 127.0.0.1:6411 forever still fails because 6411 belongs to the old container mapping. Sentinel HA requires a client that knows the service name, asks Sentinels for the current primary, validates the result, and repeats discovery after reconnect events.

2. Sentinel client discovery protocol

A Sentinel-aware client should try known Sentinel endpoints, ask SENTINEL get-master-addr-by-name, verify the returned Redis instance is actually primary, and connect. On connection loss/timeout it should rediscover rather than blindly reconnect to the cached old endpoint. Redis Sentinel also disconnects normal clients from reconfigured instances during failover so stale connections are less likely to keep using a node whose role changed.

3. redis-py: separate Sentinel and Redis connection arguments

redis-py provides a Sentinel abstraction. sentinel_kwargs apply to Sentinel control-plane connections; the ordinary username/password parameters are passed to Redis data-node connections created by master_for(). Mixing them is a common authentication bug.

Python · Sentinel-aware AtlasMart writer
import timefrom redis.sentinel import Sentinelfrom redis.exceptions import ConnectionError, TimeoutErrorSENTINELS = [    ("atlasmart-redis-ch20-sentinel-1", 26379),    ("atlasmart-redis-ch20-sentinel-2", 26379),    ("atlasmart-redis-ch20-sentinel-3", 26379),]sentinel = Sentinel(    SENTINELS,    min_other_sentinels=1,    sentinel_kwargs={        "username": "sentinel-client",        "password": "AtlasMart-Ch20-SentinelClient-Lab-Only-2026",        "socket_timeout": 0.5,    },    username="atlasmart-app",    password="AtlasMart-Ch20-App-Lab-Only-2026",    socket_timeout=0.5,    socket_connect_timeout=0.5,    decode_responses=True,)client = sentinel.master_for("atlasmart-primary")print("initial-discovery", sentinel.discover_master("atlasmart-primary"))seq = 0while seq < 120:    seq += 1    try:        value = client.incr("atlasmart:ch20:client-sequence")        current = sentinel.discover_master("atlasmart-primary")        print(seq, "ok", value, current)    except (ConnectionError, TimeoutError) as exc:        print(seq, "transient", type(exc).__name__)    time.sleep(0.25)

4. Run the client in the same routed namespace as announced names

The lab Sentinels announce Docker hostnames, so run this test client on atlasmart-redis-ch20-net. One free path is a temporary Python container. It may need internet access once to install the pinned redis-py package; if your environment is offline, install redis==8.1.0 into an existing local Python environment that can resolve/reach the same announced endpoints.

Shell · run the pinned client container
docker run --rm --name atlasmart-redis-ch20-client \  --network atlasmart-redis-ch20-net \  -v "$PWD:/work" -w /work python:3.13-slim \  sh -lc 'pip install --disable-pip-version-check redis==8.1.0 && python sentinel_client.py'# Windows PowerShell can use the same Docker image with an absolute bind path.

5. Failure test: keep the client running while the primary stops

Start the client, then stop whichever node Sentinel currently reports as primary. Record the last successful operation before failure, connection/timeout errors during the outage, the first operation against the new primary, and discovered endpoint before/after. Do not call server promotion “application recovery” until this loop resumes.

Shell · discover first, then stop only that current primary container
export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-client SENTINEL get-master-addr-by-name atlasmart-primaryunset REDISCLI_AUTHredis-cli -h 127.0.0.1 -p 26401 --user sentinel-client --raw SENTINEL get-master-addr-by-name atlasmart-primary > ch20-current-master.txtFAILED_PRIMARY=$(head -n 1 ch20-current-master.txt)echo "stopping $FAILED_PRIMARY"docker stop "$FAILED_PRIMARY"

6. DNS/NAT/container considerations

Redis Sentinel hostname support is optional and must be configured consistently. With resolve-hostnames yes Sentinel may resolve hostnames; with announce-hostnames yes it can return names to clients. Some clients or TLS certificate policies may require hostnames; others historically expected IPs. Test the exact client/version rather than assuming.

Port remapping is particularly dangerous because a Sentinel may truthfully advertise port 6379 for the Redis process while the host exposes it as 6412. The correct fix is a routable topology/address plan, not ad-hoc string replacement in application code.

7. Deliberately wrong approach: cache the first discovered primary forever

Discovery is not a one-time bootstrapping step. Redis client guidelines say reconnection should resolve the service again through Sentinel. A cached address can point to a dead node or a node that has been demoted to replica. Repair: treat Sentinel as the source of current topology at connection establishment/re-establishment and bound retry/timeouts within an application-level budget.

8. TLS and managed-service boundary

This localhost Docker lab uses no TLS; production should normally protect both application↔Redis and control-plane paths according to the deployment. Hostname announcement can become part of certificate name validation. Managed Redis products may expose their own endpoint/failover API rather than raw Sentinel, so verify provider semantics rather than assuming Sentinel ports/commands exist.

9. Observability: client trace plus Sentinel events

Correlate application timestamps with Sentinel +sdown, +odown, election, +switch-master, and replica reconfiguration logs. The gap between server promotion and first successful application write is the client-visible recovery interval. That interval includes detection, election, promotion, endpoint rediscovery, socket creation, authentication, and command retry policy.

Check your understanding

  1. Why should a client rediscover after a socket error?
  2. Why are sentinel_kwargs separate from normal Redis credentials in redis-py?
  3. Why is host port 6412 not necessarily an address Sentinel should announce?
  4. What measurement defines application recovery?
Review the answers

The old address may no longer be the primary after failover; Sentinel is the current configuration provider.

They authenticate/control connections to Sentinel itself, while the normal connection kwargs authenticate to the discovered Redis data node.

Inside the Docker topology the Redis process listens on 6379; remapped host ports are not automatically routable from Sentinel peers or network-local clients.

The first successful application operation after failure/reconnect, correlated with discovery and server failover—not merely the Sentinel switch-master timestamp.

10. Verification and reset

Verify the client discovers a different primary after failover and resumes writes. Restart the failed node and wait for Sentinel to reconfigure it as a replica before proceeding. Keep the topology for the final game day.

Summary and next step

Client Discovery Through Sentinel, Reconnect Behavior, DNS/NAT/Container Considerations is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Run a Sentinel Failover Game Day and Verify Application Recovery and Data-Loss Bounds.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.