Chapter 20 · Redis Sentinel: Monitoring, Automatic Failover, and Service Discovery
Client Discovery Through Sentinel, Reconnect Behavior, DNS/NAT/Container Considerations
Use a Sentinel-aware redis-py client to rediscover the promoted primary and measure reconnect behavior across Docker/DNS/NAT boundaries.
Learning outcomes
By the end of this lesson, you should be able to:
Explain why applications must resolve the current primary through Sentinel after reconnect/failure.
Use redis-py Sentinel discovery and
master_for() rather than hard-coding a
data-node address.
Separate Sentinel credentials from data-node credentials.
Reason about Docker/NAT/DNS/TLS name reachability as part of failover correctness.
Record actual application outage/recovery traces instead of equating server promotion with application availability.
Redis Open Source 8.10.1 using
redis:8.10.1; three Redis data nodes and three
Sentinel processes on one private Docker network; logical
database 0; AOF everysec on data nodes; Redis
Cluster is not enabled; TLS is off only because the mandatory
lab is single-host and Docker-private with host ports bound to
127.0.0.1; Redis ACLs protect both data-node and
Sentinel control connections; all passwords are disposable lab
values; no Search/JSON/vector/time-series/probabilistic feature
is required. Failure injection is confined to the named Chapter
20 topology; synthetic application fixtures use
atlasmart:ch20:*.
1. Practical problem: Sentinel can succeed while the application still fails
Suppose Sentinel promotes replica A successfully. An application
that reconnects to 127.0.0.1:6411 forever still
fails because 6411 belongs to the old container mapping.
Sentinel HA requires a client that knows the
service name, asks Sentinels for the current primary,
validates the result, and repeats discovery after reconnect
events.
2. Sentinel client discovery protocol
A Sentinel-aware client should try known Sentinel endpoints, ask
SENTINEL get-master-addr-by-name, verify the
returned Redis instance is actually primary, and connect. On
connection loss/timeout it should rediscover rather than blindly
reconnect to the cached old endpoint. Redis Sentinel also
disconnects normal clients from reconfigured instances during
failover so stale connections are less likely to keep using a
node whose role changed.
3. redis-py: separate Sentinel and Redis connection arguments
redis-py provides a Sentinel abstraction.
sentinel_kwargs apply to Sentinel control-plane
connections; the ordinary username/password parameters are
passed to Redis data-node connections created by
master_for(). Mixing them is a common
authentication bug.
import timefrom redis.sentinel import Sentinelfrom redis.exceptions import ConnectionError, TimeoutErrorSENTINELS = [ ("atlasmart-redis-ch20-sentinel-1", 26379), ("atlasmart-redis-ch20-sentinel-2", 26379), ("atlasmart-redis-ch20-sentinel-3", 26379),]sentinel = Sentinel( SENTINELS, min_other_sentinels=1, sentinel_kwargs={ "username": "sentinel-client", "password": "AtlasMart-Ch20-SentinelClient-Lab-Only-2026", "socket_timeout": 0.5, }, username="atlasmart-app", password="AtlasMart-Ch20-App-Lab-Only-2026", socket_timeout=0.5, socket_connect_timeout=0.5, decode_responses=True,)client = sentinel.master_for("atlasmart-primary")print("initial-discovery", sentinel.discover_master("atlasmart-primary"))seq = 0while seq < 120: seq += 1 try: value = client.incr("atlasmart:ch20:client-sequence") current = sentinel.discover_master("atlasmart-primary") print(seq, "ok", value, current) except (ConnectionError, TimeoutError) as exc: print(seq, "transient", type(exc).__name__) time.sleep(0.25)
4. Run the client in the same routed namespace as announced names
The lab Sentinels announce Docker hostnames, so run this test
client on atlasmart-redis-ch20-net. One free path
is a temporary Python container. It may need internet access
once to install the pinned redis-py package; if your environment
is offline, install redis==8.1.0 into an existing
local Python environment that can resolve/reach the same
announced endpoints.
docker run --rm --name atlasmart-redis-ch20-client \ --network atlasmart-redis-ch20-net \ -v "$PWD:/work" -w /work python:3.13-slim \ sh -lc 'pip install --disable-pip-version-check redis==8.1.0 && python sentinel_client.py'# Windows PowerShell can use the same Docker image with an absolute bind path.
5. Failure test: keep the client running while the primary stops
Start the client, then stop whichever node Sentinel currently reports as primary. Record the last successful operation before failure, connection/timeout errors during the outage, the first operation against the new primary, and discovered endpoint before/after. Do not call server promotion “application recovery” until this loop resumes.
export REDISCLI_AUTH=AtlasMart-Ch20-SentinelClient-Lab-Only-2026redis-cli -h 127.0.0.1 -p 26401 --user sentinel-client SENTINEL get-master-addr-by-name atlasmart-primaryunset REDISCLI_AUTHredis-cli -h 127.0.0.1 -p 26401 --user sentinel-client --raw SENTINEL get-master-addr-by-name atlasmart-primary > ch20-current-master.txtFAILED_PRIMARY=$(head -n 1 ch20-current-master.txt)echo "stopping $FAILED_PRIMARY"docker stop "$FAILED_PRIMARY"
6. DNS/NAT/container considerations
Redis Sentinel hostname support is optional and must be
configured consistently. With
resolve-hostnames yes Sentinel may resolve
hostnames; with announce-hostnames yes it can
return names to clients. Some clients or TLS certificate
policies may require hostnames; others historically expected
IPs. Test the exact client/version rather than assuming.
Port remapping is particularly dangerous because a Sentinel may truthfully advertise port 6379 for the Redis process while the host exposes it as 6412. The correct fix is a routable topology/address plan, not ad-hoc string replacement in application code.
7. Deliberately wrong approach: cache the first discovered primary forever
Discovery is not a one-time bootstrapping step. Redis client guidelines say reconnection should resolve the service again through Sentinel. A cached address can point to a dead node or a node that has been demoted to replica. Repair: treat Sentinel as the source of current topology at connection establishment/re-establishment and bound retry/timeouts within an application-level budget.
8. TLS and managed-service boundary
This localhost Docker lab uses no TLS; production should normally protect both application↔Redis and control-plane paths according to the deployment. Hostname announcement can become part of certificate name validation. Managed Redis products may expose their own endpoint/failover API rather than raw Sentinel, so verify provider semantics rather than assuming Sentinel ports/commands exist.
9. Observability: client trace plus Sentinel events
Correlate application timestamps with Sentinel
+sdown, +odown, election,
+switch-master, and replica reconfiguration logs.
The gap between server promotion and first successful
application write is the client-visible recovery interval. That
interval includes detection, election, promotion, endpoint
rediscovery, socket creation, authentication, and command retry
policy.
Check your understanding
- Why should a client rediscover after a socket error?
- Why are sentinel_kwargs separate from normal Redis credentials in redis-py?
- Why is host port 6412 not necessarily an address Sentinel should announce?
- What measurement defines application recovery?
Review the answers
The old address may no longer be the primary after failover; Sentinel is the current configuration provider.
They authenticate/control connections to Sentinel itself, while the normal connection kwargs authenticate to the discovered Redis data node.
Inside the Docker topology the Redis process listens on 6379; remapped host ports are not automatically routable from Sentinel peers or network-local clients.
The first successful application operation after failure/reconnect, correlated with discovery and server failover—not merely the Sentinel switch-master timestamp.
10. Verification and reset
Verify the client discovers a different primary after failover and resumes writes. Restart the failed node and wait for Sentinel to reconfigure it as a replica before proceeding. Keep the topology for the final game day.
Summary and next step
Client Discovery Through Sentinel, Reconnect Behavior, DNS/NAT/Container Considerations is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Run a Sentinel Failover Game Day and Verify Application Recovery and Data-Loss Bounds.