Chapter 19 · Replication, PSYNC, Backlog, Replica Reads, and Failover Foundations
Replica Reads, Staleness, replica-read-only, min-replicas Settings, and Durability Expectations
Prove read-only replica behavior and observable staleness, then evaluate min-replicas, WAIT, and WAITAOF without turning acknowledgments into stronger guarantees than Redis provides.
Learning outcomes
By the end of this lesson, you should be able to:
Classify reads by whether replica staleness is acceptable and prove a stale-read window in an isolated lab.
Verify default replica-read-only behavior
without treating it as a security boundary.
Explain min-replicas-to-write /
min-replicas-max-lag as lag-based write
admission rather than per-write synchronous replication.
Use WAIT and WAITAOF on the same
client connection and interpret their returned
acknowledgment counts correctly.
Separate replication acknowledgment, persistence acknowledgment, failover selection, backup, and application consistency.
Redis Open Source 8.10.1 using the pinned Docker
Official Image redis:8.10.1; three standalone Redis
processes on one private Docker network; primary published only
on 127.0.0.1:6401, replica A on
127.0.0.1:6402, replica B on
127.0.0.1:6403; logical database 0; AOF enabled
with appendfsync everysec on all nodes; no Sentinel
or Cluster; TLS is intentionally off because all published ports
are loopback-only; the disposable password is not a production
secret. Both primary and replicas have AOF enabled, which
permits the WAITAOF lab.
docker network create atlasmart-redis-ch19-netdocker volume create atlasmart-redis-ch19-primary-datadocker volume create atlasmart-redis-ch19-replica-a-datadocker volume create atlasmart-redis-ch19-replica-b-datadocker run -d --name atlasmart-redis-ch19-primary --network atlasmart-redis-ch19-net -p 127.0.0.1:6401:6379 -v atlasmart-redis-ch19-primary-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --repl-backlog-size 1mb --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-a --network atlasmart-redis-ch19-net -p 127.0.0.1:6402:6379 -v atlasmart-redis-ch19-replica-a-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-b --network atlasmart-redis-ch19-net -p 127.0.0.1:6403:6379 -v atlasmart-redis-ch19-replica-b-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel notice# If the named network/volumes already exist from a previous Chapter 19 lesson, reuse them rather than recreating them.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli ROLEdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli ROLEdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-b redis-cli INFO replication
1. Replica reads trade freshness for scale/fault flexibility
Redis replicas can serve reads, but asynchronous propagation creates a window in which the replica may reflect an older stream offset than the primary. That may be acceptable for a product-description page and unacceptable for a stock reservation, idempotency token, payment status, or authorization decision.
| AtlasMart read | Replica tolerance | Reason |
|---|---|---|
| Product description | Often tolerant | A short delay rarely violates an invariant. |
| Recommendation candidates | Often tolerant | Ranking freshness is usually a quality tradeoff. |
| Inventory available before reservation | Usually primary/current-state path | Stale value can oversell. |
| Payment authorization state | Usually primary/authoritative path | Stale value can duplicate or misreport a side effect. |
2. replica-read-only prevents accidental writes, not attacks
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli CONFIG GET replica-read-onlydocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli SET atlasmart:ch19:l4:should-fail x# Expected: a READONLY-style error. Exact wording may vary by release/client.
Writable replicas exist for historical compatibility but can diverge from the primary and are not recommended for ordinary designs. Keep replicas read-only and enforce independent ACL/network controls.
3. Deterministic stale-read fixture: isolate a replica but keep its process alive
Docker network isolation cleanly separates the replication link
while docker exec can still query the Redis process
through its own loopback interface. This produces an observable
stale read without host firewall changes or clock manipulation.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli SET atlasmart:ch19:l4:inventory:sku-42 v1# Wait until replica A returns v1 before continuing.docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli GET atlasmart:ch19:l4:inventory:sku-42docker network disconnect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-adocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli SET atlasmart:ch19:l4:inventory:sku-42 v2docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli GET atlasmart:ch19:l4:inventory:sku-42# Expected during isolation: replica A can still return v1 while the primary already has v2.docker network connect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-a# Verify it later converges to v2 and master_link_status returns up.
4. min-replicas settings are admission control based on recent lag
The primary can be configured to reject writes unless at least N replicas have recently acknowledged replication progress within a configured maximum lag. This bounds one failure mode, but it does not prove that a particular accepted write reached those replicas. The health decision is based on periodic replica acknowledgments.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG SET min-replicas-to-write 2docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG SET min-replicas-max-lag 2docker network disconnect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-b# Wait more than 2 seconds so replica B is no longer considered good.docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli SET atlasmart:ch19:l4:minreplicas should-be-rejected# Expected: a NOREPLICAS-style rejection because only one good replica remains.docker network connect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-bdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG SET min-replicas-to-write 0docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG SET min-replicas-max-lag 10
5. WAIT: ask for replication acknowledgment on this connection
WAIT applies to previous writes sent on the
same client connection. It waits until the
requested number of replicas acknowledge an offset at least as
new as those writes, or until the timeout. It returns the number
actually reached. A timeout is not automatically an application
failure: the write may still have happened, so the client must
classify the outcome carefully.
docker exec -it atlasmart-redis-ch19-primary redis-cliAUTH default AtlasMart-Ch19-Lab-Only-2026SET atlasmart:ch19:l4:wait:order-9001 acceptedWAIT 1 1000# Check that the returned integer is >= the required replica count.# Exit the interactive client only after the experiment.
Redis explicitly documents that WAIT does not turn Redis into a strongly consistent CP system. It improves real-world data safety but failover can still lose an acknowledged copy in some failure/configuration combinations.
6. WAITAOF: wait for AOF fsync evidence, not universal durability
WAITAOF (Redis 7.2+) can wait for the local primary
and/or replicas to report that prior writes from this connection
were fsynced to AOF. The reply is a two-integer array: local
fsync count (0 or 1) and replica fsync count. It requires AOF
for the local check and cannot be called on a replica.
SET atlasmart:ch19:l4:waitaof:order-9002 acceptedWAITAOF 1 1 1500# Expected shape: two integers [local_fsynced, replica_fsynced].# Verify both returned counts meet your requested threshold; do not infer backup or strong consistency.
7. Acknowledgment ladder: know exactly what each signal means
| Mechanism | What it can tell the client | What it does not guarantee |
|---|---|---|
| Normal write reply | Primary executed the command. | Replica application or disk fsync. |
min-replicas-* |
Enough replicas were recently considered sufficiently current for write admission. | This particular write reached them. |
WAIT |
A returned number of replicas acknowledged the connection’s prior writes. | Strong consistency, backup, guaranteed future promotion choice. |
WAITAOF |
Returned local/replica counts fsynced those prior writes to AOF. | Universal zero-loss failover/restart semantics. |
8. Deliberately wrong approach: route every correctness-critical read to replicas
Replica reads can increase read capacity, but doing so blindly shifts the consistency model of the application. A stale inventory or authorization read can violate business rules even when the replica is healthy by operational standards.
Classify reads by invariant and freshness requirement. Route correctness-critical decisions to an authoritative/current path, or add an explicit application protocol that makes the required consistency observable.
9. Verification checklist
- Replica writes fail under the default read-only policy.
- An isolated replica demonstrably returns an older AtlasMart value.
- The replica converges after network reconnection.
- The min-replicas experiment rejects writes only in the dedicated lab and settings are restored afterward.
- WAIT/WAITAOF examples keep the write and acknowledgment command on the same client connection.
- Acknowledgment counts are checked rather than assumed.
Check your understanding
- Why is a replica-read value potentially stale even with master_link_status up?
- Does min-replicas-to-write prove the accepted write reached N replicas?
- What must be true for WAIT to refer to the intended write?
- What are the two WAITAOF return numbers?
Review the answers
Replication is asynchronous; the replica can lag the primary stream while the link is healthy.
No. It gates writes using recent replica lag/ack state, not a synchronous per-write acknowledgment.
The write and WAIT must be sent in the same client connection context.
The number of local Redis instances (0/1) and replicas that fsynced the connection’s prior writes to AOF.
Production judgment and next bridge
Use replica reads only where stale data is acceptable and use acknowledgment primitives only for the exact guarantees they document. Separate those choices from backup and from automatic failover. Lesson 5 now combines these mechanisms in a controlled replication outage and recovery drill.
Summary and next step
Replica Reads, Staleness, replica-read-only, min-replicas Settings, and Durability Expectations is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Failure Drill: Break Replication, Observe Lag, Recover, and Measure Data/Availability Impact.