Chapter 19 · Replication, PSYNC, Backlog, Replica Reads, and Failover Foundations
Failure Drill: Break Replication, Observe Lag, Recover, and Measure Data/Availability Impact
Run a bounded replication failure game day: isolate a replica, keep writes flowing on the primary, observe stale state and offsets, reconnect, measure catch-up, and record the real data/availability consequences.
Learning outcomes
By the end of this lesson, you should be able to:
Run a bounded replication-link failure without stopping or reconfiguring the shared course Redis node.
Measure primary availability, replica staleness, offset divergence, reconnect path, and recovery duration.
Classify which writes were accepted, which copies were stale, and what the experiment says about data-loss bounds.
Separate manual replication recovery from Sentinel/Cluster automatic failover.
Produce a reusable AtlasMart replication game-day evidence sheet and cleanly remove all Chapter 19 resources.
Redis Open Source 8.10.1 using the pinned Docker
Official Image redis:8.10.1; three standalone Redis
processes on one private Docker network; primary published only
on 127.0.0.1:6401, replica A on
127.0.0.1:6402, replica B on
127.0.0.1:6403; logical database 0; AOF enabled
with appendfsync everysec on all nodes; no Sentinel
or Cluster; TLS is intentionally off because all published ports
are loopback-only; the disposable password is not a production
secret. This game day isolates replica A only. The primary
remains the sole write authority; there is no automatic failover
in Chapter 19.
docker network create atlasmart-redis-ch19-netdocker volume create atlasmart-redis-ch19-primary-datadocker volume create atlasmart-redis-ch19-replica-a-datadocker volume create atlasmart-redis-ch19-replica-b-datadocker run -d --name atlasmart-redis-ch19-primary --network atlasmart-redis-ch19-net -p 127.0.0.1:6401:6379 -v atlasmart-redis-ch19-primary-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --repl-backlog-size 1mb --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-a --network atlasmart-redis-ch19-net -p 127.0.0.1:6402:6379 -v atlasmart-redis-ch19-replica-a-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-b --network atlasmart-redis-ch19-net -p 127.0.0.1:6403:6379 -v atlasmart-redis-ch19-replica-b-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel notice# If the named network/volumes already exist from a previous Chapter 19 lesson, reuse them rather than recreating them.
1. Define the failure question before causing the failure
A useful failure drill is not “break Redis and see what happens.” AtlasMart asks narrower questions: Does the primary continue accepting writes while one replica is isolated? What offset gap develops? Can a direct read of the isolated replica be stale? Does reconnect use partial or full resynchronization? How long until the replica reaches the primary stream again?
Only atlasmart-redis-ch19-replica-a is
disconnected from atlasmart-redis-ch19-net. Do
not stop the host network, change firewall rules, or touch
unrelated Redis containers.
2. Capture the pre-failure baseline
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli SET atlasmart:ch19:l5:checkpoint before-failuredocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli GET atlasmart:ch19:l5:checkpointdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO stats# Record timestamps, primary offset, replica offset, sync_full/sync_partial_* counters, and connected replica count.
3. Isolate one replica and keep writing to the primary
Because replication is asynchronous and one replica is only a follower, the primary can remain available for writes while that follower is disconnected. This is availability for the primary endpoint, not automatic failover and not proof that every accepted write exists on enough copies.
docker network disconnect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-adocker exec atlasmart-redis-ch19-primary sh -lc "export REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026; for i in $(seq 1 300); do redis-cli SET atlasmart:ch19:l5:during:$i event-$i >/dev/null; done"docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli SET atlasmart:ch19:l5:checkpoint during-failuredocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli GET atlasmart:ch19:l5:checkpoint# docker exec talks to the isolated process locally: the replica should still show its pre-failure value while disconnected.
4. What the outage proves—and what it does not
| Observation | Supported conclusion | Unsupported leap |
|---|---|---|
| Primary accepted writes | Primary endpoint remained available with one replica missing. | Those writes are safe against any later failover. |
| Replica A shows old checkpoint | A read from that isolated copy is stale. | All replicas are equally stale. |
| Primary offset advanced | New replication-stream bytes were generated. | Backlog definitely covers the entire outage. |
| Replica B remains connected | Another copy may be receiving writes. | Replica B is a backup or guaranteed promotion target. |
5. Reconnect and measure recovery
Reconnect replica A and record the start time. Poll its link state and offset until it reaches the primary’s current position. The exact recovery duration depends on missing history, backlog availability, network, dataset size, write rate, and whether Redis can use partial resynchronization.
docker network connect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-a# Record a wall-clock start time outside Redis.docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO statsdocker logs --tail 160 atlasmart-redis-ch19-replica-adocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli GET atlasmart:ch19:l5:checkpoint# Stop timing when link_status is up, sync is complete, the checkpoint is current, and replica offset has caught up to the sampled primary offset.
6. Partial versus full resync changes the recovery cost
With the Chapter 19 default 1 MiB backlog and only a 300-command burst, this drill is intentionally biased toward partial resynchronization. But do not assert the path before checking logs/counters. If the backlog had wrapped, the recovery would require a full dataset transfer and the resource profile would look like Lesson 3.
Record sync_partial_ok,
sync_partial_err, sync_full, the
replica logs, and the recovered offset. Those together make the
recovery mechanism auditable.
7. Acknowledged writes and failover are separate questions
This drill keeps the primary running, so it does not demonstrate
data loss on promotion. To reason about that boundary, combine
evidence from Lesson 4: a normal write reply means the primary
executed the command; WAIT can add replica-ack
evidence; WAITAOF can add AOF-fsync evidence. Even
those mechanisms do not guarantee a particular replica will be
promoted or that every failure mode preserves the write.
Automatic failure detection, quorum/majority, replica selection, and client discovery belong to Redis Sentinel. Chapter 19 deliberately stops at replication evidence so failover policy is not conflated with data propagation.
8. Build a game-day evidence sheet
| Metric/evidence | Before | During isolation | After recovery |
|---|---|---|---|
| Primary connected replicas | record | record | record |
| Primary replication offset | record | record | record |
| Replica A offset/link state | record | stale/down | caught up/up |
| AtlasMart checkpoint value | same | primary=new / replica=old | same/new |
| sync_partial_ok / sync_full | record | unchanged during link break | record delta |
| Recovery duration | n/a | timer starts at reconnect | record measured seconds |
| Write errors at primary | record | record | record |
9. Deliberately wrong conclusion: “the drill passed, therefore RPO is zero”
A successful partial resync only proves this bounded failure fit within the backlog and recovery resources available at this moment. It does not prove zero Recovery Point Objective (RPO) for primary crashes, failover, persistence corruption, multi-node loss, operator error, or a larger disconnection window.
Define RPO and Recovery Time Objective (RTO) per failure class. Repeat drills with representative write rate, backlog size, persistence, topology, and failure mode; combine them with restore tests from Chapter 17.
10. Verification checklist
- Only replica A was isolated; the primary remained the write authority.
- The isolated replica returned demonstrably stale AtlasMart state through local docker exec.
- Primary write behavior and connected replica count were recorded during failure.
- Reconnect path was classified using offsets, logs, and sync counters.
- Recovery duration was measured rather than guessed.
- No automatic failover, strong-consistency, or backup claim was inferred from the drill.
Check your understanding
- Why can the primary remain writable when one replica is disconnected?
- What proves the isolated replica is stale?
- What determines whether reconnect is partial or full?
- Why does this drill not test automatic failover?
Review the answers
Replication is asynchronous by default and normal primary writes do not require every replica to be connected unless additional admission/acknowledgment policy is configured.
A direct read of a known key returns the earlier value while the primary has already accepted a newer value, together with offset/link evidence.
Whether the requested replication history/offset is recognized and the missing bytes are still retained in the primary backlog.
No Sentinel or Cluster control plane is present; the primary never stops being the designated write endpoint.
11. Clean up only the Chapter 19 topology
docker rm -f atlasmart-redis-ch19-replica-a atlasmart-redis-ch19-replica-b atlasmart-redis-ch19-primarydocker volume rm atlasmart-redis-ch19-replica-a-data atlasmart-redis-ch19-replica-b-data atlasmart-redis-ch19-primary-datadocker network rm atlasmart-redis-ch19-net# These commands target only the explicitly named disposable Chapter 19 resources. Do not use broad Docker prune commands.
Production judgment and next bridge
Replication operations should be rehearsed with explicit evidence: offsets, backlog coverage, sync path, resource headroom, stale-read behavior, and recovery duration. Those mechanisms now provide the foundation for Chapter 20, where Sentinel adds monitoring, failure detection, quorum/majority logic, promotion, and client service discovery.
Summary and next step
Failure Drill: Break Replication, Observe Lag, Recover, and Measure Data/Availability Impact is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Sentinel Architecture, Quorum vs Majority, Objective Down, and Failure Detection.