Chapter 19 · Replication, PSYNC, Backlog, Replica Reads, and Failover Foundations

Failure Drill: Break Replication, Observe Lag, Recover, and Measure Data/Availability Impact

Run a bounded replication failure game day: isolate a replica, keep writes flowing on the primary, observe stale state and offsets, reconnect, measure catch-up, and record the real data/availability consequences.

Advanced180–260 minutesfailure drill, lag, reconnect, recovery time, data/availability impactRedis Open Source 8.10.1Docker + redis-cli3-node isolated topologyFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

By the end of this lesson, you should be able to:

01

Run a bounded replication-link failure without stopping or reconfiguring the shared course Redis node.

02

Measure primary availability, replica staleness, offset divergence, reconnect path, and recovery duration.

03

Classify which writes were accepted, which copies were stale, and what the experiment says about data-loss bounds.

04

Separate manual replication recovery from Sentinel/Cluster automatic failover.

05

Produce a reusable AtlasMart replication game-day evidence sheet and cleanly remove all Chapter 19 resources.

Reproducible lab baseline

Redis Open Source 8.10.1 using the pinned Docker Official Image redis:8.10.1; three standalone Redis processes on one private Docker network; primary published only on 127.0.0.1:6401, replica A on 127.0.0.1:6402, replica B on 127.0.0.1:6403; logical database 0; AOF enabled with appendfsync everysec on all nodes; no Sentinel or Cluster; TLS is intentionally off because all published ports are loopback-only; the disposable password is not a production secret. This game day isolates replica A only. The primary remains the sole write authority; there is no automatic failover in Chapter 19.

Shell · create the isolated Chapter 19 topology
docker network create atlasmart-redis-ch19-netdocker volume create atlasmart-redis-ch19-primary-datadocker volume create atlasmart-redis-ch19-replica-a-datadocker volume create atlasmart-redis-ch19-replica-b-datadocker run -d --name atlasmart-redis-ch19-primary --network atlasmart-redis-ch19-net -p 127.0.0.1:6401:6379 -v atlasmart-redis-ch19-primary-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --repl-backlog-size 1mb --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-a --network atlasmart-redis-ch19-net -p 127.0.0.1:6402:6379 -v atlasmart-redis-ch19-replica-a-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-b --network atlasmart-redis-ch19-net -p 127.0.0.1:6403:6379 -v atlasmart-redis-ch19-replica-b-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel notice# If the named network/volumes already exist from a previous Chapter 19 lesson, reuse them rather than recreating them.

1. Define the failure question before causing the failure

A useful failure drill is not “break Redis and see what happens.” AtlasMart asks narrower questions: Does the primary continue accepting writes while one replica is isolated? What offset gap develops? Can a direct read of the isolated replica be stale? Does reconnect use partial or full resynchronization? How long until the replica reaches the primary stream again?

Blast radius

Only atlasmart-redis-ch19-replica-a is disconnected from atlasmart-redis-ch19-net. Do not stop the host network, change firewall rules, or touch unrelated Redis containers.

2. Capture the pre-failure baseline

Shell · evidence before isolation
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli SET atlasmart:ch19:l5:checkpoint before-failuredocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli GET atlasmart:ch19:l5:checkpointdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO stats# Record timestamps, primary offset, replica offset, sync_full/sync_partial_* counters, and connected replica count.

3. Isolate one replica and keep writing to the primary

Because replication is asynchronous and one replica is only a follower, the primary can remain available for writes while that follower is disconnected. This is availability for the primary endpoint, not automatic failover and not proof that every accepted write exists on enough copies.

Shell · bounded failure injection and write burst
docker network disconnect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-adocker exec atlasmart-redis-ch19-primary sh -lc "export REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026; for i in $(seq 1 300); do redis-cli SET atlasmart:ch19:l5:during:$i event-$i >/dev/null; done"docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli SET atlasmart:ch19:l5:checkpoint during-failuredocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli GET atlasmart:ch19:l5:checkpoint# docker exec talks to the isolated process locally: the replica should still show its pre-failure value while disconnected.

4. What the outage proves—and what it does not

Observation Supported conclusion Unsupported leap
Primary accepted writes Primary endpoint remained available with one replica missing. Those writes are safe against any later failover.
Replica A shows old checkpoint A read from that isolated copy is stale. All replicas are equally stale.
Primary offset advanced New replication-stream bytes were generated. Backlog definitely covers the entire outage.
Replica B remains connected Another copy may be receiving writes. Replica B is a backup or guaranteed promotion target.

5. Reconnect and measure recovery

Reconnect replica A and record the start time. Poll its link state and offset until it reaches the primary’s current position. The exact recovery duration depends on missing history, backlog availability, network, dataset size, write rate, and whether Redis can use partial resynchronization.

Shell · recovery evidence
docker network connect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-a# Record a wall-clock start time outside Redis.docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO statsdocker logs --tail 160 atlasmart-redis-ch19-replica-adocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli GET atlasmart:ch19:l5:checkpoint# Stop timing when link_status is up, sync is complete, the checkpoint is current, and replica offset has caught up to the sampled primary offset.

6. Partial versus full resync changes the recovery cost

With the Chapter 19 default 1 MiB backlog and only a 300-command burst, this drill is intentionally biased toward partial resynchronization. But do not assert the path before checking logs/counters. If the backlog had wrapped, the recovery would require a full dataset transfer and the resource profile would look like Lesson 3.

Record sync_partial_ok, sync_partial_err, sync_full, the replica logs, and the recovered offset. Those together make the recovery mechanism auditable.

7. Acknowledged writes and failover are separate questions

This drill keeps the primary running, so it does not demonstrate data loss on promotion. To reason about that boundary, combine evidence from Lesson 4: a normal write reply means the primary executed the command; WAIT can add replica-ack evidence; WAITAOF can add AOF-fsync evidence. Even those mechanisms do not guarantee a particular replica will be promoted or that every failure mode preserves the write.

Bridge to Chapter 20

Automatic failure detection, quorum/majority, replica selection, and client discovery belong to Redis Sentinel. Chapter 19 deliberately stops at replication evidence so failover policy is not conflated with data propagation.

8. Build a game-day evidence sheet

Metric/evidence Before During isolation After recovery
Primary connected replicas record record record
Primary replication offset record record record
Replica A offset/link state record stale/down caught up/up
AtlasMart checkpoint value same primary=new / replica=old same/new
sync_partial_ok / sync_full record unchanged during link break record delta
Recovery duration n/a timer starts at reconnect record measured seconds
Write errors at primary record record record

9. Deliberately wrong conclusion: “the drill passed, therefore RPO is zero”

A successful partial resync only proves this bounded failure fit within the backlog and recovery resources available at this moment. It does not prove zero Recovery Point Objective (RPO) for primary crashes, failover, persistence corruption, multi-node loss, operator error, or a larger disconnection window.

Repair

Define RPO and Recovery Time Objective (RTO) per failure class. Repeat drills with representative write rate, backlog size, persistence, topology, and failure mode; combine them with restore tests from Chapter 17.

10. Verification checklist

  • Only replica A was isolated; the primary remained the write authority.
  • The isolated replica returned demonstrably stale AtlasMart state through local docker exec.
  • Primary write behavior and connected replica count were recorded during failure.
  • Reconnect path was classified using offsets, logs, and sync counters.
  • Recovery duration was measured rather than guessed.
  • No automatic failover, strong-consistency, or backup claim was inferred from the drill.

Check your understanding

  1. Why can the primary remain writable when one replica is disconnected?
  2. What proves the isolated replica is stale?
  3. What determines whether reconnect is partial or full?
  4. Why does this drill not test automatic failover?
Review the answers

Replication is asynchronous by default and normal primary writes do not require every replica to be connected unless additional admission/acknowledgment policy is configured.

A direct read of a known key returns the earlier value while the primary has already accepted a newer value, together with offset/link evidence.

Whether the requested replication history/offset is recognized and the missing bytes are still retained in the primary backlog.

No Sentinel or Cluster control plane is present; the primary never stops being the designated write endpoint.

11. Clean up only the Chapter 19 topology

Shell · Chapter 19 cleanup
docker rm -f atlasmart-redis-ch19-replica-a atlasmart-redis-ch19-replica-b atlasmart-redis-ch19-primarydocker volume rm atlasmart-redis-ch19-replica-a-data atlasmart-redis-ch19-replica-b-data atlasmart-redis-ch19-primary-datadocker network rm atlasmart-redis-ch19-net# These commands target only the explicitly named disposable Chapter 19 resources. Do not use broad Docker prune commands.

Production judgment and next bridge

Replication operations should be rehearsed with explicit evidence: offsets, backlog coverage, sync path, resource headroom, stale-read behavior, and recovery duration. Those mechanisms now provide the foundation for Chapter 20, where Sentinel adds monitoring, failure detection, quorum/majority logic, promotion, and client service discovery.

Summary and next step

Failure Drill: Break Replication, Observe Lag, Recover, and Measure Data/Availability Impact is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Sentinel Architecture, Quorum vs Majority, Objective Down, and Failure Detection.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.