Chapter 19 · Replication, PSYNC, Backlog, Replica Reads, and Failover Foundations

Primary/Replica Replication, Asynchronous Semantics, Initial Sync, and Command Propagation

Build a three-node Redis replication lab and observe asynchronous command propagation, replication offsets, initial synchronization, and the difference between a replica copy and a synchronous commit.

Advanced180–260 minutesprimary/replica, async replication, INFO replication, ROLE, offsetsRedis Open Source 8.10.1Docker + redis-cli3-node isolated topologyFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

By the end of this lesson, you should be able to:

01

Describe Redis primary/replica replication as asynchronous command propagation rather than a synchronous commit protocol.

02

Identify primary and replica roles, replication IDs, offsets, link state, and initial synchronization evidence with ROLE and INFO replication.

03

Explain why an initial full sync differs from steady-state command propagation and why a replica can momentarily serve older state.

04

Distinguish replication from persistence, backup, Sentinel/Cluster failover, and strong consistency.

05

Build and verify a three-node AtlasMart replication lab without touching the shared Chapter 01 node.

Reproducible lab baseline

Redis Open Source 8.10.1 using the pinned Docker Official Image redis:8.10.1; three standalone Redis processes on one private Docker network; primary published only on 127.0.0.1:6401, replica A on 127.0.0.1:6402, replica B on 127.0.0.1:6403; logical database 0; AOF enabled with appendfsync everysec on all nodes; no Sentinel or Cluster; TLS is intentionally off because all published ports are loopback-only; the disposable password is not a production secret. All replication failure injection in Chapter 19 is confined to this topology.

1. Practical problem: two copies do not imply one synchronous write

AtlasMart wants read capacity and a better failure posture for catalog and session data. Adding replicas sounds simple, but the application team starts saying “the write is safe because we have three copies.” Redis does not make that promise merely because two replicas exist.

In Redis Open Source replication, the primary accepts the normal write workload and emits a replication stream. A replica consumes that stream and applies its effects. The primary normally replies to the application without waiting for every replica to process the write. That is the core asynchronous boundary for the entire chapter.

2. Mental model: data history plus a moving offset

A Redis primary owns a replication history identified by a replication ID. As bytes are emitted into the replication stream, a monotonically increasing replication offset advances. A replica that shares the same replication ID but has a lower offset is behind that history. Redis can therefore reason about “which history?” and “how far into it?” separately.

Signal Meaning Do not infer
role Whether this node is primary or replica Automatic failover exists
master_replid Identity of the replication history Durability on disk
master_repl_offset / replica offset Position in the command stream Application-level causal consistency
master_link_status Replica link state Replica data is necessarily current
lag Heartbeat-style lag indicator on the primary Exact per-write staleness in milliseconds

3. Build the isolated AtlasMart topology

The lab uses one primary and two replicas because a single replica cannot expose several of the tradeoffs later in the chapter. The password is intentionally disposable and visible in commands; never copy this credential model to production. For production, use ACL users, TLS/private networking, secret management, and a tested topology.

Shell · create the isolated Chapter 19 topology
docker network create atlasmart-redis-ch19-netdocker volume create atlasmart-redis-ch19-primary-datadocker volume create atlasmart-redis-ch19-replica-a-datadocker volume create atlasmart-redis-ch19-replica-b-datadocker run -d --name atlasmart-redis-ch19-primary --network atlasmart-redis-ch19-net -p 127.0.0.1:6401:6379 -v atlasmart-redis-ch19-primary-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --repl-backlog-size 1mb --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-a --network atlasmart-redis-ch19-net -p 127.0.0.1:6402:6379 -v atlasmart-redis-ch19-replica-a-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-b --network atlasmart-redis-ch19-net -p 127.0.0.1:6403:6379 -v atlasmart-redis-ch19-replica-b-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel notice# If the named network/volumes already exist from a previous Chapter 19 lesson, reuse them rather than recreating them.
Shell · verify roles and replication state
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli ROLEdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli ROLEdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-b redis-cli INFO replication

4. What happens on the first synchronization

A new replica has no usable common history yet. When partial resynchronization is unavailable, Redis performs a full synchronization. In the normal disk-based path, the primary creates an RDB snapshot in a background child, buffers newer writes while that snapshot is produced/transferred, and then streams the buffered changes. The replica replaces/loads its dataset and then continues consuming the live command stream.

Shell · inspect initial-sync evidence
docker logs --tail 120 atlasmart-redis-ch19-primarydocker logs --tail 120 atlasmart-redis-ch19-replica-adocker logs --tail 120 atlasmart-redis-ch19-replica-bdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO stats# Look for full-sync/repl-transfer evidence and the sync_full counter; exact log wording is version dependent.

5. Steady-state propagation: observe a write crossing nodes

Once synchronization is complete, normal writes are propagated as replication-stream commands. The primary keeps serving clients while replicas consume the stream. This is why the normal write latency does not automatically include replica application time.

Shell · write once, then inspect all three copies
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli HSET atlasmart:ch19:l1:inventory:sku-42 available 25 reserved 0docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli GET atlasmart:ch19:l1:markerdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli SET atlasmart:ch19:l1:marker propagated-v1docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli GET atlasmart:ch19:l1:markerdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-b redis-cli GET atlasmart:ch19:l1:markerdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replication

6. Observable offsets prove progress, not synchronous commit

After the fixture propagates, compare the primary offset with each replica offset. Equal offsets at the instant you sampled them are useful evidence that the nodes reached the same stream position. They do not retroactively prove that the application response waited for those replicas, nor do they make replication a backup.

Evidence rule

Capture the write response time separately from the later offset/read checks. Otherwise an after-the-fact replica read can be mistaken for a synchronous durability guarantee.

7. Replica reads: scaling with an explicit freshness tradeoff

Read-only replicas can offload eligible queries, but the application must classify which reads can tolerate staleness. Catalog descriptions, recommendations, and analytical reads may accept bounded lag more readily than inventory reservation or payment state. Redis does not infer this business rule for you.

A replica is read-only by default, but read-only is an accident-prevention control, not a security perimeter. ACLs, network isolation, authentication, and TLS remain separate controls.

8. Deliberately wrong approach: “replica = backup”

A replica faithfully follows destructive changes too. If the primary deletes a key, applies bad application data, or restarts empty in a dangerous no-persistence configuration, replication can propagate that state. A replica therefore improves availability/read scale but does not preserve an independent historical recovery point.

Repair

Keep persistence and tested off-host backups as separate recovery mechanisms. Chapter 17 established that durability files must be restorable; Chapter 19 adds that a live replica is not a substitute for those restore points.

9. Verification checklist

  • Primary reports role:master and two connected replicas after startup.
  • Each replica reports role:slave/role:replica semantics and master_link_status:up.
  • The AtlasMart marker eventually reads identically on both replicas.
  • Replication offsets are captured and interpreted as stream positions, not transaction commits.
  • Initial full-sync evidence and sync_full are recorded before later PSYNC experiments.

Check your understanding

  1. Why can a successful SET reply precede replica application?
  2. What does a replication ID + offset pair represent?
  3. Does an equal replica offset prove the original client waited for that replica?
  4. Why is a replica not a backup?
Review the answers

Because Redis replication is asynchronous by default; the primary normally does not wait for replica processing before replying.

A particular history of the dataset plus a position in that history.

No. It only shows the sampled stream positions had caught up by the time of observation.

It follows primary changes, including unwanted changes, and does not by itself provide an independent historical restore point.

Production judgment and next bridge

Use replicas when read scaling, redundancy, or a failover foundation justifies the extra memory/network/full-sync cost. Monitor link status, offsets, sync counters, replica memory, output buffers, and recovery duration. Do not promise zero data loss merely because replicas exist. Lesson 2 now turns the replication ID/offset model into the practical PSYNC/backlog reconnect window.

Summary and next step

Primary/Replica Replication, Asynchronous Semantics, Initial Sync, and Command Propagation is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Partial Resynchronization, Replication IDs, Backlog Sizing, and Reconnect Windows.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.