Chapter 19 · Replication, PSYNC, Backlog, Replica Reads, and Failover Foundations

Partial Resynchronization, Replication IDs, Backlog Sizing, and Reconnect Windows

Use replication IDs, offsets, backlog fields, deliberate disconnects, and sync counters to prove when PSYNC can use partial resynchronization and when Redis must perform a full synchronization.

Advanced180–260 minutesPSYNC, replication IDs, backlog, partial/full resync, reconnect windowRedis Open Source 8.10.1Docker + redis-cli3-node isolated topologyFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

By the end of this lesson, you should be able to:

01

Explain PSYNC as a request to resume a known replication history from a known offset.

02

Interpret master_replid, master_replid2, backlog first offset, backlog size, and backlog history length.

03

Estimate a reconnect window from measured replication-stream byte rate rather than dataset size alone.

04

Prove both partial and full resynchronization paths with bounded network isolation and sync counters.

05

Design backlog capacity from outage duration, write amplification, headroom, and recovery cost.

Reproducible lab baseline

Redis Open Source 8.10.1 using the pinned Docker Official Image redis:8.10.1; three standalone Redis processes on one private Docker network; primary published only on 127.0.0.1:6401, replica A on 127.0.0.1:6402, replica B on 127.0.0.1:6403; logical database 0; AOF enabled with appendfsync everysec on all nodes; no Sentinel or Cluster; TLS is intentionally off because all published ports are loopback-only; the disposable password is not a production secret. Lesson 1 topology is reused. If it is absent, run the setup block below.

Shell · create the isolated Chapter 19 topology
docker network create atlasmart-redis-ch19-netdocker volume create atlasmart-redis-ch19-primary-datadocker volume create atlasmart-redis-ch19-replica-a-datadocker volume create atlasmart-redis-ch19-replica-b-datadocker run -d --name atlasmart-redis-ch19-primary --network atlasmart-redis-ch19-net -p 127.0.0.1:6401:6379 -v atlasmart-redis-ch19-primary-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --repl-backlog-size 1mb --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-a --network atlasmart-redis-ch19-net -p 127.0.0.1:6402:6379 -v atlasmart-redis-ch19-replica-a-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-b --network atlasmart-redis-ch19-net -p 127.0.0.1:6403:6379 -v atlasmart-redis-ch19-replica-b-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel notice# If the named network/volumes already exist from a previous Chapter 19 lesson, reuse them rather than recreating them.

1. PSYNC asks two questions: “same history?” and “still retained?”

When a replica reconnects, it sends its remembered replication ID and processed offset using the internal PSYNC mechanism. The primary can perform a partial resynchronization only when that history is still recognized and the missing byte range is still retained in the replication backlog. Otherwise Redis falls back to a full synchronization.

The application should not issue PSYNC manually in normal operation; Redis servers manage it. The useful operator skill is reading the evidence that explains why a reconnect was partial or full.

2. Read the backlog as a circular history window

Shell · inspect PSYNC/backlog fields
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG GET repl-backlog-sizedocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO stats# Record master_replid, master_replid2, master_repl_offset, repl_backlog_first_byte_offset, repl_backlog_histlen, sync_partial_ok, sync_partial_err, and sync_full.
Field Operational use
repl_backlog_size Maximum configured circular backlog capacity.
repl_backlog_histlen Bytes of useful replication history currently retained.
repl_backlog_first_byte_offset Oldest stream offset still recoverable from the backlog.
master_replid2 / second_repl_offset Continuity aid after promotion/failover; not a permanent second history.

3. Backlog sizing uses stream rate, not key count

A useful first-order estimate is coverage_seconds ≈ backlog_bytes / replication_stream_bytes_per_second. Measure the offset delta under representative writes over a known interval. Payload bytes, command framing, expirations, evictions, scripts/functions, and workload bursts all affect the replication stream, so “one million keys” is not a backlog size.

Shell · measure a bounded stream-rate sample
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replication# Record master_repl_offset as OFFSET_BEFORE and the current time.docker exec atlasmart-redis-ch19-primary sh -lc "export REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026; for i in $(seq 1 400); do redis-cli SET atlasmart:ch19:l2:rate:$i value-$i >/dev/null; done"docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replication# OFFSET_AFTER - OFFSET_BEFORE divided by elapsed seconds is the observed replication-stream byte rate for this tiny fixture.

4. Partial resync experiment: short isolation inside the window

Disconnect replica A from the private Docker network, perform a small bounded write set, and reconnect it. Because the default Chapter 19 backlog is deliberately larger than this missing history, a partial resync should normally be possible. Treat the logs and counters—not this expectation—as the evidence.

Shell · bounded partial-resync drill
docker network disconnect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-adocker exec atlasmart-redis-ch19-primary sh -lc "export REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026; for i in $(seq 1 80); do redis-cli SET atlasmart:ch19:l2:partial:$i p-$i >/dev/null; done"docker network connect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-a# Give the replica a moment to reconnect, then inspect:docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO statsdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-a redis-cli INFO replicationdocker logs --tail 100 atlasmart-redis-ch19-replica-a

5. Full resync experiment: deliberately overrun a small backlog

Now reduce the primary backlog to a small bounded value, isolate replica B, and generate enough replication traffic to move the oldest retained offset beyond the replica’s last processed offset. After reconnect, Redis should be unable to satisfy that old PSYNC position and will normally perform a full synchronization.

Shell · force the history to age out of a disposable backlog
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG SET repl-backlog-size 256kbdocker network disconnect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-bdocker exec atlasmart-redis-ch19-primary sh -lc "export REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026; payload=$(printf %01024d 0); for i in $(seq 1 700); do redis-cli SET atlasmart:ch19:l2:overflow:$i $payload >/dev/null; done"docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker network connect atlasmart-redis-ch19-net atlasmart-redis-ch19-replica-b# After recovery, compare sync_full / sync_partial_err and inspect replica-B logs.docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO statsdocker logs --tail 140 atlasmart-redis-ch19-replica-bdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG SET repl-backlog-size 1mb
Expected mechanism, not guaranteed wording

Exact log messages and counter timing are version dependent. The acceptance test is that the replica reconnects and converges; use sync_full/sync_partial_*, backlog offsets, and logs together to classify the path.

6. Secondary replication IDs preserve continuity after promotion

Redis keeps a secondary replication ID/offset to support partial synchronization across certain failover histories. When a replica is promoted, it begins a new replication history but remembers the previous history long enough for other replicas to reconnect without an unnecessary full copy when their offsets are still compatible. Chapter 20 will use this idea under Sentinel; this lesson only establishes the mechanism.

7. Deliberately wrong approach: “1 MB backlog gives N minutes”

A fixed backlog size has no fixed time meaning. During a quiet period it can cover a long disconnect; during a burst it can wrap in seconds. A full sync then adds fork/COW, network, replica loading, and buffer pressure exactly when the system is already recovering from a fault.

Repair

Measure replication-stream byte rate at representative and peak write rates, choose a reconnect objective, include burst/failure headroom, and monitor actual partial/full sync outcomes.

8. Verification checklist

  • Capture replication ID/offset/backlog fields before each disconnect.
  • Show one short reconnect whose evidence is consistent with partial resynchronization.
  • Show one deliberately overflowed reconnect whose evidence is consistent with full synchronization.
  • Restore the backlog size after the experiment.
  • Do not use dataset size or key count as a substitute for replication-stream rate.

Check your understanding

  1. What two conditions make partial resynchronization possible?
  2. Why can a small write burst shorten the reconnect window dramatically?
  3. What does sync_partial_err tell you?
  4. Is a larger backlog free?
Review the answers

The replica must refer to a replication history the primary still recognizes, and the missing offset range must still exist in the backlog.

Because backlog coverage is in stream bytes; higher byte rate wraps the circular history faster.

It records failed partial synchronization attempts; it is useful alongside logs and sync_full to understand why Redis fell back to full sync.

No. It consumes memory, so it belongs in the Chapter 18 memory budget and must be balanced against the cost of full resynchronization.

Production judgment and next bridge

Backlog is recovery capacity: size it from measured stream rate and expected interruption windows, then verify with controlled reconnects. A backlog that avoids full sync can save CPU, memory, disk, and network pressure during an incident. Lesson 3 examines those full-sync resource paths directly.

Summary and next step

Partial Resynchronization, Replication IDs, Backlog Sizing, and Reconnect Windows is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Disk-Based vs Diskless Sync, Replication Buffers, Network Pressure, and Large Dataset Recovery.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.