Chapter 19 · Replication, PSYNC, Backlog, Replica Reads, and Failover Foundations
Disk-Based vs Diskless Sync, Replication Buffers, Network Pressure, and Large Dataset Recovery
Compare disk-based and diskless full synchronization, then connect replication buffers, output limits, fork/COW pressure, network transfer, and replica load time to recovery capacity planning.
Learning outcomes
By the end of this lesson, you should be able to:
Describe the disk-based full-sync path and the diskless alternative without assuming one is universally faster.
Inspect repl-diskless-sync, synchronization
logs, replication buffers, client output buffers, and
COW-related memory signals.
Force a fresh replica to perform a full synchronization under a bounded dataset.
Connect network throughput, fork cost, buffer growth, replica load time, and disk headroom to recovery duration.
Choose a sync mode from measured platform constraints rather than folklore.
Redis Open Source 8.10.1 using the pinned Docker
Official Image redis:8.10.1; three standalone Redis
processes on one private Docker network; primary published only
on 127.0.0.1:6401, replica A on
127.0.0.1:6402, replica B on
127.0.0.1:6403; logical database 0; AOF enabled
with appendfsync everysec on all nodes; no Sentinel
or Cluster; TLS is intentionally off because all published ports
are loopback-only; the disposable password is not a production
secret. This lesson may recreate replica B and its named volume
to force a clean full synchronization. Primary and replica A
remain intact.
docker network create atlasmart-redis-ch19-netdocker volume create atlasmart-redis-ch19-primary-datadocker volume create atlasmart-redis-ch19-replica-a-datadocker volume create atlasmart-redis-ch19-replica-b-datadocker run -d --name atlasmart-redis-ch19-primary --network atlasmart-redis-ch19-net -p 127.0.0.1:6401:6379 -v atlasmart-redis-ch19-primary-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --repl-backlog-size 1mb --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-a --network atlasmart-redis-ch19-net -p 127.0.0.1:6402:6379 -v atlasmart-redis-ch19-replica-a-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel noticedocker run -d --name atlasmart-redis-ch19-replica-b --network atlasmart-redis-ch19-net -p 127.0.0.1:6403:6379 -v atlasmart-redis-ch19-replica-b-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel notice# If the named network/volumes already exist from a previous Chapter 19 lesson, reuse them rather than recreating them.
1. Full sync is a resource event, not just a correctness event
When PSYNC cannot resume from backlog, Redis must transfer the dataset again. The recovery path consumes resources on both sides: the primary needs a consistent snapshot source plus buffering for concurrent writes; the network carries that snapshot and later stream bytes; the replica must receive and load the dataset. Large datasets turn this into a capacity-planning problem.
2. Disk-based full synchronization
In the traditional path, the primary forks a child to create an RDB file. New writes continue on the primary and are buffered for the replica. After the RDB is sent and loaded, buffered/live stream changes bring the replica forward. Slow storage can make the RDB creation/read path expensive, while write-heavy workloads can increase copy-on-write memory during the fork.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG GET repl-diskless-syncdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CLIENT LIST TYPE replica
3. Diskless sync changes the transfer path, not the need for headroom
With repl-diskless-sync yes, the child can stream
the RDB directly to replicas instead of using disk as an
intermediate file. This can help when disks are slow, but it
shifts emphasis toward network capacity, child lifetime, replica
coordination, and buffering. It does not eliminate fork/COW or
the cost of loading the dataset on the replica.
The repl-diskless-sync-delay setting can
intentionally wait for additional replicas so one child can
serve several arrivals. That delay can improve resource
efficiency while increasing recovery latency for the first
replica. Measure the tradeoff.
4. Build a bounded recovery dataset
docker exec atlasmart-redis-ch19-primary sh -lc "export REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026; payload=$(printf %04096d 0); for i in $(seq 1 1500); do redis-cli SET atlasmart:ch19:l3:blob:$i $payload >/dev/null; done"docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli DBSIZEdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO memory# This is intentionally small enough for a laptop lab; it does not model production full-sync duration.
5. Force a fresh full sync with diskless mode enabled
docker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG SET repl-diskless-sync yesdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli CONFIG SET repl-diskless-sync-delay 0docker rm -f atlasmart-redis-ch19-replica-bdocker volume rm atlasmart-redis-ch19-replica-b-datadocker volume create atlasmart-redis-ch19-replica-b-datadocker run -d --name atlasmart-redis-ch19-replica-b --network atlasmart-redis-ch19-net -p 127.0.0.1:6403:6379 -v atlasmart-redis-ch19-replica-b-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --requirepass AtlasMart-Ch19-Lab-Only-2026 --masterauth AtlasMart-Ch19-Lab-Only-2026 --replicaof atlasmart-redis-ch19-primary 6379 --replica-read-only yes --loglevel noticedocker logs --tail 160 atlasmart-redis-ch19-primarydocker logs --tail 160 atlasmart-redis-ch19-replica-bdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-primary redis-cli INFO statsdocker exec -e REDISCLI_AUTH=AtlasMart-Ch19-Lab-Only-2026 atlasmart-redis-ch19-replica-b redis-cli INFO replication
6. Buffers are part of the recovery budget
During replication, the primary can retain backlog and replica
output-buffer memory while the replica catches up. Redis exposes
replication-buffer memory in INFO memory, and
replica client output-buffer state is visible through
CLIENT LIST TYPE replica. Output buffer limits
prevent an indefinitely slow replica from consuming unbounded
memory; if limits are reached, Redis can close the connection,
potentially triggering another synchronization cycle.
| Pressure source | Evidence | Risk |
|---|---|---|
| Replication backlog |
mem_replication_backlog /
repl_backlog_*
|
Memory reserved to enable PSYNC. |
| Replica output/buffer growth |
CLIENT LIST TYPE replica, omem
|
Slow receiver can be disconnected at limits. |
| Fork/COW | INFO persistence COW fields |
Write-heavy fork can raise peak memory. |
| Network | net input/output counters + elapsed sync time | Long transfer extends recovery and buffering window. |
| Replica load | master_sync_in_progress, logs |
Large dataset loading can block/limit reads depending on config. |
7. Deliberately wrong approach: full sync with no spare memory/disk/network
A deployment sized only for steady-state dataset memory may succeed until a replica needs a full resync. At that moment fork/COW, backlog/output buffers, disk snapshot space, and network throughput all become concurrent demands. The recovery event can then cause OOM, slow clients, or repeated replica disconnects.
Budget recovery headroom explicitly and load-test full synchronization under representative write rate. Keep the experiment isolated; do not “test” by starving a production primary.
8. Disk-based versus diskless: choose from evidence
| Question | Disk-based may fit when… | Diskless may fit when… |
|---|---|---|
| Storage | Fast local storage and enough temporary space are available. | Disk is slow or snapshot I/O would be the bottleneck. |
| Network | Normal dataset transfer is acceptable. | Network can sustain direct RDB stream without creating new bottlenecks. |
| Replica fan-in | Sync arrival is predictable. | Coordinating several arrivals with diskless delay can reduce repeated forks. |
| Operational evidence | Measured recovery and COW are acceptable. | Measured diskless recovery is better for this platform. |
9. Verification checklist
-
Record the actual
repl-diskless-syncmode before the experiment. - Force a full sync using only the disposable replica-B volume.
-
Capture primary/replica logs,
sync_full, COW fields, replication buffers, and elapsed recovery time. - Do not extrapolate laptop transfer rate to production.
- Keep client-output-buffer limits and full-sync headroom in the capacity plan.
Check your understanding
- Does diskless sync eliminate fork/COW?
- Why can a slow replica increase primary memory pressure?
- Why might repl-diskless-sync-delay increase recovery latency?
- What should decide disk-based versus diskless mode?
Review the answers
No. It changes how the RDB snapshot is transferred; the fork/snapshot consistency work and replica load cost still matter.
The primary may retain backlog/output data while the replica is behind, subject to buffer limits.
Redis can wait intentionally so more replicas share one diskless child transfer.
Measured disk, network, fork/COW, replica load, concurrency, and recovery behavior on the target platform.
Production judgment and next bridge
Full synchronization belongs in capacity and incident planning. Test it before outages force it, and monitor recovery duration as a service-level signal. Lesson 4 now focuses on what applications should believe when they read replicas or request replication/persistence acknowledgments.
Summary and next step
Disk-Based vs Diskless Sync, Replication Buffers, Network Pressure, and Large Dataset Recovery is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Replica Reads, Staleness, replica-read-only, min-replicas Settings, and Durability Expectations.