Chapter 17 · Persistence: RDB Snapshots, AOF, fsync, Rewrites, and Crash Recovery

Choose Durability Settings from RPO, Latency, Memory, Fork Behavior, and Storage Characteristics

Build a measured durability decision framework from recovery point objective, latency, fork/COW memory, disk behavior, startup time, replication, and backup requirements.

Advanced180–250 minutesRPO, latency, memory, storage, recovery designRedis Open Source 8.10.1Docker local labFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

The final design question is not “RDB or AOF?” in isolation. AtlasMart must choose a persistence envelope that meets business recovery objectives on its actual storage and memory limits, while preserving enough headroom for fork/rewrite and separating node durability from service availability and backup.

01

Translate business RPO/RTO into persistence and backup requirements.

02

Compare no persistence, RDB, AOF policies, and combined modes without universal tuning values.

03

Budget fork/COW memory and rewrite disk headroom alongside steady-state dataset memory.

04

Define a benchmark/recovery matrix with p50/p95/p99 latency and measured restart/restore time.

05

Produce an operational acceptance checklist spanning persistence, replication, backup, security, and rollback.

Exact Chapter 17 lab boundary

The normal course baseline remains Redis Open Source 8.10.1 from pinned image redis:8.10.1. Because this chapter intentionally restarts/crashes persistence processes, destructive exercises use dedicated disposable Chapter 17 containers and named volumes rather than the shared atlasmart-redis-ch01 container. All ports bind only to 127.0.0.1; the temporary lab password is public classroom data, not a production secret; TLS is omitted only on loopback. Mandatory topology is standalone, logical DB 0, no explicit maxmemory/eviction policy. This lesson is a decision/measurement capstone; it does not mutate the shared Chapter 01 node.

1. Start with RPO and RTO, not a config snippet

Recovery point objective (RPO) is the maximum acceptable data-loss interval. Recovery time objective (RTO) is the target time to restore acceptable service. RDB save cadence, AOF fsync policy, backup frequency, startup replay time, and replication topology all contribute differently. None should be selected solely because it is a Redis default.

2. Compare the persistence modes honestly

Mode Normal durability shape Primary costs / caveats
No persistence memory only all node-local data can disappear on restart
RDB last successful snapshot fork/COW + snapshot IO; wider RPO between saves
AOF everysec recent command log; roughly one-second disaster window documented continuous writes + fsync/rewrite + replay cost
AOF always fsync every command batch higher write latency/IO; still not an off-host backup
RDB + AOF AOF used for normal startup, RDB also available pays both mechanisms; more operational surfaces to validate

3. Memory budget includes COW headroom

maxmemory, when configured, is not total process RSS and does not reserve space for forked-child COW pages. Persistence jobs can increase RSS while the parent mutates memory. Capacity planning therefore includes dataset/overhead, client/output buffers, replication/AOF buffers, integrated indexes, allocator fragmentation, and expected COW peak under write load.

redis-cli · evidence to capture during a representative test
INFO memoryINFO persistenceINFO statsLATENCY LATESTSLOWLOG GET 20# Capture current_cow_peak, rdb_last_cow_size, aof_last_cow_size, latest_fork_usec, memory RSS/fragmentation, and tail latency.

4. Disk budget includes transient rewrite and backup state

Measure active MP-AOF bytes, RDB size, rewrite growth while the child runs, backup artifacts awaiting seal/retention cleanup, filesystem free space, and storage throttling/burst limits. A dataset that “fits on disk” can still fail a rewrite because the transient peak does not.

5. Benchmark persistence with realistic writes

Run the same bounded AtlasMart workload across candidate modes using disposable keys under atlasmart:ch17:l5:bench:. Hold payloads, key distribution, pipeline depth, concurrency, client connection model, and dataset size constant. Record throughput plus p50/p95/p99 command latency, fork duration, COW peak, fsync/rewrite metrics, storage utilization, and restart/replay time. An average-only benchmark hides the pauses persistence planning is supposed to expose.

Python · result schema, not fabricated numbers
result = {  "mode": "aof-everysec",  "requests": measured_requests,  "p50_ms": percentile(latencies, 50),  "p95_ms": percentile(latencies, 95),  "p99_ms": percentile(latencies, 99),  "fork_usec": observed_fork_usec,  "cow_peak_bytes": observed_cow_peak,  "restart_seconds": observed_restart_seconds,}print(result)# Fill only from an executed test. Do not copy example numbers from another machine.

6. Crash recovery and backup restore are separate acceptance tests

Test Failure injected Acceptance evidence
Process restart clean stop/restart expected key count + logs + load time
Abrupt process loss docker kill on disposable node survived/missing acknowledged markers vs configured RPO
Persistence corruption copy damage only a duplicate artifact validator/startup behavior and documented repair decision
New-node restore fresh empty node + retained backup integrity + key/application invariants + elapsed RTO
Replica/Sentinel/Cluster failover topology-specific node/network failure availability + acknowledged-write loss bounds + client recovery

7. Replication can improve availability but does not erase persistence questions

Redis replication is asynchronous. Sentinel or Cluster can promote a replica, but the promoted replica may not contain every acknowledged primary write. Persistence determines what a node can reconstruct after restart; replication determines another copy/topology; backup protects against broader failure and historical recovery. State the guarantee of each layer independently.

8. Security and platform constraints belong in durability design

Protect persistence and backup artifacts because they contain application data. Restrict file/volume access, encrypt storage/transport where appropriate, rotate credentials independently, and include ACL/TLS/config dependencies in restore drills. Managed Redis products may expose different persistence options and responsibilities; verify service-specific RPO, restore, backup ownership, and cost rather than mapping self-managed knobs one-to-one.

9. Decision worksheet

Question Evidence required
Maximum tolerable data loss? business RPO by data class
Maximum recovery time? measured startup/replay/restore RTO
Write latency budget? p50/p95/p99 under chosen fsync policy
Memory headroom? RSS + fragmentation + COW peak + buffers
Disk headroom? active files + rewrite/backup transient peak
Historical recovery needed? retention/off-host backup policy + restore tests
HA required? replication/Sentinel/Cluster topology and failover drills
Rollback path? version-compatible artifacts/config and tested downgrade/restore procedure

10. Production judgment and chapter synthesis

Choose the simplest persistence mode that meets measured RPO/RTO without violating latency, memory, disk, and operational constraints. Document what is guaranteed, what is merely likely, and what is outside Redis. Re-test after Redis upgrades, storage-class changes, dataset growth, major schema/index changes, or topology migrations because the previous headroom/latency evidence may no longer hold.

Check your understanding

  1. What is the first input to a persistence choice?
  2. Why is maxmemory insufficient for persistence capacity planning?
  3. What latency statistics should be reported?
  4. Does Sentinel failover make AOF or backup unnecessary?
  5. When should the durability benchmark be rerun?
Review the answers

Business recovery objectives such as RPO and RTO, not a Redis default.

Fork/COW, buffers, indexes, allocator overhead and RSS can exceed the configured data-memory budget.

At least a distribution such as p50/p95/p99, with workload and environment disclosed.

No. Availability, node persistence, and backup address different failures.

After material changes in version, storage, dataset size/write rate, topology, persistence policy, or integrated-index workload.

11. Summary and bridge

Chapter 17 converted persistence from a slogan into measured recovery engineering: RDB snapshots define point-in-time recovery, AOF narrows the normal loss window according to fsync policy, rewrite needs COW/disk headroom, combined startup prefers AOF, and backup requires restore proof. Chapter 18 now examines memory management, eviction, fragmentation, and out-of-memory prevention—the headroom constraints that persistence forks already exposed.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.