Chapter 17 · Persistence: RDB Snapshots, AOF, fsync, Rewrites, and Crash Recovery
Choose Durability Settings from RPO, Latency, Memory, Fork Behavior, and Storage Characteristics
Build a measured durability decision framework from recovery point objective, latency, fork/COW memory, disk behavior, startup time, replication, and backup requirements.
Learning outcomes
The final design question is not “RDB or AOF?” in isolation. AtlasMart must choose a persistence envelope that meets business recovery objectives on its actual storage and memory limits, while preserving enough headroom for fork/rewrite and separating node durability from service availability and backup.
Translate business RPO/RTO into persistence and backup requirements.
Compare no persistence, RDB, AOF policies, and combined modes without universal tuning values.
Budget fork/COW memory and rewrite disk headroom alongside steady-state dataset memory.
Define a benchmark/recovery matrix with p50/p95/p99 latency and measured restart/restore time.
Produce an operational acceptance checklist spanning persistence, replication, backup, security, and rollback.
The normal course baseline remains Redis Open Source
8.10.1 from pinned image
redis:8.10.1. Because this chapter intentionally
restarts/crashes persistence processes, destructive exercises
use dedicated disposable Chapter 17 containers and named
volumes rather than the shared
atlasmart-redis-ch01 container. All ports bind
only to 127.0.0.1; the temporary lab password is
public classroom data, not a production secret; TLS is omitted
only on loopback. Mandatory topology is standalone, logical DB
0, no explicit maxmemory/eviction policy. This lesson is a
decision/measurement capstone; it does not mutate the shared
Chapter 01 node.
1. Start with RPO and RTO, not a config snippet
Recovery point objective (RPO) is the maximum acceptable data-loss interval. Recovery time objective (RTO) is the target time to restore acceptable service. RDB save cadence, AOF fsync policy, backup frequency, startup replay time, and replication topology all contribute differently. None should be selected solely because it is a Redis default.
2. Compare the persistence modes honestly
| Mode | Normal durability shape | Primary costs / caveats |
|---|---|---|
| No persistence | memory only | all node-local data can disappear on restart |
| RDB | last successful snapshot | fork/COW + snapshot IO; wider RPO between saves |
| AOF everysec | recent command log; roughly one-second disaster window documented | continuous writes + fsync/rewrite + replay cost |
| AOF always | fsync every command batch | higher write latency/IO; still not an off-host backup |
| RDB + AOF | AOF used for normal startup, RDB also available | pays both mechanisms; more operational surfaces to validate |
3. Memory budget includes COW headroom
maxmemory, when configured, is not total process
RSS and does not reserve space for forked-child COW pages.
Persistence jobs can increase RSS while the parent mutates
memory. Capacity planning therefore includes dataset/overhead,
client/output buffers, replication/AOF buffers, integrated
indexes, allocator fragmentation, and expected COW peak under
write load.
INFO memoryINFO persistenceINFO statsLATENCY LATESTSLOWLOG GET 20# Capture current_cow_peak, rdb_last_cow_size, aof_last_cow_size, latest_fork_usec, memory RSS/fragmentation, and tail latency.
4. Disk budget includes transient rewrite and backup state
Measure active MP-AOF bytes, RDB size, rewrite growth while the child runs, backup artifacts awaiting seal/retention cleanup, filesystem free space, and storage throttling/burst limits. A dataset that “fits on disk” can still fail a rewrite because the transient peak does not.
5. Benchmark persistence with realistic writes
Run the same bounded AtlasMart workload across candidate modes
using disposable keys under
atlasmart:ch17:l5:bench:. Hold payloads, key
distribution, pipeline depth, concurrency, client connection
model, and dataset size constant. Record throughput plus
p50/p95/p99 command latency, fork duration, COW peak,
fsync/rewrite metrics, storage utilization, and restart/replay
time. An average-only benchmark hides the pauses persistence
planning is supposed to expose.
result = { "mode": "aof-everysec", "requests": measured_requests, "p50_ms": percentile(latencies, 50), "p95_ms": percentile(latencies, 95), "p99_ms": percentile(latencies, 99), "fork_usec": observed_fork_usec, "cow_peak_bytes": observed_cow_peak, "restart_seconds": observed_restart_seconds,}print(result)# Fill only from an executed test. Do not copy example numbers from another machine.
6. Crash recovery and backup restore are separate acceptance tests
| Test | Failure injected | Acceptance evidence |
|---|---|---|
| Process restart | clean stop/restart | expected key count + logs + load time |
| Abrupt process loss | docker kill on disposable node | survived/missing acknowledged markers vs configured RPO |
| Persistence corruption copy | damage only a duplicate artifact | validator/startup behavior and documented repair decision |
| New-node restore | fresh empty node + retained backup | integrity + key/application invariants + elapsed RTO |
| Replica/Sentinel/Cluster failover | topology-specific node/network failure | availability + acknowledged-write loss bounds + client recovery |
7. Replication can improve availability but does not erase persistence questions
Redis replication is asynchronous. Sentinel or Cluster can promote a replica, but the promoted replica may not contain every acknowledged primary write. Persistence determines what a node can reconstruct after restart; replication determines another copy/topology; backup protects against broader failure and historical recovery. State the guarantee of each layer independently.
8. Security and platform constraints belong in durability design
Protect persistence and backup artifacts because they contain application data. Restrict file/volume access, encrypt storage/transport where appropriate, rotate credentials independently, and include ACL/TLS/config dependencies in restore drills. Managed Redis products may expose different persistence options and responsibilities; verify service-specific RPO, restore, backup ownership, and cost rather than mapping self-managed knobs one-to-one.
9. Decision worksheet
| Question | Evidence required |
|---|---|
| Maximum tolerable data loss? | business RPO by data class |
| Maximum recovery time? | measured startup/replay/restore RTO |
| Write latency budget? | p50/p95/p99 under chosen fsync policy |
| Memory headroom? | RSS + fragmentation + COW peak + buffers |
| Disk headroom? | active files + rewrite/backup transient peak |
| Historical recovery needed? | retention/off-host backup policy + restore tests |
| HA required? | replication/Sentinel/Cluster topology and failover drills |
| Rollback path? | version-compatible artifacts/config and tested downgrade/restore procedure |
10. Production judgment and chapter synthesis
Choose the simplest persistence mode that meets measured RPO/RTO without violating latency, memory, disk, and operational constraints. Document what is guaranteed, what is merely likely, and what is outside Redis. Re-test after Redis upgrades, storage-class changes, dataset growth, major schema/index changes, or topology migrations because the previous headroom/latency evidence may no longer hold.
Check your understanding
- What is the first input to a persistence choice?
- Why is maxmemory insufficient for persistence capacity planning?
- What latency statistics should be reported?
- Does Sentinel failover make AOF or backup unnecessary?
- When should the durability benchmark be rerun?
Review the answers
Business recovery objectives such as RPO and RTO, not a Redis default.
Fork/COW, buffers, indexes, allocator overhead and RSS can exceed the configured data-memory budget.
At least a distribution such as p50/p95/p99, with workload and environment disclosed.
No. Availability, node persistence, and backup address different failures.
After material changes in version, storage, dataset size/write rate, topology, persistence policy, or integrated-index workload.
11. Summary and bridge
Chapter 17 converted persistence from a slogan into measured recovery engineering: RDB snapshots define point-in-time recovery, AOF narrows the normal loss window according to fsync policy, rewrite needs COW/disk headroom, combined startup prefers AOF, and backup requires restore proof. Chapter 18 now examines memory management, eviction, fragmentation, and out-of-memory prevention—the headroom constraints that persistence forks already exposed.
Authoritative references
- Redis persistence
- INFO
- BGSAVE
- LASTSAVE
- SAVE
- BGREWRITEAOF
- CONFIG GET
- CONFIG SET
- Redis 8.10 release notes
- Redis 8.10 whats new
- BACKUP
- Redis replication
- Redis Sentinel
- Redis Cluster specification
- Redis ACLs
- Redis latency diagnosis
- Redis memory optimization
- Redis security
- Redis licenses
- Docker Official Redis image