Chapter 17 · Persistence: RDB Snapshots, AOF, fsync, Rewrites, and Crash Recovery
Use RDB + AOF Together, Understand Startup Recovery Order, and Validate Persistence Files
Prove startup behavior when RDB and AOF are enabled together, validate persistence artifacts safely, and separate persistence recovery from a tested backup/restore process.
Learning outcomes
AtlasMart wants compact snapshots for portable recovery points while keeping AOF for a tighter node-restart RPO. When both are enabled, operators must know which artifact wins at startup and how to validate files without turning a repair tool into accidental data deletion.
Prove normal startup prefers AOF when both RDB and AOF are enabled.
Inspect both dump.rdb and the Redis 7+ MP-AOF set.
Use redis-check-rdb / redis-check-aof as inspection tools before any repair.
Explain truncated versus corrupted AOF handling and the risk of --fix.
Separate local persistence validation from backup/restore acceptance.
The normal course baseline remains Redis Open Source
8.10.1 from pinned image
redis:8.10.1. Because this chapter intentionally
restarts/crashes persistence processes, destructive exercises
use dedicated disposable Chapter 17 containers and named
volumes rather than the shared
atlasmart-redis-ch01 container. All ports bind
only to 127.0.0.1; the temporary lab password is
public classroom data, not a production secret; TLS is omitted
only on loopback. Mandatory topology is standalone, logical DB
0, no explicit maxmemory/eviction policy. This lesson uses a
dedicated combined-persistence node on host port 6384.
1. Both mechanisms can coexist
RDB gives explicit point-in-time snapshots while AOF records subsequent writes. Redis documentation states that on a normal restart with both enabled, Redis reconstructs from AOF because it is expected to be the more complete history. This is a startup rule, not a license to ignore RDB or backup design.
docker volume create atlasmart-redis-ch17-both_datadocker run --rm -v atlasmart-redis-ch17-both_data:/data redis:8.10.1 sh -lc 'printf "%s\n" "user default off" "user academy-admin on >AtlasMart-Admin-Lab-Only-2026 ~* &* +@all" > /data/users.acl'docker run -d --name atlasmart-redis-ch17-both -p 127.0.0.1:6384:6379 -v atlasmart-redis-ch17-both_data:/data redis:8.10.1 redis-server --dir /data --aclfile /data/users.acl --appendonly yes --appendfsync everysec --save 60 1 --dbfilename dump.rdbdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin ACL WHOAMI# users.acl lives in the named volume, so academy-admin remains available after the deliberate restart/crash.
2. Create an RDB point, then a newer AOF-only write
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin SET atlasmart:ch17:l4:snapshotted v1docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin BGSAVEsleep 1docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin LASTSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin SET atlasmart:ch17:l4:after-snapshot v2docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin INFO persistence
Now dump.rdb can represent the first key, while AOF
should also include the later write. This creates an observable
difference between the two sources without corrupting either.
3. Restart and prove AOF precedence
docker restart atlasmart-redis-ch17-bothsleep 1docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin MGET atlasmart:ch17:l4:snapshotted atlasmart:ch17:l4:after-snapshotdocker logs --tail 60 atlasmart-redis-ch17-bothdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin INFO persistence
If the later key survives, that is consistent with normal AOF-first reconstruction. The log provides stronger evidence than merely noticing both files exist.
4. Inspect artifacts before any repair
docker exec atlasmart-redis-ch17-both sh -lc 'find /data -maxdepth 2 -type f -printf "%p %s bytes\n" | sort'
docker exec atlasmart-redis-ch17-both sh -lc 'test -f /data/dump.rdb && redis-check-rdb /data/dump.rdb || true'docker exec atlasmart-redis-ch17-both sh -lc 'find /data -name "*.aof" -type f -maxdepth 3 -print'# For a real corrupted AOF, make a copy before redis-check-aof --fix; fixing can discard data after the invalid region.
5. Truncation and corruption are different recovery cases
A crash can leave a partial final AOF command. With the default truncated-load behavior, modern Redis can discard a malformed tail and continue loading, accepting the associated data loss. Corruption in the middle is more serious and can abort startup. Repair tools are not magic: inspect the offset and preserve the original before modifying anything.
If you want a corruption exercise, copy the whole Chapter 17 volume to a disposable second volume/container, alter only that copy, then compare validator output and restored state. The mandatory lesson stops at safe inspection.
6. Redis 8.10 BACKUP is a separate backup workflow
Redis 8.10 adds a BACKUP command family that uses
MP-AOF-compatible BASE, INCR, and manifest artifacts and
supports a seal boundary. That is relevant because persistence
files inside the live data directory are not automatically an
off-host backup. The new command does not remove the requirement
to retain copies independently and restore-test them.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin COMMAND INFO BACKUPdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin CONFIG GET preload-file backup-sealed-ttl# Do not start a backup unless you have defined destination, retention, seal/cleanup, and restore test.
7. Startup preload is intentionally explicit
Redis 8.10 preload-file is startup-only and can
point to an RDB or AOF manifest. When configured, it bypasses
normal appenddirname/dump.rdb loading for that startup and
aborts if the selected artifact cannot load. Use it for a
controlled restore workflow, not as an ad-hoc production switch.
8. Persistence is still not backup
| Capability | Primary purpose | Common mistaken claim |
|---|---|---|
| RDB | point-in-time node persistence / portable snapshots | ‘we have dump.rdb, therefore DR is solved’ |
| AOF | lower-RPO node reconstruction | ‘everysec means no acknowledged write can be lost’ |
| Replication | availability/read scaling | ‘the replica is an independent backup’ |
| Backup | independent retained recovery artifact | ‘copy succeeded, therefore restore works’ |
| Restore drill | proof of recoverability | ‘we can test only during an incident’ |
9. Cleanup and production judgment
docker rm -f atlasmart-redis-ch17-bothdocker volume rm atlasmart-redis-ch17-both_data# Removes only the named Chapter 17 container/volume. Never use docker system prune for this lab.
For production, document normal startup source, restore-source override, artifact retention, integrity checks, encryption/access controls, off-host/off-site copies, RPO/RTO, and application reconciliation. Test both “node restart from local persistence” and “new node from backup” because they are different procedures with different failure modes.
Check your understanding
- When RDB and AOF are both enabled, what does normal Redis startup prefer?
- Should redis-check-aof --fix be run before preserving a copy?
- What does BACKUP add in Redis 8.10?
- Does a valid persistence file prove disaster recovery is complete?
- Why test a fresh-node restore separately from restart?
Review the answers
AOF, because it is expected to be the more complete dataset.
No. Preserve the original and inspect first; --fix can discard a large tail.
A node-side self-contained MP-AOF-compatible backup workflow with base, increment, manifest, and seal semantics.
No. You still need independent retention and a tested restore/application-validation workflow.
Restart may rely on local files/config that a disaster has also destroyed.
10. Summary and bridge
Combined persistence improves recovery options, but startup precedence and repair behavior must be explicit. Lesson 5 turns these mechanisms into a durability decision based on measurable RPO, latency, memory, storage, and recovery objectives.
Authoritative references
- Redis persistence
- INFO
- BGSAVE
- LASTSAVE
- SAVE
- BGREWRITEAOF
- CONFIG GET
- CONFIG SET
- Redis 8.10 release notes
- Redis 8.10 whats new
- BACKUP
- Redis replication
- Redis Sentinel
- Redis Cluster specification
- Redis ACLs
- Redis latency diagnosis
- Redis memory optimization
- Redis security
- Redis licenses
- Docker Official Redis image