Chapter 17 · Persistence: RDB Snapshots, AOF, fsync, Rewrites, and Crash Recovery

Use RDB + AOF Together, Understand Startup Recovery Order, and Validate Persistence Files

Prove startup behavior when RDB and AOF are enabled together, validate persistence artifacts safely, and separate persistence recovery from a tested backup/restore process.

Advanced180–250 minutesRDB+AOF startup order, validation, corruption handlingRedis Open Source 8.10.1Docker local labFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

AtlasMart wants compact snapshots for portable recovery points while keeping AOF for a tighter node-restart RPO. When both are enabled, operators must know which artifact wins at startup and how to validate files without turning a repair tool into accidental data deletion.

01

Prove normal startup prefers AOF when both RDB and AOF are enabled.

02

Inspect both dump.rdb and the Redis 7+ MP-AOF set.

03

Use redis-check-rdb / redis-check-aof as inspection tools before any repair.

04

Explain truncated versus corrupted AOF handling and the risk of --fix.

05

Separate local persistence validation from backup/restore acceptance.

Exact Chapter 17 lab boundary

The normal course baseline remains Redis Open Source 8.10.1 from pinned image redis:8.10.1. Because this chapter intentionally restarts/crashes persistence processes, destructive exercises use dedicated disposable Chapter 17 containers and named volumes rather than the shared atlasmart-redis-ch01 container. All ports bind only to 127.0.0.1; the temporary lab password is public classroom data, not a production secret; TLS is omitted only on loopback. Mandatory topology is standalone, logical DB 0, no explicit maxmemory/eviction policy. This lesson uses a dedicated combined-persistence node on host port 6384.

1. Both mechanisms can coexist

RDB gives explicit point-in-time snapshots while AOF records subsequent writes. Redis documentation states that on a normal restart with both enabled, Redis reconstructs from AOF because it is expected to be the more complete history. This is a startup rule, not a license to ignore RDB or backup design.

Shell · create one isolated persistence node
docker volume create atlasmart-redis-ch17-both_datadocker run --rm -v atlasmart-redis-ch17-both_data:/data redis:8.10.1 sh -lc 'printf "%s\n" "user default off" "user academy-admin on >AtlasMart-Admin-Lab-Only-2026 ~* &* +@all" > /data/users.acl'docker run -d --name atlasmart-redis-ch17-both -p 127.0.0.1:6384:6379 -v atlasmart-redis-ch17-both_data:/data redis:8.10.1 redis-server --dir /data --aclfile /data/users.acl --appendonly yes --appendfsync everysec --save 60 1 --dbfilename dump.rdbdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin ACL WHOAMI# users.acl lives in the named volume, so academy-admin remains available after the deliberate restart/crash.

2. Create an RDB point, then a newer AOF-only write

redis-cli · construct two recovery generations
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin SET atlasmart:ch17:l4:snapshotted v1docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin BGSAVEsleep 1docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin LASTSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin SET atlasmart:ch17:l4:after-snapshot v2docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin INFO persistence

Now dump.rdb can represent the first key, while AOF should also include the later write. This creates an observable difference between the two sources without corrupting either.

3. Restart and prove AOF precedence

Shell · normal restart and state check
docker restart atlasmart-redis-ch17-bothsleep 1docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin MGET atlasmart:ch17:l4:snapshotted atlasmart:ch17:l4:after-snapshotdocker logs --tail 60 atlasmart-redis-ch17-bothdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin INFO persistence

If the later key survives, that is consistent with normal AOF-first reconstruction. The log provides stronger evidence than merely noticing both files exist.

4. Inspect artifacts before any repair

Shell · locate persistence files
docker exec atlasmart-redis-ch17-both sh -lc 'find /data -maxdepth 2 -type f -printf "%p %s bytes\n" | sort'
Shell · validate copies/read-only first
docker exec atlasmart-redis-ch17-both sh -lc 'test -f /data/dump.rdb && redis-check-rdb /data/dump.rdb || true'docker exec atlasmart-redis-ch17-both sh -lc 'find /data -name "*.aof" -type f -maxdepth 3 -print'# For a real corrupted AOF, make a copy before redis-check-aof --fix; fixing can discard data after the invalid region.

5. Truncation and corruption are different recovery cases

A crash can leave a partial final AOF command. With the default truncated-load behavior, modern Redis can discard a malformed tail and continue loading, accepting the associated data loss. Corruption in the middle is more serious and can abort startup. Repair tools are not magic: inspect the offset and preserve the original before modifying anything.

Never practice corruption on the only artifact

If you want a corruption exercise, copy the whole Chapter 17 volume to a disposable second volume/container, alter only that copy, then compare validator output and restored state. The mandatory lesson stops at safe inspection.

6. Redis 8.10 BACKUP is a separate backup workflow

Redis 8.10 adds a BACKUP command family that uses MP-AOF-compatible BASE, INCR, and manifest artifacts and supports a seal boundary. That is relevant because persistence files inside the live data directory are not automatically an off-host backup. The new command does not remove the requirement to retain copies independently and restore-test them.

redis-cli · feature-gate the current server
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin COMMAND INFO BACKUPdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-both redis-cli --user academy-admin CONFIG GET preload-file backup-sealed-ttl# Do not start a backup unless you have defined destination, retention, seal/cleanup, and restore test.

7. Startup preload is intentionally explicit

Redis 8.10 preload-file is startup-only and can point to an RDB or AOF manifest. When configured, it bypasses normal appenddirname/dump.rdb loading for that startup and aborts if the selected artifact cannot load. Use it for a controlled restore workflow, not as an ad-hoc production switch.

8. Persistence is still not backup

Capability Primary purpose Common mistaken claim
RDB point-in-time node persistence / portable snapshots ‘we have dump.rdb, therefore DR is solved’
AOF lower-RPO node reconstruction ‘everysec means no acknowledged write can be lost’
Replication availability/read scaling ‘the replica is an independent backup’
Backup independent retained recovery artifact ‘copy succeeded, therefore restore works’
Restore drill proof of recoverability ‘we can test only during an incident’

9. Cleanup and production judgment

Shell · bounded cleanup
docker rm -f atlasmart-redis-ch17-bothdocker volume rm atlasmart-redis-ch17-both_data# Removes only the named Chapter 17 container/volume. Never use docker system prune for this lab.

For production, document normal startup source, restore-source override, artifact retention, integrity checks, encryption/access controls, off-host/off-site copies, RPO/RTO, and application reconciliation. Test both “node restart from local persistence” and “new node from backup” because they are different procedures with different failure modes.

Check your understanding

  1. When RDB and AOF are both enabled, what does normal Redis startup prefer?
  2. Should redis-check-aof --fix be run before preserving a copy?
  3. What does BACKUP add in Redis 8.10?
  4. Does a valid persistence file prove disaster recovery is complete?
  5. Why test a fresh-node restore separately from restart?
Review the answers

AOF, because it is expected to be the more complete dataset.

No. Preserve the original and inspect first; --fix can discard a large tail.

A node-side self-contained MP-AOF-compatible backup workflow with base, increment, manifest, and seal semantics.

No. You still need independent retention and a tested restore/application-validation workflow.

Restart may rely on local files/config that a disaster has also destroyed.

10. Summary and bridge

Combined persistence improves recovery options, but startup precedence and repair behavior must be explicit. Lesson 5 turns these mechanisms into a durability decision based on measurable RPO, latency, memory, storage, and recovery objectives.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.