Chapter 17 · Persistence: RDB Snapshots, AOF, fsync, Rewrites, and Crash Recovery
AOF Persistence: Command Logging, fsync Policies, Durability, and File Growth
Trace Append Only File persistence from acknowledged writes through buffering and fsync policies to file growth, restart replay, and realistic data-loss envelopes.
Learning outcomes
AtlasMart now wants a smaller loss window for order-processing metadata than periodic snapshots can provide. Append Only File (AOF) persistence records write commands and replays them at startup, but the application-visible reply and the storage-device durability point are still not identical.
Explain AOF command logging and Redis 7+ multi-part AOF layout.
Compare appendfsync always, everysec, and no without inventing zero-loss guarantees.
Observe AOF size/status and replay after restart.
Demonstrate an intentionally acknowledged write before abrupt process loss without fabricating the exact loss outcome.
Separate AOF persistence from backup and replication.
The normal course baseline remains Redis Open Source
8.10.1 from pinned image
redis:8.10.1. Because this chapter intentionally
restarts/crashes persistence processes, destructive exercises
use dedicated disposable Chapter 17 containers and named
volumes rather than the shared
atlasmart-redis-ch01 container. All ports bind
only to 127.0.0.1; the temporary lab password is
public classroom data, not a production secret; TLS is omitted
only on loopback. Mandatory topology is standalone, logical DB
0, no explicit maxmemory/eviction policy. This lesson uses a
dedicated AOF-only node on host port 6382.
1. AOF records state-changing commands for replay
With appendonly yes, Redis appends state-changing
operations in Redis protocol form so startup can replay them.
Since Redis 7, AOF is multi-part: a manifest identifies at most
one base file plus one or more incremental AOF files. Treat
these filenames as current-version implementation behavior, not
an application API.
docker volume create atlasmart-redis-ch17-aof_datadocker run --rm -v atlasmart-redis-ch17-aof_data:/data redis:8.10.1 sh -lc 'printf "%s\n" "user default off" "user academy-admin on >AtlasMart-Admin-Lab-Only-2026 ~* &* +@all" > /data/users.acl'docker run -d --name atlasmart-redis-ch17-aof -p 127.0.0.1:6382:6379 -v atlasmart-redis-ch17-aof_data:/data redis:8.10.1 redis-server --dir /data --aclfile /data/users.acl --appendonly yes --appendfsync everysec --save docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin ACL WHOAMI# users.acl lives in the named volume, so academy-admin remains available after the deliberate restart/crash.
2. The fsync policy defines a durability/latency tradeoff
| appendfsync | Mechanism | Failure window / cost |
|---|---|---|
| always | fsync after each batch of appended writes before replies | Strongest local fsync policy; highest latency/IO cost |
| everysec | background fsync about once per second | Fast default; disaster can lose roughly a second of acknowledged writes |
| no | Redis does not call fsync; OS decides flush timing | Lowest Redis fsync cost; widest/OS-dependent loss window |
everysec does not mean zero loss. Redis
documentation explicitly allows approximately one second of
loss in a disaster, and slow storage can also influence
latency behavior.
3. Make the AOF observable
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin SET atlasmart:ch17:l2:order:1001 paiddocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INCR atlasmart:ch17:l2:sequencedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin HSET atlasmart:ch17:l2:payment:1001 state captureddocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin CONFIG GET appendonly appendfsync appenddirname
docker exec atlasmart-redis-ch17-aof sh -lc 'find /data -maxdepth 2 -type f -printf "%p %s bytes\n" | sort'# Expect an appendonly directory with a manifest and current base/incremental components; exact generation numbers can vary.
4. Acknowledged write before crash: record, do not predetermine
To teach the loss window honestly, write a unique marker and immediately kill the dedicated process. On some runs the marker may already have been fsynced; on others it may not. The lesson therefore records the result instead of promising deterministic loss. Repeat many isolated trials only if you are measuring a real storage environment.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin SET atlasmart:ch17:l2:crash-marker acknowledged-before-killdocker kill atlasmart-redis-ch17-aofdocker start atlasmart-redis-ch17-aofsleep 1docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin GET atlasmart:ch17:l2:crash-markerdocker logs --tail 50 atlasmart-redis-ch17-aof# Record survived/missing; either outcome is compatible with everysec depending on timing/storage.
5. Replay evidence is not a backup certificate
After restart, INFO persistence, logs, and key
values prove what this node reconstructed from its AOF set. They
do not prove an independent copy exists. For Redis 7+ MP-AOF,
copying the live append directory during a rewrite without a
consistency procedure can yield an invalid backup set.
If you need an AOF backup, follow a consistency procedure
around rewrites or use the Redis 8.10
BACKUP family. Backup success still requires an
actual restore test.
6. File growth is expected
AOF logs operations, so repeated changes to the same logical
value can grow the log far beyond the current dataset
representation. aof_current_size and
aof_base_size are useful evidence. Rewrite compacts
that history into a new base representation; Lesson 3 measures
that process and its temporary resource demand.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin SET atlasmart:ch17:l2:counter 0docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INCR atlasmart:ch17:l2:counterdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INCR atlasmart:ch17:l2:counterdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INCR atlasmart:ch17:l2:counterdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INCR atlasmart:ch17:l2:counterdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INCR atlasmart:ch17:l2:counterdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INCR atlasmart:ch17:l2:counterdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INCR atlasmart:ch17:l2:counterdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INCR atlasmart:ch17:l2:counterdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch17-aof redis-cli --user academy-admin INFO persistence
7. Integrity handling: inspect before repair
Redis can tolerate some truncated AOF tails depending on
aof-load-truncated. For genuine corruption,
redis-check-aof can inspect and optionally repair,
but --fix can discard everything after the first
invalid region. Always copy the artifact first and understand
the proposed loss before mutation.
This lesson does not hex-edit or truncate the only persistence artifact. Corruption repair is practiced later only on a disposable copy, never on an authoritative file.
8. Cleanup
docker rm -f atlasmart-redis-ch17-aofdocker volume rm atlasmart-redis-ch17-aof_data# Removes only the named Chapter 17 container/volume. Never use docker system prune for this lab.
9. Production judgment
Choose AOF when a smaller RPO than snapshots is worth additional disk writes, restart replay cost, rewrite planning, and fsync latency. Benchmark with the actual filesystem/container/storage class, persistence policy, payload sizes, write rate, pipeline depth, and concurrency. Track AOF write/rewrite status, fsync delays, current/base size, COW memory, disk free space, and startup replay time.
Check your understanding
- Does everysec guarantee every acknowledged write survives a power loss?
- Why can AOF be much larger than the dataset?
- What is MP-AOF?
- Should redis-check-aof --fix be the first action on the only copy?
- Does successful AOF replay prove you have a backup?
Review the answers
No. Redis documents an approximately one-second possible loss window.
It records write history until rewrite compacts it.
Redis 7+ uses a manifest plus a base file and incremental AOF files.
No. Preserve a copy, inspect without --fix, understand the truncation/loss, then decide.
No; it proves this node can reconstruct from its local persistence set.
10. Summary and bridge
AOF narrows the normal loss window but introduces fsync and log-growth costs. Lesson 3 examines the rewrite that bounds that growth and the temporary disk/COW headroom it needs.
Authoritative references
- Redis persistence
- INFO
- BGSAVE
- LASTSAVE
- SAVE
- BGREWRITEAOF
- CONFIG GET
- CONFIG SET
- Redis 8.10 release notes
- Redis 8.10 whats new
- BACKUP
- Redis replication
- Redis Sentinel
- Redis Cluster specification
- Redis ACLs
- Redis latency diagnosis
- Redis memory optimization
- Redis security
- Redis licenses
- Docker Official Redis image