Chapter 18 · Memory Management, Eviction, Expiration, Fragmentation, and OOM Prevention
Big Keys, Hot Keys, Large Collections, Lazy Freeing, Active Defrag, and Memory Diagnostics
Diagnose big and hot keys, large collections, asynchronous freeing, allocator fragmentation, and active defragmentation without turning diagnostics into destructive tuning.
Learning outcomes
AtlasMart is within its total memory budget, yet one shard/client path experiences latency spikes. Aggregate memory can look healthy while a single giant collection, hot key, slow free, or fragmented allocator causes operational pain.
Distinguish a big key from a hot key and explain why either can concentrate cost.
Use MEMORY USAGE plus redis-cli keyspace analyzers without KEYS/production-wide blocking scans.
Explain DEL versus UNLINK and observe lazyfree_pending_objects.
Interpret MEMORY DOCTOR, allocator metrics, and MEMORY PURGE cautiously.
Treat active defragmentation as measured CPU/memory tradeoff, not a blanket fix.
The course baseline remains Redis Open Source
8.10.1 from pinned image
redis:8.10.1. Memory-pressure experiments use
dedicated disposable Chapter 18 containers, named volumes, and
loopback-only host ports 6391–6395 so the shared Chapter 01
lab is never pushed toward eviction or out-of-memory behavior.
The temporary academy-admin password is
classroom-only. TLS is omitted only on loopback. Mandatory
topology is standalone, logical DB 0. Each lesson states its
own persistence, maxmemory, eviction, and container-memory
settings. Lesson 4 uses a 64 MiB maxmemory / allkeys-lfu node
so --hotkeys is available. Persistence is disabled. The
container limit is 192 MiB; all data is disposable.
1. Big and hot are different failure shapes
| Shape | What it means | Typical pressure |
|---|---|---|
| Big key | One key contains many elements or many bytes. | Long O(N) operations, large replies, slow deletion/freeing, network and memory spikes. |
| Hot key | One key receives disproportionate request rate. | CPU/event-loop concentration, one Cluster slot/shard becoming the bottleneck. |
| Many small keys | Large cardinality with per-key/object overhead. | Metadata/hashtable overhead and allocator pressure. |
| Fragmentation | Allocator active/resident pages exceed current allocations. | RSS remains high even after logical data shrinks. |
A key can be both big and hot, but the remediation differs. Splitting a 10-million-member collection addresses size; local caching/read replication or redesign may address read hotness.
2. Build a bounded diagnostic dataset
docker volume create atlasmart-redis-ch18-diagnostics_datadocker run --rm -v atlasmart-redis-ch18-diagnostics_data:/data redis:8.10.1 sh -lc 'printf "%s\n" "user default off" "user academy-admin on >AtlasMart-Admin-Lab-Only-2026 ~* &* +@all" > /data/users.acl'docker run -d --name atlasmart-redis-ch18-diagnostics --memory 192m -p 127.0.0.1:6394:6379 -v atlasmart-redis-ch18-diagnostics_data:/data redis:8.10.1 redis-server --dir /data --aclfile /data/users.acl --save '' --appendonly no --maxmemory 64mb --maxmemory-policy allkeys-lfudocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin ACL WHOAMI
import os, redisHOST = "127.0.0.1"PORT = 6394r = redis.Redis(host=HOST, port=PORT, username="academy-admin", password="AtlasMart-Admin-Lab-Only-2026", decode_responses=False)assert r.ping()pipe = r.pipeline(transaction=False)for i in range(1500): pipe.set(f"atlasmart:ch18:l4:small:{i}", b"s"*1024)pipe.execute()r.set("atlasmart:ch18:l4:large:string", b"L" * (2 * 1024 * 1024))for start in range(0, 20000, 1000): r.rpush("atlasmart:ch18:l4:large:list", *[f"item-{i}" for i in range(start,start+1000)])for _ in range(500): r.get("atlasmart:ch18:l4:small:0")print("done")
3. Measure individual keys before redesigning them
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY USAGE atlasmart:ch18:l4:large:stringdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin STRLEN atlasmart:ch18:l4:large:stringdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY USAGE atlasmart:ch18:l4:large:list SAMPLES 0docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin LLEN atlasmart:ch18:l4:large:listdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin OBJECT FREQ atlasmart:ch18:l4:small:0
MEMORY USAGE estimates nested aggregate memory
using five samples by default; SAMPLES 0 examines
all nested elements and can therefore be more expensive on a
genuinely huge collection. Use it deliberately.
4. Use redis-cli SCAN-based analyzers, not KEYS
--bigkeys, --memkeys, and newer
--keystats scan the keyspace incrementally.
--hotkeys is available only when the configured
policy is LFU.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin --memkeys --pattern "atlasmart:ch18:l4:*" --count 100docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin --bigkeys --pattern "atlasmart:ch18:l4:*" --count 100docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin --hotkeys --pattern "atlasmart:ch18:l4:*" --count 100
These modes use SCAN rather than KEYS, but scanning a large production keyspace still consumes CPU/network and may traverse the full dataset. Rate-limit and schedule diagnostics according to your environment.
5. UNLINK moves reclamation off the main deletion path
DEL synchronously frees the value.
UNLINK removes keys from the keyspace and lets
background lazy-free workers reclaim memory. That can reduce
main-thread deletion cost for large values, but it does not make
the work disappear; CPU and memory remain consumed until
reclamation catches up.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin UNLINK atlasmart:ch18:l4:large:listdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin EXISTS atlasmart:ch18:l4:large:listdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memory# Compare lazyfree_pending_objects and lazyfreed_objects. Do not expect RSS to drop immediately.
6. MEMORY DOCTOR is a hint engine, not a diagnosis substitute
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY DOCTORdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY STATSdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin CONFIG GET activedefrag active-defrag-ignore-bytes active-defrag-threshold-lower active-defrag-threshold-upper active-defrag-cycle-min active-defrag-cycle-max
A small lab may return that there is too little data to diagnose, which is a valid result. Production decisions should use absolute fragmentation bytes, allocator ratios, RSS history, workload phase, and latency—not a single advisory string.
7. Active defragmentation trades CPU for allocator compaction
When supported and enabled, Redis active defragmentation moves
allocations so allocator pages can become reclaimable.
active_defrag_running exposes current
activity/target CPU percentage. Enabling or tuning it without
measured fragmentation can add CPU work without meaningful
benefit.
This mandatory lab inspects active-defrag configuration and metrics but does not force it on. If you test it, do so only on an isolated load-representative node, record allocator fragmentation bytes and p50/p95/p99 latency before/after, and restore the original configuration.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin MEMORY PURGEdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch18-diagnostics redis-cli --user academy-admin INFO memory# MEMORY PURGE is allocator-dependent and can itself be slow; it is not a periodic production maintenance recipe.
8. Big/hot-key repair patterns
| Problem | Safer design direction | Tradeoff to measure |
|---|---|---|
| Large list/hash/zset | Partition by bounded business dimension or time bucket. | More keys/coordination and changed access patterns. |
| Hot read key | Application-local cache, replica reads where staleness is acceptable, or data fan-out. | Freshness and invalidation complexity. |
| Hot write key | Shard the logical counter/workload if semantics allow, or redesign contention point. | Aggregation/ordering complexity. |
| Slow deletion | UNLINK or lazyfree configuration after measurement. | Background CPU/memory backlog. |
| Fragmentation | Allow allocator reuse; consider defrag/purge only when supported and justified. | CPU spikes and incomplete RSS reduction. |
9. Verification checklist
-
Identify key size using both logical length/cardinality and
MEMORY USAGE. - Use SCAN-based analyzers with a bounded prefix in the lab.
- Confirm
--hotkeysrequires LFU policy. - Observe lazyfree metrics after UNLINK.
- Do not force active defrag merely because mem_fragmentation_ratio is high.
- Record absolute fragmentation bytes and tail latency before any tuning.
docker rm -f atlasmart-redis-ch18-diagnosticsdocker volume rm atlasmart-redis-ch18-diagnostics_data# Removes only this named Chapter 18 node and volume.
Check your understanding
- What is the difference between a big key and a hot key?
- Why can UNLINK improve latency without reducing total work?
- Why is --hotkeys policy-dependent?
- Should active defrag be enabled whenever RSS exceeds used_memory?
Review the answers
A big key is large in bytes/elements; a hot key is accessed disproportionately often. One key can be both.
It removes the key immediately but shifts value reclamation to background lazy-free workers; CPU/memory work still occurs.
It relies on LFU frequency metadata, so redis-cli supports it only with volatile-lfu or allkeys-lfu.
No. First distinguish allocator fragmentation from other RSS overhead and measure absolute bytes and latency; active defrag costs CPU.
Production judgment and next bridge
Memory pathologies are about shape and traffic as much as total bytes. Lesson 5 assembles the measurements from Chapters 17–18 into one process/container capacity budget that remains valid during forks, replication/failover, and index growth.
Summary and next step
Big Keys, Hot Keys, Large Collections, Lazy Freeing, Active Defrag, and Memory Diagnostics is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Build a Memory Budget Including Data, Indexes, Replication Buffers, Fork Headroom, and Failover.
Authoritative references
- Redis key eviction
- Redis memory optimization
- INFO
- MEMORY STATS
- MEMORY USAGE
- MEMORY DOCTOR
- MEMORY PURGE
- MEMORY MALLOC-STATS
- OBJECT FREQ
- OBJECT IDLETIME
- UNLINK
- Redis CLI keyspace analysis
- Redis latency monitoring
- Redis persistence
- Redis replication
- Redis Sentinel
- Redis Cluster specification
- Redis Search and query
- Redis 8.10 release notes
- Redis 8.10 whats new
- Redis ACLs
- Redis security
- Redis licenses
- Docker Official Redis image