Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone
INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, MONITOR Caveats, and Exported Metrics
Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.
Redis Open Source 8.10.1 using the pinned
redis:8.10.1 image. The observability/benchmark
node is atlasmart-redis-ch24 on
127.0.0.1:6441, standalone topology, logical
database 0, AOF everysec plus RDB save rules,
maxmemory 0/noeviction unless a
bounded experiment says otherwise, named
academy-admin and atlasmart-app ACL
users, and fixture prefix atlasmart:ch24:*. TLS is
off only on this loopback-local disposable node; the capstone
security acceptance criteria reuse Chapter 22 TLS/ACL guidance.
Python examples target redis==8.1.0.
Search/JSON/vector/time-series/probabilistic features are
optional and must be included in capacity accounting only when
the chosen AtlasMart architecture actually uses them.
Learning outcomes
This lesson turns INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, MONITOR Caveats, and Exported Metrics into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.
Explain the mechanisms and terminology behind INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, MONITOR Caveats, and Exported Metrics.
Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.
Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.
Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.
Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.
1. Problem: the alert says “Redis latency,” but which layer is slow?
AtlasMart checkout reports a p99 response-time regression. Redis may be executing commands slowly, the client may be queued behind its own connection pool, the network may be delayed, a persistence fork may pause the process, or the application may simply be missing its cache and waiting on another database. Observability means collecting enough independent evidence to distinguish those mechanisms. A service-level objective (SLO) is the target reliability or latency level for a service; an error budget is the amount of allowed failure relative to that objective.
2. Evidence hierarchy: start with cheap snapshots
Use server metadata before intrusive tools. INFO is
sectioned and machine-readable;
commandstats reports calls and execution time by
command, latencystats reports command latency
percentiles, memory exposes allocator/process
state, replication exposes role/offsets, and
keyspace exposes logical-database counts. Redis
8.8+ also adds slow-log aggregate fields to command statistics.
These are server-side observations; they do not include the
application network round trip.
docker volume create atlasmart-redis-ch24-datadocker run -d --name atlasmart-redis-ch24 --restart no -p 127.0.0.1:6441:6379 -v atlasmart-redis-ch24-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --save 300 10 --latency-tracking yes --requirepass AtlasMart-Ch24-Bootstrap-Lab-Only-2026docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Bootstrap-Lab-Only-2026 atlasmart-redis-ch24 redis-cli ACL SETUSER academy-admin on ">AtlasMart-Ch24-Admin-Lab-Only-2026" "~*" "&*" +@alldocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Bootstrap-Lab-Only-2026 atlasmart-redis-ch24 redis-cli ACL SETUSER atlasmart-app on ">AtlasMart-Ch24-App-Lab-Only-2026" "~atlasmart:ch24:*" "&atlasmart:ch24:*" +@read +@write +@connection -@dangerousdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Bootstrap-Lab-Only-2026 atlasmart-redis-ch24 redis-cli ACL SETUSER default offdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO server
3. Capture a before-change observability bundle
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO serverdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO statsdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO commandstatsdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO latencystatsdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO keyspace
Expected evidence includes redis_version:8.10.1,
standalone role, command counters, latency percentiles,
used-memory/RSS values, and DB 0 key counts. A clean snapshot
proves the node state at capture time; it does not prove the
absence of short spikes between samples.
4. SLOWLOG answers “which commands blocked Redis?”
The Slow Log records commands whose
server execution time crosses
slowlog-log-slower-than. Redis explicitly excludes
client I/O from that duration, so a slow application request
with no Slow Log entry can still be caused by network, client
queues, serialization, or another upstream service. Save the
original threshold before changing it and restore it after the
lab.
OLD_SLOW=$(docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --raw --user academy-admin CONFIG GET slowlog-log-slower-than | tail -1)docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET slowlog-log-slower-than 0docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-App-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user atlasmart-app SET atlasmart:ch24:l1:probe valuedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin SLOWLOG GET 5docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET slowlog-log-slower-than "$OLD_SLOW"
5. LATENCY captures server events that command statistics alone can miss
Redis latency monitoring is disabled by default when
latency-monitor-threshold is 0. When enabled, the
framework records named event classes such as command, fork,
eviction, expiration, and AOF-related events.
LATENCY LATEST returns samples;
LATENCY DOCTOR turns those samples into a
human-readable diagnosis. Choose the threshold from the workload
SLO, not a universal number.
OLD_LAT=$(docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --raw --user academy-admin CONFIG GET latency-monitor-threshold | tail -1)docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET latency-monitor-threshold 1docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY LATESTdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY DOCTORdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET latency-monitor-threshold "$OLD_LAT"
6. MEMORY and CLIENT LIST explain pressure around the command engine
MEMORY STATS complements INFO memory,
while CLIENT LIST exposes connection identity,
age/idle time, authenticated user, RESP version, query buffers,
output buffers, and library metadata where the client supplies
it. A large omem or tot-mem can reveal
slow consumers even when commands themselves are fast.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin MEMORY STATSdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin MEMORY DOCTORdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CLIENT LIST
7. MONITOR is an emergency microscope, not a dashboard
Leaving MONITOR running on production because it shows every command. Redis documents MONITOR as an @admin/@slow/@dangerous debugging command and shows that even one MONITOR client can materially reduce throughput. It can also expose application keys and command arguments. Prefer INFO, SLOWLOG, latency tracking, tracing, and sampled application telemetry. If MONITOR is ever used, scope it to a disposable environment or a tightly bounded incident procedure.
MONITOR also does not log every administrative command, and sensitive authentication data is specially handled. Absence from MONITOR is therefore not evidence that no administrative/security event occurred.
8. Export metrics without confusing collection with interpretation
A production metrics system normally scrapes a supported Redis/managed-service exporter. The free local lab instead demonstrates a small JSON snapshot so the evidence pipeline stays reproducible without adding another daemon. The important design is stable metric names, timestamps, labels such as node/role, and a retention policy—not the choice of dashboard product.
import json, time, redisr = redis.Redis(host="127.0.0.1", port=6441, username="academy-admin", password="AtlasMart-Ch24-Admin-Lab-Only-2026", decode_responses=True)info = r.info()payload = { "captured_at_unix": time.time(), "redis_version": info.get("redis_version"), "used_memory": info.get("used_memory"), "used_memory_rss": info.get("used_memory_rss"), "connected_clients": info.get("connected_clients"), "instantaneous_ops_per_sec": info.get("instantaneous_ops_per_sec"), "keyspace_hits": info.get("keyspace_hits"), "keyspace_misses": info.get("keyspace_misses"), "evicted_keys": info.get("evicted_keys"),}print(json.dumps(payload, indent=2, sort_keys=True))
9. Evidence matrix: what each tool proves
| Evidence | Strong at | Does not prove |
|---|---|---|
| INFO commandstats / latencystats | Aggregate server command activity and tracked percentiles | Application round-trip or business success |
| SLOWLOG | Commands with slow server execution | Network/client delay |
| LATENCY | Named server latency events over threshold | Complete request traces |
| MEMORY | Redis allocator/process memory evidence | Container/host capacity by itself |
| CLIENT LIST | Connection/buffer/client identity state | Why an upstream caller is slow without correlation |
| MONITOR | Real-time processed command stream | Safe continuous observability or complete security audit |
Check your understanding
- Why can an application request be slow while SLOWLOG is empty?
- What does INFO latencystats add to commandstats?
- Why is MONITOR unsuitable as a normal dashboard source?
- What must an exported metric include besides a value?
Review the answers
SLOWLOG excludes network/client I/O and only records server execution above its threshold.
Per-command tracked latency percentiles, while commandstats focuses on calls/execution totals and related counters.
It streams every processed command, is intrusive, can expose workload details, and Redis documents a significant performance cost.
At minimum a stable meaning plus timestamp and node/topology labels so it can be correlated and compared.
10. Production judgment and bridge
Start incidents with low-cost evidence, correlate client and server time, and increase intrusiveness only when the hypothesis requires it. Tail latency, memory/buffer growth, persistence events, replication state, security identity, and workload shape belong in the same timeline. The next lesson decomposes latency mechanisms rather than treating every spike as a Redis command problem.
Summary and next step
INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, MONITOR Caveats, and Exported Metrics is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics.
Authoritative references
docker rm -f atlasmart-redis-ch24 2>/dev/null || truedocker volume rm atlasmart-redis-ch24-data 2>/dev/null || true# Do not run docker system prune; remove only resources named by this chapter.