Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone
Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics
Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.
Redis Open Source 8.10.1 using the pinned
redis:8.10.1 image. The observability/benchmark
node is atlasmart-redis-ch24 on
127.0.0.1:6441, standalone topology, logical
database 0, AOF everysec plus RDB save rules,
maxmemory 0/noeviction unless a
bounded experiment says otherwise, named
academy-admin and atlasmart-app ACL
users, and fixture prefix atlasmart:ch24:*. TLS is
off only on this loopback-local disposable node; the capstone
security acceptance criteria reuse Chapter 22 TLS/ACL guidance.
Python examples target redis==8.1.0.
Search/JSON/vector/time-series/probabilistic features are
optional and must be included in capacity accounting only when
the chosen AtlasMart architecture actually uses them.
Learning outcomes
This lesson turns Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.
Explain the mechanisms and terminology behind Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics.
Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.
Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.
Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.
Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.
1. Problem: p99 jumps every few minutes, but average Redis time looks normal
AtlasMart sees periodic checkout spikes while p50 remains healthy. Averages hide rare pauses, and Redis cannot execute faster than the scheduling baseline of the operating environment. Diagnose from the bottom up: intrinsic platform latency, Redis event latency, command execution, fork/persistence work, CPU pressure, storage stalls, network round-trip, and client-side queueing.
2. Intrinsic latency is a platform floor, not a Redis request benchmark
redis-cli --intrinsic-latency N does not connect to
Redis. Redis documentation requires running it on the machine
that runs Redis because it measures how long the process is
descheduled by the kernel/hypervisor. In this Docker learning
path, running it inside the Redis container shares the host
kernel but also reflects container/cgroup constraints; record
that limitation next to the result.
docker exec atlasmart-redis-ch24 redis-cli --intrinsic-latency 5# Record max latency and test location. Do not compare a laptop result directly with a production VM.
3. Separate client round trip from server execution
Measure both. redis-cli --latency observes
round-trip behavior from the client location; SLOWLOG and INFO
latencystats observe server execution. The difference can
include network delay, TCP/TLS overhead, client connection-pool
waits, serialization, scheduler delays, or proxy/load-balancer
time.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin --latency -i 0.1# Stop with Ctrl-C after a short bounded observation window.
4. LATENCY DOCTOR turns recorded events into hypotheses
Enable latency monitoring at a threshold related to the SLO,
reproduce a bounded event, then use LATENCY LATEST,
LATENCY HISTORY, and LATENCY DOCTOR.
The doctor is advisory evidence, not a root-cause oracle:
correlate it with SLOWLOG, persistence state, CPU/RSS, disk, and
client timing.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET latency-monitor-threshold 1docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY LATESTdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY DOCTOR
5. Fork pauses and copy-on-write must be correlated with persistence
RDB snapshots and AOF rewrites can fork a child process. Fork
latency and copy-on-write (COW) memory depend on dataset size,
write rate, kernel/virtualization behavior, and memory headroom.
Use INFO persistence before and after a bounded
BGSAVE, then correlate any
fork latency event instead of declaring that
persistence is always slow.
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin BGSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY HISTORY fork
6. Big commands are workload-size problems, not command-name folklore
Command complexity is documented per command, but actual latency depends on collection cardinality, reply size, CPU, memory locality, and concurrent work. A bounded large reply can also be slow for the client while SLOWLOG remains small because output transfer is not included in command execution time. This distinction is why both server and client timing are required.
import time, redisr = redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")key = "atlasmart:ch24:l2:list"r.delete(key)payload = b"x" * 256pipe = r.pipeline(transaction=False)for _ in range(20000): pipe.rpush(key, payload)pipe.execute()for stop in (99, 999, 9999, 19999): t0 = time.perf_counter_ns(); rows = r.lrange(key, 0, stop); dt = (time.perf_counter_ns()-t0)/1e6 print(stop+1, "items", round(dt,3), "ms", "bytes", sum(map(len, rows)))
7. CPU, RSS, disk, and network are context, not Redis counters
Redis cannot report host CPU steal time, Docker throttling,
filesystem queue depth, or a physical-network fault from
INFO alone. Capture platform context at the same
timestamps. In the local lab,
docker stats --no-stream gives container
CPU/memory/network/block-I/O counters. In production use the
platform’s supported host/container metrics and keep them
correlated with Redis samples.
docker stats --no-stream atlasmart-redis-ch24docker logs --tail 100 atlasmart-redis-ch24
8. A latency decision tree
| Observation | Next evidence | Likely boundary |
|---|---|---|
| Client p99 high; SLOWLOG/latencystats low | Network timing, pool wait, CLIENT LIST, application traces | Outside Redis command execution |
| SLOWLOG entries high | Command/data complexity, key cardinality, CPU | Command execution |
| LATENCY fork spikes | INFO persistence, RSS/COW, host scheduling | Persistence/fork |
| Intrinsic latency high | Host/hypervisor/cgroup metrics | Platform floor |
| Output buffers grow | CLIENT LIST omem/oll, slow consumer | Client/network backpressure |
| AOF-related latency events | Disk/fsync metrics, appendfsync policy | Storage/durability tradeoff |
9. Deliberately wrong: tune Redis before proving the layer
Changing maxmemory, disabling persistence, raising timeouts, or adding replicas because “Redis is slow” can hide the symptom while damaging durability or capacity. Capture an evidence bundle first, reproduce if safe, change one mechanism, and verify the same measurements after the change.
Check your understanding
- Why must intrinsic latency be measured on the Redis host?
- Why can LRANGE of a large reply feel slow without a corresponding huge SLOWLOG duration?
- What should accompany a fork latency event?
- Why is a timeout increase not a latency fix?
Review the answers
It measures kernel/hypervisor scheduling delay local to that execution environment, not a network request.
Slow Log excludes client I/O and reply transfer.
INFO persistence plus RSS/COW and host scheduling/memory evidence.
It changes how long callers wait; it does not remove the underlying latency mechanism.
10. Production judgment and bridge
Treat latency as a layered budget. Preserve durability and correctness while testing hypotheses, and report p50/p95/p99/max rather than averages alone. The next lesson turns this diagnostic discipline into representative load and capacity tests.
Summary and next step
Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency.
Authoritative references
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET latency-monitor-threshold 0docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-App-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user atlasmart-app UNLINK atlasmart:ch24:l2:list