Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone

Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics

Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.

Advanced210–320 minuteslatency doctor, intrinsic latency, fork, CPU, network, storageRedis Open Source 8.10.1redis-py 8.1.0 where Python is usedDocker + redis-cli + PythonStandalone evidence node · DB 0AOF everysec + RDB · maxmemory 0/noeviction baselineNamed ACL users · TLS off only on loopbackReuses Sentinel/Cluster labs for failover acceptanceFree/local-firstLast reviewed: September 6, 2026
Reproducible Chapter 24 baseline

Redis Open Source 8.10.1 using the pinned redis:8.10.1 image. The observability/benchmark node is atlasmart-redis-ch24 on 127.0.0.1:6441, standalone topology, logical database 0, AOF everysec plus RDB save rules, maxmemory 0/noeviction unless a bounded experiment says otherwise, named academy-admin and atlasmart-app ACL users, and fixture prefix atlasmart:ch24:*. TLS is off only on this loopback-local disposable node; the capstone security acceptance criteria reuse Chapter 22 TLS/ACL guidance. Python examples target redis==8.1.0. Search/JSON/vector/time-series/probabilistic features are optional and must be included in capacity accounting only when the chosen AtlasMart architecture actually uses them.

Learning outcomes

This lesson turns Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.

01

Explain the mechanisms and terminology behind Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics.

02

Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.

03

Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.

04

Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.

05

Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.

1. Problem: p99 jumps every few minutes, but average Redis time looks normal

AtlasMart sees periodic checkout spikes while p50 remains healthy. Averages hide rare pauses, and Redis cannot execute faster than the scheduling baseline of the operating environment. Diagnose from the bottom up: intrinsic platform latency, Redis event latency, command execution, fork/persistence work, CPU pressure, storage stalls, network round-trip, and client-side queueing.

2. Intrinsic latency is a platform floor, not a Redis request benchmark

redis-cli --intrinsic-latency N does not connect to Redis. Redis documentation requires running it on the machine that runs Redis because it measures how long the process is descheduled by the kernel/hypervisor. In this Docker learning path, running it inside the Redis container shares the host kernel but also reflects container/cgroup constraints; record that limitation next to the result.

Shell · bounded intrinsic latency sample
docker exec atlasmart-redis-ch24 redis-cli --intrinsic-latency 5# Record max latency and test location. Do not compare a laptop result directly with a production VM.

3. Separate client round trip from server execution

Measure both. redis-cli --latency observes round-trip behavior from the client location; SLOWLOG and INFO latencystats observe server execution. The difference can include network delay, TCP/TLS overhead, client connection-pool waits, serialization, scheduler delays, or proxy/load-balancer time.

Shell · client-side latency sample
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin --latency -i 0.1# Stop with Ctrl-C after a short bounded observation window.

4. LATENCY DOCTOR turns recorded events into hypotheses

Enable latency monitoring at a threshold related to the SLO, reproduce a bounded event, then use LATENCY LATEST, LATENCY HISTORY, and LATENCY DOCTOR. The doctor is advisory evidence, not a root-cause oracle: correlate it with SLOWLOG, persistence state, CPU/RSS, disk, and client timing.

Shell · inspect the latency event framework
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET latency-monitor-threshold 1docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY LATESTdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY DOCTOR

5. Fork pauses and copy-on-write must be correlated with persistence

RDB snapshots and AOF rewrites can fork a child process. Fork latency and copy-on-write (COW) memory depend on dataset size, write rate, kernel/virtualization behavior, and memory headroom. Use INFO persistence before and after a bounded BGSAVE, then correlate any fork latency event instead of declaring that persistence is always slow.

Shell · bounded fork evidence
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin BGSAVEdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO persistencedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY HISTORY fork

6. Big commands are workload-size problems, not command-name folklore

Command complexity is documented per command, but actual latency depends on collection cardinality, reply size, CPU, memory locality, and concurrent work. A bounded large reply can also be slow for the client while SLOWLOG remains small because output transfer is not included in command execution time. This distinction is why both server and client timing are required.

Python · bounded big-reply comparison
import time, redisr = redis.Redis(host="127.0.0.1", port=6441, username="atlasmart-app", password="AtlasMart-Ch24-App-Lab-Only-2026")key = "atlasmart:ch24:l2:list"r.delete(key)payload = b"x" * 256pipe = r.pipeline(transaction=False)for _ in range(20000): pipe.rpush(key, payload)pipe.execute()for stop in (99, 999, 9999, 19999):    t0 = time.perf_counter_ns(); rows = r.lrange(key, 0, stop); dt = (time.perf_counter_ns()-t0)/1e6    print(stop+1, "items", round(dt,3), "ms", "bytes", sum(map(len, rows)))

7. CPU, RSS, disk, and network are context, not Redis counters

Redis cannot report host CPU steal time, Docker throttling, filesystem queue depth, or a physical-network fault from INFO alone. Capture platform context at the same timestamps. In the local lab, docker stats --no-stream gives container CPU/memory/network/block-I/O counters. In production use the platform’s supported host/container metrics and keep them correlated with Redis samples.

Shell · platform context snapshot
docker stats --no-stream atlasmart-redis-ch24docker logs --tail 100 atlasmart-redis-ch24

8. A latency decision tree

Observation Next evidence Likely boundary
Client p99 high; SLOWLOG/latencystats low Network timing, pool wait, CLIENT LIST, application traces Outside Redis command execution
SLOWLOG entries high Command/data complexity, key cardinality, CPU Command execution
LATENCY fork spikes INFO persistence, RSS/COW, host scheduling Persistence/fork
Intrinsic latency high Host/hypervisor/cgroup metrics Platform floor
Output buffers grow CLIENT LIST omem/oll, slow consumer Client/network backpressure
AOF-related latency events Disk/fsync metrics, appendfsync policy Storage/durability tradeoff

9. Deliberately wrong: tune Redis before proving the layer

Failure pattern

Changing maxmemory, disabling persistence, raising timeouts, or adding replicas because “Redis is slow” can hide the symptom while damaging durability or capacity. Capture an evidence bundle first, reproduce if safe, change one mechanism, and verify the same measurements after the change.

Check your understanding

  1. Why must intrinsic latency be measured on the Redis host?
  2. Why can LRANGE of a large reply feel slow without a corresponding huge SLOWLOG duration?
  3. What should accompany a fork latency event?
  4. Why is a timeout increase not a latency fix?
Review the answers

It measures kernel/hypervisor scheduling delay local to that execution environment, not a network request.

Slow Log excludes client I/O and reply transfer.

INFO persistence plus RSS/COW and host scheduling/memory evidence.

It changes how long callers wait; it does not remove the underlying latency mechanism.

10. Production judgment and bridge

Treat latency as a layered budget. Preserve durability and correctness while testing hypotheses, and report p50/p95/p99/max rather than averages alone. The next lesson turns this diagnostic discipline into representative load and capacity tests.

Summary and next step

Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with redis-benchmark and Custom Load Tests: Pipeline Depth, Payload, Durability, Skew, and Tail Latency.

Authoritative references

Shell · restore temporary latency setting and fixture
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET latency-monitor-threshold 0docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-App-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user atlasmart-app UNLINK atlasmart:ch24:l2:list

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.