Chapter 24 · Observability, Latency, Benchmarking, Capacity, Backup, Upgrades, and Capstone

INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, MONITOR Caveats, and Exported Metrics

Measure per-key concentration before choosing local caches, replica reads, key redesign, or sharding, and connect hot-key relief to consistency and write concentration.

Advanced210–320 minutesINFO, latency, SLOWLOG, MEMORY, CLIENT LIST, MONITOR, metricsRedis Open Source 8.10.1redis-py 8.1.0 where Python is usedDocker + redis-cli + PythonStandalone evidence node · DB 0AOF everysec + RDB · maxmemory 0/noeviction baselineNamed ACL users · TLS off only on loopbackReuses Sentinel/Cluster labs for failover acceptanceFree/local-firstLast reviewed: September 6, 2026
Reproducible Chapter 24 baseline

Redis Open Source 8.10.1 using the pinned redis:8.10.1 image. The observability/benchmark node is atlasmart-redis-ch24 on 127.0.0.1:6441, standalone topology, logical database 0, AOF everysec plus RDB save rules, maxmemory 0/noeviction unless a bounded experiment says otherwise, named academy-admin and atlasmart-app ACL users, and fixture prefix atlasmart:ch24:*. TLS is off only on this loopback-local disposable node; the capstone security acceptance criteria reuse Chapter 22 TLS/ACL guidance. Python examples target redis==8.1.0. Search/JSON/vector/time-series/probabilistic features are optional and must be included in capacity accounting only when the chosen AtlasMart architecture actually uses them.

Learning outcomes

This lesson turns INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, MONITOR Caveats, and Exported Metrics into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.

01

Explain the mechanisms and terminology behind INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, MONITOR Caveats, and Exported Metrics.

02

Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.

03

Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.

04

Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.

05

Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.

1. Problem: the alert says “Redis latency,” but which layer is slow?

AtlasMart checkout reports a p99 response-time regression. Redis may be executing commands slowly, the client may be queued behind its own connection pool, the network may be delayed, a persistence fork may pause the process, or the application may simply be missing its cache and waiting on another database. Observability means collecting enough independent evidence to distinguish those mechanisms. A service-level objective (SLO) is the target reliability or latency level for a service; an error budget is the amount of allowed failure relative to that objective.

2. Evidence hierarchy: start with cheap snapshots

Use server metadata before intrusive tools. INFO is sectioned and machine-readable; commandstats reports calls and execution time by command, latencystats reports command latency percentiles, memory exposes allocator/process state, replication exposes role/offsets, and keyspace exposes logical-database counts. Redis 8.8+ also adds slow-log aggregate fields to command statistics. These are server-side observations; they do not include the application network round trip.

Shell · start the disposable Chapter 24 node
docker volume create atlasmart-redis-ch24-datadocker run -d --name atlasmart-redis-ch24 --restart no -p 127.0.0.1:6441:6379 -v atlasmart-redis-ch24-data:/data redis:8.10.1 redis-server --appendonly yes --appendfsync everysec --save 300 10 --latency-tracking yes --requirepass AtlasMart-Ch24-Bootstrap-Lab-Only-2026docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Bootstrap-Lab-Only-2026 atlasmart-redis-ch24 redis-cli ACL SETUSER academy-admin on ">AtlasMart-Ch24-Admin-Lab-Only-2026" "~*" "&*" +@alldocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Bootstrap-Lab-Only-2026 atlasmart-redis-ch24 redis-cli ACL SETUSER atlasmart-app on ">AtlasMart-Ch24-App-Lab-Only-2026" "~atlasmart:ch24:*" "&atlasmart:ch24:*" +@read +@write +@connection -@dangerousdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Bootstrap-Lab-Only-2026 atlasmart-redis-ch24 redis-cli ACL SETUSER default offdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO server

3. Capture a before-change observability bundle

Shell · low-cost Redis evidence
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO serverdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO statsdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO commandstatsdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO latencystatsdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO memorydocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO replicationdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin INFO keyspace

Expected evidence includes redis_version:8.10.1, standalone role, command counters, latency percentiles, used-memory/RSS values, and DB 0 key counts. A clean snapshot proves the node state at capture time; it does not prove the absence of short spikes between samples.

4. SLOWLOG answers “which commands blocked Redis?”

The Slow Log records commands whose server execution time crosses slowlog-log-slower-than. Redis explicitly excludes client I/O from that duration, so a slow application request with no Slow Log entry can still be caused by network, client queues, serialization, or another upstream service. Save the original threshold before changing it and restore it after the lab.

Shell · bounded Slow Log experiment
OLD_SLOW=$(docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --raw --user academy-admin CONFIG GET slowlog-log-slower-than | tail -1)docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET slowlog-log-slower-than 0docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-App-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user atlasmart-app SET atlasmart:ch24:l1:probe valuedocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin SLOWLOG GET 5docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET slowlog-log-slower-than "$OLD_SLOW"

5. LATENCY captures server events that command statistics alone can miss

Redis latency monitoring is disabled by default when latency-monitor-threshold is 0. When enabled, the framework records named event classes such as command, fork, eviction, expiration, and AOF-related events. LATENCY LATEST returns samples; LATENCY DOCTOR turns those samples into a human-readable diagnosis. Choose the threshold from the workload SLO, not a universal number.

Shell · enable, inspect, and restore latency monitoring
OLD_LAT=$(docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --raw --user academy-admin CONFIG GET latency-monitor-threshold | tail -1)docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET latency-monitor-threshold 1docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY LATESTdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin LATENCY DOCTORdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CONFIG SET latency-monitor-threshold "$OLD_LAT"

6. MEMORY and CLIENT LIST explain pressure around the command engine

MEMORY STATS complements INFO memory, while CLIENT LIST exposes connection identity, age/idle time, authenticated user, RESP version, query buffers, output buffers, and library metadata where the client supplies it. A large omem or tot-mem can reveal slow consumers even when commands themselves are fast.

Shell · memory and client evidence
docker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin MEMORY STATSdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin MEMORY DOCTORdocker exec -e REDISCLI_AUTH=AtlasMart-Ch24-Admin-Lab-Only-2026 atlasmart-redis-ch24 redis-cli --user academy-admin CLIENT LIST

7. MONITOR is an emergency microscope, not a dashboard

Deliberately wrong approach

Leaving MONITOR running on production because it shows every command. Redis documents MONITOR as an @admin/@slow/@dangerous debugging command and shows that even one MONITOR client can materially reduce throughput. It can also expose application keys and command arguments. Prefer INFO, SLOWLOG, latency tracking, tracing, and sampled application telemetry. If MONITOR is ever used, scope it to a disposable environment or a tightly bounded incident procedure.

MONITOR also does not log every administrative command, and sensitive authentication data is specially handled. Absence from MONITOR is therefore not evidence that no administrative/security event occurred.

8. Export metrics without confusing collection with interpretation

A production metrics system normally scrapes a supported Redis/managed-service exporter. The free local lab instead demonstrates a small JSON snapshot so the evidence pipeline stays reproducible without adding another daemon. The important design is stable metric names, timestamps, labels such as node/role, and a retention policy—not the choice of dashboard product.

Python · export a compact evidence snapshot
import json, time, redisr = redis.Redis(host="127.0.0.1", port=6441, username="academy-admin", password="AtlasMart-Ch24-Admin-Lab-Only-2026", decode_responses=True)info = r.info()payload = {    "captured_at_unix": time.time(),    "redis_version": info.get("redis_version"),    "used_memory": info.get("used_memory"),    "used_memory_rss": info.get("used_memory_rss"),    "connected_clients": info.get("connected_clients"),    "instantaneous_ops_per_sec": info.get("instantaneous_ops_per_sec"),    "keyspace_hits": info.get("keyspace_hits"),    "keyspace_misses": info.get("keyspace_misses"),    "evicted_keys": info.get("evicted_keys"),}print(json.dumps(payload, indent=2, sort_keys=True))

9. Evidence matrix: what each tool proves

Evidence Strong at Does not prove
INFO commandstats / latencystats Aggregate server command activity and tracked percentiles Application round-trip or business success
SLOWLOG Commands with slow server execution Network/client delay
LATENCY Named server latency events over threshold Complete request traces
MEMORY Redis allocator/process memory evidence Container/host capacity by itself
CLIENT LIST Connection/buffer/client identity state Why an upstream caller is slow without correlation
MONITOR Real-time processed command stream Safe continuous observability or complete security audit

Check your understanding

  1. Why can an application request be slow while SLOWLOG is empty?
  2. What does INFO latencystats add to commandstats?
  3. Why is MONITOR unsuitable as a normal dashboard source?
  4. What must an exported metric include besides a value?
Review the answers

SLOWLOG excludes network/client I/O and only records server execution above its threshold.

Per-command tracked latency percentiles, while commandstats focuses on calls/execution totals and related counters.

It streams every processed command, is intrusive, can expose workload details, and Redis documents a significant performance cost.

At minimum a stable meaning plus timestamp and node/topology labels so it can be correlated and compared.

10. Production judgment and bridge

Start incidents with low-cost evidence, correlate client and server time, and increase intrusiveness only when the hypothesis requires it. Tail latency, memory/buffer growth, persistence events, replication state, security identity, and workload shape belong in the same timeline. The next lesson decomposes latency mechanisms rather than treating every spike as a Redis command problem.

Summary and next step

INFO, LATENCY, SLOWLOG, MEMORY, CLIENT LIST, MONITOR Caveats, and Exported Metrics is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Latency Doctor, Intrinsic Latency, Fork Pauses, Big Commands, Network, CPU, and Storage Diagnostics.

Authoritative references

Shell · bounded cleanup
docker rm -f atlasmart-redis-ch24 2>/dev/null || truedocker volume rm atlasmart-redis-ch24-data 2>/dev/null || true# Do not run docker system prune; remove only resources named by this chapter.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.