Chapter 15 · Pipelining, Batching, Client Libraries, and Client-Side Caching
Benchmark a Chatty Workflow Before and After Pipelining with Realistic Payloads
Benchmark one realistic AtlasMart read workflow sequentially and with non-transactional pipelines using identical payloads, warmup, request sets, and p50/p95/p99 evidence.
Learning outcomes
Chapter 15 ends with a controlled performance decision. AtlasMart's product-card renderer needs four independent values per request. We will benchmark the same deterministic request set sequentially and with a non-transactional pipeline, keeping dataset, payload sizes, key skew, persistence, connection model, and concurrency constant.
Build a reproducible chatty read fixture with realistic payload sizes and deterministic key selection.
Separate warmup from measured runs and record p50/p95/p99 request latency plus throughput.
Compare sequential and pipeline modes with identical Redis commands/results and verify result equality.
Use Redis INFO/CLIENT evidence to contextualize measurements without inventing server-side causality.
Decide whether pipelining is worth the complexity and recognize when one native multi-key command is simpler.
All Chapter 15 mandatory labs reuse the disposable Chapter 01
environment: Redis Open Source 8.10.1 from pinned
Docker image redis:8.10.1, container
atlasmart-redis-ch01, standalone topology, host
endpoint 127.0.0.1:6379, TLS disabled only
because traffic remains on loopback, default ACL user
disabled, named academy-admin and
atlasmart-app users, logical database 0, AOF with
appendfsync everysec plus RDB snapshots,
persistent /data, and no explicit
maxmemory/eviction policy. Python work pins
redis-py 8.1.0. Chapter-specific fixtures stay
under atlasmart:ch15:.
1. Define the workload before timing it
A benchmark is only useful when the workload is explicit. Each synthetic product stores four String keys: compact identity, pricing JSON, inventory JSON, and merchandising text. Values are intentionally hundreds of bytes rather than one-character toy values. A fixed random seed chooses product IDs with a hot-set skew so both modes read the same sequence.
| Dimension | Chapter 15 value | Why record it |
|---|---|---|
| Products | 200 | Controls key cardinality |
| Keys per product request | 4 GET commands | Defines chattiness |
| Payload | ~0.5–1 KiB per key | Affects network/parser/reply buffers |
| Measured requests | 600 after warmup | Enough for percentiles without a stress test |
| Concurrency | 1 | Isolates round-trip shape first |
| Persistence | AOF everysec + RDB | Keeps course baseline consistent |
| Network | 127.0.0.1 loopback | Means RTT is unusually small versus remote deployment |
2. Build the deterministic fixture with bounded pipeline chunks
import json, redisr = redis.Redis(host="127.0.0.1", port=6379, username="atlasmart-app", password="AtlasMart-App-Lab-Only-2026", decode_responses=True)for start in range(0, 200, 50): p = r.pipeline(transaction=False) for i in range(start, start + 50): base = f"atlasmart:ch15:l5:product:{i}" p.set(base+":identity", json.dumps({"id":i,"sku":f"SKU-{i:04d}","name":"AtlasMart Camera "+str(i)})) p.set(base+":price", json.dumps({"currency":"USD","list":199.99,"sale":179.99,"taxClass":"standard","pad":"p"*350})) p.set(base+":inventory", json.dumps({"available":25+(i%11),"warehouse":"WH-A","reserved":i%5,"pad":"i"*350})) p.set(base+":merch", json.dumps({"badge":"featured","copy":"m"*700})) p.execute()print("fixture_ready", r.exists("atlasmart:ch15:l5:product:0:identity"))
3. Benchmark sequential and pipeline modes fairly
The harness precomputes the request sequence once. Each mode
reads exactly the same four keys in the same order and asserts
identical returned values. The pipeline explicitly uses
transaction=False; otherwise redis-py's default
transaction would contaminate the comparison.
import math, random, time, redisr = redis.Redis(host="127.0.0.1", port=6379, username="atlasmart-app", password="AtlasMart-App-Lab-Only-2026", decode_responses=True, max_connections=8, client_name="atlasmart-ch15-bench")rng = random.Random(20260906)requests = [rng.randrange(20) if rng.random() < 0.80 else rng.randrange(20, 200) for _ in range(650)]def keys(i): base = f"atlasmart:ch15:l5:product:{i}" return [base+":identity", base+":price", base+":inventory", base+":merch"]def seq(i): return [r.get(k) for k in keys(i)]def piped(i): p = r.pipeline(transaction=False) for k in keys(i): p.get(k) return p.execute()def pct(v, q): v = sorted(v); return v[min(len(v)-1, math.ceil(q*len(v))-1)]baseline = [seq(i) for i in requests[:25]]assert all(piped(i) == expected for i, expected in zip(requests[:25], baseline))def run(fn): for i in requests[:50]: fn(i) # warmup, excluded lat=[]; t0=time.perf_counter() for i in requests[50:]: q0=time.perf_counter_ns(); fn(i); lat.append((time.perf_counter_ns()-q0)/1_000_000) elapsed=time.perf_counter()-t0 n=len(lat) return {"requests":n,"req_s":round(n/elapsed,1),"p50_ms":round(pct(lat,.50),3),"p95_ms":round(pct(lat,.95),3),"p99_ms":round(pct(lat,.99),3)}print("sequential", run(seq))print("pipeline", run(piped))# Record actual output. Loopback may show a small gain—or client overhead may dominate.
The correct result is the one your machine measures. On 127.0.0.1 the RTT is tiny; a pipeline can help, tie, or even lose for small groups depending on client/parser overhead. Do not replace measured output with a marketing claim.
4. Record Redis-side context around the run
Before and after a benchmark, record server/client state so
later comparisons are explainable. INFO stats can
show total command and network-byte counters;
INFO commandstats shows per-command call/usec
statistics; CLIENT LIST identifies connection
state. None of these counters directly says “pipeline was
faster”—they contextualize the client latency result.
INFO serverINFO statsINFO commandstatsINFO clientsCLIENT LIST TYPE normalSLOWLOG GET 20# Save before/after snapshots if you want to attribute changes. Other concurrent workloads contaminate the counters.
5. Interpret p50, p95, p99—not only average throughput
Median (p50) describes a typical request, while p95/p99 expose the slow tail that users and upstream timeouts often experience. Throughput tells how many request groups completed per second. A pipeline may improve throughput while increasing request-level p99 if batches or client scheduling create queueing; all dimensions matter.
| Metric | Question answered |
|---|---|
| req/s | How much grouped work completed per second? |
| p50 | What did a typical product-card request observe? |
| p95/p99 | How bad was the slow tail? |
| Redis commandstats | How many GETs and server CPU time were recorded? |
| net input/output bytes | How much Redis protocol traffic moved? |
6. The comparison still has boundaries
This benchmark is single-client, loopback, warm dataset, no explicit maxmemory eviction, standalone, and AOF everysec. It does not predict WAN/TLS latency, Cluster routing, Sentinel failover, replica reads, high concurrency, cold CPU caches, persistence rewrite activity, or noisy-neighbor behavior. Re-run in a representative environment before choosing a production batch strategy.
7. Pipelining may not be the simplest optimization
If a native multi-key command expresses the operation directly,
use it as another candidate. Four independent String
GETs can often be represented by
MGET when keys are known together and, in Cluster,
share the required slot constraints. A native command can reduce
both network turns and client command construction. Do not force
pipelining where Redis already has a simpler primitive.
MGET atlasmart:ch15:l5:product:42:identity atlasmart:ch15:l5:product:42:price atlasmart:ch15:l5:product:42:inventory atlasmart:ch15:l5:product:42:merch# Compare semantics/topology eligibility before benchmarking MGET as an additional candidate.
8. Failure/misuse: benchmark drift can manufacture a result
A common bad benchmark uses tiny payloads in one mode, larger
payloads in another, changes key distribution, skips warmup for
one mode, or accidentally uses pipeline() with the
redis-py transactional default. Any of those differences can
dominate the comparison. Keep fixtures immutable between modes
and assert result equality.
Same server version, topology, persistence, client version, connection model, command set, payload bytes, key sequence/skew, concurrency, warmup, measurement window, and background workload—or clearly disclose the difference.
9. Optional third experiment: repeated hot reads with client-side caching
After the pipeline comparison is complete, you may run Lesson
4's RESP3 CacheConfig client against the same
request sequence. That is a different optimization: it removes
server reads on local cache hits rather than merely batching
them. Keep it a separately labeled mode and report cache
hit/miss/reconnect behavior; do not merge it into the pipeline
numbers.
10. Verification and cleanup
- Both modes used the same 600 measured request IDs after the same-size warmup.
- Each request returned four identical values and an assertion checked sample equality.
- The pipeline mode used
transaction=False. - Actual p50/p95/p99 and requests/second were printed rather than pre-filled.
- Redis INFO/CLIENT context was recorded and its attribution limits were stated.
import redisr = redis.Redis(host="127.0.0.1", port=6379, username="atlasmart-app", password="AtlasMart-App-Lab-Only-2026")for start in range(0, 200, 25): names=[] for i in range(start, start+25): base=f"atlasmart:ch15:l5:product:{i}" names += [base+":identity",base+":price",base+":inventory",base+":merch"] r.unlink(*names)print("remaining_sample", r.exists("atlasmart:ch15:l5:product:0:identity"))
11. Production judgment and chapter synthesis
Use the simplest optimization that meets the measured objective. First reduce unnecessary commands and choose better native data structures/commands. Then consider bounded pipelines when RTT dominates. Size pools and timeouts to the workload, classify retry ambiguity, and add client-side caching only where a freshness contract makes it safe. Monitor p95/p99, pool waits, reconnect/retry counts, Redis client/output buffers, network bytes, and cache invalidations/hit ratio. Re-test after version, topology, TLS, payload, persistence, or traffic changes.
Check your understanding
- Why is loopback a weak predictor of remote pipeline benefit?
- Why assert sequential and pipeline results are identical?
- What redis-py flag is essential for a batching-only benchmark?
- Why report p99 as well as requests/second?
- When might MGET be preferable?
Review the answers
Its RTT is extremely small, so network wait contributes much less than on remote/TLS paths.
Performance is irrelevant if the optimization changes behavior or result ordering.
pipeline(transaction=False).
Throughput can improve while a slow tail violates application deadlines/SLOs.
When the same known String keys can be read with one native command and topology/slot constraints allow it.
12. Chapter summary and bridge
Chapter 15 connected Redis client performance to the whole request path: RTT, pipeline buffering, bounded batches, connection models, timeouts, ambiguous retries, RESP3 invalidations, local-cache freshness, and measurement discipline. Chapter 16 turns to Pub/Sub and keyspace notifications, where a different boundary becomes central: delivery signals are ephemeral and must not be confused with durable business state.
Authoritative references
- Redis Open Source 8.10 release notes
- Redis pipelining
- Using Redis commands
- Redis transactions
- redis-py guide
- redis-py pipelines and transactions
- redis-py production usage
- redis-py connection guide
- redis-py asynchronous usage
- redis-py error handling
- Connection pools and multiplexing
- Client-side caching introduction
- Client-side caching reference
- CLIENT TRACKING
- CLIENT TRACKINGINFO
- CLIENT CACHING
- CLIENT LIST
- CLIENT INFO
- CLIENT ID
- CLIENT SETNAME
- MONITOR
- INFO
- SLOWLOG
- Redis latency diagnosis
- Redis security
- Redis ACLs
- Redis Cluster specification
- Redis persistence
- Redis replication
- redis-py documentation
- redis-py 8.1.0 on PyPI
- Redis licenses