Chapter 15 · Pipelining, Batching, Client Libraries, and Client-Side Caching

Batch Size, Memory, Backpressure, and Latency/Throughput Tradeoffs

Choose bounded pipeline batches from client/server memory, response size, queueing, backpressure, throughput, and tail-latency evidence rather than maximizing batch size.

Advanced170–230 minutesBatch sizing, buffers, and backpressureRedis Open Source 8.10.1redis-py 8.1.0Free/local-firstLast reviewed: September 6, 2026

Learning outcomes

After Lesson 1, AtlasMart can batch commands correctly. The next failure mode is operational: treating “larger batch” as synonymous with “better performance.” Every queued request and reply occupies memory and increases the amount of work one application request waits to complete. This lesson measures those tradeoffs with bounded synthetic data.

01

Explain batch size as a latency/throughput/memory tradeoff rather than a universal tuning constant.

02

Measure redis-py client-side command-queue memory before execute() with deterministic payloads.

03

Observe Redis client output-buffer evidence using a bounded slow-reader fixture and interpret zero/nonzero results cautiously.

04

Benchmark several batch sizes with identical total work and report throughput plus p50/p95/p99 batch latency.

05

Apply chunking and backpressure so producers cannot create unbounded pending Redis work.

Exact lab baseline

All Chapter 15 mandatory labs reuse the disposable Chapter 01 environment: Redis Open Source 8.10.1 from pinned Docker image redis:8.10.1, container atlasmart-redis-ch01, standalone topology, host endpoint 127.0.0.1:6379, TLS disabled only because traffic remains on loopback, default ACL user disabled, named academy-admin and atlasmart-app users, logical database 0, AOF with appendfsync everysec plus RDB snapshots, persistent /data, and no explicit maxmemory/eviction policy. Python work pins redis-py 8.1.0. Chapter-specific fixtures stay under atlasmart:ch15:.

1. Batch size creates a queueing boundary

A batch of 1 minimizes “time waiting for the rest of the batch” but pays many round trips. A very large batch amortizes RTT aggressively but holds more request/reply state, creates a larger unit of retry/diagnosis, and can monopolize application work for longer. The useful operating point depends on RTT, payload size, concurrency, parser cost, Redis CPU, output buffers, persistence, and the latency objective.

Increase batch size Likely benefit Likely cost
From 1 to tens Fewer RTT waits Some queueing delay
Tens to hundreds More throughput on chatty workloads More client/reply memory
Hundreds to thousands RTT amortization may flatten Larger tail latency/failure scope
Unbounded No reliable extra benefit Memory exhaustion and backpressure collapse

2. Measure the client queue before sending anything

redis-py queues pipeline commands in the application process until execute(). The following fixture uses tracemalloc to compare Python allocation peaks while queuing different numbers of 512-byte values. It deliberately does not execute the pipeline, so it isolates client-side queued-command cost.

Python · client-side pipeline queue memory
import gc, tracemalloc, redisr = redis.Redis(host="127.0.0.1", port=6379, username="atlasmart-app", password="AtlasMart-App-Lab-Only-2026", decode_responses=False)payload = b"x" * 512for n in (10, 100, 1000, 5000):    gc.collect()    tracemalloc.start()    pipe = r.pipeline(transaction=False)    for i in range(n):        pipe.set(f"atlasmart:ch15:l2:queued:{i}", payload)    current, peak = tracemalloc.get_traced_memory()    tracemalloc.stop()    pipe.reset()  # drop queued commands without sending them    print(n, "peak_bytes", peak)# Record your measured peaks; Python allocator/client version affect exact numbers.
What this proves

Memory increases with queued client work, but tracemalloc measures Python allocations, not complete process RSS and not Redis server memory. Measure both layers when capacity matters.

3. A slow reader can make reply buffering visible

Redis must send one reply for every command. To make that normally brief state observable, this bounded experiment creates a raw TCP client, authenticates it, names it, sends many GET requests for a 4 KiB value, and intentionally delays reading. A second administrative connection inspects the named client. Kernel socket buffering may absorb some or all replies, so output-buffer fields can still be zero on a fast machine; that is evidence about the environment, not proof that buffering does not exist.

Python · bounded delayed-reader experiment
import socket, time, redisHOST, PORT = "127.0.0.1", 6379KEY = "atlasmart:ch15:l2:blob"admin = redis.Redis(host=HOST, port=PORT, username="academy-admin", password="AtlasMart-Admin-Lab-Only-2026", decode_responses=True)app = redis.Redis(host=HOST, port=PORT, username="atlasmart-app", password="AtlasMart-App-Lab-Only-2026")app.set(KEY, b"z" * 4096)def resp(*parts):    out = [f"*{len(parts)}\r\n".encode()]    for part in parts:        b = str(part).encode()        out += [f"${len(b)}\r\n".encode(), b, b"\r\n"]    return b"".join(out)sock = socket.create_connection((HOST, PORT), timeout=2)sock.setsockopt(socket.SOL_SOCKET, socket.SO_RCVBUF, 4096)wire = resp("AUTH", "atlasmart-app", "AtlasMart-App-Lab-Only-2026") + resp("CLIENT", "SETNAME", "atlasmart-ch15-slow-reader")wire += b"".join(resp("GET", KEY) for _ in range(2000))sock.sendall(wire)time.sleep(0.20)matches = [c for c in admin.client_list() if c.get("name") == "atlasmart-ch15-slow-reader"]print(matches)sock.close()app.unlink(KEY)# Inspect obl/oll/omem when present. Exact values are environment-dependent.
Safety bound

The fixture queues about 8 MiB of reply payload at most and then closes the dedicated connection. Do not scale this up on a shared or production server.

4. Compare batch sizes with identical total work

The benchmark below writes the same number of keys with the same 512-byte payload for every batch size. It reports throughput and the distribution of batch completion latency. Larger batches naturally contain more commands, so compare both throughput and latency—not latency alone.

Python · bounded batch-size benchmark
import math, statistics, time, redisr = redis.Redis(host="127.0.0.1", port=6379, username="atlasmart-app", password="AtlasMart-App-Lab-Only-2026", decode_responses=False)N, payload = 2000, b"p" * 512def pct(values, q):    values = sorted(values)    return values[min(len(values)-1, math.ceil(q*len(values))-1)]for batch in (1, 10, 50, 250, 1000):    lat = []    t_all = time.perf_counter()    for start in range(0, N, batch):        p = r.pipeline(transaction=False)        for i in range(start, min(start + batch, N)):            p.set(f"atlasmart:ch15:l2:b{batch}:{i}", payload)        t0 = time.perf_counter()        p.execute()        lat.append((time.perf_counter()-t0)*1000)    elapsed = time.perf_counter() - t_all    print({"batch":batch,"ops_s":round(N/elapsed,1),"batch_p50_ms":round(pct(lat,.50),3),"batch_p95_ms":round(pct(lat,.95),3),"batch_p99_ms":round(pct(lat,.99),3)})    for start in range(0, N, 500):        r.unlink(*(f"atlasmart:ch15:l2:b{batch}:{i}" for i in range(start, min(start+500, N))))# Warm the environment separately if you want stable comparative runs; do not invent expected percentiles.

5. Backpressure means the producer is not allowed to outrun the consumer forever

If an application can enqueue Redis work faster than it can send/read replies, memory becomes the queue. Backpressure is the rule that caps in-flight work and makes producers wait, reject, shed, or persist work elsewhere when capacity is full. A bounded chunk generator is the simplest synchronous form.

Python · bounded chunking instead of one giant pipeline
def chunks(items, size):    batch = []    for item in items:        batch.append(item)        if len(batch) == size:            yield batch            batch = []    if batch:        yield batchfor batch in chunks(range(10000), 250):    pipe = r.pipeline(transaction=False)    for i in batch:        pipe.set(f"atlasmart:ch15:l2:chunk:{i}", "x")    pipe.execute()# At most 250 commands are queued in this client pipeline at a time.

6. Giant batches are also a timeout and retry problem

A larger batch takes longer to complete and creates a larger ambiguous unit if the connection fails after Redis processed some or all commands but before the client received every reply. Do not automatically retry a giant non-idempotent pipeline. Bound the batch and classify commands by retry safety before choosing recovery behavior.

Deliberately wrong approach

“If execute() times out, just resend the whole 50,000-command pipeline” can duplicate non-idempotent effects. Retry policy belongs to the workload contract, not to the pipeline abstraction.

7. Server memory is more than maxmemory

The Chapter 01 lab has no explicit maxmemory setting. Even when a production deployment does, client query/output buffers and allocator/process overhead are separate from the logical dataset budget. A batch-sizing experiment should therefore record Redis process memory, client/output state, payload size, and concurrent clients—not just MEMORY USAGE for keys.

redis-cli · record server and client context
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin INFO serverdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin INFO clientsdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin CONFIG GET appendonly appendfsync maxmemory maxmemory-policydocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin ACL WHOAMI# Record redis_version, persistence, client count, maxmemory policy, and ACL identity before interpreting measurements.

8. Select a measured operating range

Do not choose the batch with the maximum operations/second automatically. If batch 1000 adds little throughput over 250 but substantially increases p99 batch completion and memory, 250 may be the safer operating point. Conversely, a remote high-RTT path may justify larger batches. Record the curve and choose against your service-level objective (SLO), memory budget, and failure semantics.

9. Verification and cleanup

  • Client queue memory was measured before network execution.
  • The delayed-reader test used one named disposable connection and a bounded reply volume.
  • Every batch-size run performed identical total writes and payload sizes.
  • Both throughput and p50/p95/p99 batch latency were recorded.
  • Cleanup used explicit prefixes/keys rather than FLUSHDB.
redis-cli · remove bounded chunk fixture if you ran it
SCAN 0 MATCH atlasmart:ch15:l2:chunk:* COUNT 500# Iterate SCAN to cursor 0 and UNLINK only returned atlasmart:ch15:l2:chunk:* keys; do not use KEYS/FLUSHDB on shared data.

10. Production judgment

Batching is appropriate when per-command RTT dominates and operations can tolerate the chosen failure unit. Measure realistic payloads, concurrency, TLS/network distance, persistence, parsing, and reply sizes. Add bounded queues/semaphores in application code, enforce request deadlines, and alert on rising client output memory, timeouts, queue depth, and p99. Do not transfer one environment's “optimal batch size” into another.

Check your understanding

  1. Why can a larger batch improve throughput?
  2. Why can the same larger batch worsen p99?
  3. What does tracemalloc prove in the lab?
  4. What is backpressure?
  5. Why is a zero CLIENT LIST omem reading not proof of zero buffering?
Review the answers

It amortizes round-trip/socket overhead over more commands.

More commands/replies must finish before the batch completes, increasing queueing and failure scope.

Python-side allocation growth while commands are queued; it does not measure complete process or Redis memory.

A bound that prevents producers from creating unlimited in-flight work when downstream capacity is saturated.

Kernel/socket buffers may have absorbed replies before inspection; the observation is timing/environment dependent.

11. Summary and next step

Batch size is a queueing and memory control, not a trophy number. Lesson 3 moves one layer outward to connection management: pools versus multiplexing, pool exhaustion, timeout budgets, reconnects, and retry ambiguity.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.