The percentile tail reveals stalls that the mean—and sometimes the load generator—can hide.

Tail Latency, Coordinated Omission, Warmup, GC, Network, Disk, and Client-Pool Effects

Measure the latency distribution users can actually experience, including open-loop offered load, coordinated omission, warmup, garbage collection, disk/network stalls, pool saturation, and backpressure.

Advanced125–165 minutesTail-latency labPython 3.13+ · standard libraryVendor-neutral · free/local mandatory pathLast reviewed: August 2026
01

Interpret p50, p95, p99, and p99.9 as distribution summaries and explain why fan-out makes tail latency especially important.

02

Demonstrate coordinated omission by comparing closed-loop request issuance with open-loop scheduled arrivals during a server stall.

03

Account for warmup, runtime/GC pauses, disk and network queues, client-pool waits, backpressure, and test-environment disclosure.

04

Design latency tests that preserve arrival intent, classify timeouts/errors, and separate client queueing from server processing.

1. A fast median does not protect a distributed request from a slow leaf

AtlasMart's catalog page may fan out to several shards or derived services. The page completes when the required components complete, so a rare slow component can become common at the aggregate request level. Tail latency is the high-percentile portion of the response-time distribution. p50 summarizes the median; p95/p99/p99.9 expose increasingly rare stalls. Report the sample count and observation window because p99.9 from a few hundred samples is not stable evidence.

Latency has components: time waiting for a client connection, network transit, coordinator queueing, replica/storage service, retries, and response transit. Runtime garbage collection, page faults, compaction/checkpoints, disk queueing, noisy neighbors, retransmissions, and connection-pool limits can create pauses. Warmup changes cache/JIT/page-cache state. Backpressure intentionally slows or rejects offered work to keep queues bounded.

2. Coordinated omission: when the generator politely waits through the outage

A closed-loop benchmark often sends a request, waits for the response, then sends the next. If the server pauses for 500 ms, the client also pauses offering work. That behavior can omit the queueing delay that would affect users or upstream services whose arrivals continue on schedule. An open-loop or scheduled-load test preserves intended arrival times and measures lateness relative to those schedules.

This does not make closed-loop tests useless. They answer a different question: “what latency does one client experience when it self-throttles after every response?” The mistake is interpreting that distribution as the latency under an independent arrival process. HdrHistogram calls out correction for known coordinated omission, but measurement at the right boundary is preferable when possible.

Timeouts belong in the latency story

Dropping timed-out requests from the histogram makes the surviving distribution look better as overload grows. Record timeout/rejection counts and the elapsed client wait. Decide whether a timed-out request may still have executed, which also affects retries and correctness.

3. Deliberately wrong approach: closed-loop p99=5 ms, therefore the 500 ms pause was harmless

The lab models a 5 ms server that pauses from t=1000 to 1500 ms. The closed-loop client records only one request above 100 ms; because it stops sending while blocked, its p99 remains 5 ms. The scheduled open-loop stream keeps arrivals every 10 ms and records 81 requests above 100 ms, with p99=455 ms.

4. AtlasMart lab: make coordinated omission observable

Mandatory lab environment

Python 3.13+ standard library only. This is a discrete-event teaching model: no real sleep, GC, disk, network, or database is measured.

python · AtlasMart deterministic simulation
from statistics import mean

def pct(xs, p):
    s=sorted(xs)
    rank=max(0,min(len(s)-1,int((p/100)*len(s)+.999999)-1))
    return s[rank]

def show(name, xs):
    print(name, "n=",len(xs), "mean=",round(mean(xs),1),
          "p50=",pct(xs,50), "p95=",pct(xs,95), "p99=",pct(xs,99), "p99.9=",pct(xs,99.9))

# One server: normal service is 5 ms. It pauses from t=1000 to 1500 ms.
# Closed loop sends the next request only after the previous response.
closed=[]
t=0
for _ in range(1000):
    start=t
    if 1000 <= start < 1500:
        finish=1505
    else:
        finish=start+5
    closed.append(finish-start)
    t=finish

# Open loop schedules one request every 10 ms regardless of earlier completion.
open_loop=[]
server_free=0
for scheduled in range(0,10000,10):
    if 1000 <= max(scheduled,server_free) < 1500:
        server_free=1500
    start=max(scheduled,server_free)
    finish=start+5
    open_loop.append(finish-scheduled)
    server_free=finish

print("SYNTHETIC ENVIRONMENT: 5 ms service, 500 ms pause, no network/disk noise")
show("closed-loop observed", closed)
show("open-loop scheduled", open_loop)
print("\nclosed-loop requests >100ms:", sum(x>100 for x in closed))
print("open-loop requests >100ms:", sum(x>100 for x in open_loop))
print("coordinated omission: the closed loop stops offering work during the stall, so it omits queueing latencies real arrivals would experience")
print("production benchmarks must also disclose warmup, GC/runtime, disk, network, pool size, backpressure, and failure state")
Expected evidence

Closed-loop p99 stays 5 ms even though one 505 ms stall occurred, while the open-loop p95 reaches 255 ms and p99 455 ms because scheduled requests accumulate behind the pause. This proves the load-generation model changes the observed distribution.

5. Production judgment: disclose the entire latency experiment

Document CPU/RAM/storage/network, region/zone distance, OS/runtime, database/client versions, connection pool sizes, TLS, dataset/cache state, warmup duration, offered rate, read/write mix, durability/consistency, timeout/retry policy, and maintenance/failure state. Capture client-side histograms with sufficient dynamic range and precision. Correlate with server queues, GC/runtime pauses, disk p99, network loss/retransmit, and pool wait.

Do not “correct” coordinated omission mechanically unless the expected arrival interval is actually known and the correction matches the experiment. Prefer a harness that models the intended arrival process. The next lesson converts these measured service costs and storage footprints into a capacity plan that must still pass after a node or zone is unavailable.

Check your understanding

  1. Why can p50 be excellent while users still complain?
  2. What is coordinated omission?
  3. Why should timeout observations not simply be discarded?
  4. How do warmup and cache state affect comparisons?
  5. What should a latency benchmark disclose besides percentiles?
Review the answers

1. A median describes only the midpoint; a minority of requests can experience very large stalls, especially when distributed fan-out makes slow components more likely to affect whole requests.

2. A measurement bias where the generator stops or slows issuing work while the system is slow, omitting delays that would have affected independently scheduled arrivals.

3. They are user-visible failures and often represent the worst latency region; excluding them biases the distribution and can hide overload.

4. Cold caches/JIT/page cache and later warmed state can have different latency/throughput; a test must define and repeat its warmup policy.

5. Sample count, environment, arrival model, workload, pool/timeout/retry settings, guarantees, warmup, maintenance/failure state, errors, and relevant resource/queue metrics.

References

Foundational claims use primary research or current official documentation where practical. Product references are implementation anchors only; the mandatory labs are vendor-neutral.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.