Latency Percentiles, Throughput, Errors, Saturation, and Bottleneck Diagnosis: Core Concepts and Mental Model
Chapter 28 turned p95 into policy. Chapter 29 asks the harder question: why did p95 change? JMeter supplies symptoms; a root-cause claim requires correlated system evidence and a controlled experiment.
Learning objectives
- Connect offered-load pressure to queueing/service time, latency distributions, achieved throughput and errors.
- Define p50/p90/p95/p99, throughput, active threads, queue depth/wait and saturation.
- Separate JMeter elapsed/latency/connect time from pacing and target service time.
- Form competing generator-versus-SUT hypotheses and validate one variable at a time.
1. One graph is a symptom, not root cause
p95 can rise because of a server queue, database pool, generator CPU/GC, DNS/connect behavior, payload changes, assertion/parser cost, or mixed labels. Diagnosis must explain multiple signals together.
127.0.0.1:8029. Discovery profiles are 2/4/8 threads
×30 loops, 50 ms pacing; maximum single run=240 samples and ≤20
seconds. No production/shared target, real credentials, paid
telemetry, remote engines or OS/JVM tuning.
2. Mental model
Diagnosis starts with a workload and ends with a controlled falsification test; graphs are evidence, not root-cause labels.
flowchart TD O[Configured concurrency + pacing] --> Q[Queueing + service work] Q --> L[Latency distribution] Q --> T[Achieved throughput] Q --> E[Errors] G[Generator CPU/GC/network/disk] --> L G --> T S[SUT CPU/memory/DB/network/queue] --> Q L --> H[Competing hypotheses] T --> H E --> H G --> H S --> H H --> X[One-variable rerun]
Configured concurrency/pacing creates pressure. Work can queue before service. JMeter elapsed time includes the client-visible request path; throughput/errors describe achieved work. Generator and SUT telemetry locate pressure. The one-variable rerun tests whether the leading hypothesis predicts the observed response.
3. Metrics
| Metric | Meaning | Use |
|---|---|---|
| p50 | Median elapsed time. | Typical behavior. |
| p90/p95/p99 | Tail percentiles. | Expose queue/stall tail hidden by averages. |
| Throughput | Completed samples/time. | Plateau under increasing pressure is a saturation clue. |
| Successful throughput | Successful completed samples/time. | Prevents quick errors from looking like useful capacity. |
| Error rate | Failed samples/samples. | Shows rejection/timeouts/correctness failure. |
| Active threads | JMeter threads active when samples finish. | Correlates closed-model concurrency with response behavior. |
| Queue wait/depth | Target requests waiting for constrained resource. | Direct contention evidence. |
| Saturation | Resource at effective limit; extra pressure yields delay/plateau/errors. | Needs resource evidence, not latency alone. |
4. Timing boundaries
JMeter elapsed time runs from just before request send until the last response byte. Latency ends at the first response; connect time covers connection establishment including SSL. A Constant Timer delays sampler execution and is not target service time.
5. Offered versus achieved
This lab uses a closed Thread Group: increasing threads while holding pacing constant increases offered pressure, but arrivals depend on response time. Achieved throughput is measured from JTL. Chapter30 will examine coordinated omission and other validity consequences.
6. Saturation knee
A defensible knee often combines diminishing throughput gains, rising tail latency/queue wait, a constrained resource pinned near its limit, possible errors, and a generator that is still healthy. No one signal alone proves the bottleneck.
7. Read-only inspection first
& "$env:JMETER_HOME\bin\jmeter.bat" -v\njava -version\nGet-Content .\config\bottleneck.properties\nGet-FileHash .\plans\bottleneck.jmx,.\config\bottleneck.properties -Algorithm SHA256\ncurl.exe --fail --silent http://127.0.0.1:8029/health
Prove JMeter/Java, workload, sample label, result schema, fixture capacity/service delay and target authorization before changing anything.
8. Evidence contract
Every diagnosis retains the source CSV JTL,
matching jmeter.log, dashboard, exact JMX/properties,
configured-versus-achieved workload, generator-health snapshot and
target queue/service events before any correction. The mandatory lab
stays on 127.0.0.1:8029 and uses
no real credentials.
9. DevOps connection
Bottleneck diagnosis is experimental reasoning: JMeter supplies client symptoms, but root cause requires correlated system evidence and controlled changes.
Knowledge check
Why can p95 rise without proving server CPU is the bottleneck?
Queue, DB, network, generator, mixed-label and other mechanisms can all raise elapsed time; root cause needs correlated evidence.
What does throughput plateau + queue wait growth suggest?
A constrained target resource may be saturating if generator health remains acceptable; validate with a causal change.
Why is timer delay not target latency?
It occurs before the sampler request timing boundary.
Why filter by label?
Mixing different operations can produce a percentile that describes neither operation.
What is the closed-model limitation?
Arrival rate feeds back from response time rather than being independently fixed.
Official references and version notes
- Apache JMeter downloads — JMeter 5.6.3 and Java 8+ requirement.
- Dashboard Report — CSV requirements, percentile settings, response-time/throughput/error/thread graphs and estimator caveat.
- Properties Reference — result-save fields, thread counts and report-generator properties.
- JMeter Glossary — elapsed, latency and connect-time definitions.
Rechecked against current primary documentation on 2026-09-05. Mandatory runtime: Apache JMeter 5.6.3, Java 17, no third-party plugin. Dashboard percentiles default to 90/95/99 and are configurable. JMeter warns dashboard percentile estimates can differ from Aggregate Report, especially for small/wide samples; therefore this chapter audits p50/p90/p95/p99 from raw CSV JTL with a documented nearest-rank method and retains the dashboard as corroborating evidence. The lab uses a standard closed Thread Group, so “offered load” means configured concurrency plus pacing pressure, not an independent fixed open-arrival RPS.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.