Latency Percentiles, Throughput, Errors, Saturation, and Bottleneck Diagnosis: Diagnostics, Failure Modes, and Production Practices
Most diagnosis errors promote correlation to root cause too early. Preserve first-failure evidence, verify scope/workload, check generator health, inspect SUT telemetry, then run the smallest discriminating experiment.
Learning objectives
- Apply a preserve-first diagnostic sequence to failure modes involving Latency Percentiles, Throughput, Errors, Saturation, and Bottleneck Diagnosis.
- Separate plan/configuration, generator, protocol/network, target, CI/container, and distributed-engine causes before changing settings.
- Reproduce a failure with the smallest authorized workload and retain the original JTL, jmeter.log, and supporting evidence.
- Reject shortcuts such as blanket retries, disabled verification, unbounded load increases, silent global overrides, or deleting first-failure artifacts.
- Verify the least-invasive correction under the original bounded workload before declaring the problem resolved.
1. Preserve-first sequence
Diagnosis starts with a workload and ends with a controlled falsification test; graphs are evidence, not root-cause labels.
flowchart TD E[Preserve JTL/log/target/generator] --> V[Confirm versions] V --> C[Confirm JMX/data/properties/CLI/target] C --> S[Validate scope/labels/workload] S --> P[Inspect protocol/session/data] P --> G[Inspect generator JVM/OS/network] G --> T[Inspect SUT CPU/memory/DB/network/queue] T --> H[Competing hypotheses] H --> X[One-variable rerun]
127.0.0.1:8029.
Keep JTL, matching jmeter.log, dashboard, target
events, generator snapshots and exact inputs.
2. Averages masking tail latency
Average95 ms can coexist with p99800 ms. Show the distribution and exact label/time window before tuning.
3. High CPU called root cause
CPU95% is a saturation symptom, not automatically the causal bottleneck. Ask whether throughput/queue/service behavior moves with it and whether a CPU-capacity/work reduction shifts the knee under the same workload.
4. Generator saturation mistaken for target latency
If JMeter CPU/GC/network/disk saturates while target queue/resource use stays low, target capacity claims are invalid. Fix/scale the injector first.
5. Ignoring queue depth/DB limits
Server CPU may be modest while requests wait for a DB pool/thread pool/semaphore/lock/downstream rate limiter. Queue/pool wait can be more discriminating than CPU.
6. Comparing different workloads
4 threads/50ms pacing versus 8 threads/0ms pacing cannot isolate a server-version effect. Hash inputs and compare configured/achieved load.
7. Intentionally broken: mixed labels
Health 5 ms ×900
Work 300 ms ×100
A combined percentile describes the mixture. Keep raw JTL and repair
the analysis filter to Work; do not rewrite samples.
8. Fast errors hide useful throughput loss
Quick queue-timeout responses can keep total samples/s high while successful RPS collapses. Always show both throughput and error distribution.
9. Causal symptom table
| Pattern | Leading hypothesis | Discriminator |
|---|---|---|
| p95↑ + RPS plateau + queue↑ + in-service=cap + generator healthy | Target worker/pool limit | Increase only capacity; queue/tail should drop. |
| p95↑ + RPS plateau + JMeter CPU/GC saturated + target queue low | Generator bottleneck | Reduce generator cost/add injector capacity. |
| p99↑ + average stable + rare target GC/queue spikes | Episodic tail event | Timeline correlation + repeated controlled run. |
| Errors↑ fast + total RPS stable + success RPS↓ | Rejection/timeout saturation | Statuses + queue/resource limit. |
| Different threads/timers/labels | Invalid comparison | Restore identical workload/scope. |
10. Shortcuts to reject
- Do not add blanket retries or arbitrary sleeps.
- Do not assign giant heaps without generator evidence.
- Do not mass-disable evidence without measured cost.
- Do not use global property hacks.
- Do not disable TLS/RMI verification.
- Do not increase load unboundedly.
- Do not delete result files after failure.
- Do not change multiple SUT variables before validation.
Knowledge check
Can high server CPU alone establish root cause?
No; correlate other signals and validate with a controlled CPU-related experiment.
Why separate successful throughput?
Fast failures can make total completion rate look healthy.
How repair mixed-label analysis?
Keep raw JTL, filter the intended label/transaction and recompute.
What points toward generator saturation?
JMeter resource saturation with target resources below limits and no response to target-capacity change.
Why are different thread/timer settings invalid for a causal comparison?
Workload changed simultaneously, so attribution is confounded.
Official references and version notes
- Apache JMeter downloads — JMeter 5.6.3 and Java 8+ requirement.
- Dashboard Report — CSV requirements, percentile settings, response-time/throughput/error/thread graphs and estimator caveat.
- Properties Reference — result-save fields, thread counts and report-generator properties.
- JMeter Glossary — elapsed, latency and connect-time definitions.
Rechecked against current primary documentation on 2026-09-05. Mandatory runtime: Apache JMeter 5.6.3, Java 17, no third-party plugin. Dashboard percentiles default to 90/95/99 and are configurable. JMeter warns dashboard percentile estimates can differ from Aggregate Report, especially for small/wide samples; therefore this chapter audits p50/p90/p95/p99 from raw CSV JTL with a documented nearest-rank method and retains the dashboard as corroborating evidence. The lab uses a standard closed Thread Group, so “offered load” means configured concurrency plus pacing pressure, not an independent fixed open-arrival RPS.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.