Chapter 29Lesson 04~235 minutes

Latency Percentiles, Throughput, Errors, Saturation, and Bottleneck Diagnosis: Diagnostics, Failure Modes, and Production Practices

Most diagnosis errors promote correlation to root cause too early. Preserve first-failure evidence, verify scope/workload, check generator health, inspect SUT telemetry, then run the smallest discriminating experiment.

Learning objectives

  • Apply a preserve-first diagnostic sequence to failure modes involving Latency Percentiles, Throughput, Errors, Saturation, and Bottleneck Diagnosis.
  • Separate plan/configuration, generator, protocol/network, target, CI/container, and distributed-engine causes before changing settings.
  • Reproduce a failure with the smallest authorized workload and retain the original JTL, jmeter.log, and supporting evidence.
  • Reject shortcuts such as blanket retries, disabled verification, unbounded load increases, silent global overrides, or deleting first-failure artifacts.
  • Verify the least-invasive correction under the original bounded workload before declaring the problem resolved.

1. Preserve-first sequence

Preserve-first diagnosis

Diagnosis starts with a workload and ends with a controlled falsification test; graphs are evidence, not root-cause labels.

flowchart TD
E[Preserve JTL/log/target/generator] --> V[Confirm versions]
V --> C[Confirm JMX/data/properties/CLI/target]
C --> S[Validate scope/labels/workload]
S --> P[Inspect protocol/session/data]
P --> G[Inspect generator JVM/OS/network]
G --> T[Inspect SUT CPU/memory/DB/network/queue]
T --> H[Competing hypotheses]
H --> X[One-variable rerun]
All runnable diagnosis remains ≤240 samples/run on 127.0.0.1:8029. Keep JTL, matching jmeter.log, dashboard, target events, generator snapshots and exact inputs.

2. Averages masking tail latency

Average95 ms can coexist with p99800 ms. Show the distribution and exact label/time window before tuning.

3. High CPU called root cause

CPU95% is a saturation symptom, not automatically the causal bottleneck. Ask whether throughput/queue/service behavior moves with it and whether a CPU-capacity/work reduction shifts the knee under the same workload.

4. Generator saturation mistaken for target latency

If JMeter CPU/GC/network/disk saturates while target queue/resource use stays low, target capacity claims are invalid. Fix/scale the injector first.

5. Ignoring queue depth/DB limits

Server CPU may be modest while requests wait for a DB pool/thread pool/semaphore/lock/downstream rate limiter. Queue/pool wait can be more discriminating than CPU.

6. Comparing different workloads

4 threads/50ms pacing versus 8 threads/0ms pacing cannot isolate a server-version effect. Hash inputs and compare configured/achieved load.

7. Intentionally broken: mixed labels

Health  5 ms ×900
Work  300 ms ×100

A combined percentile describes the mixture. Keep raw JTL and repair the analysis filter to Work; do not rewrite samples.

8. Fast errors hide useful throughput loss

Quick queue-timeout responses can keep total samples/s high while successful RPS collapses. Always show both throughput and error distribution.

9. Causal symptom table

Pattern Leading hypothesis Discriminator
p95↑ + RPS plateau + queue↑ + in-service=cap + generator healthy Target worker/pool limit Increase only capacity; queue/tail should drop.
p95↑ + RPS plateau + JMeter CPU/GC saturated + target queue low Generator bottleneck Reduce generator cost/add injector capacity.
p99↑ + average stable + rare target GC/queue spikes Episodic tail event Timeline correlation + repeated controlled run.
Errors↑ fast + total RPS stable + success RPS↓ Rejection/timeout saturation Statuses + queue/resource limit.
Different threads/timers/labels Invalid comparison Restore identical workload/scope.
Security boundary: the runnable diagnosis uses synthetic loopback traffic and no real credentials.

10. Shortcuts to reject

  • Do not add blanket retries or arbitrary sleeps.
  • Do not assign giant heaps without generator evidence.
  • Do not mass-disable evidence without measured cost.
  • Do not use global property hacks.
  • Do not disable TLS/RMI verification.
  • Do not increase load unboundedly.
  • Do not delete result files after failure.
  • Do not change multiple SUT variables before validation.

Knowledge check

Can high server CPU alone establish root cause?

Why separate successful throughput?

How repair mixed-label analysis?

What points toward generator saturation?

Why are different thread/timer settings invalid for a causal comparison?

Next lesson

Checkpoint: discriminate competing hypotheses

Lesson 5 applies the diagnostic sequence to a bounded checkpoint that distinguishes generator, network, and SUT causes before making a capacity claim.

Official references and version notes

Version and compatibility note

Rechecked against current primary documentation on 2026-09-05. Mandatory runtime: Apache JMeter 5.6.3, Java 17, no third-party plugin. Dashboard percentiles default to 90/95/99 and are configurable. JMeter warns dashboard percentile estimates can differ from Aggregate Report, especially for small/wide samples; therefore this chapter audits p50/p90/p95/p99 from raw CSV JTL with a documented nearest-rank method and retains the dashboard as corroborating evidence. The lab uses a standard closed Thread Group, so “offered load” means configured concurrency plus pacing pressure, not an independent fixed open-arrival RPS.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.