Performance Testing Foundations: Load, Stress, Spike, Soak, and Capacity: Diagnostics, Failure Modes, and Production Practices
Diagnose performance evidence without hiding the original cause. The workflow begins with raw artifacts and target authorization, then separates workload mistakes, protocol/setup failures, injector limits, and target-side saturation.
Learning objectives
- Apply a preserve-first diagnostic sequence before changing configuration.
- Explain why virtual-user count, average latency, and a single throughput number can mislead.
-
Diagnose a safe loopback connection failure from JTL and
jmeter.log. - Distinguish target latency from injector CPU/GC, listener/result overhead, DNS/network behavior, and server saturation.
- Reject unsafe troubleshooting shortcuts such as public-target experiments, blanket security disablement, and unbounded workload increases.
- Repair the smallest failing layer and rerun the smallest controlled workload.
1. Preserve evidence before touching the test
When a run surprises you, do not immediately increase heap, add
retries, delete results, or rerun at a different load. First
preserve the exact JMX, data files, properties, CLI command, JTL,
jmeter.log, generator snapshot, SUT telemetry, target
build/environment identity, and authorization context.
Use this diagram together with the surrounding prose: the arrows represent ownership and evidence flow, not decorative sequencing.
flowchart TD E[Preserve JTL, jmeter.log, JMX, run record] --> V[Confirm JMeter / Java / plugin versions] V --> A[Confirm authorized target and exact CLI inputs] A --> C[Validate tree scope + resolved config] C --> P[Inspect protocol/session/data state] P --> G[Inspect generator JVM/OS/network] G --> S[Inspect SUT telemetry] S --> X[Check distributed/CI/container layer if present] X --> F[Apply smallest correction] F --> R[Rerun smallest controlled workload]
2. Failure mode: “20 users means 20 requests/second”
Imagine a closed test with 20 threads and no think time. If each iteration takes about 250 ms, the idealized upper bound could be roughly 80 iterations/second before overhead and bottlenecks; if response time rises to one second, the same thread pool may only complete around 20 iterations/second. The thread count stayed constant while achieved rate changed dramatically.
The fix is not a new flag. The fix is to state whether the experiment controls concurrency or arrival rate, then measure the achieved workload. Later chapters add pacing and explicit workload components; Chapter 01 establishes the interpretation discipline.
3. Failure mode: average-only reporting
A run can show an acceptable average while a minority of samples suffer severe tail latency. Always pair the central tendency with percentiles, errors, sample counts, workload level, and time windows. If only a small number of samples exist, say that the percentile is unstable rather than displaying a precise number with false confidence.
4. Intentionally broken local example: wrong port, preserved evidence
Start the Chapter 01 fixture on 127.0.0.1:8000. Copy
the JMX and change only the sampler port to 8001. Keep
the run bounded at ten samples. The target remains loopback-only, so
the failure is safe.
Run the broken copy into a new result directory. Expected evidence
is a set of unsuccessful samples and a connection failure such as a
Java ConnectException. The exact message is
platform/JDK dependent, so inspect rather than hard-code it.
jmeter -n -t chapter01-wrong-port.jmx -l results/wrong-port.jtl -j results/wrong-port-jmeter.log
Diagnosis:
- Preserve both files.
-
Re-run
curl --fail --silent http://127.0.0.1:8000/healthto prove the fixture is healthy on the authorized port. - Inspect the JMX and confirm the sampler points to 8001.
- Repair only that port.
- Run a new artifact set; do not overwrite the failed run.
This is a configuration/target-address failure, not evidence that the target is slow or lacks capacity.
5. Failure mode: zero think time without justification
Interactive users usually do not submit a new action immediately after every response. A zero-pause closed loop can create a much higher request rate than the business workflow. Conversely, machine-to-machine APIs may legitimately have little think time. The mistake is not “zero is always wrong”; the mistake is using zero without tying it to a demand model.
7. Failure mode: claiming target capacity from a saturated injector
Suppose throughput plateaus while the JMeter host reaches sustained CPU saturation, long GC pauses, socket exhaustion, or network saturation. The generator is no longer a neutral source of offered load. Increasing thread count can make measurement worse because the injector spends more time scheduling, allocating, logging, or waiting on its own resources.
| Signal | Possible interpretation | What to verify before blaming target |
|---|---|---|
| Generator CPU near saturation | Injector may be compute-bound. | JVM/process CPU, plan scripting, listeners, host competition. |
| Heap/GC pressure | Allocation/result retention may be limiting. | Heap usage, GC pauses, retained response data, heavy listeners. |
| Socket/port errors | Host/network limits may be reached. | File descriptors, ephemeral ports, connection reuse, OS errors. |
| Network maxed | Injector cannot offer more bytes. | NIC throughput, packet loss, result-transfer traffic. |
| Target telemetry also saturated | Target may truly be limiting. | Correlate timestamps and verify generator still has headroom. |
8. Failure mode: changing workload and environment together
If run A uses build 101 at 10 requests/second and run B uses build 102 at 20 requests/second, a latency difference cannot be attributed cleanly to the code change. The performance result changed under two major factors. Keep the workload constant for regression comparison, or explicitly design a factorial experiment and analyze it as such.
9. Causal timing checklist
- Sampler elapsed time: what the client observed for the request/response operation.
- Timer/pacing delay: intentional waiting that shapes user behavior and should not be confused with server response time.
- Generator CPU/GC: client-side work that can delay request issuance or result processing.
- DNS/connect/network: path and connection establishment effects distinct from application processing.
- Listener/save-service cost: evidence collection can consume injector resources.
- Server saturation: queues, CPU, pools, locks, downstreams, and databases require SUT telemetry.
10. Production practices that preserve diagnostic value
Use unique run IDs, immutable result directories, stable sampler labels, explicit environment/build identities, time-synchronized hosts, conservative preflight, and a smallest-workload reproduction path. Avoid blanket retries, giant sleeps, arbitrary heap increases, result deletion, broad property hacks, disabling TLS/RMI security, or uncontrolled load increases. Those actions can erase the evidence needed to understand the first failure.
Knowledge check
A 20-thread test completes 70 requests/second. Is JMeter broken because 70 is not 20?
No. Thread count is concurrency, not a fixed request rate. Throughput emerges from iteration duration, pacing, target behavior, and generator capacity.
What artifact should be deleted first when a run fails?
None. Preserve the first-failure JTL, jmeter.log, JMX, inputs, and environment evidence before cleanup.
The wrong-port run shows ConnectException. What is the least destructive correction?
Verify the authorized fixture is healthy on 127.0.0.1:8000, correct only the sampler port, and rerun into a new artifact set.
Why is a saturated injector a validity problem?
Because it may be unable to produce the intended workload or may add client-side delay, so observed throughput/latency cannot be attributed cleanly to target capacity.
Why is changing build and workload in the same comparison weak evidence?
Because two major factors changed, so the observed difference cannot be causally attributed to one of them.
Official references and version notes
- Apache JMeter downloads — current production-release and Java requirement baseline.
- Getting Started — GUI authoring, CLI load execution, Java requirements, CLI flags, and operational guidance.
- Best Practices — load-generation and result-collection practices.
- Component Reference — Thread Group and current Open Model Thread Group status.
- HTML Dashboard Report — result-report terminology and percentile-oriented reporting.
Version-sensitive statements were rechecked against current Apache
JMeter primary documentation on 2026-09-04. The production
download is Apache JMeter 5.6.3 and requires Java 8+. The
mandatory Chapter 01 executable lab uses Java 17 as a pinned local
lab choice, no third-party plugins, no distributed engines, and
only the loopback target 127.0.0.1:8000. Apache
guidance requires CLI mode for actual load execution; GUI mode is
limited to construction and bounded debugging. The Open Model
Thread Group is discussed conceptually only and remains marked
experimental in the current Component Reference.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.