Performance Testing Foundations: Load, Stress, Spike, Soak, and Capacity: Configuration, Design Patterns, and Trade-Offs
Choose a performance-test design that matches the question. This lesson treats workload model, threshold policy, percentile reporting, repetition, target safety, and generator cost as explicit engineering trade-offs rather than JMeter checkbox choices.
Learning objectives
- Choose among load, stress, spike, soak, and capacity experiments based on the decision being made.
- Explain closed/concurrency-based versus open/arrival-rate workload thinking.
- Describe the current experimental status of JMeter's Open Model Thread Group and choose a stable fallback.
- Separate exploratory characterization from an enforceable performance gate.
- Use percentiles and error rates without hiding the underlying sample population and time window.
- Design repeated experiments that change one important factor at a time.
1. The design starts with the decision
A load test for a release gate and a stress test for architecture discovery may use some of the same samplers, but they should not use the same acceptance logic. A gate asks a binary reviewable question against a stable baseline. A characterization run asks how the system behaves across a range and may intentionally cross a failure boundary.
| Decision | Experiment design | Why |
|---|---|---|
| Approve a release for expected demand | Representative load with predeclared p95/error objectives and injector headroom. | The gate should match the operational demand it protects. |
| Find the onset of degradation | Controlled stress steps with server and generator telemetry. | You need a response curve, not a single pass/fail point. |
| Test recovery from a flash crowd | Spike with a fast increase, short hold, and observed recovery. | The transient and recovery are the phenomenon of interest. |
| Look for leaks or queue drift | Long soak at a representative workload. | Time-dependent defects need duration. |
| Estimate a capacity boundary | Repeated levels near the SLO boundary with controlled environment and replication. | The result is environment-specific and uncertainty matters. |
2. Closed-user and open-arrival mental models
In a closed workload, a fixed pool of virtual users repeatedly executes work. When responses slow down, those users wait longer and naturally offer less new work. In an open model, new arrivals are scheduled independently of how long earlier requests take; the system may accumulate concurrency as response times rise.
Use this diagram together with the surrounding prose: the arrows represent ownership and evidence flow, not decorative sequencing.
flowchart TB subgraph Closed[Closed / concurrency-oriented] U[Fixed user pool] --> W[Request] W --> R[Wait for response] R --> Z[Think / pacing] Z --> W end subgraph Open[Open / arrival-oriented] A[Arrival schedule] --> Q1[New request] A --> Q2[New request] A --> Q3[New request] Q1 --> S[System] Q2 --> S Q3 --> S end
Neither model is universally “more realistic.” Choose the model that resembles how demand reaches the real system. Human sessions often have closed-user aspects; queues, events, or externally scheduled API calls can be closer to open arrivals.
3. Exploratory characterization versus enforceable gate
An exploratory run can answer, “Where does the error curve become unacceptable?” A release gate must answer, “Did this version satisfy a rule that was defined before the run?” Mixing the two creates moving thresholds: engineers see the result, then choose a threshold that happens to pass.
| Property | Exploratory run | Enforceable gate |
|---|---|---|
| Primary goal | Learn behavior. | Make a repeatable release decision. |
| Workload | May sweep several levels. | Fixed/versioned profile. |
| Thresholds | May be descriptive. | Predeclared and reviewable. |
| Environment | Can be a controlled lab. | Must be stable enough for comparison. |
| Result | Curves, hypotheses, next experiments. | Pass/fail plus raw evidence and diagnostic context. |
| Change policy | Experiment can evolve quickly. | Baseline/threshold changes require governance. |
4. Average versus percentiles: summarize without erasing the distribution
Consider ten elapsed times: 40, 41, 41, 42, 42, 43, 44, 45, 48, and 300 ms. The average is 72.6 ms. That average is mathematically correct but hides the one severe tail observation. A percentile is also only a summary, yet a high percentile makes tail behavior harder to hide.
Percentile definitions and estimators can vary between tools. When comparing results, keep the tool/version and population consistent. Do not compare a p95 computed over one set of labels and success rules with a p95 from a differently filtered population and call the difference a regression.
5. One run is an observation; repeated controlled runs support a claim
JIT compilation, filesystem cache, application cache, DNS, connection reuse, background jobs, GC, shared infrastructure, and many other sources create run-to-run variation. A robust comparison keeps important factors fixed, separates warm-up from steady state when relevant, and repeats the experiment enough to understand normal variance.
For the Chapter 01 local fixture, a useful teaching sequence is
three tiny runs with identical JMX and fixture delay. The aim is not
capacity; it is to observe whether the measured elapsed times are
reasonably consistent. Then change only
delay_ms and repeat. The causal story is much stronger
because one major factor changed.
6. Keep configuration layers separate
| Layer | Examples | Owner / concern |
|---|---|---|
| JMeter test plan | Thread Group, sampler, timers, assertions | Workload and protocol behavior. |
| JMeter runtime | JMeter properties, result save configuration | Engine behavior and evidence volume. |
| Java/JVM | Java version, heap, GC | Injector runtime capability. |
| Host OS/network | CPU, sockets, DNS, interface bandwidth | Load-generator constraints and path effects. |
| System under test | Build, DB pool, caches, replicas | Target capacity and correctness. |
| Plugin/driver | Third-party components and dependencies | Compatibility and distributed synchronization. |
| CI/container/orchestrator | Workspace, image, replicas, network | Execution environment, not JMeter semantics. |
A performance experiment becomes uninterpretable when engineers change several of these layers and attribute the result to only one of them.
7. Worked decision: a synthetic service has an expected 10 requests/second peak
Suppose a dedicated local/private test environment expects roughly 10 requests/second during normal operation and the release requirement is p95 elapsed time below 200 ms with less than 1% failed samples. A useful first gate design targets that representative load and verifies injector headroom. A stress experiment is separate: perhaps step above 10 requests/second to learn where the SLO stops holding. Do not use the stress ceiling as the normal-load gate.
If the real production demand arrives independently of response completion, open-arrival reasoning may better represent it. If the application is mostly a fixed population of interactive sessions, a closed user model with realistic think time may be a better approximation. The choice comes from demand semantics, not from which JMeter component is easiest to configure.
8. Generator cost and validity trade-offs
Higher traffic is not automatically a better test. More threads increase JVM, socket, memory, result, and target cost. More result fields can increase disk/network overhead. Longer tests improve time coverage but cost more environment time. Distributed generation can extend load but adds synchronization and network evidence complexity.
The design goal is the smallest controlled experiment that can answer the question with adequate confidence. Chapter 26 will cover generator sizing in depth; the invariant begins now: measure injector health before believing target-capacity conclusions.
9. Free/local design path
Every mandatory concept in this chapter can be practiced with the loopback Python fixture, Apache JMeter, the JDK, raw JTL files, and OS resource monitors. Managed load clouds, commercial observability, shared staging, and Kubernetes are not prerequisites. If a later production architecture uses those systems, the same evidence contract still applies.
Knowledge check
Why can a closed workload reduce its offered rate when the target slows down?
The fixed user pool waits longer for responses, so each user begins new iterations less frequently.
Is JMeter's Open Model Thread Group currently a stable component guarantee?
No. The current Apache Component Reference marks it experimental and warns that it may change.
Why should a release gate threshold be defined before the run?
So the decision rule is auditable and is not chosen retrospectively to fit the observed result.
Can two p95 values be compared safely if their sample populations use different labels or filters?
Not without reconciling those populations. Percentile comparison is meaningful only when workload, population, time window, and calculation semantics are aligned.
Why repeat a controlled experiment?
To estimate normal run-to-run variation and reduce the chance that a one-off environmental effect is mistaken for a product change.
Official references and version notes
- Apache JMeter downloads — current production-release and Java requirement baseline.
- Getting Started — GUI authoring, CLI load execution, Java requirements, CLI flags, and operational guidance.
- Best Practices — load-generation and result-collection practices.
- Component Reference — Thread Group and current Open Model Thread Group status.
- HTML Dashboard Report — result-report terminology and percentile-oriented reporting.
Version-sensitive statements were rechecked against current Apache
JMeter primary documentation on 2026-09-04. The production
download is Apache JMeter 5.6.3 and requires Java 8+. The
mandatory Chapter 01 executable lab uses Java 17 as a pinned local
lab choice, no third-party plugins, no distributed engines, and
only the loopback target 127.0.0.1:8000. Apache
guidance requires CLI mode for actual load execution; GUI mode is
limited to construction and bounded debugging. The Open Model
Thread Group is discussed conceptually only and remains marked
experimental in the current Component Reference.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.