Chapter 01Lesson 01~100 minutes

Performance Testing Foundations: Load, Stress, Spike, Soak, and Capacity: Core Concepts and Mental Model

Build the performance-engineering mental model that must exist before a JMeter test plan can produce trustworthy evidence: business risk, workload intent, target authorization, achieved load, client/server observations, and experimental validity.

Workload modelPercentilesCapacitySLOValidity

Learning objectives

  • Translate a business performance concern into a measurable workload hypothesis.
  • Distinguish concurrency, arrival rate, throughput, response time, percentiles, errors, saturation, and capacity.
  • Classify load, stress, spike, soak/endurance, and capacity questions without treating them as synonyms.
  • Separate configured workload, achieved workload, load-generator health, SUT state, and result evidence.
  • Write explicit authorization, stop conditions, and acceptance criteria before generating traffic.
  • Explain why one successful JMeter run cannot by itself prove capacity or root cause.

1. Current baseline and chapter boundary

Apache JMeter is a Java application that can generate protocol-level traffic and record sample results. The current production download is JMeter 5.6.3, which requires Java 8 or newer. This course treats JMeter as a load-generation and measurement tool, not as a browser, an APM platform, a security scanner, or a capacity oracle.

Chapter 01 deliberately begins with the experiment, not with menus. Chapter 02 will teach Java setup, installation, the JMeter architecture, and the first authored test plan. Here, the goal is to know what question a future plan must answer and what evidence makes the answer credible.

Authorization boundary: load testing is inherently disruptive. Every executable exercise in this chapter is restricted to 127.0.0.1 and uses conservative traffic ceilings. Do not redirect the examples to public, employer, customer, shared staging, identity, commerce, or third-party systems.

2. Performance testing starts with business risk, not threads

“Run 500 users” is not yet a performance question. A useful question names the behavior that matters, the workload that should exercise it, the environment in which the result is valid, and the condition that would count as success or failure. For example: “Can the synthetic checkout service sustain 20 completed iterations per second for ten minutes in the dedicated test environment while p95 elapsed response time remains below 250 ms and the error rate remains below 0.5%, without saturating the injector?”

That statement is stronger because every noun can be measured. It identifies a workload, a time window, two service objectives, and an injector-health constraint. It also avoids claiming that JMeter itself knows the cause of a slowdown. JMeter observes client-side behavior; server-side telemetry must explain what the target was doing.

3. Mental model: from performance question to validity judgment

Performance experiment evidence flow

Use this diagram together with the surrounding prose: the arrows represent ownership and evidence flow, not decorative sequencing.

flowchart TD
B[Business risk / SLO] --> H[Workload hypothesis]
H --> M[Concurrency or arrival model]
M --> L[Controlled load shape]
L --> G[JMeter load generator]
G --> P[Protocol requests]
P --> S[System under test]
S --> R[Responses]
R --> J[JTL + jmeter.log]
S --> T[SUT telemetry]
G --> GH[Generator health]
J --> V[Validity judgment]
T --> V
GH --> V
V --> D[Decision]

The key word is judgment. A graph does not automatically become a conclusion. If the load generator is CPU-bound, if the workload changed between runs, if the target build changed, or if the observed samples include warm-up effects that one run excluded and another included, then a numeric difference may not represent a product regression.

4. Vocabulary: configured work, observed work, and response

This chapter uses elapsed response time when discussing the end-to-end time recorded for a JMeter sample. JMeter also has a field named Latency with more specific semantics; later reporting chapters separate those fields explicitly. Do not use “latency” as an unqualified synonym for every timing field.

Term Meaning in this chapter What it is not
Virtual user / thread One concurrent scenario executor in a closed-user model. A promise of one request per second.
Concurrency How many scenario executions can be in progress at once. Arrival rate or throughput.
Arrival rate How often new work is offered to the system over time. The same thing as active users.
Throughput Completed samples/transactions per time unit. Configured load; it is an observed result.
Elapsed response time Time from sampler start until the response is completed. JMeter's narrower Latency field; keep the names explicit.
Percentile A rank in the observed latency distribution, such as p95. A guarantee that every request is below the value.
Error rate Fraction of samples that fail transport/protocol/assertion criteria. Only HTTP 5xx; correctness can fail in other ways.
Saturation A constrained resource is near or at its effective limit. High CPU alone; interpret with other evidence.
Capacity Highest defined workload that still satisfies explicit correctness/performance criteria in a stated environment. A number inferred from one uncontrolled laptop run.
SLO A target for a measured service-level indicator under stated conditions. A metric with no workload/environment definition.

A useful closed-system intuition is that active users cannot issue new iterations while they are still waiting for the previous iteration and any think time. Therefore, the same thread count can produce radically different throughput when response times or pacing change. That is why “100 threads = 100 requests/second” is false.

5. Load, stress, spike, soak, and capacity answer different questions

Question type Primary question Typical load shape What a useful result says
Load Does the service satisfy objectives at expected traffic? Representative expected workload, usually with warm-up and steady state. Behavior at the stated normal workload.
Stress How does behavior degrade as offered load exceeds normal expectations? Controlled steps or ramps beyond expected demand. Where errors/latency/saturation emerge and how recovery behaves.
Spike Can the service absorb a sudden transient increase? Fast increase, short hold, controlled recovery. Transient resilience and recovery, not long-term capacity.
Soak / endurance Does behavior drift or leak over a long steady run? Stable representative workload for an extended period. Time-dependent problems such as leaks, queue growth, or resource exhaustion.
Capacity What maximum workload still meets explicit acceptance criteria? Repeated controlled levels near the boundary. A bounded environment-specific threshold with uncertainty, not a universal product limit.

These labels describe experimental intent, not special magic modes. The same basic JMeter components can participate in several experiment types, but the schedule, duration, evidence, and interpretation must match the question. A two-minute “soak test” is contradictory because it cannot expose long-duration drift.

6. State boundaries: what JMeter knows and what it cannot know

State/evidence Owned by Examples Why you must separate it
Workload definition Test design / JMX users, loops, ramp, timers, arrival schedule Configured intent can differ from achieved load.
Generator state JMeter JVM + host OS CPU, heap/GC, sockets, disk, network A saturated injector can make the target look slower than it is.
Protocol/session state Sampler/client layer HTTP connections, cookies, tokens Session behavior changes request cost and correctness.
Target state System under test threads, queues, DB pool, CPU, cache This is where server-side bottlenecks live.
Raw result evidence JTL + jmeter.log sample timestamps, elapsed time, success, errors Preserve first-failure evidence before summary/report transformation.
Server telemetry SUT observability CPU, memory, queue depth, DB metrics Client latency alone rarely proves root cause.
Validity record Experiment owner versions, workload, environment, anomalies Without it, comparisons can be numerically precise but scientifically invalid.

When a result worsens, begin by locating which state changed. A new application build, different dataset, injector JVM, network path, or workload profile can all alter results. Good performance engineering records these dependencies before the run so diagnosis does not become guesswork afterward.

7. Acceptance criteria need workload context

An SLO such as “p95 below 200 ms” is incomplete when used as a test gate unless you also state which transaction population, workload level, environment, warm-up/steady-state window, and correctness condition it applies to. A fast error response is not success. Likewise, a lower average can coexist with a much worse tail.

For a beginner-friendly synthetic example, use this charter:

Charter field Concrete Chapter 01 value
Target Loopback-only fixture at http://127.0.0.1:8000.
Question Can a tiny local baseline complete correctly without injector distress?
Configured workload 2 threads, 2-second ramp-up, 5 loops each: at most 10 /work samples.
Correctness No failed samples; fixture remains healthy before and after.
Timing evidence Record p50 and p95 elapsed response time, but do not treat ten samples as a stable production percentile estimate.
Generator guard Observe Java/host CPU and memory; abort if the workstation becomes unstable.
Stop condition Stop immediately if the target is not loopback, errors appear unexpectedly, or resource use becomes unsafe.
Validity limit The run is a teaching baseline only; it does not establish service capacity.

8. Read-only preflight before any load exists

Before generating traffic, prove the intended target and runtime assumptions. The first checks are intentionally non-destructive:

java -version
jmeter -v

On Windows PowerShell, when JMeter is not on PATH, the equivalent version check is:

java -version
& "$env:JMETER_HOME\bin\jmeter.bat" -v

For this chapter, record the values rather than “fixing” anything yet. If JMeter is not installed, complete the conceptual work and perform the reproducible installation in Chapter 02 before executing the load lab.

9. Common foundation mistakes

  • Equating users with request rate. Threads constrain concurrency; achieved rate emerges from scenario time, timers, response time, errors, and capacity.
  • Reporting averages alone. Averages can hide tail behavior that affects a meaningful fraction of users.
  • Choosing thresholds after seeing the result. That converts an acceptance criterion into retrospective storytelling.
  • Ignoring injector health. A saturated load generator invalidates strong claims about target capacity.
  • Running against an uncontrolled public target. That is unsafe and creates uncontrollable network/environment noise.
  • Changing workload and environment together. You lose causal interpretability because multiple factors moved at once.

10. Hands-on thinking lab: classify the question before choosing a tool

No traffic is generated in this lesson. Classify each question and write the missing evidence requirement:

  1. “Can normal weekday traffic meet p95 and error objectives?” → load.
  2. “At what offered load do queues, errors, or latency become unacceptable?” → stress/capacity characterization.
  3. “What happens if traffic jumps from 5 to 50 requests/second in a few seconds?” → spike.
  4. “Does memory or queue depth drift during six hours of representative traffic?” → soak/endurance.
  5. For each case, add generator-health evidence and SUT telemetry before claiming the target is the limiting factor.

11. Why this matters in DevOps

A performance result can become release evidence only when it is reproducible and reviewable. That means the workload, environment, target build, versions, data, stop criteria, raw results, generator health, and server-side observations must travel together. The rest of this course turns that operating model into JMeter components and automation, but the evidence contract starts here.

Knowledge check

Why can 50 JMeter threads not be translated directly into 50 requests per second?

What is the difference between configured load and achieved load?

Why is p95 usually more informative than an average alone?

A laptop reaches 100% CPU while throughput stops increasing. Can you claim the target has reached capacity?

What is the first safety question before a load test?

Next lesson

From experiment charter to one bounded loopback run

Lesson 2 creates a disposable local service, performs target and runtime preflight, runs a deliberately tiny JMeter plan in CLI mode, and separates configured workload from observed evidence.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current Apache JMeter primary documentation on 2026-09-04. The production download is Apache JMeter 5.6.3 and requires Java 8+. The mandatory Chapter 01 executable lab uses Java 17 as a pinned local lab choice, no third-party plugins, no distributed engines, and only the loopback target 127.0.0.1:8000. Apache guidance requires CLI mode for actual load execution; GUI mode is limited to construction and bounded debugging. The Open Model Thread Group is discussed conceptually only and remains marked experimental in the current Component Reference.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.