Chapter 30Lesson 01~195 minutes

Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design: Core Concepts and Mental Model

Chapter 29 showed how correlated JMeter and target evidence can support a bottleneck hypothesis. But even a beautifully correlated result can be wrong if the experiment accidentally includes cold startup, compares different datasets, uses a traffic model that stops sending work during stalls, or declares a regression from one noisy run. This chapter puts measurement validity before threshold enforcement.

ValidityWarm-upCoordinated omissionNoiseExperimental design

Learning objectives

  • Design a performance experiment from question and hypothesis to uncertainty-aware conclusion.
  • Define warm-up, measurement window, repetition, noise, open/closed workload and coordinated omission.
  • Separate configured load, achieved load and the traffic that would have arrived under another workload model.
  • Inspect versions, environment, data, timing and run variation before changing a threshold.
  • Explain why validity precedes CI gate enforcement.

1. The practical problem: a measured difference can be an artifact

Suppose branch B is 12 ms slower than branch A in one run. That difference could be product code, JVM/JIT warm-up, a cold cache, different test data, background CPU noise, a runner change, clock skew, an arrival-model artifact, or normal run-to-run variation. A valid experiment controls or measures these competing explanations before calling the difference a regression.

Mandatory boundary: all executable work is local at 127.0.0.1:8030. Normal measurement = 4 threads ×25 loops =100 samples/run; warm-up =20 samples; arrival-oriented demonstration ≤60 requests at 10/s for 6 s. No public/shared target, real credential, remote engine, paid telemetry, OS/JVM tuning or production failure injection is required.

2. Mental model: the experiment is a chain of assumptions

Experimental-validity workflow

Validity is an experiment property, not a dashboard checkbox. Every arrow below changes what a measured difference can legitimately mean.

flowchart TD
Q[Question] --> H[Hypothesis + predicted direction]
H --> E[Controlled environment + versions/data]
E --> W[Workload model
closed or arrival-oriented]
W --> U[Warm-up]
U --> M[Measurement window]
M --> R[Repeated runs]
R --> C[Engineering/statistical comparison]
C --> K[Conclusion + uncertainty]
V[Generator/target validity evidence] --> C
N[Noise/background state] --> C

The question defines what decision the test should support. A hypothesis predicts what metric should change. Environment/data/tool versions define comparability. The workload model determines how arrivals react to response time. Warm-up removes intentionally non-representative cold behavior. A predeclared measurement window defines which samples count. Repetitions reveal variation. Only then do percentiles/means/throughput support a conclusion—and the conclusion must still list uncertainty.

3. New terms

Term Beginner definition Why it matters
Warm-up Traffic intentionally excluded from the comparison while JVM/cache/connection/service state reaches the intended operating regime. Prevents startup/JIT/cache costs from masquerading as steady-state latency.
Measurement window Predeclared interval/samples whose results count toward the comparison. Prevents post-hoc cherry-picking.
Repetition Re-running the same protocol/environment/workload as an independent comparison unit. Shows run-to-run variation/noise.
Closed workload Each virtual user sends a next request only after the prior request plus think/pacing completes. Arrival rate falls when responses stall; coordinated omission risk.
Open/arrival-oriented workload Arrivals are scheduled independently of previous response completion. Can preserve intended arrival pressure during stalls; requires enough injector concurrency.
Coordinated omission The generator omits requests precisely while the SUT is slow because future requests wait for prior ones. Observed latency distribution can underrepresent what scheduled arrivals would have experienced.
Noise Uncontrolled or intentionally modeled variation: scheduler, GC, antivirus, background jobs, network, cache, DB, CI neighbors. Can obscure or mimic a product difference.
Validity threat A plausible alternative explanation for the result. Must be controlled, measured or acknowledged.

4. State that must be frozen or observed

Boundary State
Generator JMeter/Java, CPU/GC/network/disk, thread group, timer/rate schedule, plugins.
Tree/scope Warm-up versus measurement labels, controllers, assertions, listeners and sample filters.
Variables/properties/data Resolved target/build/data/seed/window/run-id values and identical datasets.
Protocol/session Keep-alive, DNS/TLS, cookies/tokens, cache/session reuse and connection warm-up.
Target Build/config, base latency, cache/JIT state, background jobs, service dependencies.
Results JTL timestamp/elapsed/latency/connect/success/thread fields, dashboard, log and analyzer version.
Trust/safety Authorized target, fake credentials and bounded failure/noise injection.
Validity Configured/achieved sample counts, model limits, generator headroom, repetitions and remaining threats.

5. Closed versus arrival-oriented interpretation

In the stable JMeter Thread Group, a thread waits for its current sample to finish before proceeding, so the realized arrival rate depends on response time. That is often correct for “N concurrent users.” It is not the same question as “the system must receive 10 arrivals/s even during a 700 ms stall.” The latter is arrival-oriented and must schedule work independently of previous completion.

JMeter 5.6.3 includes an experimental Open Model Thread Group that can express schedules such as rate(10/sec) random_arrivals(6 sec). Because the component is explicitly experimental, this chapter's mandatory coordinated-omission demonstration uses a small Python standard-library scheduler as a stable fallback; the JMeter Open Model path is optional.

6. Timing boundaries and warm-up

JMeter elapsed/latency/connect fields begin at sampler execution—not during a Constant Timer. Warm-up can occur in multiple layers: Java JIT, DNS/TLS/connection pool, target cache, database pages, application code paths and generated data. Do not assume “30 seconds” is universally enough. Inspect time series and target state, choose a bounded warm-up rule, and keep it identical between A/B conditions.

7. Repetition and uncertainty

One run estimates one realization. Three repetitions are still a small sample, but they expose whether a 15 ms difference is large relative to run variation. This chapter reports median, sample standard deviation, coefficient of variation and a simple t-based interval for the mean. The interval is an uncertainty illustration, not a magical significance certificate.

8. Read-only/non-destructive inspection first

& "$env:JMETER_HOME\bin\jmeter.bat" -v
java -version

Get-FileHash `
  .\plans\warmup.jmx, .\plans\measure.jmx, `
  .\config\validity.properties, .\data\* `
  -Algorithm SHA256

Get-Content .\experiment\protocol.json
Get-Content .\results\*\summary.json -ErrorAction SilentlyContinue

Also inspect generator process/JVM state and target /health. Do not change warm-up, data, runtime or thresholds until the actual executed state is proven.

Evidence baseline: every controlled run in this chapter uses Apache JMeter 5.6.3 with Java 17 and retains the raw CSV JTL plus the matching jmeter.log before any filtering or comparison. The lab uses no real credentials.

9. DevOps connection

Performance engineering becomes credible when a team can distinguish a product regression from cold startup, workload-model bias, a noisy CI runner or random variation. The same evidence discipline improves release gates, incident analysis, capacity experiments and optimization reviews.

Knowledge check

Why does measurement validity come before threshold enforcement?

What is coordinated omission in a closed workload?

Does warm-up mean discarding whatever early samples look slow?

Why repeat runs?

Why is Open Model optional here?

Next lesson

Build the experiment protocol

Lesson 2 runs with/without warm-up, adds deterministic noise, calculates run variation, marks windows and demonstrates coordinated omission against the same local fixture.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current primary documentation on 2026-09-05. The mandatory runtime remains Apache JMeter 5.6.3 with Java 17 and no third-party plugin. JMeter 5.6.3 requires Java 8+; the current 5.6.x changes page recommends Java 17 or later. The Open Model Thread Group remains explicitly experimental; the mandatory course path therefore uses the stable standard Thread Group for closed-model measurements and a bounded Python-standard-library fixed-arrival probe to demonstrate coordinated omission. An optional Open Model schedule such as rate(10/sec) random_arrivals(6 sec) is shown only as a current 5.6.3 comparison. All important comparisons use raw CSV JTL plus a documented label/window filter; dashboards remain corroborating evidence. JMeter elapsed time begins immediately before sending the request and ends after the last response byte; latency ends at the first response, and connect time covers connection establishment including SSL handshake.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.