Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design: Core Concepts and Mental Model
Chapter 29 showed how correlated JMeter and target evidence can support a bottleneck hypothesis. But even a beautifully correlated result can be wrong if the experiment accidentally includes cold startup, compares different datasets, uses a traffic model that stops sending work during stalls, or declares a regression from one noisy run. This chapter puts measurement validity before threshold enforcement.
Learning objectives
- Design a performance experiment from question and hypothesis to uncertainty-aware conclusion.
- Define warm-up, measurement window, repetition, noise, open/closed workload and coordinated omission.
- Separate configured load, achieved load and the traffic that would have arrived under another workload model.
- Inspect versions, environment, data, timing and run variation before changing a threshold.
- Explain why validity precedes CI gate enforcement.
1. The practical problem: a measured difference can be an artifact
Suppose branch B is 12 ms slower than branch A in one run. That difference could be product code, JVM/JIT warm-up, a cold cache, different test data, background CPU noise, a runner change, clock skew, an arrival-model artifact, or normal run-to-run variation. A valid experiment controls or measures these competing explanations before calling the difference a regression.
127.0.0.1:8030. Normal measurement = 4 threads ×25
loops =100 samples/run; warm-up =20 samples; arrival-oriented
demonstration ≤60 requests at 10/s for 6 s. No public/shared target,
real credential, remote engine, paid telemetry, OS/JVM tuning or
production failure injection is required.
2. Mental model: the experiment is a chain of assumptions
Validity is an experiment property, not a dashboard checkbox. Every arrow below changes what a measured difference can legitimately mean.
flowchart TD Q[Question] --> H[Hypothesis + predicted direction] H --> E[Controlled environment + versions/data] E --> W[Workload model closed or arrival-oriented] W --> U[Warm-up] U --> M[Measurement window] M --> R[Repeated runs] R --> C[Engineering/statistical comparison] C --> K[Conclusion + uncertainty] V[Generator/target validity evidence] --> C N[Noise/background state] --> C
The question defines what decision the test should support. A hypothesis predicts what metric should change. Environment/data/tool versions define comparability. The workload model determines how arrivals react to response time. Warm-up removes intentionally non-representative cold behavior. A predeclared measurement window defines which samples count. Repetitions reveal variation. Only then do percentiles/means/throughput support a conclusion—and the conclusion must still list uncertainty.
3. New terms
| Term | Beginner definition | Why it matters |
|---|---|---|
| Warm-up | Traffic intentionally excluded from the comparison while JVM/cache/connection/service state reaches the intended operating regime. | Prevents startup/JIT/cache costs from masquerading as steady-state latency. |
| Measurement window | Predeclared interval/samples whose results count toward the comparison. | Prevents post-hoc cherry-picking. |
| Repetition | Re-running the same protocol/environment/workload as an independent comparison unit. | Shows run-to-run variation/noise. |
| Closed workload | Each virtual user sends a next request only after the prior request plus think/pacing completes. | Arrival rate falls when responses stall; coordinated omission risk. |
| Open/arrival-oriented workload | Arrivals are scheduled independently of previous response completion. | Can preserve intended arrival pressure during stalls; requires enough injector concurrency. |
| Coordinated omission | The generator omits requests precisely while the SUT is slow because future requests wait for prior ones. | Observed latency distribution can underrepresent what scheduled arrivals would have experienced. |
| Noise | Uncontrolled or intentionally modeled variation: scheduler, GC, antivirus, background jobs, network, cache, DB, CI neighbors. | Can obscure or mimic a product difference. |
| Validity threat | A plausible alternative explanation for the result. | Must be controlled, measured or acknowledged. |
4. State that must be frozen or observed
| Boundary | State |
|---|---|
| Generator | JMeter/Java, CPU/GC/network/disk, thread group, timer/rate schedule, plugins. |
| Tree/scope | Warm-up versus measurement labels, controllers, assertions, listeners and sample filters. |
| Variables/properties/data | Resolved target/build/data/seed/window/run-id values and identical datasets. |
| Protocol/session | Keep-alive, DNS/TLS, cookies/tokens, cache/session reuse and connection warm-up. |
| Target | Build/config, base latency, cache/JIT state, background jobs, service dependencies. |
| Results | JTL timestamp/elapsed/latency/connect/success/thread fields, dashboard, log and analyzer version. |
| Trust/safety | Authorized target, fake credentials and bounded failure/noise injection. |
| Validity | Configured/achieved sample counts, model limits, generator headroom, repetitions and remaining threats. |
5. Closed versus arrival-oriented interpretation
In the stable JMeter Thread Group, a thread waits for its current sample to finish before proceeding, so the realized arrival rate depends on response time. That is often correct for “N concurrent users.” It is not the same question as “the system must receive 10 arrivals/s even during a 700 ms stall.” The latter is arrival-oriented and must schedule work independently of previous completion.
JMeter 5.6.3 includes an
experimental Open Model Thread Group that can
express schedules such as
rate(10/sec) random_arrivals(6 sec). Because the
component is explicitly experimental, this chapter's mandatory
coordinated-omission demonstration uses a small Python
standard-library scheduler as a stable fallback; the JMeter Open
Model path is optional.
6. Timing boundaries and warm-up
JMeter elapsed/latency/connect fields begin at sampler execution—not during a Constant Timer. Warm-up can occur in multiple layers: Java JIT, DNS/TLS/connection pool, target cache, database pages, application code paths and generated data. Do not assume “30 seconds” is universally enough. Inspect time series and target state, choose a bounded warm-up rule, and keep it identical between A/B conditions.
7. Repetition and uncertainty
One run estimates one realization. Three repetitions are still a small sample, but they expose whether a 15 ms difference is large relative to run variation. This chapter reports median, sample standard deviation, coefficient of variation and a simple t-based interval for the mean. The interval is an uncertainty illustration, not a magical significance certificate.
8. Read-only/non-destructive inspection first
& "$env:JMETER_HOME\bin\jmeter.bat" -v
java -version
Get-FileHash `
.\plans\warmup.jmx, .\plans\measure.jmx, `
.\config\validity.properties, .\data\* `
-Algorithm SHA256
Get-Content .\experiment\protocol.json
Get-Content .\results\*\summary.json -ErrorAction SilentlyContinue
Also inspect generator process/JVM state and target
/health. Do not change warm-up, data, runtime or
thresholds until the actual executed state is proven.
jmeter.log before any filtering or comparison. The lab
uses no real credentials.
9. DevOps connection
Performance engineering becomes credible when a team can distinguish a product regression from cold startup, workload-model bias, a noisy CI runner or random variation. The same evidence discipline improves release gates, incident analysis, capacity experiments and optimization reviews.
Knowledge check
Why does measurement validity come before threshold enforcement?
A threshold applied to an invalid or incomparable experiment can block/pass changes for the wrong reason.
What is coordinated omission in a closed workload?
Future requests are not generated while a previous request is slow, so some bad experiences that fixed arrivals would have seen are omitted.
Does warm-up mean discarding whatever early samples look slow?
No. Warm-up must be defined before comparison using a reproducible rule tied to expected cold-state behavior.
Why repeat runs?
To quantify environmental/run variation and judge whether an A/B difference is large relative to that variation.
Why is Open Model optional here?
JMeter 5.6.3 documents it as experimental; the mandatory path uses stable Thread Group plus a bounded portable arrival scheduler.
Official references and version notes
- Apache JMeter downloads — current JMeter 5.6.3 and Java 8+ requirement.
- JMeter current changes — Java 17+ recommendation for the 5.6.x line.
- Open Model Thread Group — currently experimental and schedule-driven.
- JMeter Dashboard Report — CSV requirements, report windows/percentiles and estimator caveats.
- JMeter Glossary — elapsed, latency and connect-time definitions.
- JMeter Listeners / result fields — timestamps, elapsed, success, active-thread counts and other JTL fields.
Version-sensitive statements were rechecked against current
primary documentation on 2026-09-05. The mandatory runtime remains
Apache JMeter 5.6.3 with Java 17 and no
third-party plugin. JMeter 5.6.3 requires Java 8+; the current
5.6.x changes page recommends Java 17 or later. The
Open Model Thread Group remains explicitly experimental; the mandatory course path therefore uses the stable standard
Thread Group for closed-model measurements and a bounded
Python-standard-library fixed-arrival probe to demonstrate
coordinated omission. An optional Open Model schedule such as
rate(10/sec) random_arrivals(6 sec) is shown only as
a current 5.6.3 comparison. All important comparisons use raw CSV
JTL plus a documented label/window filter; dashboards remain
corroborating evidence. JMeter elapsed time begins immediately
before sending the request and ends after the last response byte;
latency ends at the first response, and connect time covers
connection establishment including SSL handshake.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.