Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design: Diagnostics, Failure Modes, and Production Practices
When an experiment is invalid, the correct repair is usually to preserve the evidence, identify the confounder, restore a comparable protocol, and rerun the smallest bounded test. Deleting “bad” samples or tuning until green destroys the information needed to understand the failure.
Learning objectives
- Apply a preserve-first diagnostic sequence to failure modes involving Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design.
- Separate plan/configuration, generator, protocol/network, target, CI/container, and distributed-engine causes before changing settings.
- Reproduce a failure with the smallest authorized workload and retain the original JTL, jmeter.log, and supporting evidence.
- Reject shortcuts such as blanket retries, disabled verification, unbounded load increases, silent global overrides, or deleting first-failure artifacts.
- Verify the least-invasive correction under the original bounded workload before declaring the problem resolved.
1. Preserve-first diagnostic sequence
Validity is an experiment property, not a dashboard checkbox. Every arrow below changes what a measured difference can legitimately mean.
flowchart TD E[Preserve JTL + jmeter.log + engine + target evidence] --> V[Confirm JMeter/Java/plugin/tool versions] V --> C[Confirm JMX/data/properties/CLI + authorized target] C --> S[Validate scope/labels/resolved variables + warm-up/window/model] S --> P[Inspect protocol/session/data/cache state] P --> G[Inspect generator JVM/OS/network] G --> T[Inspect SUT build/cache/DB/noise/telemetry] T --> D[Inspect distributed/CI/container/time-sync state] D --> H[Identify validity threat / competing cause] H --> F[Least-destructive protocol correction] F --> R[Small comparable rerun]
jmeter.log, target
events, generator snapshots, build/data hashes, warm-up/window
definition and run order. Do not delete the run that exposed the
problem.
2. Failure mode: startup/JIT warm-up unknowingly included
Broken protocol starts a fresh service and immediately compares its first 100 samples with a long-lived baseline. Target events show the first 20 requests paid an intentional cold penalty.
Repair: keep the contaminated JTL, restart both conditions, run the same 20-request warm-up, then measure. Do not simply remove the first 20 rows after seeing them unless that exclusion was the predeclared protocol.
3. Failure mode: datasets/build configurations differ
A uses cached 1 KiB records; B uses uncached 5 MiB records or a debug build. The performance difference may be real, but it is not attributable to the code change under test.
Repair: hash/manifest build artifacts, data/config, feature flags and dependency versions. Re-run with equivalent datasets/config or explicitly change the question to test those differences.
4. Failure mode: declaring significance from one run
One A/B pair cannot show whether a 10 ms difference is larger than normal run variation. Repeat under the same protocol; report variation/effect magnitude. Avoid saying “statistically significant” unless the assumptions/test/design actually support that claim.
5. Failure mode: coordinated omission hides stalls
A one-thread closed test aims for “10 requests/s,” but when one request stalls for 750 ms the next requests do not exist. The JTL might contain only one slow sample during a wall-clock incident. A fixed-arrival scheduler still launches several arrivals that experience the bad window.
Repair the workload model for the performance question. Do not post-process nonexistent requests into the JTL and call them measured samples. Document whether you used a closed user population or an independent arrival schedule.
6. Failure mode: clock skew across telemetry
JMeter and target metrics appear offset by five seconds, so GC/queue spikes seem unrelated. In distributed/CI/container tests record timezone/UTC, NTP/clock offsets or correlate using run IDs/sequence markers. Do not manually slide graphs until they “look correlated” without recording the offset evidence.
7. Failure mode: tuning until the test passes
Changing warm-up, timeout, heap, threshold, thread count and dataset after every failure until the latest run is green creates a different experiment. Keep a protocol/version history. A product fix changes the product; a measurement fix changes the protocol and normally requires re-baselining both A and B.
8. Intentionally broken example: discard unexplained outliers
# anti-pattern
samples = [x for x in samples if x < 500] # "500ms looks abnormal"
report_p95(samples)
If 500+ ms samples coincide with the fixture's logged periodic stall, they are exactly the phenomenon under study. Preserve them. Exclude a point only for a documented invalidity cause defined consistently (for example, target process crashed and the run is invalid), not because it harms the metric.
9. Background-noise diagnosis
Compare generator CPU/GC/disk/network and SUT process/background state across repetitions. If one run contains an OS update/backup/CI-neighbor event, mark it as a validity threat. The correct action may be to rerun the entire predefined repetition—not silently delete a few slow samples.
10. Causal table
| Symptom | Validity threat | Repair |
|---|---|---|
| Early samples much slower only on fresh B process | Warm-up confounder | Apply identical predeclared warm-up to both. |
| A/B dataset hashes differ | Data confounder | Restore equivalent data or change the experiment question. |
| A=80/B=90 once, repeats overlap widely | Run variation | More controlled repetitions; report uncertainty. |
| Closed arrival rate falls during stall | Coordinated omission | Use/compare arrival-oriented model for fixed-arrival question. |
| JMeter/target spikes offset in time | Clock skew | Measure/synchronize offsets; use run/sequence correlation. |
| Threshold/workload changed after failure | Tuning-until-pass | Freeze protocol; rerun both conditions. |
| Only slow rows deleted | Outlier bias | Restore raw evidence; justify exclusions by validity cause. |
127.0.0.1:8030 with Apache JMeter 5.6.3/Java 17.
Keep the original JTL and matching jmeter.log together
with target/generator evidence before any protocol correction. No real
credentials are used.
11. Safety-sensitive boundaries
Sustained traffic, credentials, DB/message/API mutation, recorder certs, RMI engines, CI secrets, containers, OS/JVM tuning and failure injection require authorization and rollback. This chapter uses fake/local data and deterministic delay only. Never experiment on uncontrolled public/production systems.
12. Troubleshooting shortcuts to reject
- Do not add blanket retries or arbitrary long sleeps.
- Do not assign giant heaps without generator evidence.
- Do not mass-disable evidence/listeners without measured cost.
- Do not use global property hacks.
- Do not disable TLS/RMI verification.
- Do not increase workload unboundedly.
- Do not delete result files or unexplained outliers.
- Do not weaken thresholds/warm-up rules until a run passes.
Knowledge check
Why preserve a cold-contaminated JTL?
It proves the validity problem and lets you compare the corrected protocol without hiding the original cause.
Can you call one A/B pair statistically significant?
Not responsibly without an appropriate statistical design/assumptions; first quantify repetition variation.
How should clock skew be repaired?
Measure/synchronize/document offsets or correlate via run/sequence markers—never visually shift graphs without evidence.
Why are unexplained slow samples not disposable?
They may be the real tail behavior; deleting them biases the distribution.
What indicates coordinated omission rather than merely a slow server?
Configured fixed-demand intent but closed generator arrivals fall precisely while responses stall; arrival-oriented comparison reveals omitted demand.
Official references and version notes
- Apache JMeter downloads — current JMeter 5.6.3 and Java 8+ requirement.
- JMeter current changes — Java 17+ recommendation for the 5.6.x line.
- Open Model Thread Group — currently experimental and schedule-driven.
- JMeter Dashboard Report — CSV requirements, report windows/percentiles and estimator caveats.
- JMeter Glossary — elapsed, latency and connect-time definitions.
- JMeter Listeners / result fields — timestamps, elapsed, success, active-thread counts and other JTL fields.
Version-sensitive statements were rechecked against current
primary documentation on 2026-09-05. The mandatory runtime remains
Apache JMeter 5.6.3 with Java 17 and no
third-party plugin. JMeter 5.6.3 requires Java 8+; the current
5.6.x changes page recommends Java 17 or later. The
Open Model Thread Group remains explicitly experimental; the mandatory course path therefore uses the stable standard
Thread Group for closed-model measurements and a bounded
Python-standard-library fixed-arrival probe to demonstrate
coordinated omission. An optional Open Model schedule such as
rate(10/sec) random_arrivals(6 sec) is shown only as
a current 5.6.3 comparison. All important comparisons use raw CSV
JTL plus a documented label/window filter; dashboards remain
corroborating evidence. JMeter elapsed time begins immediately
before sending the request and ends after the last response byte;
latency ends at the first response, and connect time covers
connection establishment including SSL handshake.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.