Chapter 30Lesson 04~240 minutes

Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design: Diagnostics, Failure Modes, and Production Practices

When an experiment is invalid, the correct repair is usually to preserve the evidence, identify the confounder, restore a comparable protocol, and rerun the smallest bounded test. Deleting “bad” samples or tuning until green destroys the information needed to understand the failure.

Cold-start contaminationDataset driftOne-run claimsClock skewOutlier discipline

Learning objectives

  • Apply a preserve-first diagnostic sequence to failure modes involving Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design.
  • Separate plan/configuration, generator, protocol/network, target, CI/container, and distributed-engine causes before changing settings.
  • Reproduce a failure with the smallest authorized workload and retain the original JTL, jmeter.log, and supporting evidence.
  • Reject shortcuts such as blanket retries, disabled verification, unbounded load increases, silent global overrides, or deleting first-failure artifacts.
  • Verify the least-invasive correction under the original bounded workload before declaring the problem resolved.

1. Preserve-first diagnostic sequence

Invalid-experiment diagnosis

Validity is an experiment property, not a dashboard checkbox. Every arrow below changes what a measured difference can legitimately mean.

flowchart TD
E[Preserve JTL + jmeter.log + engine + target evidence] --> V[Confirm JMeter/Java/plugin/tool versions]
V --> C[Confirm JMX/data/properties/CLI + authorized target]
C --> S[Validate scope/labels/resolved variables + warm-up/window/model]
S --> P[Inspect protocol/session/data/cache state]
P --> G[Inspect generator JVM/OS/network]
G --> T[Inspect SUT build/cache/DB/noise/telemetry]
T --> D[Inspect distributed/CI/container/time-sync state]
D --> H[Identify validity threat / competing cause]
H --> F[Least-destructive protocol correction]
F --> R[Small comparable rerun]
All runnable diagnosis stays local and bounded. Preserve first-failure JTL, matching jmeter.log, target events, generator snapshots, build/data hashes, warm-up/window definition and run order. Do not delete the run that exposed the problem.

2. Failure mode: startup/JIT warm-up unknowingly included

Broken protocol starts a fresh service and immediately compares its first 100 samples with a long-lived baseline. Target events show the first 20 requests paid an intentional cold penalty.

Repair: keep the contaminated JTL, restart both conditions, run the same 20-request warm-up, then measure. Do not simply remove the first 20 rows after seeing them unless that exclusion was the predeclared protocol.

3. Failure mode: datasets/build configurations differ

A uses cached 1 KiB records; B uses uncached 5 MiB records or a debug build. The performance difference may be real, but it is not attributable to the code change under test.

Repair: hash/manifest build artifacts, data/config, feature flags and dependency versions. Re-run with equivalent datasets/config or explicitly change the question to test those differences.

4. Failure mode: declaring significance from one run

One A/B pair cannot show whether a 10 ms difference is larger than normal run variation. Repeat under the same protocol; report variation/effect magnitude. Avoid saying “statistically significant” unless the assumptions/test/design actually support that claim.

5. Failure mode: coordinated omission hides stalls

A one-thread closed test aims for “10 requests/s,” but when one request stalls for 750 ms the next requests do not exist. The JTL might contain only one slow sample during a wall-clock incident. A fixed-arrival scheduler still launches several arrivals that experience the bad window.

Repair the workload model for the performance question. Do not post-process nonexistent requests into the JTL and call them measured samples. Document whether you used a closed user population or an independent arrival schedule.

6. Failure mode: clock skew across telemetry

JMeter and target metrics appear offset by five seconds, so GC/queue spikes seem unrelated. In distributed/CI/container tests record timezone/UTC, NTP/clock offsets or correlate using run IDs/sequence markers. Do not manually slide graphs until they “look correlated” without recording the offset evidence.

7. Failure mode: tuning until the test passes

Changing warm-up, timeout, heap, threshold, thread count and dataset after every failure until the latest run is green creates a different experiment. Keep a protocol/version history. A product fix changes the product; a measurement fix changes the protocol and normally requires re-baselining both A and B.

8. Intentionally broken example: discard unexplained outliers

# anti-pattern
samples = [x for x in samples if x < 500]  # "500ms looks abnormal"
report_p95(samples)

If 500+ ms samples coincide with the fixture's logged periodic stall, they are exactly the phenomenon under study. Preserve them. Exclude a point only for a documented invalidity cause defined consistently (for example, target process crashed and the run is invalid), not because it harms the metric.

9. Background-noise diagnosis

Compare generator CPU/GC/disk/network and SUT process/background state across repetitions. If one run contains an OS update/backup/CI-neighbor event, mark it as a validity threat. The correct action may be to rerun the entire predefined repetition—not silently delete a few slow samples.

10. Causal table

Symptom Validity threat Repair
Early samples much slower only on fresh B process Warm-up confounder Apply identical predeclared warm-up to both.
A/B dataset hashes differ Data confounder Restore equivalent data or change the experiment question.
A=80/B=90 once, repeats overlap widely Run variation More controlled repetitions; report uncertainty.
Closed arrival rate falls during stall Coordinated omission Use/compare arrival-oriented model for fixed-arrival question.
JMeter/target spikes offset in time Clock skew Measure/synchronize offsets; use run/sequence correlation.
Threshold/workload changed after failure Tuning-until-pass Freeze protocol; rerun both conditions.
Only slow rows deleted Outlier bias Restore raw evidence; justify exclusions by validity cause.
Diagnostic evidence baseline: all runnable diagnosis stays on 127.0.0.1:8030 with Apache JMeter 5.6.3/Java 17. Keep the original JTL and matching jmeter.log together with target/generator evidence before any protocol correction. No real credentials are used.

11. Safety-sensitive boundaries

Sustained traffic, credentials, DB/message/API mutation, recorder certs, RMI engines, CI secrets, containers, OS/JVM tuning and failure injection require authorization and rollback. This chapter uses fake/local data and deterministic delay only. Never experiment on uncontrolled public/production systems.

12. Troubleshooting shortcuts to reject

  • Do not add blanket retries or arbitrary long sleeps.
  • Do not assign giant heaps without generator evidence.
  • Do not mass-disable evidence/listeners without measured cost.
  • Do not use global property hacks.
  • Do not disable TLS/RMI verification.
  • Do not increase workload unboundedly.
  • Do not delete result files or unexplained outliers.
  • Do not weaken thresholds/warm-up rules until a run passes.

Knowledge check

Why preserve a cold-contaminated JTL?

Can you call one A/B pair statistically significant?

How should clock skew be repaired?

Why are unexplained slow samples not disposable?

What indicates coordinated omission rather than merely a slow server?

Next lesson

Checkpoint: execute a controlled A/B experiment

Lesson 5 freezes the protocol, warms fresh A/B targets, repeats each measurement three times, quantifies variation and documents coordinated-omission and remaining threats.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current primary documentation on 2026-09-05. The mandatory runtime remains Apache JMeter 5.6.3 with Java 17 and no third-party plugin. JMeter 5.6.3 requires Java 8+; the current 5.6.x changes page recommends Java 17 or later. The Open Model Thread Group remains explicitly experimental; the mandatory course path therefore uses the stable standard Thread Group for closed-model measurements and a bounded Python-standard-library fixed-arrival probe to demonstrate coordinated omission. An optional Open Model schedule such as rate(10/sec) random_arrivals(6 sec) is shown only as a current 5.6.3 comparison. All important comparisons use raw CSV JTL plus a documented label/window filter; dashboards remain corroborating evidence. JMeter elapsed time begins immediately before sending the request and ends after the last response byte; latency ends at the first response, and connect time covers connection establishment including SSL handshake.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.