Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design: Guided Hands-On Workflow
The fixture has four controlled effects: a base service time, a cold-start penalty on the first 20 requests, deterministic per-run jitter/background noise, and a wall-clock periodic stall used only for coordinated-omission demonstration. Each effect is explicit in target events so learners can distinguish it from JMeter timing.
Learning objectives
- Execute a bounded hands-on workflow for Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design using the course's current runtime and authorized local or synthetic resources.
- Build and verify the concrete lab artifacts step by step instead of treating configuration snippets as isolated examples.
- Preserve the JTL, jmeter.log, target, generator, and configuration evidence required by the workflow before interpreting results.
- Distinguish configured state from achieved behavior, and stop when safety, count, environment, or generator-validity conditions are not met.
- Explain how the completed workflow prepares the configuration and trade-off analysis in the next lesson.
1. Safety / experiment ceilings
127.0.0.1:8030.
Measurement run =4×25=100 samples; warm-up=20; three repeats per
condition in the checkpoint. Coordinated-omission demonstration =60
scheduled requests maximum. Abort on count mismatch, non-loopback
traffic, fixture crash, generator saturation, missing
JTL/events/log, or any unplanned change to data/runtime/workload
between conditions.
2. Create the deterministic validity fixture
fixtures/validity_fixture.py:
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from urllib.parse import urlparse, parse_qs
import argparse, hashlib, json, re, threading, time
FIXTURE_VERSION = "prompt30-validity-fixture-v1"
SAFE = re.compile(r"^[A-Za-z0-9_.-]{1,64}$")
lock = threading.Lock()
event_log = None
base_ms = 55
warmup_count = 20
warmup_extra_ms = 90
jitter_ms = 10
noise_every = 17
noise_ms = 25
stall_start_after_ms = 1000
stall_period_ms = 2500
stall_width_ms = 600
stall_extra_ms = 700
started_mono = time.monotonic()
state = {"work_requests": 0, "co_requests": 0}
def stable_jitter(run_id, seq):
if jitter_ms <= 0:
return 0
raw = hashlib.sha256(f"{run_id}:{seq}".encode()).digest()
return int.from_bytes(raw[:4], "big") % (jitter_ms + 1)
def write_event(event):
if event_log is None:
return
with lock:
with event_log.open("a", encoding="utf-8") as h:
h.write(json.dumps(event, sort_keys=True) + "\n")
def snapshot():
with lock:
return {
"fixture_version": FIXTURE_VERSION,
"base_ms": base_ms,
"warmup_count": warmup_count,
"warmup_extra_ms": warmup_extra_ms,
"jitter_ms": jitter_ms,
"noise_every": noise_every,
"noise_ms": noise_ms,
"stall_start_after_ms": stall_start_after_ms,
"stall_period_ms": stall_period_ms,
"stall_width_ms": stall_width_ms,
"stall_extra_ms": stall_extra_ms,
**state,
}
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def send_json(self, status, payload):
raw = json.dumps(payload, sort_keys=True).encode()
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(raw)))
self.send_header("X-Fixture-Version", FIXTURE_VERSION)
self.end_headers()
self.wfile.write(raw)
def do_GET(self):
req_start_ms = int(time.time() * 1000)
parsed = urlparse(self.path)
if parsed.path == "/health":
self.send_json(200, {"status": "ok", "state": snapshot()})
return
if parsed.path not in ("/work", "/co"):
self.send_json(404, {"status": "not_found"})
return
q = parse_qs(parsed.query)
run_id = q.get("run_id", [""])[0]
seq_raw = q.get("seq", [""])[0]
if not SAFE.fullmatch(run_id):
self.send_json(400, {"status": "invalid_run_id"})
return
try:
seq = int(seq_raw)
except ValueError:
self.send_json(400, {"status": "invalid_seq"})
return
if parsed.path == "/work":
with lock:
state["work_requests"] += 1
ordinal = state["work_requests"]
cold = ordinal <= warmup_count
jitter = stable_jitter(run_id, seq)
noisy = noise_every > 0 and ordinal % noise_every == 0
delay = base_ms + jitter + (warmup_extra_ms if cold else 0) + (noise_ms if noisy else 0)
time.sleep(delay / 1000.0)
ended = int(time.time() * 1000)
self.send_json(200, {
"status": "ok", "run_id": run_id, "seq": seq,
"base_ms": base_ms, "cold": cold, "jitter_ms": jitter,
"noise_active": noisy, "configured_delay_ms": delay,
})
write_event({
"ts_ms": ended, "operation": "work", "status": 200,
"run_id": run_id, "seq": seq, "ordinal": ordinal,
"base_ms": base_ms, "cold": cold, "jitter_ms": jitter,
"noise_active": noisy, "configured_delay_ms": delay,
"service_wall_ms": ended - req_start_ms,
})
return
# /co: wall-clock periodic stall for coordinated-omission demonstration.
with lock:
state["co_requests"] += 1
ordinal = state["co_requests"]
since_start_ms = int((time.monotonic() - started_mono) * 1000)
stall_active = False
if since_start_ms >= stall_start_after_ms and stall_period_ms > 0:
phase = (since_start_ms - stall_start_after_ms) % stall_period_ms
stall_active = phase < stall_width_ms
jitter = stable_jitter(run_id, seq)
delay = base_ms + jitter + (stall_extra_ms if stall_active else 0)
time.sleep(delay / 1000.0)
ended = int(time.time() * 1000)
self.send_json(200, {
"status": "ok", "run_id": run_id, "seq": seq,
"stall_active": stall_active,
"configured_delay_ms": delay,
})
write_event({
"ts_ms": ended, "operation": "co", "status": 200,
"run_id": run_id, "seq": seq, "ordinal": ordinal,
"stall_active": stall_active,
"configured_delay_ms": delay,
"service_wall_ms": ended - req_start_ms,
})
def log_message(self, format, *args):
return
def main():
p = argparse.ArgumentParser()
p.add_argument("--host", default="127.0.0.1")
p.add_argument("--port", type=int, default=8030)
p.add_argument("--base-ms", type=int, default=55)
p.add_argument("--warmup-count", type=int, default=20)
p.add_argument("--warmup-extra-ms", type=int, default=90)
p.add_argument("--jitter-ms", type=int, default=10)
p.add_argument("--noise-every", type=int, default=17)
p.add_argument("--noise-ms", type=int, default=25)
p.add_argument("--stall-start-after-ms", type=int, default=1000)
p.add_argument("--stall-period-ms", type=int, default=2500)
p.add_argument("--stall-width-ms", type=int, default=600)
p.add_argument("--stall-extra-ms", type=int, default=700)
p.add_argument("--log", required=True)
args = p.parse_args()
global event_log, base_ms, warmup_count, warmup_extra_ms, jitter_ms
global noise_every, noise_ms, stall_start_after_ms, stall_period_ms
global stall_width_ms, stall_extra_ms, started_mono
base_ms = args.base_ms
warmup_count = args.warmup_count
warmup_extra_ms = args.warmup_extra_ms
jitter_ms = args.jitter_ms
noise_every = args.noise_every
noise_ms = args.noise_ms
stall_start_after_ms = args.stall_start_after_ms
stall_period_ms = args.stall_period_ms
stall_width_ms = args.stall_width_ms
stall_extra_ms = args.stall_extra_ms
started_mono = time.monotonic()
event_log = Path(args.log).resolve()
event_log.parent.mkdir(parents=True, exist_ok=True)
event_log.write_text("", encoding="utf-8")
print(f"fixture_version={FIXTURE_VERSION}", flush=True)
print(f"listen=http://{args.host}:{args.port}", flush=True)
print(json.dumps(snapshot(), sort_keys=True), flush=True)
ThreadingHTTPServer((args.host, args.port), Handler).serve_forever()
if __name__ == "__main__":
main()
/work models warm-up/jitter/noise.
/co models a wall-clock stall: closed clients stop
generating future work while blocked; the arrival probe continues
scheduling at its fixed rate.
3. Result/runtime properties
config/validity.properties:
target.host=127.0.0.1
target.port=8030
threads=4
loops=25
pacing.ms=50
connect.timeout.ms=500
response.timeout.ms=2500
jmeter.httpsampler=HttpClient4
httpclient4.retrycount=0
jmeter.save.saveservice.output_format=csv
jmeter.save.saveservice.print_field_names=true
jmeter.save.saveservice.timestamp_format=ms
jmeter.save.saveservice.time=true
jmeter.save.saveservice.label=true
jmeter.save.saveservice.response_code=true
jmeter.save.saveservice.response_message=true
jmeter.save.saveservice.thread_name=true
jmeter.save.saveservice.successful=true
jmeter.save.saveservice.bytes=true
jmeter.save.saveservice.sent_bytes=true
jmeter.save.saveservice.thread_counts=true
jmeter.save.saveservice.latency=true
jmeter.save.saveservice.connect_time=true
jmeter.save.saveservice.assertion_results_failure_message=true
jmeter.save.saveservice.response_data=false
jmeter.save.saveservice.response_data.on_error=false
jmeter.save.saveservice.samplerData=false
jmeter.save.saveservice.responseHeaders=false
jmeter.save.saveservice.requestHeaders=false
jmeter.save.saveservice.url=false
jmeter.reportgenerator.overall_granularity=2000
jmeter.reportgenerator.aggregate_rpt_pct1=90
jmeter.reportgenerator.aggregate_rpt_pct2=95
jmeter.reportgenerator.aggregate_rpt_pct3=99
These fields make window filtering and generator-validity checks auditable. JTL stores no response body/header.
4. Author two stable JMX plans
Warm-up plan:
Test Plan
├── HTTP Request Defaults -> 127.0.0.1:8030, HttpClient4
└── Thread Group — Warmup
1 thread × 20 loops
├── Counter -> SEQ
└── HTTP Request — Warmup
GET /work?run_id=${__P(run.id,warmup)}&seq=${SEQ}
Use KeepAlive=checked
├── Constant Timer 20 ms
└── Response Assertion: HTTP code = 200
Measurement plan:
Test Plan
├── HTTP Request Defaults
│ host=${__P(target.host,127.0.0.1)}
│ port=${__P(target.port,8030)}
│ implementation=HttpClient4
└── Thread Group — Measure
${__P(threads,4)} threads × ${__P(loops,25)} loops
├── Counter -> SEQ (per user)
└── HTTP Request — Measure
GET /work?run_id=${__P(run.id,p30)}&seq=${SEQ}
Use KeepAlive=checked
├── Constant Timer ${__P(pacing.ms,50)} ms
└── Response Assertion: HTTP code = 200
Warm-up uses a separate JTL/label and is never merged into the measurement distribution. Measurement uses identical four-thread/25-loop behavior in A and B.
5. First experiment: measure cold startup accidentally
Start variant A fresh:
python .\fixtures\validity_fixture.py `
--host 127.0.0.1 --port 8030 `
--base-ms 55 --warmup-count 20 --warmup-extra-ms 90 `
--jitter-ms 10 --noise-every 17 --noise-ms 25 `
--log .\results\no-warmup-target.jsonl
Immediately run measure.jmx without a warm-up. The
first ~20 target requests carry an extra 90 ms. Preserve the
JTL/dashboard/log; this is intentionally contaminated evidence, not
something to delete.
6. Repeat with an explicit warm-up
Restart the same variant A configuration with a fresh event log. Run
warmup.jmx once (20 requests) and then
measure.jmx. Now the cold penalty is consumed outside
the comparison window. Predict: full measurement p95 should fall
while base/jitter/noise configuration remains identical.
7. Mark/compute explicit measurement windows
tools/analyze_window.py:
import argparse, csv, json, math
from pathlib import Path
def nearest_rank(values, pct):
data = sorted(values)
if not data:
return 0
return data[max(1, math.ceil(len(data) * pct / 100.0)) - 1]
def main():
p = argparse.ArgumentParser()
p.add_argument("--jtl", required=True)
p.add_argument("--label", default="Measure")
p.add_argument("--out", required=True)
p.add_argument("--window-start-s", type=float, default=0.0)
p.add_argument("--window-end-s", type=float)
args = p.parse_args()
rows = list(csv.DictReader(Path(args.jtl).open(newline="", encoding="utf-8")))
rows = [r for r in rows if r.get("label") == args.label]
if not rows:
raise SystemExit("no matching rows")
first_ts = min(int(r["timeStamp"]) for r in rows)
selected = []
for r in rows:
offset_s = (int(r["timeStamp"]) - first_ts) / 1000.0
if offset_s < args.window_start_s:
continue
if args.window_end_s is not None and offset_s >= args.window_end_s:
continue
selected.append(r)
elapsed = [int(float(r["elapsed"])) for r in selected]
failures = sum(r.get("success", "").lower() != "true" for r in selected)
if selected:
start = min(int(r["timeStamp"]) for r in selected)
end = max(int(r["timeStamp"]) + int(float(r["elapsed"])) for r in selected)
span_s = max((end - start) / 1000.0, 0.001)
else:
span_s = 0.0
result = {
"label": args.label,
"source_rows": len(rows),
"selected_rows": len(selected),
"excluded_rows": len(rows) - len(selected),
"window_start_s": args.window_start_s,
"window_end_s": args.window_end_s,
"p50_ms": nearest_rank(elapsed, 50),
"p90_ms": nearest_rank(elapsed, 90),
"p95_ms": nearest_rank(elapsed, 95),
"p99_ms": nearest_rank(elapsed, 99),
"mean_ms": round(sum(elapsed) / len(elapsed), 3) if elapsed else 0,
"error_rate_pct": round(100.0 * failures / len(selected), 3) if selected else 100.0,
"throughput_rps": round(len(selected) / span_s, 3) if span_s else 0.0,
}
Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps(result, indent=2))
if __name__ == "__main__":
main()
Example full measurement:
python .\tools\analyze_window.py `
--jtl .\results\a1\results.jtl `
--label Measure `
--window-start-s 0 `
--out .\results\a1\summary.json
For a fixed-duration study you can predeclare, for example,
window-start-s=2/window-end-s=10. Never
decide the cut after seeing which interval makes A look faster.
8. Introduce synthetic background noise deliberately
Restart variant A with a stronger but still bounded noise pattern
such as --noise-every 7 --noise-ms 60, then repeat the
same measurement three times. Because noise is known and logged, you
can see p95/CV widen without inventing a product regression. Restore
the normal 17/25 noise settings before the checkpoint.
9. Quantify repetition variation
tools/summarize_repetitions.py:
import argparse, json, math, statistics
from pathlib import Path
T975 = {
1: 12.706, 2: 4.303, 3: 3.182, 4: 2.776, 5: 2.571,
6: 2.447, 7: 2.365, 8: 2.306, 9: 2.262, 10: 2.228
}
def summarize(values):
n = len(values)
mean = statistics.mean(values)
median = statistics.median(values)
if n >= 2:
sd = statistics.stdev(values)
cv = 100.0 * sd / mean if mean else 0.0
t = T975.get(n - 1, 1.96)
margin = t * sd / math.sqrt(n)
ci = [mean - margin, mean + margin]
else:
sd = cv = 0.0
ci = [mean, mean]
return {
"n": n, "mean": round(mean, 3), "median": round(median, 3),
"sample_sd": round(sd, 3), "cv_pct": round(cv, 3),
"min": min(values), "max": max(values),
"approx_95pct_t_ci_of_mean": [round(ci[0], 3), round(ci[1], 3)]
}
def main():
p = argparse.ArgumentParser()
p.add_argument("summaries", nargs="+")
p.add_argument("--metric", default="p95_ms")
p.add_argument("--out", required=True)
args = p.parse_args()
docs = [json.loads(Path(x).read_text(encoding="utf-8")) for x in args.summaries]
values = [float(d[args.metric]) for d in docs]
result = {
"metric": args.metric,
"runs": [{"path": p, "value": v} for p, v in zip(args.summaries, values)],
"variation": summarize(values),
"warning": "With only three repetitions the t interval is intentionally wide; use it as an uncertainty illustration, not proof of a population parameter."
}
Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps(result, indent=2))
if __name__ == "__main__":
main()
python .\tools\summarize_repetitions.py `
.\results\a1\summary.json `
.\results\a2\summary.json `
.\results\a3\summary.json `
--metric p95_ms `
--out .\results\a-p95-variation.json
Report median, mean, SD, CV, min/max and the deliberately cautious t-based interval. With n=3 the interval can be wide; that itself teaches why a single run is weak evidence.
10. Coordinated omission with a deterministic wall-clock stall
For this small demonstration restart the fixture with
--warmup-count 0, base≈50 ms and the default periodic
stall. Run the closed JMeter plan:
Test Plan
├── HTTP Request Defaults -> 127.0.0.1:8030, HttpClient4
└── Thread Group — Closed CO demonstration
1 thread × 60 loops
├── Counter -> SEQ
└── HTTP Request — CO-Closed
GET /co?run_id=co-closed&seq=${SEQ}
Use KeepAlive=checked
├── Constant Timer 50 ms
└── Response Assertion: HTTP code = 200
Interpretation:
- without stalls, ~50 ms service + 50 ms timer gives roughly 10 cycles/s;
- during a long response the single closed user cannot issue the intended next requests;
- the missing arrivals are coordinated with the stall.
Then run the fixed-arrival fallback:
import argparse, concurrent.futures, json, time, urllib.request
from pathlib import Path
def request(url, intended_s):
start = time.monotonic()
try:
with urllib.request.urlopen(url, timeout=3) as r:
status = r.status
r.read()
except Exception:
status = 0
end = time.monotonic()
return {
"intended_s": round(intended_s, 4),
"actual_start_s": round(start, 4),
"elapsed_ms": round((end - start) * 1000, 3),
"status": status
}
def main():
p = argparse.ArgumentParser()
p.add_argument("--url-base", default="http://127.0.0.1:8030/co")
p.add_argument("--rate", type=float, default=10.0)
p.add_argument("--duration", type=float, default=6.0)
p.add_argument("--run-id", default="co-arrival")
p.add_argument("--out", required=True)
args = p.parse_args()
count = int(args.rate * args.duration)
if count > 120:
raise SystemExit("safety ceiling: max 120 scheduled requests")
interval = 1.0 / args.rate
t0 = time.monotonic()
futures = []
with concurrent.futures.ThreadPoolExecutor(max_workers=32) as ex:
for i in range(count):
intended = t0 + i * interval
sleep_s = intended - time.monotonic()
if sleep_s > 0:
time.sleep(sleep_s)
url = f"{args.url_base}?run_id={args.run_id}&seq={i+1}"
futures.append(ex.submit(request, url, intended - t0))
results = [f.result() for f in futures]
result = {
"rate_per_s": args.rate,
"duration_s": args.duration,
"scheduled_requests": count,
"completed": len(results),
"slow_over_500ms": sum(r["elapsed_ms"] > 500 for r in results),
"p95_ms": sorted(r["elapsed_ms"] for r in results)[max(0, int(0.95 * len(results)) - 1)],
"results": results
}
Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps({k:v for k,v in result.items() if k != "results"}, indent=2))
if __name__ == "__main__":
main()
python .\tools\arrival_probe.py `
--url-base http://127.0.0.1:8030/co `
--rate 10 --duration 6 `
--run-id co-arrival `
--out .\results\co-arrival.json
Expected qualitative difference: when one closed thread hits a ~700 ms stall it cannot issue the next scheduled requests, so fewer samples arrive during the bad window. The arrival scheduler continues launching at 10/s, causing several requests to observe the stall. The closed result can therefore underrepresent the latency distribution for a fixed-arrival demand question.
11. Optional current JMeter Open Model comparison
In JMeter 5.6.3 you may author an Open Model Thread Group (experimental) with a bounded profile such as:
rate(10/sec) random_arrivals(6 sec)
Keep this optional and label the exact JMeter version. Validate generated/achieved sample count and generator headroom because an arrival schedule can require more concurrent threads during stalls.
jmeter.log,
dashboard, target JSONL, window summary and generator snapshot. No
real credentials are used.
12. Generator-health check
$Proc = Get-CimInstance Win32_Process |
Where-Object { $_.Name -eq "java.exe" -and $_.CommandLine -like "*ApacheJMeter.jar*" } |
Select-Object -First 1
if ($Proc) {
$Pid = $Proc.ProcessId
Get-Process -Id $Pid |
Select-Object Id,CPU,WorkingSet64,PrivateMemorySize64,Handles,Threads
jcmd $Pid GC.heap_info
jstat -gcutil $Pid 1000 5
}
If the injector saturates, arrival scheduling or latency comparisons can become invalid. Preserve CPU/heap/GC evidence rather than assuming local loopback means “free.”
13. Challenge
You must compare a checkout service that receives approximately 50 external arrivals/s regardless of response time. Would four closed JMeter users be a sufficient primary workload model?
Not necessarily. A closed user model reduces arrivals during slow responses. Prefer an arrival-oriented model/schedule with enough generator headroom, then explicitly compare achieved arrivals with the configured schedule.
Knowledge check
What changes between no-warmup and warmup experiments?
Only whether the documented 20-request cold phase is consumed before the measurement JTL; target/runtime/workload settings remain the same.
Why log synthetic noise?
It makes an intentional validity factor observable instead of an unexplained outlier.
Why keep separate warm-up and measurement JTLs?
It prevents accidental mixing of cold samples into the performance distribution.
What demonstrates coordinated omission here?
The fixed-arrival scheduler keeps sending during wall-clock stalls while the one-thread closed client cannot.
Why can Open Model need more generator capacity?
Independent arrivals continue during slow responses, increasing concurrent in-flight work instead of self-throttling.
Official references and version notes
- Apache JMeter downloads — current JMeter 5.6.3 and Java 8+ requirement.
- JMeter current changes — Java 17+ recommendation for the 5.6.x line.
- Open Model Thread Group — currently experimental and schedule-driven.
- JMeter Dashboard Report — CSV requirements, report windows/percentiles and estimator caveats.
- JMeter Glossary — elapsed, latency and connect-time definitions.
- JMeter Listeners / result fields — timestamps, elapsed, success, active-thread counts and other JTL fields.
Version-sensitive statements were rechecked against current
primary documentation on 2026-09-05. The mandatory runtime remains
Apache JMeter 5.6.3 with Java 17 and no
third-party plugin. JMeter 5.6.3 requires Java 8+; the current
5.6.x changes page recommends Java 17 or later. The
Open Model Thread Group remains explicitly experimental; the mandatory course path therefore uses the stable standard
Thread Group for closed-model measurements and a bounded
Python-standard-library fixed-arrival probe to demonstrate
coordinated omission. An optional Open Model schedule such as
rate(10/sec) random_arrivals(6 sec) is shown only as
a current 5.6.3 comparison. All important comparisons use raw CSV
JTL plus a documented label/window filter; dashboards remain
corroborating evidence. JMeter elapsed time begins immediately
before sending the request and ends after the last response byte;
latency ends at the first response, and connect time covers
connection establishment including SSL handshake.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.