Chapter 30Lesson 02~340 minutes

Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design: Guided Hands-On Workflow

The fixture has four controlled effects: a base service time, a cold-start penalty on the first 20 requests, deterministic per-run jitter/background noise, and a wall-clock periodic stall used only for coordinated-omission demonstration. Each effect is explicit in target events so learners can distinguish it from JMeter timing.

Warm-up protocolRepeated runsSynthetic noiseMeasurement windowCoordinated omission

Learning objectives

  • Execute a bounded hands-on workflow for Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design using the course's current runtime and authorized local or synthetic resources.
  • Build and verify the concrete lab artifacts step by step instead of treating configuration snippets as isolated examples.
  • Preserve the JTL, jmeter.log, target, generator, and configuration evidence required by the workflow before interpreting results.
  • Distinguish configured state from achieved behavior, and stop when safety, count, environment, or generator-validity conditions are not met.
  • Explain how the completed workflow prepares the configuration and trade-off analysis in the next lesson.

1. Safety / experiment ceilings

Target only 127.0.0.1:8030. Measurement run =4×25=100 samples; warm-up=20; three repeats per condition in the checkpoint. Coordinated-omission demonstration =60 scheduled requests maximum. Abort on count mismatch, non-loopback traffic, fixture crash, generator saturation, missing JTL/events/log, or any unplanned change to data/runtime/workload between conditions.

2. Create the deterministic validity fixture

fixtures/validity_fixture.py:

from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from urllib.parse import urlparse, parse_qs
import argparse, hashlib, json, re, threading, time

FIXTURE_VERSION = "prompt30-validity-fixture-v1"
SAFE = re.compile(r"^[A-Za-z0-9_.-]{1,64}$")

lock = threading.Lock()
event_log = None
base_ms = 55
warmup_count = 20
warmup_extra_ms = 90
jitter_ms = 10
noise_every = 17
noise_ms = 25
stall_start_after_ms = 1000
stall_period_ms = 2500
stall_width_ms = 600
stall_extra_ms = 700
started_mono = time.monotonic()
state = {"work_requests": 0, "co_requests": 0}

def stable_jitter(run_id, seq):
    if jitter_ms <= 0:
        return 0
    raw = hashlib.sha256(f"{run_id}:{seq}".encode()).digest()
    return int.from_bytes(raw[:4], "big") % (jitter_ms + 1)

def write_event(event):
    if event_log is None:
        return
    with lock:
        with event_log.open("a", encoding="utf-8") as h:
            h.write(json.dumps(event, sort_keys=True) + "\n")

def snapshot():
    with lock:
        return {
            "fixture_version": FIXTURE_VERSION,
            "base_ms": base_ms,
            "warmup_count": warmup_count,
            "warmup_extra_ms": warmup_extra_ms,
            "jitter_ms": jitter_ms,
            "noise_every": noise_every,
            "noise_ms": noise_ms,
            "stall_start_after_ms": stall_start_after_ms,
            "stall_period_ms": stall_period_ms,
            "stall_width_ms": stall_width_ms,
            "stall_extra_ms": stall_extra_ms,
            **state,
        }

class Handler(BaseHTTPRequestHandler):
    protocol_version = "HTTP/1.1"

    def send_json(self, status, payload):
        raw = json.dumps(payload, sort_keys=True).encode()
        self.send_response(status)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(raw)))
        self.send_header("X-Fixture-Version", FIXTURE_VERSION)
        self.end_headers()
        self.wfile.write(raw)

    def do_GET(self):
        req_start_ms = int(time.time() * 1000)
        parsed = urlparse(self.path)

        if parsed.path == "/health":
            self.send_json(200, {"status": "ok", "state": snapshot()})
            return

        if parsed.path not in ("/work", "/co"):
            self.send_json(404, {"status": "not_found"})
            return

        q = parse_qs(parsed.query)
        run_id = q.get("run_id", [""])[0]
        seq_raw = q.get("seq", [""])[0]
        if not SAFE.fullmatch(run_id):
            self.send_json(400, {"status": "invalid_run_id"})
            return
        try:
            seq = int(seq_raw)
        except ValueError:
            self.send_json(400, {"status": "invalid_seq"})
            return

        if parsed.path == "/work":
            with lock:
                state["work_requests"] += 1
                ordinal = state["work_requests"]

            cold = ordinal <= warmup_count
            jitter = stable_jitter(run_id, seq)
            noisy = noise_every > 0 and ordinal % noise_every == 0
            delay = base_ms + jitter + (warmup_extra_ms if cold else 0) + (noise_ms if noisy else 0)
            time.sleep(delay / 1000.0)

            ended = int(time.time() * 1000)
            self.send_json(200, {
                "status": "ok", "run_id": run_id, "seq": seq,
                "base_ms": base_ms, "cold": cold, "jitter_ms": jitter,
                "noise_active": noisy, "configured_delay_ms": delay,
            })
            write_event({
                "ts_ms": ended, "operation": "work", "status": 200,
                "run_id": run_id, "seq": seq, "ordinal": ordinal,
                "base_ms": base_ms, "cold": cold, "jitter_ms": jitter,
                "noise_active": noisy, "configured_delay_ms": delay,
                "service_wall_ms": ended - req_start_ms,
            })
            return

        # /co: wall-clock periodic stall for coordinated-omission demonstration.
        with lock:
            state["co_requests"] += 1
            ordinal = state["co_requests"]

        since_start_ms = int((time.monotonic() - started_mono) * 1000)
        stall_active = False
        if since_start_ms >= stall_start_after_ms and stall_period_ms > 0:
            phase = (since_start_ms - stall_start_after_ms) % stall_period_ms
            stall_active = phase < stall_width_ms

        jitter = stable_jitter(run_id, seq)
        delay = base_ms + jitter + (stall_extra_ms if stall_active else 0)
        time.sleep(delay / 1000.0)
        ended = int(time.time() * 1000)

        self.send_json(200, {
            "status": "ok", "run_id": run_id, "seq": seq,
            "stall_active": stall_active,
            "configured_delay_ms": delay,
        })
        write_event({
            "ts_ms": ended, "operation": "co", "status": 200,
            "run_id": run_id, "seq": seq, "ordinal": ordinal,
            "stall_active": stall_active,
            "configured_delay_ms": delay,
            "service_wall_ms": ended - req_start_ms,
        })

    def log_message(self, format, *args):
        return

def main():
    p = argparse.ArgumentParser()
    p.add_argument("--host", default="127.0.0.1")
    p.add_argument("--port", type=int, default=8030)
    p.add_argument("--base-ms", type=int, default=55)
    p.add_argument("--warmup-count", type=int, default=20)
    p.add_argument("--warmup-extra-ms", type=int, default=90)
    p.add_argument("--jitter-ms", type=int, default=10)
    p.add_argument("--noise-every", type=int, default=17)
    p.add_argument("--noise-ms", type=int, default=25)
    p.add_argument("--stall-start-after-ms", type=int, default=1000)
    p.add_argument("--stall-period-ms", type=int, default=2500)
    p.add_argument("--stall-width-ms", type=int, default=600)
    p.add_argument("--stall-extra-ms", type=int, default=700)
    p.add_argument("--log", required=True)
    args = p.parse_args()

    global event_log, base_ms, warmup_count, warmup_extra_ms, jitter_ms
    global noise_every, noise_ms, stall_start_after_ms, stall_period_ms
    global stall_width_ms, stall_extra_ms, started_mono

    base_ms = args.base_ms
    warmup_count = args.warmup_count
    warmup_extra_ms = args.warmup_extra_ms
    jitter_ms = args.jitter_ms
    noise_every = args.noise_every
    noise_ms = args.noise_ms
    stall_start_after_ms = args.stall_start_after_ms
    stall_period_ms = args.stall_period_ms
    stall_width_ms = args.stall_width_ms
    stall_extra_ms = args.stall_extra_ms
    started_mono = time.monotonic()

    event_log = Path(args.log).resolve()
    event_log.parent.mkdir(parents=True, exist_ok=True)
    event_log.write_text("", encoding="utf-8")

    print(f"fixture_version={FIXTURE_VERSION}", flush=True)
    print(f"listen=http://{args.host}:{args.port}", flush=True)
    print(json.dumps(snapshot(), sort_keys=True), flush=True)
    ThreadingHTTPServer((args.host, args.port), Handler).serve_forever()

if __name__ == "__main__":
    main()

/work models warm-up/jitter/noise. /co models a wall-clock stall: closed clients stop generating future work while blocked; the arrival probe continues scheduling at its fixed rate.

3. Result/runtime properties

config/validity.properties:

target.host=127.0.0.1
target.port=8030
threads=4
loops=25
pacing.ms=50
connect.timeout.ms=500
response.timeout.ms=2500

jmeter.httpsampler=HttpClient4
httpclient4.retrycount=0

jmeter.save.saveservice.output_format=csv
jmeter.save.saveservice.print_field_names=true
jmeter.save.saveservice.timestamp_format=ms
jmeter.save.saveservice.time=true
jmeter.save.saveservice.label=true
jmeter.save.saveservice.response_code=true
jmeter.save.saveservice.response_message=true
jmeter.save.saveservice.thread_name=true
jmeter.save.saveservice.successful=true
jmeter.save.saveservice.bytes=true
jmeter.save.saveservice.sent_bytes=true
jmeter.save.saveservice.thread_counts=true
jmeter.save.saveservice.latency=true
jmeter.save.saveservice.connect_time=true
jmeter.save.saveservice.assertion_results_failure_message=true
jmeter.save.saveservice.response_data=false
jmeter.save.saveservice.response_data.on_error=false
jmeter.save.saveservice.samplerData=false
jmeter.save.saveservice.responseHeaders=false
jmeter.save.saveservice.requestHeaders=false
jmeter.save.saveservice.url=false

jmeter.reportgenerator.overall_granularity=2000
jmeter.reportgenerator.aggregate_rpt_pct1=90
jmeter.reportgenerator.aggregate_rpt_pct2=95
jmeter.reportgenerator.aggregate_rpt_pct3=99

These fields make window filtering and generator-validity checks auditable. JTL stores no response body/header.

4. Author two stable JMX plans

Warm-up plan:

Test Plan
├── HTTP Request Defaults -> 127.0.0.1:8030, HttpClient4
└── Thread Group — Warmup
    1 thread × 20 loops
    ├── Counter -> SEQ
    └── HTTP Request — Warmup
        GET /work?run_id=${__P(run.id,warmup)}&seq=${SEQ}
        Use KeepAlive=checked
        ├── Constant Timer 20 ms
        └── Response Assertion: HTTP code = 200

Measurement plan:

Test Plan
├── HTTP Request Defaults
│   host=${__P(target.host,127.0.0.1)}
│   port=${__P(target.port,8030)}
│   implementation=HttpClient4
└── Thread Group — Measure
    ${__P(threads,4)} threads × ${__P(loops,25)} loops
    ├── Counter -> SEQ (per user)
    └── HTTP Request — Measure
        GET /work?run_id=${__P(run.id,p30)}&seq=${SEQ}
        Use KeepAlive=checked
        ├── Constant Timer ${__P(pacing.ms,50)} ms
        └── Response Assertion: HTTP code = 200

Warm-up uses a separate JTL/label and is never merged into the measurement distribution. Measurement uses identical four-thread/25-loop behavior in A and B.

5. First experiment: measure cold startup accidentally

Start variant A fresh:

python .\fixtures\validity_fixture.py `
  --host 127.0.0.1 --port 8030 `
  --base-ms 55 --warmup-count 20 --warmup-extra-ms 90 `
  --jitter-ms 10 --noise-every 17 --noise-ms 25 `
  --log .\results\no-warmup-target.jsonl

Immediately run measure.jmx without a warm-up. The first ~20 target requests carry an extra 90 ms. Preserve the JTL/dashboard/log; this is intentionally contaminated evidence, not something to delete.

6. Repeat with an explicit warm-up

Restart the same variant A configuration with a fresh event log. Run warmup.jmx once (20 requests) and then measure.jmx. Now the cold penalty is consumed outside the comparison window. Predict: full measurement p95 should fall while base/jitter/noise configuration remains identical.

7. Mark/compute explicit measurement windows

tools/analyze_window.py:

import argparse, csv, json, math
from pathlib import Path

def nearest_rank(values, pct):
    data = sorted(values)
    if not data:
        return 0
    return data[max(1, math.ceil(len(data) * pct / 100.0)) - 1]

def main():
    p = argparse.ArgumentParser()
    p.add_argument("--jtl", required=True)
    p.add_argument("--label", default="Measure")
    p.add_argument("--out", required=True)
    p.add_argument("--window-start-s", type=float, default=0.0)
    p.add_argument("--window-end-s", type=float)
    args = p.parse_args()

    rows = list(csv.DictReader(Path(args.jtl).open(newline="", encoding="utf-8")))
    rows = [r for r in rows if r.get("label") == args.label]
    if not rows:
        raise SystemExit("no matching rows")

    first_ts = min(int(r["timeStamp"]) for r in rows)
    selected = []
    for r in rows:
        offset_s = (int(r["timeStamp"]) - first_ts) / 1000.0
        if offset_s < args.window_start_s:
            continue
        if args.window_end_s is not None and offset_s >= args.window_end_s:
            continue
        selected.append(r)

    elapsed = [int(float(r["elapsed"])) for r in selected]
    failures = sum(r.get("success", "").lower() != "true" for r in selected)
    if selected:
        start = min(int(r["timeStamp"]) for r in selected)
        end = max(int(r["timeStamp"]) + int(float(r["elapsed"])) for r in selected)
        span_s = max((end - start) / 1000.0, 0.001)
    else:
        span_s = 0.0

    result = {
        "label": args.label,
        "source_rows": len(rows),
        "selected_rows": len(selected),
        "excluded_rows": len(rows) - len(selected),
        "window_start_s": args.window_start_s,
        "window_end_s": args.window_end_s,
        "p50_ms": nearest_rank(elapsed, 50),
        "p90_ms": nearest_rank(elapsed, 90),
        "p95_ms": nearest_rank(elapsed, 95),
        "p99_ms": nearest_rank(elapsed, 99),
        "mean_ms": round(sum(elapsed) / len(elapsed), 3) if elapsed else 0,
        "error_rate_pct": round(100.0 * failures / len(selected), 3) if selected else 100.0,
        "throughput_rps": round(len(selected) / span_s, 3) if span_s else 0.0,
    }
    Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
    print(json.dumps(result, indent=2))

if __name__ == "__main__":
    main()

Example full measurement:

python .\tools\analyze_window.py `
  --jtl .\results\a1\results.jtl `
  --label Measure `
  --window-start-s 0 `
  --out .\results\a1\summary.json

For a fixed-duration study you can predeclare, for example, window-start-s=2/window-end-s=10. Never decide the cut after seeing which interval makes A look faster.

8. Introduce synthetic background noise deliberately

Restart variant A with a stronger but still bounded noise pattern such as --noise-every 7 --noise-ms 60, then repeat the same measurement three times. Because noise is known and logged, you can see p95/CV widen without inventing a product regression. Restore the normal 17/25 noise settings before the checkpoint.

9. Quantify repetition variation

tools/summarize_repetitions.py:

import argparse, json, math, statistics
from pathlib import Path

T975 = {
    1: 12.706, 2: 4.303, 3: 3.182, 4: 2.776, 5: 2.571,
    6: 2.447, 7: 2.365, 8: 2.306, 9: 2.262, 10: 2.228
}

def summarize(values):
    n = len(values)
    mean = statistics.mean(values)
    median = statistics.median(values)
    if n >= 2:
        sd = statistics.stdev(values)
        cv = 100.0 * sd / mean if mean else 0.0
        t = T975.get(n - 1, 1.96)
        margin = t * sd / math.sqrt(n)
        ci = [mean - margin, mean + margin]
    else:
        sd = cv = 0.0
        ci = [mean, mean]
    return {
        "n": n, "mean": round(mean, 3), "median": round(median, 3),
        "sample_sd": round(sd, 3), "cv_pct": round(cv, 3),
        "min": min(values), "max": max(values),
        "approx_95pct_t_ci_of_mean": [round(ci[0], 3), round(ci[1], 3)]
    }

def main():
    p = argparse.ArgumentParser()
    p.add_argument("summaries", nargs="+")
    p.add_argument("--metric", default="p95_ms")
    p.add_argument("--out", required=True)
    args = p.parse_args()

    docs = [json.loads(Path(x).read_text(encoding="utf-8")) for x in args.summaries]
    values = [float(d[args.metric]) for d in docs]
    result = {
        "metric": args.metric,
        "runs": [{"path": p, "value": v} for p, v in zip(args.summaries, values)],
        "variation": summarize(values),
        "warning": "With only three repetitions the t interval is intentionally wide; use it as an uncertainty illustration, not proof of a population parameter."
    }
    Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
    print(json.dumps(result, indent=2))

if __name__ == "__main__":
    main()
python .\tools\summarize_repetitions.py `
  .\results\a1\summary.json `
  .\results\a2\summary.json `
  .\results\a3\summary.json `
  --metric p95_ms `
  --out .\results\a-p95-variation.json

Report median, mean, SD, CV, min/max and the deliberately cautious t-based interval. With n=3 the interval can be wide; that itself teaches why a single run is weak evidence.

10. Coordinated omission with a deterministic wall-clock stall

For this small demonstration restart the fixture with --warmup-count 0, base≈50 ms and the default periodic stall. Run the closed JMeter plan:

Test Plan
├── HTTP Request Defaults -> 127.0.0.1:8030, HttpClient4
└── Thread Group — Closed CO demonstration
    1 thread × 60 loops
    ├── Counter -> SEQ
    └── HTTP Request — CO-Closed
        GET /co?run_id=co-closed&seq=${SEQ}
        Use KeepAlive=checked
        ├── Constant Timer 50 ms
        └── Response Assertion: HTTP code = 200

Interpretation:
- without stalls, ~50 ms service + 50 ms timer gives roughly 10 cycles/s;
- during a long response the single closed user cannot issue the intended next requests;
- the missing arrivals are coordinated with the stall.

Then run the fixed-arrival fallback:

import argparse, concurrent.futures, json, time, urllib.request
from pathlib import Path

def request(url, intended_s):
    start = time.monotonic()
    try:
        with urllib.request.urlopen(url, timeout=3) as r:
            status = r.status
            r.read()
    except Exception:
        status = 0
    end = time.monotonic()
    return {
        "intended_s": round(intended_s, 4),
        "actual_start_s": round(start, 4),
        "elapsed_ms": round((end - start) * 1000, 3),
        "status": status
    }

def main():
    p = argparse.ArgumentParser()
    p.add_argument("--url-base", default="http://127.0.0.1:8030/co")
    p.add_argument("--rate", type=float, default=10.0)
    p.add_argument("--duration", type=float, default=6.0)
    p.add_argument("--run-id", default="co-arrival")
    p.add_argument("--out", required=True)
    args = p.parse_args()

    count = int(args.rate * args.duration)
    if count > 120:
        raise SystemExit("safety ceiling: max 120 scheduled requests")

    interval = 1.0 / args.rate
    t0 = time.monotonic()
    futures = []
    with concurrent.futures.ThreadPoolExecutor(max_workers=32) as ex:
        for i in range(count):
            intended = t0 + i * interval
            sleep_s = intended - time.monotonic()
            if sleep_s > 0:
                time.sleep(sleep_s)
            url = f"{args.url_base}?run_id={args.run_id}&seq={i+1}"
            futures.append(ex.submit(request, url, intended - t0))
        results = [f.result() for f in futures]

    result = {
        "rate_per_s": args.rate,
        "duration_s": args.duration,
        "scheduled_requests": count,
        "completed": len(results),
        "slow_over_500ms": sum(r["elapsed_ms"] > 500 for r in results),
        "p95_ms": sorted(r["elapsed_ms"] for r in results)[max(0, int(0.95 * len(results)) - 1)],
        "results": results
    }
    Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
    print(json.dumps({k:v for k,v in result.items() if k != "results"}, indent=2))

if __name__ == "__main__":
    main()
python .\tools\arrival_probe.py `
  --url-base http://127.0.0.1:8030/co `
  --rate 10 --duration 6 `
  --run-id co-arrival `
  --out .\results\co-arrival.json

Expected qualitative difference: when one closed thread hits a ~700 ms stall it cannot issue the next scheduled requests, so fewer samples arrive during the bad window. The arrival scheduler continues launching at 10/s, causing several requests to observe the stall. The closed result can therefore underrepresent the latency distribution for a fixed-arrival demand question.

11. Optional current JMeter Open Model comparison

In JMeter 5.6.3 you may author an Open Model Thread Group (experimental) with a bounded profile such as:

rate(10/sec) random_arrivals(6 sec)

Keep this optional and label the exact JMeter version. Validate generated/achieved sample count and generator headroom because an arrival schedule can require more concurrent threads during stalls.

Run evidence: use Apache JMeter 5.6.3/Java 17 and keep each measurement JTL with its matching jmeter.log, dashboard, target JSONL, window summary and generator snapshot. No real credentials are used.

12. Generator-health check

$Proc = Get-CimInstance Win32_Process |
  Where-Object { $_.Name -eq "java.exe" -and $_.CommandLine -like "*ApacheJMeter.jar*" } |
  Select-Object -First 1

if ($Proc) {
  $Pid = $Proc.ProcessId
  Get-Process -Id $Pid |
    Select-Object Id,CPU,WorkingSet64,PrivateMemorySize64,Handles,Threads
  jcmd $Pid GC.heap_info
  jstat -gcutil $Pid 1000 5
}

If the injector saturates, arrival scheduling or latency comparisons can become invalid. Preserve CPU/heap/GC evidence rather than assuming local loopback means “free.”

13. Challenge

You must compare a checkout service that receives approximately 50 external arrivals/s regardless of response time. Would four closed JMeter users be a sufficient primary workload model?

Not necessarily. A closed user model reduces arrivals during slow responses. Prefer an arrival-oriented model/schedule with enough generator headroom, then explicitly compare achieved arrivals with the configured schedule.

Knowledge check

What changes between no-warmup and warmup experiments?

Why log synthetic noise?

Why keep separate warm-up and measurement JTLs?

What demonstrates coordinated omission here?

Why can Open Model need more generator capacity?

Next lesson

Choose experiment structure deliberately

Lesson 3 compares warm-up duration, fixed duration/iterations, open/closed models, long runs/repetitions and baseline-control/canary comparisons.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current primary documentation on 2026-09-05. The mandatory runtime remains Apache JMeter 5.6.3 with Java 17 and no third-party plugin. JMeter 5.6.3 requires Java 8+; the current 5.6.x changes page recommends Java 17 or later. The Open Model Thread Group remains explicitly experimental; the mandatory course path therefore uses the stable standard Thread Group for closed-model measurements and a bounded Python-standard-library fixed-arrival probe to demonstrate coordinated omission. An optional Open Model schedule such as rate(10/sec) random_arrivals(6 sec) is shown only as a current 5.6.3 comparison. All important comparisons use raw CSV JTL plus a documented label/window filter; dashboards remain corroborating evidence. JMeter elapsed time begins immediately before sending the request and ends after the last response byte; latency ends at the first response, and connect time covers connection establishment including SSL handshake.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.