Checkpoint Lab — Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design
The checkpoint tests a known engineering difference instead of hunting for one. Variant A has base service time 55 ms; variant B is intentionally 70 ms. Everything else—warm-up, jitter/noise, JMeter plan, threads, loops, pacing, result fields, generator, analyzer and measurement window—must remain identical. The task is to detect the known +15 ms change without overstating certainty.
Learning objectives
- Complete the chapter checkpoint for Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design as one reviewable, bounded experiment.
- State the workload, predictions, acceptance criteria, authorization boundary, and abort conditions before execution.
- Reconcile configured versus achieved work with JTL, jmeter.log, target evidence, and generator validity before making a conclusion.
- Produce an evidence packet that records the exact inputs, results, diagnosis or gate outcome, and any material limitations.
- Perform cleanup or rollback and explain how the checkpoint evidence hands off to the next chapter or operating practice.
1. Exact assumptions and ceilings
| Item | Checkpoint |
|---|---|
| JMeter/Java | Apache JMeter 5.6.3 / Java17 / no third-party plugin. |
| Fixture/tooling | Python3 standard library only. |
| Target | 127.0.0.1:8030. |
| Variant A | base_ms=55. |
| Variant B | base_ms=70 (+15 ms known change). |
| Shared target settings | warmup_count20, warmup_extra90, jitter≤10, every17th request +25ms noise. |
| Warm-up | 20 requests on every fresh A/B fixture before measurement. |
| Measurement | 4 threads ×25 loops =100 samples/run, 50 ms pacing. |
| Repetitions | 3 A +3 B measurement runs; no result overwrite. |
| Single-run ceiling | 100 measured samples; ≤20 seconds. |
| Primary comparison | Measure-label p95 plus median/variation; also error/RPS/count/generator evidence. |
| CO demonstration | separate ≤60-arrival teaching experiment; not mixed into A/B metric. |
2. Write/freeze the experiment protocol before running
Your protocol must state:
- Question: can the experiment detect the known +15 ms base-service change?
- H0 engineering interpretation: A/B difference is not clearly distinguishable from run variation.
- H1 engineering interpretation: B shows consistently higher measurement p95/median than A by an amount materially larger than within-condition variation.
- Warm-up: 20 requests excluded from measurement.
- Measurement: exactly 100 Measure samples/run; same JMX/properties.
- Repetitions: three per condition; preferably alternate/order the conditions if host drift is a concern.
- Primary metric: p95 nearest-rank on label Measure; secondary p50/RPS/error.
- Validity gates: exact counts, generator health, same hashes/build/data/noise config.
- Known limitations: n=3, loopback, closed model, synthetic service and coordinated omission.
3. Predictions before execution
P1: fresh no-warm-up measurement would be contaminated by the first20 +90 ms samples; the checkpoint warm-up should remove that effect from Measure JTL.
P2: B p50/p95 should generally exceed A because only base service increases 55→70 ms.
P3: A and B each show small repeat variation due to deterministic run-id jitter/noise; the known effect should be interpreted relative to this variation.
P4: configured/achieved samples remain100, errors≈0 and generator health stays acceptable; otherwise the comparison is invalid.
4. Variant A protocol
Start fresh A:
python .\fixtures\validity_fixture.py `
--host 127.0.0.1 --port 8030 `
--base-ms 55 `
--warmup-count 20 --warmup-extra-ms 90 `
--jitter-ms 10 --noise-every 17 --noise-ms 25 `
--log .\results\a-target.jsonl
Run the warm-up JMX once; preserve its own JTL/log. Then run Measure
three times with run IDs a1, a2,
a3, each to a separate results directory. Analyze each
Measure JTL with window-start-s=0.
5. Variant B protocol
Stop A, start a fresh B process with the exact same command except
--base-ms 70 and a new event log. Run the exact same
20-request warm-up, then b1/b2/b3
measurement runs. Do not change
JMeter/Java/properties/JMX/noise/jitter/timers.
6. Quantify within-condition variation
python .\tools\summarize_repetitions.py `
.\results\a1\summary.json .\results\a2\summary.json .\results\a3\summary.json `
--metric p95_ms --out .\results\a-p95.json
python .\tools\summarize_repetitions.py `
.\results\b1\summary.json .\results\b2\summary.json .\results\b3\summary.json `
--metric p95_ms --out .\results\b-p95.json
Report p95 median/mean/SD/CV/range and approximate t interval. Repeat for p50 if useful. Do not call the n=3 interval definitive significance evidence.
7. Compare A and B as engineering evidence
| Evidence | A | B | Interpretation |
|---|---|---|---|
| Base target setting | 55 ms | 70 ms | Known +15 ms causal change. |
| Warm-up | 20 requests | 20 requests | Same cold-state treatment. |
| Measurement/run | 100 Measure samples | 100 | Same configured/achieved work. |
| p50/p95 repeats | record 3 values | record 3 values | Compare effect with within-condition spread. |
| Errors/RPS | record | record | Detect confounding failure/throughput differences. |
| Generator CPU/GC | valid | valid | Exclude injector saturation. |
| Noise/jitter | same config | same | Controlled validity factor. |
8. Noise observation and run order
Record whether background host state changed during the six runs. If thermal/CI/OS drift is plausible, alternate A/B on fresh processes in a second experiment rather than always A then B. Do not cherry-pick “quiet” runs after observing results.
9. Coordinated-omission validity threat
The checkpoint's 4-thread closed workload asks about a fixed concurrent-user population. It does not prove the result for a fixed external arrival rate. Use Lesson2's separate wall-clock-stall + 10/s arrival probe to demonstrate how a closed model can omit demand during stalls. State this threat even if the A/B effect is detected cleanly.
10. Required evidence packet
| Artifact | Required content |
|---|---|
| Experiment protocol | Question/hypothesis, primary metric, warm-up/window, repetitions, stop rules. |
| Environment manifest | JMeter/Java/Python, OS/host notes, JMX/properties hashes and fixture settings. |
| Warm-up evidence | 20-sample warm-up JTL/log for every fresh A/B fixture. |
| Repeated measurements | a1-a3/b1-b3 JTL + matching jmeter.log + dashboards + summary.json. |
| Variation | A/B p95 repeat summary with median/SD/CV/range/t interval. |
| Target evidence | A/B event logs proving base55/base70, cold flags/noise/jitter. |
| Generator evidence | CPU/heap/GC snapshot from representative measurement. |
| Load-model note | Closed 4-user model; optional/teaching arrival-oriented comparison kept separate. |
| Validity threats | Small n, loopback, warm-up assumptions, synthetic noise, CO, time drift. |
11. Conclusion template
12. Verification checklist
- Only 127.0.0.1:8030.
- JMeter5.6.3 / Java17 / no plugin.
- Same JMX/properties/analyzer and target noise/jitter settings.
- Exactly20 warm-up requests per fresh target.
- Exactly100 Measure samples per a1-a3/b1-b3.
- Only base_ms55→70 changes.
- A/B p50/p95/repetition variation retained.
- Generator health and target events retained.
- No unexplained outlier deletion or tuning-until-pass.
- Coordinated omission and other validity threats documented.
13. Cleanup / rollback
- Stop fixture and confirm port8030 is free.
- Keep protocol, A/B event logs, all JTL/log/dashboard/summary and manifests until review.
- Rollback is start variant A again with base_ms55; no global OS/JVM setting changed.
- Delete disposable results only after retention policy.
- No production/shared target, credential, remote engine, cloud resource or system-wide tuning was modified.
14. Production operating-model addition and Chapter31 bridge
Chapter30 adds an experimental-validity contract: question/hypothesis, environment/data/version manifest, explicit closed/open workload interpretation, predeclared warm-up and measurement window, configured-versus-achieved load, repeated runs, variation/effect-size comparison, noise/clock/generator evidence, coordinated-omission assessment, no post-hoc outlier/protocol tuning, and uncertainty disclosure are required before performance thresholds or regression claims are trusted.
Chapter31 moves to Security, Data Protection, Authorization, and Safe Load-Test Boundaries, where this experimental discipline is combined with least privilege, sensitive-data controls, target authorization, secret handling and blast-radius limits.
Knowledge check
What is the known causal A/B change?
Only base service time changes from55 to70 ms; all workload/warm-up/noise/runtime settings remain fixed.
Why warm each fresh fixture before measuring?
To keep the known first20-request cold penalty outside the A/B measurement window.
Why are three repetitions still not definitive proof?
n=3 exposes variation but produces wide uncertainty and cannot remove model/environment validity limits.
If one B run has only70 samples, can it be included because its p95 matches the others?
No. Configured-versus-achieved mismatch makes that run invalid for the planned comparison.
What does Chapter31 add next?
Authorization, secrets/data protection and safe operational boundaries around the now-valid experiment design.
Official references and version notes
- Apache JMeter downloads — current JMeter 5.6.3 and Java 8+ requirement.
- JMeter current changes — Java 17+ recommendation for the 5.6.x line.
- Open Model Thread Group — currently experimental and schedule-driven.
- JMeter Dashboard Report — CSV requirements, report windows/percentiles and estimator caveats.
- JMeter Glossary — elapsed, latency and connect-time definitions.
- JMeter Listeners / result fields — timestamps, elapsed, success, active-thread counts and other JTL fields.
Version-sensitive statements were rechecked against current
primary documentation on 2026-09-05. The mandatory runtime remains
Apache JMeter 5.6.3 with Java 17 and no
third-party plugin. JMeter 5.6.3 requires Java 8+; the current
5.6.x changes page recommends Java 17 or later. The
Open Model Thread Group remains explicitly experimental; the mandatory course path therefore uses the stable standard
Thread Group for closed-model measurements and a bounded
Python-standard-library fixed-arrival probe to demonstrate
coordinated omission. An optional Open Model schedule such as
rate(10/sec) random_arrivals(6 sec) is shown only as
a current 5.6.3 comparison. All important comparisons use raw CSV
JTL plus a documented label/window filter; dashboards remain
corroborating evidence. JMeter elapsed time begins immediately
before sending the request and ends after the last response byte;
latency ends at the first response, and connect time covers
connection establishment including SSL handshake.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.