Chapter 30Lesson 05~340 minutes

Checkpoint Lab — Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design

The checkpoint tests a known engineering difference instead of hunting for one. Variant A has base service time 55 ms; variant B is intentionally 70 ms. Everything else—warm-up, jitter/noise, JMeter plan, threads, loops, pacing, result fields, generator, analyzer and measurement window—must remain identical. The task is to detect the known +15 ms change without overstating certainty.

CheckpointA/B3 repetitionsWarm-upValidity threats

Learning objectives

  • Complete the chapter checkpoint for Test Validity, Coordinated Omission, Warm-Up, Noise, and Experimental Design as one reviewable, bounded experiment.
  • State the workload, predictions, acceptance criteria, authorization boundary, and abort conditions before execution.
  • Reconcile configured versus achieved work with JTL, jmeter.log, target evidence, and generator validity before making a conclusion.
  • Produce an evidence packet that records the exact inputs, results, diagnosis or gate outcome, and any material limitations.
  • Perform cleanup or rollback and explain how the checkpoint evidence hands off to the next chapter or operating practice.

1. Exact assumptions and ceilings

Item Checkpoint
JMeter/Java Apache JMeter 5.6.3 / Java17 / no third-party plugin.
Fixture/tooling Python3 standard library only.
Target 127.0.0.1:8030.
Variant A base_ms=55.
Variant B base_ms=70 (+15 ms known change).
Shared target settings warmup_count20, warmup_extra90, jitter≤10, every17th request +25ms noise.
Warm-up 20 requests on every fresh A/B fixture before measurement.
Measurement 4 threads ×25 loops =100 samples/run, 50 ms pacing.
Repetitions 3 A +3 B measurement runs; no result overwrite.
Single-run ceiling 100 measured samples; ≤20 seconds.
Primary comparison Measure-label p95 plus median/variation; also error/RPS/count/generator evidence.
CO demonstration separate ≤60-arrival teaching experiment; not mixed into A/B metric.
Abort: non-loopback target, JTL count ≠100, target/fixture failure, generator saturation, different hashes/config between A/B, missing warm-up on a fresh fixture, unplanned noise/service setting change, missing evidence, or any threshold/protocol change after seeing results.

2. Write/freeze the experiment protocol before running

Your protocol must state:

  1. Question: can the experiment detect the known +15 ms base-service change?
  2. H0 engineering interpretation: A/B difference is not clearly distinguishable from run variation.
  3. H1 engineering interpretation: B shows consistently higher measurement p95/median than A by an amount materially larger than within-condition variation.
  4. Warm-up: 20 requests excluded from measurement.
  5. Measurement: exactly 100 Measure samples/run; same JMX/properties.
  6. Repetitions: three per condition; preferably alternate/order the conditions if host drift is a concern.
  7. Primary metric: p95 nearest-rank on label Measure; secondary p50/RPS/error.
  8. Validity gates: exact counts, generator health, same hashes/build/data/noise config.
  9. Known limitations: n=3, loopback, closed model, synthetic service and coordinated omission.

3. Predictions before execution

P1: fresh no-warm-up measurement would be contaminated by the first20 +90 ms samples; the checkpoint warm-up should remove that effect from Measure JTL.

P2: B p50/p95 should generally exceed A because only base service increases 55→70 ms.

P3: A and B each show small repeat variation due to deterministic run-id jitter/noise; the known effect should be interpreted relative to this variation.

P4: configured/achieved samples remain100, errors≈0 and generator health stays acceptable; otherwise the comparison is invalid.

4. Variant A protocol

Start fresh A:

python .\fixtures\validity_fixture.py `
  --host 127.0.0.1 --port 8030 `
  --base-ms 55 `
  --warmup-count 20 --warmup-extra-ms 90 `
  --jitter-ms 10 --noise-every 17 --noise-ms 25 `
  --log .\results\a-target.jsonl

Run the warm-up JMX once; preserve its own JTL/log. Then run Measure three times with run IDs a1, a2, a3, each to a separate results directory. Analyze each Measure JTL with window-start-s=0.

5. Variant B protocol

Stop A, start a fresh B process with the exact same command except --base-ms 70 and a new event log. Run the exact same 20-request warm-up, then b1/b2/b3 measurement runs. Do not change JMeter/Java/properties/JMX/noise/jitter/timers.

6. Quantify within-condition variation

python .\tools\summarize_repetitions.py `
  .\results\a1\summary.json .\results\a2\summary.json .\results\a3\summary.json `
  --metric p95_ms --out .\results\a-p95.json

python .\tools\summarize_repetitions.py `
  .\results\b1\summary.json .\results\b2\summary.json .\results\b3\summary.json `
  --metric p95_ms --out .\results\b-p95.json

Report p95 median/mean/SD/CV/range and approximate t interval. Repeat for p50 if useful. Do not call the n=3 interval definitive significance evidence.

7. Compare A and B as engineering evidence

Evidence A B Interpretation
Base target setting 55 ms 70 ms Known +15 ms causal change.
Warm-up 20 requests 20 requests Same cold-state treatment.
Measurement/run 100 Measure samples 100 Same configured/achieved work.
p50/p95 repeats record 3 values record 3 values Compare effect with within-condition spread.
Errors/RPS record record Detect confounding failure/throughput differences.
Generator CPU/GC valid valid Exclude injector saturation.
Noise/jitter same config same Controlled validity factor.

8. Noise observation and run order

Record whether background host state changed during the six runs. If thermal/CI/OS drift is plausible, alternate A/B on fresh processes in a second experiment rather than always A then B. Do not cherry-pick “quiet” runs after observing results.

9. Coordinated-omission validity threat

The checkpoint's 4-thread closed workload asks about a fixed concurrent-user population. It does not prove the result for a fixed external arrival rate. Use Lesson2's separate wall-clock-stall + 10/s arrival probe to demonstrate how a closed model can omit demand during stalls. State this threat even if the A/B effect is detected cleanly.

10. Required evidence packet

Artifact Required content
Experiment protocol Question/hypothesis, primary metric, warm-up/window, repetitions, stop rules.
Environment manifest JMeter/Java/Python, OS/host notes, JMX/properties hashes and fixture settings.
Warm-up evidence 20-sample warm-up JTL/log for every fresh A/B fixture.
Repeated measurements a1-a3/b1-b3 JTL + matching jmeter.log + dashboards + summary.json.
Variation A/B p95 repeat summary with median/SD/CV/range/t interval.
Target evidence A/B event logs proving base55/base70, cold flags/noise/jitter.
Generator evidence CPU/heap/GC snapshot from representative measurement.
Load-model note Closed 4-user model; optional/teaching arrival-oriented comparison kept separate.
Validity threats Small n, loopback, warm-up assumptions, synthetic noise, CO, time drift.

11. Conclusion template

Example form—not values to copy: “A and B used JMeter5.6.3/Java17 with identical 4×25 Measure workload, 50 ms pacing, target noise/jitter and a 20-request warm-up on each fresh target. Three independent measurement windows per condition each produced exactly100 samples with acceptable generator health. Variant B differed only in base service time 55→70 ms. B's p50/p95 values were consistently higher than A and the between-condition shift was larger than the observed within-condition spread, supporting detection of the known latency change in this local protocol. However n=3 yields wide uncertainty, the target is loopback/synthetic, and the closed workload can exhibit coordinated omission for fixed-arrival questions; therefore this experiment does not establish a production SLO regression threshold.”

12. Verification checklist

  • Only 127.0.0.1:8030.
  • JMeter5.6.3 / Java17 / no plugin.
  • Same JMX/properties/analyzer and target noise/jitter settings.
  • Exactly20 warm-up requests per fresh target.
  • Exactly100 Measure samples per a1-a3/b1-b3.
  • Only base_ms55→70 changes.
  • A/B p50/p95/repetition variation retained.
  • Generator health and target events retained.
  • No unexplained outlier deletion or tuning-until-pass.
  • Coordinated omission and other validity threats documented.

13. Cleanup / rollback

  1. Stop fixture and confirm port8030 is free.
  2. Keep protocol, A/B event logs, all JTL/log/dashboard/summary and manifests until review.
  3. Rollback is start variant A again with base_ms55; no global OS/JVM setting changed.
  4. Delete disposable results only after retention policy.
  5. No production/shared target, credential, remote engine, cloud resource or system-wide tuning was modified.

14. Production operating-model addition and Chapter31 bridge

Chapter30 adds an experimental-validity contract: question/hypothesis, environment/data/version manifest, explicit closed/open workload interpretation, predeclared warm-up and measurement window, configured-versus-achieved load, repeated runs, variation/effect-size comparison, noise/clock/generator evidence, coordinated-omission assessment, no post-hoc outlier/protocol tuning, and uncertainty disclosure are required before performance thresholds or regression claims are trusted.

Chapter31 moves to Security, Data Protection, Authorization, and Safe Load-Test Boundaries, where this experimental discipline is combined with least privilege, sensitive-data controls, target authorization, secret handling and blast-radius limits.

Knowledge check

What is the known causal A/B change?

Why warm each fresh fixture before measuring?

Why are three repetitions still not definitive proof?

If one B run has only70 samples, can it be included because its p95 matches the others?

What does Chapter31 add next?

Next chapter

Security, Data Protection, Authorization, and Safe Load-Test Boundaries

Chapter31 applies explicit authorization and data/trust controls to production performance-testing workflows.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current primary documentation on 2026-09-05. The mandatory runtime remains Apache JMeter 5.6.3 with Java 17 and no third-party plugin. JMeter 5.6.3 requires Java 8+; the current 5.6.x changes page recommends Java 17 or later. The Open Model Thread Group remains explicitly experimental; the mandatory course path therefore uses the stable standard Thread Group for closed-model measurements and a bounded Python-standard-library fixed-arrival probe to demonstrate coordinated omission. An optional Open Model schedule such as rate(10/sec) random_arrivals(6 sec) is shown only as a current 5.6.3 comparison. All important comparisons use raw CSV JTL plus a documented label/window filter; dashboards remain corroborating evidence. JMeter elapsed time begins immediately before sending the request and ends after the last response byte; latency ends at the first response, and connect time covers connection establishment including SSL handshake.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.