Checkpoint Lab — Performance Governance, Baselines, SLOs, Regression Budgets, and Reporting
The checkpoint produces a complete miniature governance packet for two controlled builds. The expected outcome is intentionally not all green: Build B should pass the absolute service objective but fail the relative p95 budget. A seven-day exception demonstrates that business approval can be auditable without rewriting technical evidence.
Learning objectives
- Complete the chapter checkpoint for Performance Governance, Baselines, SLOs, Regression Budgets, and Reporting as one reviewable, bounded experiment.
- State the workload, predictions, acceptance criteria, authorization boundary, and abort conditions before execution.
- Reconcile configured versus achieved work with JTL, jmeter.log, target evidence, and generator validity before making a conclusion.
- Produce an evidence packet that records the exact inputs, results, diagnosis or gate outcome, and any material limitations.
- Perform cleanup or rollback and explain how the checkpoint evidence hands off to the next chapter or operating practice.
1. Exact assumptions and ceilings
| Item | Checkpoint |
|---|---|
| Runtime | Apache JMeter5.6.3 / Java17 / Python3 stdlib / no third-party plugin. |
| Authorized target | 127.0.0.1:8034 only. |
| Baseline build | build-1.0.0, base55 ms, jitter0..5 ms, no synthetic errors. |
| Current build | build-1.1.0-demo, base64 ms; all other fixture/JMeter settings equal. |
| Load/run | 2 threads×10 loops=20 samples;40 ms pacing. |
| Repetitions | 3 baseline +3 current. |
| Target ceiling | 70 requests/process; expected60 measured requests/build. |
| Policy | Absolute p95≤90ms/error≤1%/throughput≥10; relative p95≤+12%, p50≤+15%, error≤+0.5pp, throughput drop≤20%. |
| Exception | Only for actual failed metric(s); owner+approver+issue+reason; maximum7 days. |
2. Predictions before execution
P1: all six runs will achieve20 samples and zero errors with valid generator state.
P2: Build B p50/p95 will be higher because only base service delay changes55→64 ms; absolute p95 should remain below90 ms.
P3: the relative p95 change should exceed the12% budget, so the technical gate should FAIL even though the absolute SLO passes.
P4: an approved scoped seven-day exception should produce release disposition ALLOW_WITH_EXCEPTION while gate.json remains FAIL.
3. Preflight/freeze the governance inputs
& "$env:JMETER_HOME\bin\jmeter.bat" -v
java -version
Get-FileHash `
.\plans\governed-workload.jmx, `
.\config\governance.properties, `
.\policy\performance-policy.json `
-Algorithm SHA256
Get-Content .\policy\performance-policy.json
Record JMeter/Java, hashes, workload version, environment ID and service authorization before starting either build.
4. Build A: repeat, analyze, aggregate, baseline
- Start build-1.0.0 fixture with base55/jitter5/max70.
- Run a1/a2/a3 with identical JMX/properties; separate JTL/log/dashboard directories.
- Analyze each: configured20, achieved20, generator-valid true.
- Aggregate medians and preserve all run values.
-
Create
baseline-v1.jsonwith owner/reason/policy/source hashes. - Stop Build A after60 measured requests.
5. Build B: repeat under identical protocol
- Start build-1.1.0-demo with base64; keep jitter/error/ceiling/port identical.
- Run b1/b2/b3 with the same workload/environment/analysis formula.
- Analyze and aggregate; preserve target JSONL/JTL/jmeter.log/dashboard/generator notes.
- Verify60 measured target requests and no rejections/errors.
6. Technical gate
Run evaluate_gate.py. Require a valid comparison.
Expected pattern:
| Check | Expected qualitative result |
|---|---|
| Absolute p95≤90ms | PASS |
| Absolute error≤1% | PASS |
| Absolute throughput≥10rps | PASS |
| Relative p50≤+15% | Likely PASS; use actual evidence. |
| Relative p95≤+12% | FAIL in designed scenario. |
| Relative error≤+0.5pp | PASS |
| Relative throughput drop≤20% | PASS under normal local generator state. |
If the actual deterministic timing differs slightly, trust the generated gate output rather than forcing this table. The lab's purpose is transparent policy evaluation, not manufacturing a specific percentage.
7. Simulate one justified exception
If and only if relative_p95_regression_pct actually
failed, create P34-EXC-001 with owner, approver, PERF-134, rationale
and7-day expiry. If another metric also failed, either include it
with explicit rationale or BLOCK; a partial exception cannot cover
unapproved failures.
Record why this demonstration exception exists: the team accepts a known short-lived local regression while remediation is tracked, but does not promote the slower build into the baseline automatically.
8. Release decision remains two-layered
Run the decision tool. The packet should preserve:
technical_gate_status: FAIL
release_disposition: ALLOW_WITH_EXCEPTION
exception_id: P34-EXC-001
reason: approved_unexpired_scoped_exception
After expiry, the same gate/exception pair evaluates to BLOCK. A future baseline promotion requires a separate reviewed baseline-v2 event.
9. Append trend and build stakeholder report
Run build_report.py. The trend row must include build,
baseline version, policy/workload/environment IDs, median metrics,
technical gate, release disposition, exception ID and report hash.
The Markdown report must state limitations and round values without
losing raw JSON/JTL authority.
10. Prove the decision remains auditable later
Close/reopen the packet without rerunning JMeter. Verify:
- baseline-v1 still identifies build-1.0.0 and source/policy hashes;
- gate.json still says FAIL and names the exact failed metric(s);
- exception.json names build/baseline/owner/approver/issue/scope/expiry;
- decision.json explains ALLOW_WITH_EXCEPTION;
- trend history contains both the technical and release outcomes;
- report.json hashes match the preserved inputs;
- raw JTL/jmeter.log directories still support reanalysis during retention.
11. Negative governance test: make comparison INVALID
Copy the current aggregate and change its environment ID to
different-runner (do not alter the original). Evaluate
this copy. Require status INVALID/environment_mismatch. Attempting
to create a performance exception for this INVALID result must be
rejected. This proves the gate distinguishes regression from
comparability failure.
12. Verification checklist
- JMeter5.6.3/Java17/no plugin.
- Only127.0.0.1:8034.
- 20 configured=20 achieved each run;3 repetitions/build.
- Same workload/environment/formula/result schema.
- Only target base_ms55→64 changes.
- Baseline immutable/versioned with owner/reason/hashes.
- Technical gate output preserved independently of exception.
- Exception scope/owner/approver/issue/expiry valid.
- Trend/report contain gate + release disposition + limitations.
- Invalid environment mismatch is BLOCK_INVALID_RUN, not exception-eligible.
13. Cleanup / rollback
- Stop fixture and confirm port8034 closes.
- Keep governance packet/trend/raw evidence for the lab's retention window.
- Rollback policy means continue using baseline-v1/policy-v1; do not delete current FAIL/exception history.
- After raw-retention expiry, keep compact trend/baseline/policy/decision history according to policy.
- No production/shared target, credential, remote engine, cloud resource or system-wide setting was changed.
14. Production operating-model addition and Chapter35 bridge
Chapter34 adds a performance-governance contract: service objectives, approved workload/environment identity, immutable versioned baselines, transparent repeated-run reducers, absolute and relative regression budgets, fail-closed validity checks, immutable technical gate outputs, scoped expiring exceptions, trend history, stakeholder reporting, ownership/change control and raw-evidence retention are required for durable release policy.
Chapter35 is the capstone: Design and Operate a Production Performance-Testing Program. It combines workload modeling, correlation/data, CLI/distributed execution, telemetry, diagnosis, validity, safety, extensions, troubleshooting and governance into one end-to-end operating program.
Knowledge check
What should remain FAIL after an approved exception?
The technical gate; only release disposition becomes ALLOW_WITH_EXCEPTION.
Can you create an exception for environment_mismatch INVALID?
No. Repair the comparison and rerun; exceptions cover valid failed metrics, not invalid evidence.
Why not promote Build B into baseline-v1 after accepting its exception?
That would silently erase the regression. A new baseline version requires a separate reviewed change event.
What proves the historical decision later?
Immutable policy/baseline/gate/exception/decision/report hashes plus trend row and retained raw evidence.
Why does Chapter35 need Chapter34?
A production testing program needs governance around all earlier technical practices so results drive consistent, auditable operational decisions.
Official references and version notes
- Apache JMeter downloads — current stable JMeter 5.6.3 and Java 8+ requirement.
- JMeter current changes — Java 17+ recommendation for the 5.6.x line.
- JMeter Getting Started — non-GUI/CLI execution and result/log flags.
- JMeter Dashboard Report — CSV requirements, p90/p95/p99 defaults, report generation and percentile-estimator caveat.
- JMeter Properties Reference — result-save fields, aggregate percentile properties and reporting properties.
- JMeter Remote Testing — same-plan fan-out, exact JMeter parity, Java/data requirements and controller overhead.
Version-sensitive behavior was rechecked against current primary
documentation on 2026-09-06. Mandatory runtime:
Apache JMeter 5.6.3, Java 17, no third-party
plugin, Python 3 standard library only. Meaningful runs use CLI
with raw CSV JTL plus matching jmeter.log; HTML
dashboards are corroborating evidence. JMeter's dashboard defaults
to configurable p90/p95/p99 and can estimate percentiles
differently from other reports, especially with few samples. For
governance math this chapter therefore uses one explicit
nearest-rank formula over raw, label-filtered CSV JTL and stores
that formula/version in every summary. The mandatory local policy
requires three repetitions per build, exact workload/environment
identity, configured-versus-achieved sample checks and
generator-validity notes before a gate can be evaluated.
Remote/cloud/paid CI is optional only; if later used, every
environment/engine must carry the same workload/policy identity
and valid runtime evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.