Chapter 34Lesson 04~250 minutes

Performance Governance, Baselines, SLOs, Regression Budgets, and Reporting: Diagnostics, Failure Modes, and Production Practices

Governance fails when policy state becomes mutable or comparison identity disappears. These are technical defects: they can make a release gate look deterministic while actually changing the question after every run.

Baseline driftEnvironment mismatchPost-hoc thresholdException expiryFalse precision

Learning objectives

  • Apply a preserve-first diagnostic sequence to failure modes involving Performance Governance, Baselines, SLOs, Regression Budgets, and Reporting.
  • Separate plan/configuration, generator, protocol/network, target, CI/container, and distributed-engine causes before changing settings.
  • Reproduce a failure with the smallest authorized workload and retain the original JTL, jmeter.log, and supporting evidence.
  • Reject shortcuts such as blanket retries, disabled verification, unbounded load increases, silent global overrides, or deleting first-failure artifacts.
  • Verify the least-invasive correction under the original bounded workload before declaring the problem resolved.

1. Preserve-first governance diagnostic sequence

Governance-failure diagnosis

Governance does not replace measurement validity. It records which valid experiment, workload, environment, baseline and policy produced a release decision.

flowchart TD
P[Preserve first-failure JTL / jmeter.log / engine / target evidence] --> V[Confirm JMeter / Java / plugin / tool versions]
V --> I[Confirm exact JMX / data / properties / CLI / authorized target]
I --> S[Validate tree scope + resolved variables/properties]
S --> D[Inspect protocol/session/data]
D --> G[Inspect generator JVM/OS/network]
G --> T[Inspect SUT telemetry]
T --> C[Inspect CI/container/distributed state]
C --> Q[Inspect policy/baseline/workload/environment/exception identity]
Q --> F[Least-destructive correction]
F --> R[Small controlled rerun]
R --> O[Re-evaluate immutable gate]
Runnable diagnosis: use only 127.0.0.1:8034, one20-sample run or the normal3-repeat set. Never “debug” governance by increasing production load, deleting evidence, or weakening safety/security boundaries.

2. Failure mode: baseline silently drifts

Broken process stores baseline as baseline.json and overwrites it after every successful build. A sequence 55→58→61→64 ms may pass each small comparison even though64 ms is16% worse than55 ms.

Repair: baseline is immutable/versioned. Promotion creates baseline-v2 with owner/reason/source hashes while baseline-v1 remains in history.

3. Failure mode: comparing different environments

A baseline ran on a dedicated runner; current ran on a shared laptop/container with different CPU/Java/network limits. A relative percentage is computed anyway. The correct technical result is INVALID, not FAIL/PASS. Restore normalized environment or explicitly establish a new reviewed protocol/baseline.

4. Intentionally broken example: threshold chosen after seeing result

# anti-pattern: current p95 is 77ms, so "policy" becomes 80ms
$current = Get-Content .\runs\current\aggregate.json | ConvertFrom-Json
$policy.absolute_slo.p95_ms_max = [math]::Ceiling($current.metrics.p95_ms.median + 3)
$policy | ConvertTo-Json -Depth 10 | Set-Content .\policy\performance-policy.json

This is not calibration; it is outcome-driven policy mutation. Preserve the original policy hash, restore it, and rerun/evaluate. A legitimate policy revision is reviewed/versioned for future comparisons and must not retroactively rewrite the failed gate.

5. Failure mode: one metric decides everything

A build improves p95 because 8% of requests fail quickly. A p95-only gate might PASS. A governed composite includes error rate and workload validity; failed requests cannot be silently excluded unless the policy explicitly defines the service objective that way.

6. Failure mode: exceptions never expire

An exception stored as “approved=true” with no expiry/owner/issue becomes a permanent hidden threshold change. Require maximum lifetime, exact failed-metric scope, owner, approver, remediation issue, baseline/build identity, and automatic BLOCK after expiry.

7. Failure mode: missing workload version

Baseline=2×10 with40 ms pacing; current=8 threads with no timer. The same label exists, so a dashboard comparison is made. Without workload identity, the numbers answer different questions. Workload version/hash must be a first-class comparison key.

8. Failure mode: precision beyond validity

A report says “p95 regressed12.4378% with confidence” from20 samples/run and three repetitions. The calculation may be numerically correct, but the wording implies unsupported precision. Keep raw values; round stakeholder summaries and explicitly disclose sample/repetition/model/environment limitations.

9. Causal performance interpretation still applies

Observed change Competing cause to rule out
p95↑ Timer/scope mismatch, generator CPU/GC, listener/save-service cost, DNS/connect/TLS, target saturation.
Throughput↓ Closed-model response-time effect, generator saturation, errors/retries, target limit.
Errors↑ Assertion/data/session mismatch, dependency/target errors, rate/safety ceiling.
Trend step change Runner/JMeter/Java/environment/workload/baseline version change.
Remote trend degradation Engine count/result transfer/JAR/data/network parity change.

Governance should reference Chapter33 evidence when a gate failure may be infrastructure/generator-invalid rather than a product regression.

10. Troubleshooting shortcuts to reject

  • Do not add blanket retries or arbitrary sleeps to make a gate pass.
  • Do not assign giant heaps without generator evidence.
  • Do not mass-disable listeners/evidence without measured cost.
  • Do not use global property hacks or silently rewrite policy/baseline files.
  • Do not disable TLS/RMI verification.
  • Do not experiment on uncontrolled production/public targets.
  • Do not increase workload unboundedly to “prove” a regression.
  • Do not delete first-failure JTL/log/policy/gate evidence.

11. CI/container/distributed production practice

A CI gate must persist the policy/baseline/workload/environment identity and raw/summary artifacts, not just return exit1. Containers/runners should record image/resource identity. Distributed tests require per-engine parity and configured-versus-achieved aggregation. Managed trend systems are optional; the local CSV demonstrates the minimum auditable schema.

12. Security-sensitive governance artifacts

Governance packets may contain environment topology, build identifiers, performance vulnerabilities, target names and exception rationale. Keep secrets/real customer data out of JTL/report by design; restrict raw artifact access/retention as needed. Exceptions are approvals, not credentials, but approver identity/history should be integrity-protected in real systems.

13. Broken comparison repair exercise

Suppose current aggregate has environment_ids=["runner-v2"] while baseline requires runner-v1. The gate should emit INVALID and list environment mismatch. Do not create an exception. The fix is to rerun current on runner-v1 or create a reviewed new environment protocol and baseline. Only after a valid normalized run should regression math execute.

Knowledge check

Why is silently advancing the baseline dangerous?

Can an exception authorize an INVALID run?

Why is missing workload version a technical defect?

How can one latency metric pass while service quality worsens?

Why keep the original policy hash after a threshold revision?

Next lesson

Checkpoint: create an auditable packet

Lesson5 executes the two-build scenario, records the technical FAIL, approves one scoped exception, appends trend history, and verifies the packet later.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against current primary documentation on 2026-09-06. Mandatory runtime: Apache JMeter 5.6.3, Java 17, no third-party plugin, Python 3 standard library only. Meaningful runs use CLI with raw CSV JTL plus matching jmeter.log; HTML dashboards are corroborating evidence. JMeter's dashboard defaults to configurable p90/p95/p99 and can estimate percentiles differently from other reports, especially with few samples. For governance math this chapter therefore uses one explicit nearest-rank formula over raw, label-filtered CSV JTL and stores that formula/version in every summary. The mandatory local policy requires three repetitions per build, exact workload/environment identity, configured-versus-achieved sample checks and generator-validity notes before a gate can be evaluated. Remote/cloud/paid CI is optional only; if later used, every environment/engine must carry the same workload/policy identity and valid runtime evidence.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.