Performance Governance, Baselines, SLOs, Regression Budgets, and Reporting: Guided Hands-On Workflow
The guided workflow deliberately produces a nuanced result. Build A is the approved baseline at55 ms synthetic service time. Build B uses64 ms: it should remain inside the absolute90 ms p95 SLO but exceed the stricter12% relative p95 regression budget. The technical gate should therefore FAIL, then a separately generated seven-day exception can produce ALLOW_WITH_EXCEPTION without changing the gate.
Learning objectives
- Execute a bounded hands-on workflow for Performance Governance, Baselines, SLOs, Regression Budgets, and Reporting using the course's current runtime and authorized local or synthetic resources.
- Build and verify the concrete lab artifacts step by step instead of treating configuration snippets as isolated examples.
- Preserve the JTL, jmeter.log, target, generator, and configuration evidence required by the workflow before interpreting results.
- Distinguish configured state from achieved behavior, and stop when safety, count, environment, or generator-validity conditions are not met.
- Explain how the completed workflow prepares the configuration and trade-off analysis in the next lesson.
1. Local governance workspace
p34-governance/
├── fixtures/governance_fixture.py
├── plans/governed-workload.jmx
├── config/governance.properties
├── policy/performance-policy.json
├── baselines/
├── runs/
│ ├── baseline/{a1,a2,a3}/
│ └── current/{b1,b2,b3}/
├── tools/
│ ├── analyze_run.py
│ ├── aggregate_runs.py
│ ├── create_baseline.py
│ ├── evaluate_gate.py
│ ├── create_exception.py
│ ├── decide_release.py
│ └── build_report.py
├── trend/history.csv
└── packet/
127.0.0.1:8034; 2 threads×10 loops=20 samples/run;
three repetitions/build; max60 measured target requests/build and70
fixture ceiling; 40 ms timer; no retries/plugins/real credentials.
Abort on any count≠20, target rejection, non-loopback target,
environment/workload mismatch, generator invalidity, missing
JTL/log, or policy/baseline mutation after results are viewed.
2. Controlled local service with explicit build identity
fixtures/governance_fixture.py:
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from urllib.parse import urlparse, parse_qs
import argparse, hashlib, json, re, threading, time
FIXTURE_VERSION = "prompt34-governance-fixture-v1"
SAFE = re.compile(r"^[A-Za-z0-9_.:-]{1,96}$")
lock = threading.Lock()
event_log = None
build_id = "unset"
base_ms = 55
jitter_ms = 5
error_every = 0
max_requests = 70
state = {"requests": 0, "successes": 0, "errors": 0, "rejections": 0}
def jitter(run_id, seq):
if jitter_ms <= 0:
return 0
raw = hashlib.sha256(f"{build_id}|{run_id}|{seq}".encode()).digest()
return int.from_bytes(raw[:4], "big") % (jitter_ms + 1)
def snapshot():
with lock:
return {
"fixture_version": FIXTURE_VERSION,
"build_id": build_id,
"base_ms": base_ms,
"jitter_ms": jitter_ms,
"error_every": error_every,
"max_requests": max_requests,
**state,
}
def write_event(event):
with lock:
with event_log.open("a", encoding="utf-8") as h:
h.write(json.dumps(event, sort_keys=True) + "\n")
class Handler(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
def send_json(self, status, payload):
raw = json.dumps(payload, sort_keys=True).encode()
self.send_response(status)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(raw)))
self.send_header("X-Fixture-Version", FIXTURE_VERSION)
self.send_header("X-Build-Id", build_id)
self.end_headers()
self.wfile.write(raw)
def do_GET(self):
parsed = urlparse(self.path)
if parsed.path == "/health":
self.send_json(200, {"status": "ok", "state": snapshot()})
return
if parsed.path == "/stats":
self.send_json(200, {"status": "ok", "state": snapshot()})
return
if parsed.path != "/work":
self.send_json(404, {"status": "not_found"})
return
q = parse_qs(parsed.query)
run_id = q.get("run_id", [""])[0]
seq = q.get("seq", [""])[0]
if not SAFE.fullmatch(run_id) or not SAFE.fullmatch(seq):
self.send_json(400, {"status": "invalid_metadata"})
return
with lock:
state["requests"] += 1
n = state["requests"]
if n > max_requests:
with lock:
state["rejections"] += 1
self.send_json(429, {"status": "sample_ceiling"})
return
delay = base_ms + jitter(run_id, seq)
time.sleep(delay / 1000.0)
fail = error_every > 0 and n % error_every == 0
status = 503 if fail else 200
with lock:
state["errors" if fail else "successes"] += 1
ended = int(time.time() * 1000)
self.send_json(status, {
"status": "synthetic_error" if fail else "ok",
"build_id": build_id,
"run_id": run_id,
"seq": seq,
"configured_delay_ms": delay,
})
write_event({
"ts_ms": ended,
"operation": "work",
"status": status,
"build_id": build_id,
"run_id": run_id,
"seq": seq,
"configured_delay_ms": delay,
"request_ordinal": n,
})
def log_message(self, format, *args):
return
def main():
p = argparse.ArgumentParser()
p.add_argument("--host", default="127.0.0.1")
p.add_argument("--port", type=int, default=8034)
p.add_argument("--build-id", required=True)
p.add_argument("--base-ms", type=int, required=True)
p.add_argument("--jitter-ms", type=int, default=5)
p.add_argument("--error-every", type=int, default=0)
p.add_argument("--max-requests", type=int, default=70)
p.add_argument("--log", required=True)
args = p.parse_args()
if args.host != "127.0.0.1":
raise SystemExit("fixture may bind only to loopback")
if not 1 <= args.base_ms <= 200:
raise SystemExit("base-ms must be 1..200")
if not 0 <= args.jitter_ms <= 20:
raise SystemExit("jitter-ms must be 0..20")
if not 0 <= args.error_every <= 1000:
raise SystemExit("error-every out of range")
if not 1 <= args.max_requests <= 70:
raise SystemExit("max-requests must be 1..70")
global event_log, build_id, base_ms, jitter_ms, error_every, max_requests
build_id = args.build_id
base_ms = args.base_ms
jitter_ms = args.jitter_ms
error_every = args.error_every
max_requests = args.max_requests
event_log = Path(args.log).resolve()
event_log.parent.mkdir(parents=True, exist_ok=True)
event_log.write_text("", encoding="utf-8")
print(f"fixture_version={FIXTURE_VERSION}", flush=True)
print(f"listen=http://127.0.0.1:{args.port}", flush=True)
print(json.dumps(snapshot(), sort_keys=True), flush=True)
ThreadingHTTPServer((args.host, args.port), Handler).serve_forever()
if __name__ == "__main__":
main()
Only build_id and base_ms change between
A/B. Jitter is deterministic from build/run/sequence, so repetition
variation is visible but bounded. /health//stats
expose build/config/request counts; target JSONL proves the active
build and configured delay.
3. Write policy before measurements
policy/performance-policy.json:
{
"schema_version": 1,
"policy_id": "p34-local-policy-v1",
"service": "p34-governance-fixture",
"owner": "performance-working-group",
"workload_version": "p34-workload-v1",
"environment_id": "local-loopback-java17-jmeter563",
"required_repetitions": 3,
"required_samples_per_run": 20,
"absolute_slo": {
"p95_ms_max": 90,
"error_rate_pct_max": 1.0,
"throughput_rps_min": 10.0
},
"relative_budget": {
"p50_regression_pct_max": 15.0,
"p95_regression_pct_max": 12.0,
"error_rate_increase_pp_max": 0.5,
"throughput_drop_pct_max": 20.0
},
"comparison": {
"summary_reducer": "median_across_run_summaries",
"percentile_algorithm": "nearest_rank_v1",
"sample_label": "GovernedWork"
},
"baseline_change": {
"mode": "fixed_versioned",
"requires_review": true,
"reason_required": true
},
"exception_policy": {
"requires_owner": true,
"requires_approver": true,
"requires_issue": true,
"max_days": 7,
"scope_must_name_failed_metric": true
},
"retention": {
"raw_local_days": 7,
"trend_summary": "indefinite_for_lab_history"
}
}
The absolute SLO and relative budget are declared before B is measured. A baseline update requires review/reason; exceptions can last at most seven days and must name failed metrics.
4. Author the governed workload
config/governance.properties:
target.host=127.0.0.1
target.port=8034
threads=2
loops=10
pacing.ms=40
connect.timeout.ms=500
response.timeout.ms=1500
jmeter.httpsampler=HttpClient4
httpclient4.retrycount=0
jmeter.save.saveservice.output_format=csv
jmeter.save.saveservice.print_field_names=true
jmeter.save.saveservice.timestamp_format=ms
jmeter.save.saveservice.time=true
jmeter.save.saveservice.label=true
jmeter.save.saveservice.response_code=true
jmeter.save.saveservice.response_message=true
jmeter.save.saveservice.thread_name=true
jmeter.save.saveservice.successful=true
jmeter.save.saveservice.bytes=true
jmeter.save.saveservice.sent_bytes=true
jmeter.save.saveservice.thread_counts=true
jmeter.save.saveservice.latency=true
jmeter.save.saveservice.connect_time=true
jmeter.save.saveservice.assertion_results_failure_message=true
jmeter.save.saveservice.response_data=false
jmeter.save.saveservice.response_data.on_error=false
jmeter.save.saveservice.samplerData=false
jmeter.save.saveservice.responseHeaders=false
jmeter.save.saveservice.requestHeaders=false
jmeter.save.saveservice.url=false
JMX tree:
Test Plan
├── HTTP Request Defaults
│ host=${__P(target.host,127.0.0.1)}
│ port=${__P(target.port,8034)}
│ implementation=HttpClient4
│ connect timeout=${__P(connect.timeout.ms,500)}
│ response timeout=${__P(response.timeout.ms,1500)}
└── Thread Group — governed workload
threads=${__P(threads,2)}
loops=${__P(loops,10)}
Action after Sampler error=Continue
├── Counter -> SEQ (per user)
└── HTTP Request — GovernedWork
GET /work?run_id=${__P(run.id,p34)}&seq=T${__threadNum}:${SEQ}
├── Constant Timer ${__P(pacing.ms,40)} ms
└── Response Assertion: response code = 200
Per-run contract:
- exactly 2×10 = 20 configured samples
- label GovernedWork
- same workload_version p34-workload-v1
- raw CSV JTL + matching jmeter.log
- 3 repetitions per build
Constant Timer pacing is not part of JMeter elapsed time. The closed two-thread model means slower response can reduce achieved throughput; the report states that limitation rather than treating throughput as a pure service-rate measure.
5. Analyze each raw JTL transparently
tools/analyze_run.py:
import argparse, csv, json, math
from pathlib import Path
def nearest_rank(values, pct):
data = sorted(values)
if not data:
return 0
rank = max(1, math.ceil(len(data) * pct / 100.0))
return data[rank - 1]
def main():
p = argparse.ArgumentParser()
p.add_argument("--jtl", required=True)
p.add_argument("--label", default="GovernedWork")
p.add_argument("--build-id", required=True)
p.add_argument("--run-id", required=True)
p.add_argument("--workload-version", default="p34-workload-v1")
p.add_argument("--environment-id", default="local-loopback-java17-jmeter563")
p.add_argument("--configured-samples", type=int, default=20)
p.add_argument("--generator-valid", choices=["true","false"], default="true")
p.add_argument("--out", required=True)
args = p.parse_args()
rows = list(csv.DictReader(Path(args.jtl).open(newline="", encoding="utf-8")))
rows = [r for r in rows if r.get("label") == args.label]
elapsed = [int(float(r["elapsed"])) for r in rows]
failures = sum(r.get("success","").lower() != "true" for r in rows)
if rows:
start = min(int(r["timeStamp"]) for r in rows)
end = max(int(r["timeStamp"]) + int(float(r["elapsed"])) for r in rows)
span_s = max((end - start) / 1000.0, 0.001)
else:
span_s = 0.0
result = {
"schema_version": 1,
"summary_algorithm": "nearest_rank_v1",
"build_id": args.build_id,
"run_id": args.run_id,
"workload_version": args.workload_version,
"environment_id": args.environment_id,
"sample_label": args.label,
"configured_samples": args.configured_samples,
"achieved_samples": len(rows),
"generator_valid": args.generator_valid == "true",
"p50_ms": nearest_rank(elapsed, 50),
"p95_ms": nearest_rank(elapsed, 95),
"p99_ms": nearest_rank(elapsed, 99),
"error_rate_pct": round(100.0 * failures / len(rows), 4) if rows else 100.0,
"throughput_rps": round(len(rows) / span_s, 4) if span_s else 0.0
}
result["valid"] = (
result["generator_valid"]
and result["achieved_samples"] == result["configured_samples"]
)
Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps(result, indent=2))
if __name__ == "__main__":
main()
The script filters one label, requires configured-vs-achieved count evidence, stores generator validity, and calculates p50/p95/p99 with nearest-rank. This avoids silently mixing a dashboard estimator into baseline math.
6. Reduce three repetitions by median
tools/aggregate_runs.py:
import argparse, json, statistics
from pathlib import Path
METRICS = ("p50_ms","p95_ms","p99_ms","error_rate_pct","throughput_rps")
def main():
p = argparse.ArgumentParser()
p.add_argument("summaries", nargs="+")
p.add_argument("--build-id", required=True)
p.add_argument("--out", required=True)
args = p.parse_args()
docs = [json.loads(Path(x).read_text(encoding="utf-8")) for x in args.summaries]
workload_versions = sorted({d["workload_version"] for d in docs})
envs = sorted({d["environment_id"] for d in docs})
valid = all(d["valid"] for d in docs)
aggregate = {
"schema_version": 1,
"build_id": args.build_id,
"repetitions": len(docs),
"run_ids": [d["run_id"] for d in docs],
"workload_versions": workload_versions,
"environment_ids": envs,
"all_runs_valid": valid,
"reducer": "median_across_run_summaries",
"metrics": {}
}
for metric in METRICS:
values = [float(d[metric]) for d in docs]
aggregate["metrics"][metric] = {
"values": values,
"median": round(statistics.median(values), 4),
"min": min(values),
"max": max(values)
}
Path(args.out).write_text(json.dumps(aggregate, indent=2), encoding="utf-8")
print(json.dumps(aggregate, indent=2))
if __name__ == "__main__":
main()
The governance comparison uses the median of the three run
summaries. Each run remains visible in values/min/max;
the reducer never deletes an inconvenient repetition.
7. Execute Build A baseline runs
Start Build A:
python .\fixtures\governance_fixture.py `
--host 127.0.0.1 --port 8034 `
--build-id build-1.0.0 `
--base-ms 55 --jitter-ms 5 `
--error-every 0 --max-requests 70 `
--log .\runs\baseline\target-events.jsonl
Run JMeter three times with
run.id=a1/a2/a3, each to a
separate JTL/log/dashboard. Analyze each as build-1.0.0; require20
achieved samples and generator-valid=true. Stop the fixture after60
measured requests.
8. Create immutable baseline-v1
Aggregate a1-a3, then use tools/create_baseline.py:
import argparse, hashlib, json
from datetime import datetime, timezone
from pathlib import Path
def sha256(path):
h = hashlib.sha256()
with Path(path).open("rb") as f:
for chunk in iter(lambda: f.read(1024*1024), b""):
h.update(chunk)
return h.hexdigest()
def main():
p = argparse.ArgumentParser()
p.add_argument("--aggregate", required=True)
p.add_argument("--policy", required=True)
p.add_argument("--baseline-version", required=True)
p.add_argument("--owner", required=True)
p.add_argument("--reason", required=True)
p.add_argument("--out", required=True)
args = p.parse_args()
agg = json.loads(Path(args.aggregate).read_text(encoding="utf-8"))
policy = json.loads(Path(args.policy).read_text(encoding="utf-8"))
if not agg["all_runs_valid"]:
raise SystemExit("cannot baseline invalid runs")
if agg["repetitions"] < policy["required_repetitions"]:
raise SystemExit("insufficient repetitions")
if agg["workload_versions"] != [policy["workload_version"]]:
raise SystemExit("workload version mismatch")
if agg["environment_ids"] != [policy["environment_id"]]:
raise SystemExit("environment mismatch")
result = {
"schema_version": 1,
"baseline_version": args.baseline_version,
"baseline_build_id": agg["build_id"],
"created_at_utc": datetime.now(timezone.utc).isoformat(),
"owner": args.owner,
"change_reason": args.reason,
"policy_id": policy["policy_id"],
"policy_sha256": sha256(args.policy),
"workload_version": policy["workload_version"],
"environment_id": policy["environment_id"],
"reducer": agg["reducer"],
"metrics": {k:v["median"] for k,v in agg["metrics"].items()},
"source_aggregate_sha256": sha256(args.aggregate)
}
Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps(result, indent=2))
if __name__ == "__main__":
main()
python .\tools\create_baseline.py `
--aggregate .\runs\baseline\aggregate.json `
--policy .\policy\performance-policy.json `
--baseline-version baseline-v1 `
--owner performance-working-group `
--reason "Initial approved local governance baseline" `
--out .\baselines\baseline-v1.json
Baseline creation refuses invalid runs, insufficient repetitions, workload mismatch or environment mismatch.
9. Execute Build B current runs
Restart the same fixture as Build B:
python .\fixtures\governance_fixture.py `
--host 127.0.0.1 --port 8034 `
--build-id build-1.1.0-demo `
--base-ms 64 --jitter-ms 5 `
--error-every 0 --max-requests 70 `
--log .\runs\current\target-events.jsonl
Repeat b1-b3 with the exact same JMX/properties,
threads/loops/pacing, workload version, environment ID, summary
formula and generator-valid evidence. Aggregate into
runs/current/aggregate.json.
10. Evaluate absolute + relative gate
tools/evaluate_gate.py:
import argparse, hashlib, json
from datetime import datetime, timezone
from pathlib import Path
def pct_change(current, baseline):
if baseline == 0:
return 0.0 if current == 0 else float("inf")
return 100.0 * (current - baseline) / baseline
def pct_drop(current, baseline):
if baseline == 0:
return 0.0
return 100.0 * (baseline - current) / baseline
def sha256(path):
h = hashlib.sha256()
with Path(path).open("rb") as f:
for chunk in iter(lambda: f.read(1024*1024), b""):
h.update(chunk)
return h.hexdigest()
def main():
p = argparse.ArgumentParser()
p.add_argument("--policy", required=True)
p.add_argument("--baseline", required=True)
p.add_argument("--current", required=True)
p.add_argument("--out", required=True)
args = p.parse_args()
policy = json.loads(Path(args.policy).read_text(encoding="utf-8"))
baseline = json.loads(Path(args.baseline).read_text(encoding="utf-8"))
current = json.loads(Path(args.current).read_text(encoding="utf-8"))
invalid_reasons = []
if not current["all_runs_valid"]:
invalid_reasons.append("one_or_more_current_runs_invalid")
if current["repetitions"] < policy["required_repetitions"]:
invalid_reasons.append("insufficient_repetitions")
if current["workload_versions"] != [baseline["workload_version"]]:
invalid_reasons.append("workload_version_mismatch")
if current["environment_ids"] != [baseline["environment_id"]]:
invalid_reasons.append("environment_mismatch")
if baseline["policy_id"] != policy["policy_id"]:
invalid_reasons.append("policy_id_mismatch")
cur = {k:v["median"] for k,v in current["metrics"].items()}
base = baseline["metrics"]
checks = []
def add(name, observed, limit, passed, kind):
checks.append({"name":name,"kind":kind,"observed":round(observed,4),"limit":limit,"pass":bool(passed)})
slo = policy["absolute_slo"]
add("absolute_p95_ms", cur["p95_ms"], slo["p95_ms_max"], cur["p95_ms"] <= slo["p95_ms_max"], "absolute_slo")
add("absolute_error_rate_pct", cur["error_rate_pct"], slo["error_rate_pct_max"], cur["error_rate_pct"] <= slo["error_rate_pct_max"], "absolute_slo")
add("absolute_throughput_rps", cur["throughput_rps"], slo["throughput_rps_min"], cur["throughput_rps"] >= slo["throughput_rps_min"], "absolute_slo")
budget = policy["relative_budget"]
p50_reg = pct_change(cur["p50_ms"], base["p50_ms"])
p95_reg = pct_change(cur["p95_ms"], base["p95_ms"])
err_pp = cur["error_rate_pct"] - base["error_rate_pct"]
thr_drop = pct_drop(cur["throughput_rps"], base["throughput_rps"])
add("relative_p50_regression_pct", p50_reg, budget["p50_regression_pct_max"], p50_reg <= budget["p50_regression_pct_max"], "relative_budget")
add("relative_p95_regression_pct", p95_reg, budget["p95_regression_pct_max"], p95_reg <= budget["p95_regression_pct_max"], "relative_budget")
add("relative_error_rate_increase_pp", err_pp, budget["error_rate_increase_pp_max"], err_pp <= budget["error_rate_increase_pp_max"], "relative_budget")
add("relative_throughput_drop_pct", thr_drop, budget["throughput_drop_pct_max"], thr_drop <= budget["throughput_drop_pct_max"], "relative_budget")
status = "INVALID" if invalid_reasons else ("PASS" if all(c["pass"] for c in checks) else "FAIL")
result = {
"schema_version": 1,
"evaluated_at_utc": datetime.now(timezone.utc).isoformat(),
"policy_id": policy["policy_id"],
"baseline_version": baseline["baseline_version"],
"baseline_build_id": baseline["baseline_build_id"],
"current_build_id": current["build_id"],
"status": status,
"invalid_reasons": invalid_reasons,
"checks": checks,
"failed_metrics": [c["name"] for c in checks if not c["pass"]],
"input_sha256": {
"policy": sha256(args.policy),
"baseline": sha256(args.baseline),
"current": sha256(args.current)
}
}
Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps(result, indent=2))
raise SystemExit(0 if status == "PASS" else (2 if status == "FAIL" else 3))
if __name__ == "__main__":
main()
python .\tools\evaluate_gate.py `
--policy .\policy\performance-policy.json `
--baseline .\baselines\baseline-v1.json `
--current .\runs\current\aggregate.json `
--out .\packet\gate.json
$GateExit=$LASTEXITCODE
# Expected in this controlled demo: exit 2 = technical FAIL.
The expected qualitative result is: absolute p95/error/throughput PASS; relative p50/error/throughput likely PASS; relative p95 regression exceeds12% and FAILS. If environment/workload/count validity fails, the gate returns INVALID instead—do not create a performance exception for an invalid comparison.
11. Create one scoped temporary exception
tools/create_exception.py:
import argparse, json
from datetime import datetime, timedelta, timezone
from pathlib import Path
def main():
p = argparse.ArgumentParser()
p.add_argument("--gate", required=True)
p.add_argument("--exception-id", required=True)
p.add_argument("--owner", required=True)
p.add_argument("--approver", required=True)
p.add_argument("--issue", required=True)
p.add_argument("--metric", action="append", required=True)
p.add_argument("--reason", required=True)
p.add_argument("--days", type=int, default=7)
p.add_argument("--out", required=True)
args = p.parse_args()
gate = json.loads(Path(args.gate).read_text(encoding="utf-8"))
if gate["status"] != "FAIL":
raise SystemExit("exception may only be created for a technical FAIL")
if not 1 <= args.days <= 7:
raise SystemExit("lab policy allows 1..7 days")
failed = set(gate["failed_metrics"])
scope = set(args.metric)
if not scope or not scope.issubset(failed):
raise SystemExit("exception scope must be a subset of failed metrics")
now = datetime.now(timezone.utc)
result = {
"schema_version": 1,
"exception_id": args.exception_id,
"status": "APPROVED_TEMPORARY",
"created_at_utc": now.isoformat(),
"expires_at_utc": (now + timedelta(days=args.days)).isoformat(),
"owner": args.owner,
"approver": args.approver,
"issue": args.issue,
"reason": args.reason,
"scope_failed_metrics": sorted(scope),
"technical_gate_status": gate["status"],
"baseline_version": gate["baseline_version"],
"current_build_id": gate["current_build_id"],
"does_not_modify_gate": True
}
Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps(result, indent=2))
if __name__ == "__main__":
main()
python .\tools\create_exception.py `
--gate .\packet\gate.json `
--exception-id P34-EXC-001 `
--owner service-owner `
--approver release-reviewer `
--issue PERF-134 `
--metric relative_p95_regression_pct `
--reason "Known local demonstration regression; remediation tracked in PERF-134" `
--days 7 `
--out .\packet\exception.json
The exception is generated from the actual failed gate and expires relative to creation time. It cannot cover a metric that did not fail.
12. Keep technical gate separate from release disposition
tools/decide_release.py:
import argparse, json
from datetime import datetime, timezone
from pathlib import Path
def parse_iso(value):
return datetime.fromisoformat(value.replace("Z","+00:00"))
def main():
p = argparse.ArgumentParser()
p.add_argument("--gate", required=True)
p.add_argument("--exception")
p.add_argument("--out", required=True)
args = p.parse_args()
gate = json.loads(Path(args.gate).read_text(encoding="utf-8"))
result = {
"technical_gate_status": gate["status"],
"release_disposition": None,
"exception_id": None,
"reason": None
}
if gate["status"] == "PASS":
result.update(release_disposition="ALLOW", reason="technical_gate_pass")
elif gate["status"] == "INVALID":
result.update(release_disposition="BLOCK_INVALID_RUN", reason="comparison_not_valid")
else:
if not args.exception:
result.update(release_disposition="BLOCK", reason="technical_gate_fail_no_exception")
else:
exc = json.loads(Path(args.exception).read_text(encoding="utf-8"))
now = datetime.now(timezone.utc)
valid = (
exc.get("status") == "APPROVED_TEMPORARY"
and exc.get("technical_gate_status") == "FAIL"
and exc.get("baseline_version") == gate["baseline_version"]
and exc.get("current_build_id") == gate["current_build_id"]
and set(gate["failed_metrics"]).issubset(set(exc.get("scope_failed_metrics", [])))
and parse_iso(exc["expires_at_utc"]) > now
)
if valid:
result.update(release_disposition="ALLOW_WITH_EXCEPTION", exception_id=exc["exception_id"], reason="approved_unexpired_scoped_exception")
else:
result.update(release_disposition="BLOCK", reason="exception_invalid_expired_or_incomplete")
Path(args.out).write_text(json.dumps(result, indent=2), encoding="utf-8")
print(json.dumps(result, indent=2))
if __name__ == "__main__":
main()
python .\tools\decide_release.py `
--gate .\packet\gate.json `
--exception .\packet\exception.json `
--out .\packet\decision.json
Expected: technical_gate_status=FAIL and
release_disposition=ALLOW_WITH_EXCEPTION. If the
exception expires, does not match the build/baseline, or fails to
cover all failed metrics, disposition becomes BLOCK.
13. Build machine/human report and append trend
tools/build_report.py:
import argparse, csv, hashlib, json
from datetime import datetime, timezone
from pathlib import Path
def sha256(path):
h = hashlib.sha256()
with Path(path).open("rb") as f:
for chunk in iter(lambda: f.read(1024*1024), b""):
h.update(chunk)
return h.hexdigest()
def main():
p = argparse.ArgumentParser()
p.add_argument("--policy", required=True)
p.add_argument("--baseline", required=True)
p.add_argument("--current", required=True)
p.add_argument("--gate", required=True)
p.add_argument("--decision", required=True)
p.add_argument("--exception")
p.add_argument("--trend", required=True)
p.add_argument("--json-out", required=True)
p.add_argument("--md-out", required=True)
args = p.parse_args()
policy = json.loads(Path(args.policy).read_text())
baseline = json.loads(Path(args.baseline).read_text())
current = json.loads(Path(args.current).read_text())
gate = json.loads(Path(args.gate).read_text())
decision = json.loads(Path(args.decision).read_text())
exception = json.loads(Path(args.exception).read_text()) if args.exception else None
cur = {k:v["median"] for k,v in current["metrics"].items()}
packet = {
"generated_at_utc": datetime.now(timezone.utc).isoformat(),
"policy_id": policy["policy_id"],
"baseline_version": baseline["baseline_version"],
"baseline_build_id": baseline["baseline_build_id"],
"current_build_id": current["build_id"],
"workload_version": policy["workload_version"],
"environment_id": policy["environment_id"],
"technical_gate_status": gate["status"],
"failed_metrics": gate["failed_metrics"],
"release_disposition": decision["release_disposition"],
"exception": exception,
"current_medians": cur,
"hashes": {
"policy": sha256(args.policy),
"baseline": sha256(args.baseline),
"current": sha256(args.current),
"gate": sha256(args.gate),
"decision": sha256(args.decision)
},
"limitations": [
"local loopback synthetic service",
"three repetitions only",
"closed two-thread workload",
"nearest-rank percentiles over 20 samples per run",
"stakeholder report rounds values; raw JSON/JTL retains source precision"
]
}
Path(args.json_out).write_text(json.dumps(packet, indent=2), encoding="utf-8")
rows = []
for check in gate["checks"]:
rows.append(f"| {check['name']} | {check['observed']:.2f} | {check['limit']} | {'PASS' if check['pass'] else 'FAIL'} |")
exc_line = "None"
if exception:
exc_line = f"{exception['exception_id']} (expires {exception['expires_at_utc']}; scope: {', '.join(exception['scope_failed_metrics'])})"
md = f"""# Performance governance report
- Policy: `{policy['policy_id']}`
- Baseline: `{baseline['baseline_version']}` / build `{baseline['baseline_build_id']}`
- Current build: `{current['build_id']}`
- Workload: `{policy['workload_version']}`
- Environment: `{policy['environment_id']}`
- Technical gate: **{gate['status']}**
- Release disposition: **{decision['release_disposition']}**
- Exception: {exc_line}
## Gate checks
| Check | Observed | Limit | Result |
|---|---:|---:|---|
{chr(10).join(rows)}
## Current median-of-run summaries
- p50: {cur['p50_ms']:.1f} ms
- p95: {cur['p95_ms']:.1f} ms
- p99: {cur['p99_ms']:.1f} ms
- error rate: {cur['error_rate_pct']:.2f}%
- throughput: {cur['throughput_rps']:.2f} req/s
## Interpretation
The technical gate is never rewritten by an exception. If an exception is approved, the release disposition changes while the original failed metric, baseline version, current build, owner/approver, issue and expiry remain auditable.
## Validity limitations
- Local loopback synthetic service.
- Three repetitions only.
- Closed two-thread workload; response time affects achieved throughput.
- 20 samples/run, so tail percentiles are coarse.
- Rounded stakeholder values are for communication, not recalculation; raw JTL/JSON remain authoritative.
"""
Path(args.md_out).write_text(md, encoding="utf-8")
trend_path = Path(args.trend)
new_file = not trend_path.exists()
trend_path.parent.mkdir(parents=True, exist_ok=True)
with trend_path.open("a", newline="", encoding="utf-8") as f:
w = csv.writer(f)
if new_file:
w.writerow(["generated_at_utc","build_id","baseline_version","policy_id","workload_version","environment_id","p50_ms","p95_ms","error_rate_pct","throughput_rps","technical_gate","release_disposition","exception_id","report_json_sha256"])
w.writerow([
packet["generated_at_utc"], current["build_id"], baseline["baseline_version"], policy["policy_id"],
policy["workload_version"], policy["environment_id"], cur["p50_ms"], cur["p95_ms"],
cur["error_rate_pct"], cur["throughput_rps"], gate["status"], decision["release_disposition"],
decision.get("exception_id"), sha256(args.json_out)
])
print(json.dumps(packet, indent=2))
if __name__ == "__main__":
main()
python .\tools\build_report.py `
--policy .\policy\performance-policy.json `
--baseline .\baselines\baseline-v1.json `
--current .\runs\current\aggregate.json `
--gate .\packet\gate.json `
--decision .\packet\decision.json `
--exception .\packet\exception.json `
--trend .\trend\history.csv `
--json-out .\packet\report.json `
--md-out .\packet\report.md
The report rounds stakeholder values, keeps input hashes, stores validity limitations, and appends one immutable trend row containing both technical gate and release disposition.
14. Challenge
A current run shows p95 8% slower than baseline but comes from a different environment ID with faster CPU and different JMeter properties. Should the gate PASS because8% is inside the12% budget?
No. The comparison is INVALID before percentage math becomes authoritative. Reproduce current/baseline under the approved environment/workload identity or create a reviewed new baseline protocol.
Knowledge check
Why create policy before Build B?
Post-hoc thresholds can be chosen to force the desired result and are not governance.
What does baseline-v1 store besides metrics?
Build/workload/environment identity, owner/reason, policy hash, source aggregate hash, reducer and timestamp.
Can an exception turn gate.json from FAIL to PASS?
No. It changes release disposition only; gate.json stays immutable.
What causes INVALID rather than FAIL?
Comparison-contract failures such as workload/environment mismatch, invalid runs, insufficient repetitions or policy mismatch.
Why append both gate and release disposition to trend history?
Future reviewers must distinguish measured policy outcome from a human-approved exception decision.
Official references and version notes
- Apache JMeter downloads — current stable JMeter 5.6.3 and Java 8+ requirement.
- JMeter current changes — Java 17+ recommendation for the 5.6.x line.
- JMeter Getting Started — non-GUI/CLI execution and result/log flags.
- JMeter Dashboard Report — CSV requirements, p90/p95/p99 defaults, report generation and percentile-estimator caveat.
- JMeter Properties Reference — result-save fields, aggregate percentile properties and reporting properties.
- JMeter Remote Testing — same-plan fan-out, exact JMeter parity, Java/data requirements and controller overhead.
Version-sensitive behavior was rechecked against current primary
documentation on 2026-09-06. Mandatory runtime:
Apache JMeter 5.6.3, Java 17, no third-party
plugin, Python 3 standard library only. Meaningful runs use CLI
with raw CSV JTL plus matching jmeter.log; HTML
dashboards are corroborating evidence. JMeter's dashboard defaults
to configurable p90/p95/p99 and can estimate percentiles
differently from other reports, especially with few samples. For
governance math this chapter therefore uses one explicit
nearest-rank formula over raw, label-filtered CSV JTL and stores
that formula/version in every summary. The mandatory local policy
requires three repetitions per build, exact workload/environment
identity, configured-versus-achieved sample checks and
generator-validity notes before a gate can be evaluated.
Remote/cloud/paid CI is optional only; if later used, every
environment/engine must carry the same workload/policy identity
and valid runtime evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.