CI/CD Performance Gates in GitHub Actions, GitLab CI, and Jenkins: Core Concepts and Mental Model
Chapter 27 made the injector reproducible. Earlier chapters made its result collection, dashboard, telemetry, distributed load and generator capacity auditable. CI/CD adds one new responsibility: convert validated evidence into reviewable policy without confusing the health of the JMeter command with the performance of the application.
Learning objectives
- Explain pipeline job → pinned JMeter → JTL/dashboard → evaluator → CI status/artifacts.
- Separate engine failure, sample/assertion failure, invalid workload and SLO regression.
- Define absolute SLO, baseline and regression budget.
- Inspect runtime, workload, target, policy and artifacts before changing a gate.
- Keep provider syntax thin while the common policy remains portable.
1. Why JMeter exit alone is not a performance gate
A CLI run can complete successfully and produce a complete JTL while p95 violates an application objective. A process can also fail before useful samples exist. Therefore “JMeter returned zero” answers only whether the engine/command path completed normally; it does not answer whether the application met a latency/error objective.
127.0.0.1:8028. Each required run is 5 threads ×20
loops =100 samples with 100 ms pacing and ≤25 seconds duration. No
production/shared target, real credential, paid CI, cloud cluster,
remote engine or unbounded retry is required.
Never substitute a public/shared/production endpoint.
2. Mental model
The CI provider orchestrates a portable performance policy; it does not replace the JMeter engine, raw result evidence, target telemetry, or the evaluator.
flowchart TD P[Pipeline/local launcher] --> V[Java17 + JMeter5.6.3] V --> J[JMeter CLI + JMX/properties] J --> T[Authorized target] J --> R[JTL + jmeter.log + dashboard] T --> E[Target events] R --> G[Threshold evaluator] E --> G S[Reviewed SLO / optional baseline] --> G J --> X[Engine exit] G --> Q[Gate exit] X --> L[Launcher classification] Q --> L L --> C[CI job status] R --> A[Always-retained evidence] E --> A G --> A
JMeter owns load execution and raw results; the target independently proves achieved work. The evaluator validates the sample set before applying a documented p95/error policy. The launcher keeps engine and gate exit codes separate and CI treats any nonzero final code as blocking. Artifacts are retained independently of status.
3. Four independent signals
| Signal | Example | Interpretation |
|---|---|---|
| Engine/command | Bad JMX, Java missing, dashboard path invalid. | Execution infrastructure failed. |
| Sample/assertion |
HTTP 500 or failed assertion gives
success=false.
|
Protocol/business request failed. |
| Validity | Only 73/100 samples reach JTL/target. | Do not judge SLO; workload was not achieved. |
| SLO/regression | 100 valid samples, 0% errors, p95=181 ms with max 120 ms. | Valid experiment; application policy failed. |
4. Policy vocabulary
An absolute SLO is a fixed requirement such as p95≤120 ms and errors≤1%. A baseline is an approved reference result/environment. A regression budget is the allowed degradation from that approved reference. Baselines are governance artifacts: do not auto-update them after a failure.
5. Percentile formula is part of policy
The evaluator computes nearest-rank p95 directly from CSV JTL. The HTML dashboard remains human diagnostic evidence. This avoids scraping rendered pages and makes the exact gate algorithm auditable even when another JMeter report computes percentiles with a different estimator.
6. Exit-code contract
| Exit | Meaning |
|---|---|
| 0 | Valid run and enabled policies pass. |
| 10 | Valid run; SLO/regression policy fails. |
| 11 | Invalid run: configured and achieved evidence disagree. |
| 20 | JMeter engine/command failed. |
| 21 | Evaluator failed unexpectedly. |
| 22 | Preflight/runtime/fixture/version setup failed. |
7. Failure evidence must survive
Retain raw JTL, matching jmeter.log, generated
dashboard, gate/launcher JSON, policy/input manifest, provider
console and target events. A performance failure without evidence
becomes an expensive retry request instead of an actionable
diagnosis.
8. Read-only inspection before policy changes
java -version
"$JMETER_HOME/bin/jmeter" -v
python3 --version
python3 tools/evaluate_gate.py --help
python3 tools/run_performance_gate.py --help
cat policy/slo.json
cat policy/baseline.json
git diff -- plans/ config/ policy/ tools/
Also inspect the provider job trigger, timeout, runner/agent class, concurrency and artifact-on-failure logic. Do not lower thresholds or add retries before proving which state changed.
9. State inventory
| Boundary | Questions |
|---|---|
| Pipeline | What trigger/cadence/timeout/runner and artifact policy will execute? |
| JMeter | Exact Java/JMeter/plugin versions, CLI flags and input hashes? |
| Load | Threads/loops/timers and configured sample ceiling? |
| Protocol/session | Target identity, assertions, retries, keep-alive, auth/session state? |
| Target | Ephemeral/private or shared; independent events/telemetry? |
| Results | CSV schema, dashboard, logs and artifact retention? |
| Policy | SLO formula, threshold, optional reviewed baseline/budget? |
| Security | No secrets in CLI/JMX/logs; least-privilege CI secret scope if later needed? |
| Validity | Configured versus achieved work plus generator/runner headroom? |
10. DevOps connection
Performance gates are policy code. They should be deterministic, reviewable, cheap enough for the selected cadence and able to preserve evidence on failure. GitHub Actions, GitLab CI and Jenkins should orchestrate the same portable command rather than reimplementing thresholds three times.
Knowledge check
JMeter exits zero but p95 is 180 ms against a 120 ms SLO. Pass?
No. Engine execution succeeded; application performance policy failed.
Why does exit 11 exist?
To mark an invalid experiment such as a JTL/target count mismatch before SLO interpretation.
Why compute the gate from raw JTL?
It keeps one transparent formula independent of dashboard rendering and provider UI.
Why upload artifacts on failure?
First-failure evidence is needed to distinguish engine, target, runner, workload and policy causes.
Why must baseline updates be reviewed?
Otherwise the reference can drift and normalize regressions without explicit policy approval.
Official references and version notes
- Apache JMeter downloads — JMeter 5.6.3 and Java 8+ requirement.
- JMeter execution guidance — GUI for authoring/debug and CLI/non-GUI for load.
-
JMeter Dashboard Report
—
-e -o/-g -oand dashboard CSV requirements. -
GitHub Actions build/test examples
— artifact upload with
always()after test failure. -
GitHub setup-java
— stable Java setup action; example uses
v5with Temurin 17. -
GitLab CI YAML reference
—
artifacts:when: alwayssemantics. -
Jenkins Declarative Pipeline syntax
—
post/always, timeout, agent behavior. -
Jenkins tests and artifacts
—
archiveArtifactsin an always-run post block.
Statements were rechecked against current primary documentation on
2026-09-05. The mandatory runtime is
Apache JMeter 5.6.3 with Java 17; no third-party
JMeter plugin is required. The provider-neutral provisioner
verifies Apache's published SHA-512 before installing JMeter.
Meaningful execution stays in CLI mode and produces the HTML
dashboard using -e -o. The gate evaluator is
Python-standard-library only and calculates p95 from raw CSV JTL
using an explicitly documented nearest-rank method. The dashboard
remains diagnostic evidence rather than the policy parser.
GitHub's minimal pattern uses actions/checkout@v6,
stable actions/setup-java@v5 with Temurin 17, and
actions/upload-artifact@v4 guarded by
always(). GitLab explicitly sets
artifacts:when: always, and Jenkins archives in
Declarative post { always { ... } }. Provider
YAML/Groovy stays thin; performance formulas and exit
classification live in the shared launcher/evaluator.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.