Chapter 28Lesson 04~230 minutes

CI/CD Performance Gates in GitHub Actions, GitLab CI, and Jenkins: Diagnostics, Failure Modes, and Production Practices

A failed CI job can mean application regression, invalid workload, broken JMeter setup, noisy target, starved runner, artifact-policy error or provider syntax failure. Preserve first-failure evidence and isolate the layer before changing load or threshold.

Exit-code trapFlaky environmentArtifact failureBaseline driftSecret safety

Learning objectives

  • Apply a preserve-first diagnostic sequence to failure modes involving CI/CD Performance Gates in GitHub Actions, GitLab CI, and Jenkins.
  • Separate plan/configuration, generator, protocol/network, target, CI/container, and distributed-engine causes before changing settings.
  • Reproduce a failure with the smallest authorized workload and retain the original JTL, jmeter.log, and supporting evidence.
  • Reject shortcuts such as blanket retries, disabled verification, unbounded load increases, silent global overrides, or deleting first-failure artifacts.
  • Verify the least-invasive correction under the original bounded workload before declaring the problem resolved.

1. Preserve first-failure evidence

Runnable diagnosis remains ≤100 samples on 127.0.0.1:8028. Preserve JTL, matching jmeter.log, dashboard, launcher/gate JSON, policy/baseline, provider console/config, runner/JVM state and target events. Do not delete the failed run or increase traffic.

2. Diagnostic sequence

CI gate failure diagnosis

The CI provider orchestrates a portable performance policy; it does not replace the JMeter engine, raw result evidence, target telemetry, or the evaluator.

flowchart TD
E[Preserve JTL/log/gate/provider/target evidence] --> V[Confirm JMeter/Java/plugin/evaluator/provider versions]
V --> C[Confirm JMX/data/properties/CLI + authorized target]
C --> S[Validate tree scope/resolved properties/configured load/policy]
S --> P[Inspect protocol/session/data/assertion state]
P --> G[Inspect runner JVM/CPU/network/disk/container state]
G --> T[Inspect target telemetry/environment noise]
T --> CI[Inspect timeout/concurrency/artifacts/secrets/exit propagation]
CI --> F[Least-destructive correction]
F --> R[Small controlled rerun]

3. Intentionally broken: JMeter exit code treated as SLO

jmeter -n -t plans/ci-gate.jmx -l results.jtl -e -o dashboard
if [ $? -eq 0 ]; then
  echo "PERFORMANCE PASSED"
fi

Against degraded mode this can print PASS even though valid p95 is ~180 ms. Repair by preserving evidence and applying the common evaluator. Do not bury CI policy inside JSR223 or provider-specific JMX branches.

4. Flaky shared environment

If identical code alternates pass/fail while target telemetry shows unrelated deployments/noisy neighbors, retry-until-green samples the noise. Isolate/reserve the target, move the suite to scheduled/informational, or keep PR blocking on the deterministic fixture.

5. Huge load on every commit

A long, high-thread capacity test on each push burns runner capacity and repeatedly stresses shared systems. Cadence is a safety boundary. Keep PR smoke small; schedule/load-gate larger suites under explicit authorization and concurrency controls.

6. Success-only artifacts

Broken GitHub: upload step lacks always(). Broken GitLab: default success-only artifacts. Broken Jenkins: archive occurs only inside a successful stage. Repair with each provider's unconditional/finally mechanism and deliberately test a failing gate.

7. Baseline drift

# anti-pattern
if gate_failed:
    baseline["p95_ms"] = current_p95
    save_and_commit(baseline)

The next run passes because policy moved. Baseline updates require a reviewed reference/environment and preserved old/new evidence.

8. Secrets in CI logs

Do not put real tokens/passwords in echoed CLI -J values, JMX or artifacts. Use provider secret stores/least-privilege contexts and avoid printing values. If a real secret leaks, rotate/revoke it. This lab uses no real credentials.

9. Provider-specific logic inside JMX

A JSR223 sampler checking GITHUB_ACTIONS or constructing GitLab/Jenkins artifact paths couples orchestration to load generation. Keep JMX neutral; providers only invoke the common launcher and store its output.

10. Runner saturation

A smaller/throttled CI runner can inflate latency or cap throughput while the SUT is unchanged. Chapter 26 still applies: inspect runner CPU/GC/disk/network and classify the experiment invalid if generator constraints prevent the intended workload.

11. Causal table

Observation Likely layer Evidence
Engine=0, gate=10, 100/100 valid Application SLO gate.json/JTL/target timing.
Engine=20, partial/no JTL JMeter/runtime/JMX launcher/jmeter.log/console/target events.
Same commit alternates Shared target/runner noise Target/runner telemetry and reservations.
Fast p95 but 62/100 Invalid workload/timeout JTL/target counts and provider timeout.
No artifacts on failure CI artifact condition Provider config/job metadata.
Threshold suddenly easier Policy/baseline drift Git diff/approval/history.

12. Shortcuts to reject

  • Do not add blanket retries or arbitrary long sleeps.
  • Do not add giant heaps without runner/JVM evidence.
  • Do not mass-disable evidence/listeners without evidence of their cost.
  • Do not use global property hacks to fake success.
  • Do not disable TLS/RMI verification.
  • Do not reproduce on production/public targets.
  • Do not increase workload while validity is unresolved.
  • Do not delete failed JTL/log/dashboard/gate/provider evidence.

Knowledge check

Engine zero and gate ten means what?

Why avoid retry-until-green?

How do you test artifact policy?

What if a real token appears in CI logs?

Why keep provider APIs out of JMX?

Next lesson

Checkpoint: fail, preserve, restore, pass

Lesson 5 verifies the exact gate semantics and then maps the common command into all three CI providers.

Official references and version notes

Version and compatibility note

Statements were rechecked against current primary documentation on 2026-09-05. The mandatory runtime is Apache JMeter 5.6.3 with Java 17; no third-party JMeter plugin is required. The provider-neutral provisioner verifies Apache's published SHA-512 before installing JMeter. Meaningful execution stays in CLI mode and produces the HTML dashboard using -e -o. The gate evaluator is Python-standard-library only and calculates p95 from raw CSV JTL using an explicitly documented nearest-rank method. The dashboard remains diagnostic evidence rather than the policy parser. GitHub's minimal pattern uses actions/checkout@v6, stable actions/setup-java@v5 with Temurin 17, and actions/upload-artifact@v4 guarded by always(). GitLab explicitly sets artifacts:when: always, and Jenkins archives in Declarative post { always { ... } }. Provider YAML/Groovy stays thin; performance formulas and exit classification live in the shared launcher/evaluator.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.