Chapter 33Lesson 03~195 minutes

Testing and Quality Pipelines: JUnit, Coverage, SonarQube, Selenium, JMeter, and Quality Gates: Configuration, Design Choices, and Tradeoffs

Turn quality checks into explicit contracts: decide what fails immediately, what is merely ingested, what waits asynchronously, what is quarantined, and which evidence is authoritative.

DesignThresholdsAsync gatesFlaky testsPlugin governanceTradeoffs

Learning objectives

  • Choose deliberately between test-command failure and report-driven policy.
  • Compare webhook-backed external gates with synchronous/polling designs.
  • Separate independent evidence producers instead of building one opaque “quality” stage.
  • Define a defensible flaky-test policy that preserves failures.
  • Review plugin/tool version, security and trust implications before adoption.

1. Design principle: evidence first, policy second

Good quality architecture lets engineers answer two questions independently: what did the tool observe? and what did our delivery policy decide? When execution, ingestion and policy are fused into one shell command, every failure looks identical and teams start fixing symptoms instead of causes.

2. Decision table

Decision Option A Option B Prefer when
Test failure Fail immediately on runner exit Capture exit, publish report, then fail Use capture-then-publish when you need retained failure evidence; never swallow the failure.
Coverage Runner/CI script threshold Jenkins Coverage quality gate Script gate is portable; Jenkins gate gives native history/UI. Keep one authoritative policy owner.
Sonar gate Scanner waits synchronously Webhook-backed waitForQualityGate Webhook wait avoids holding executors and retains external task identity.
Pipeline shape One giant quality stage Separate execution/ingestion/gate stages Separate stages make ownership and first failure clearer.
Flakiness Blind retry Quarantine + history + root-cause Retry only narrowly for known transient infrastructure failure; preserve first attempt.
Browser/load evidence Always blocking Risk-tiered/advisory then blocking Block when environment/stability and ownership justify it; otherwise publish evidence without false certainty.

3. Test command versus report threshold

The test runner is authoritative for execution errors such as import failures, crashes or assertions. The JUnit publisher is authoritative for the parsed XML state Jenkins retained. A coverage publisher evaluates metrics from a different report. Do not configure all three to mutate the build result in contradictory ways without a documented precedence.

A practical pattern is: capture runner status → publish reports → evaluate/report native gates → apply one final policy stage. This makes “runner crashed,” “report missing,” “tests failed,” and “coverage threshold missed” independently visible.

4. Coverage gate design

Choose metric, baseline and consequence explicitly. Whole-project line coverage, branch coverage and changed-code coverage answer different questions. A 1% change in a generated file should not necessarily block the same way as untested security-sensitive new logic.

Control Question to answer Risk if vague
Metric LINE, BRANCH, MUTATION, etc.? Teams optimize a number without knowing what it represents.
Baseline Whole project, changed files, modified lines? Legacy code dominates or new regressions hide in averages.
Threshold Why this value? Arbitrary percentage becomes cargo-cult policy.
Consequence Inform, UNSTABLE, FAIL? A gate silently changes release behavior.
Owner Jenkins, build tool, or external analyzer? Two sources disagree and developers cannot predict outcome.

5. Synchronous versus webhook-backed quality gates

Polling from an agent wastes executor capacity and can lose context across restarts. The maintained SonarQube Jenkins integration documents a webhook-backed waitForQualityGate step. It records the analysis task through withSonarQubeEnv, then waits for server state without needing to keep a node allocated.

6. One giant stage versus independent evidence

// Hard to diagnose: everything collapses into one shell status.
stage('Quality') {
  steps { sh './test && ./coverage && ./sonar && ./browser && ./load' }
}

// Easier to reason about: distinct evidence and policy transitions.
stage('Unit test execution') { /* runner + raw reports */ }
stage('Jenkins report ingestion') { /* junit + recordCoverage */ }
stage('Static analysis submission') { /* Sonar task */ }
stage('External quality gate') { /* waitForQualityGate */ }
stage('Browser smoke') { /* controlled target */ }
stage('Load smoke') { /* controlled target */ }
stage('Quality policy') { /* aggregate reviewed outcomes */ }

More stages are not automatically better. Separate only when the states, ownership, retry behavior or evidence differ.

7. Flaky tests: retry is not a diagnosis

A retry can be appropriate for a narrowly understood infrastructure transient—for example, a disposable browser startup race with preserved first-attempt logs. It is not a generic wrapper around failing assertions. Track flaky history, owner, quarantine expiry and first-failure evidence.

Pattern Why it fails Safer design
retry(3) around all tests Deterministic regressions look intermittent and runtime triples. No blanket retry; preserve failure and classify cause.
Ignore known flaky test forever Permanent blind spot. Quarantine with owner, issue, expiry and visible non-blocking evidence.
Rerun only failures and report final attempt Erases first failure frequency. Publish initial and rerun attempts separately.
Mark all UI failures flaky Hides real browser/product regressions. Capture screenshot/network/browser evidence and classify.

8. Plugin and tool governance

Quality plugins execute inside Jenkins or deeply influence build status, so treat them as platform dependencies.

Component Dated baseline Governance note
JUnit plugin 1425.v9c7318dca_96d Requires Jenkins 2.541.3; review dependencies and prior advisories before upgrade.
Coverage plugin 3.3373.v2dc29d8d7a_d0 Requires Jenkins 2.555.3; current releases supersede older vulnerable versions.
SonarQube Scanner plugin 2.19.0 Requires Jenkins 2.479.3; 2.18.3 and earlier are listed with a stored-XSS warning on the current plugin page.
Old Quality Gates plugin 2.5 Avoid for this course: current plugin page reports unresolved credentials-transmitted-in-plain-text issue and it is a decade old.
Performance plugin 1018.ve732b_08d5600 Optional for JMeter report ingestion; raw JTL remains primary evidence.
Selenium 4.49 External browser automation; pin language binding/browser assumptions in the test environment.
JMeter 5.6.3 External load tool; run CLI mode for load tests and control target identity.

9. Trust boundaries for quality jobs

Untrusted repository code should not gain broad Sonar administration tokens, privileged browsers, internal-only network access or credentials simply because it is “just testing.” Browser and load agents can reach targets; scanner tokens can read/write project analysis; report parsers process repository-controlled files.

  • Separate untrusted PR agents from privileged internal quality agents.
  • Scope Sonar tokens to the necessary project/operation.
  • Do not expose production credentials to Selenium tests from forks.
  • Guard JMeter targets and concurrency parameters.
  • Review generated HTML/XML/logs before publishing them broadly if they may contain sensitive URLs/data.

10. Worked design: protected branch release candidate

  1. Unit tests run on a disposable quality-lab agent and always emit JUnit XML.
  2. Coverage is measured in the same source checkout and parsed by Jenkins.
  3. Sonar analysis uses a project-scoped credential on a trusted branch only.
  4. The Pipeline releases the analysis agent before waiting on the quality gate.
  5. Selenium smoke targets a disposable staging URL with no production credentials.
  6. A bounded JMeter smoke load targets the same staging environment with an explicit low concurrency ceiling.
  7. The release policy requires tests PASS, report ingestion PASS, coverage gate PASS and Sonar gate OK; Selenium/JMeter may start advisory and become blocking only after stability objectives are defined.

11. Feedback latency and cost

Putting every browser matrix and load test on every commit can increase queue time and cloud/browser cost without improving useful feedback. Keep fast deterministic unit/static checks early; schedule or risk-select expensive evidence. Measure queue wait, tool runtime, external analysis latency and report-ingestion time before increasing executors or parallelism.

12. Migration from legacy publishers

Do not replace a deprecated or vulnerable quality plugin merely by installing its successor. Inventory existing job configuration, historical result URLs, API consumers, pipeline steps, thresholds and SCM checks first. Migrate a representative job, compare evidence, document new semantics, then remove the old plugin only after dependency checks and a recoverable controller snapshot.

Next

Diagnose failure by layer

Lesson 4 intentionally breaks exit-code handling, report paths, gate waits, targets and retry behavior so you can repair the causal layer without erasing evidence.

Knowledge check

Answer before revealing the explanation.

1. Why can both the build tool and Jenkins coverage gates be a problem?

2. When is retry reasonable?

3. Why avoid the legacy Quality Gates plugin in this chapter?

4. What is the main benefit of a webhook-backed gate wait?

5. Why might Selenium/JMeter begin as advisory?

Official references and version notes

Test-report schemas, Jenkins plugins, SonarQube integration and browser/load-test tooling evolve. Prefer current primary documentation.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.