Test Reports, Coverage, JUnit, Code Quality, Browser Performance, Accessibility, and Pipeline Feedback: Configuration, Design Choices, and Tradeoffs
Choose deliberately between blocking and advisory checks, monolithic and sharded reports, UI summaries and raw evidence, coverage thresholds and contextual trends, with explicit tier and retention tradeoffs.
Learning objectives
- Choose when a quality check should block and when it should remain advisory.
- Design per-shard versus consolidated reports based on the report type's actual aggregation semantics.
- Use UI summaries for fast review while preserving raw artifacts for investigation and audit.
- Interpret coverage as contextual evidence rather than a single target that can be optimized blindly.
- Account for tier, retention, artifact access, child-pipeline boundaries, and parser timing in report architecture.
1. Start from the evidence contract
Before selecting a widget or threshold, write down the contract: which source SHA is measured, which tool/version generated the result, what file/schema GitLab parses, what tool exit status means, who owns a failure, how long raw evidence remains available, and whether the result is advisory or blocking. This avoids building policy around UI behavior that may have different tier or aggregation semantics.
Keep CI_PIPELINE_SOURCE and
CI_COMMIT_SHA in the evidence packet so the report
cannot be separated from the pipeline context that generated it.
2. Blocking versus advisory
| Design | Use when | Benefits | Risks / controls |
|---|---|---|---|
| Blocking tool exit | A deterministic failure means delivery must stop now. | Simple and visible; job status carries policy. | Flaky or noisy tools can halt delivery. Define ownership and retry policy. |
| Advisory report | Finding requires human/contextual judgment. | Provides signal without unnecessary blockage. | Can be ignored. Define review expectations and escalation. |
| Blocking threshold | A numeric/quality boundary has a defensible denominator and stable tool semantics. | Automates agreed quality floor. | Threshold gaming and abrupt cliffs. Preserve raw metric context and trend. |
| Approval/policy layer | A reviewer or governance rule should interpret evidence. | Separates evidence generation from authorization. | Tier/role complexity; must preserve decision audit trail. |
Do not use allow_failure: true reflexively to “make the
pipeline green.” If a tool is advisory, say so explicitly and
preserve the finding. If it is blocking, do not neutralize its exit
code merely to keep a widget visible—report artifacts are uploaded
independently of job success.
4. UI summary versus raw artifact
The UI is optimized for review speed: failed test names, coverage deltas, changed-line annotations, quality findings. Raw evidence is optimized for reconstruction: full XML/JSON, tool version, parser input, checksums, and sometimes screenshots or HTML. Use both when the evidence matters.
However, storing the same report under both
artifacts:reports and artifacts:paths has
storage implications. Define a retention window and access level
appropriate to the data. If a report contains sensitive URLs or
stack traces, restrict artifact access rather than assuming the MR
widget is the only exposure path.
5. Coverage threshold versus trend and context
Coverage is a measurement of exercised code according to a specific tool and denominator. It is not a direct measurement of correctness. A project can raise coverage while removing meaningful assertions, excluding difficult files, or testing trivial branches. Treat a threshold as one guardrail among several.
| Question | Evidence to inspect |
|---|---|
| Did total coverage change? | The percentage extracted by the exact regex/tool version. |
| Which changed lines are untested? | Cobertura/JaCoCo diff annotations tied to the MR SHA. |
| Did the denominator change? | Raw coverage report/configuration and source inclusion/exclusion settings. |
| Are critical paths tested? | Test names/assertions and domain-specific test plan, not percentage alone. |
| Did a shard disappear? | Per-shard report inventory and expected partition count. |
6. Tier and presentation choices
| Feature | Free path | Higher-tier extension |
|---|---|---|
| Unit test reports | JUnit ingestion, pipeline Tests, MR test summary. | Higher tiers add surrounding governance/features but are not required for core ingestion. |
| Coverage | Percentage + diff visualization. | Coverage-Check approval behavior is a higher-tier governance extension. |
| Code Quality | Import results and see MR report. | Pipeline Code Quality view: Premium/Ultimate; MR changes annotations/project summary: Ultimate-specific. |
| Accessibility | Available on Free/Premium/Ultimate. | No paid requirement for the core report widget. |
| Browser Performance | Faithful raw/synthetic performance evidence only in mandatory course path. | GitLab browser-performance MR report is Premium/Ultimate. |
7. Child pipelines and report boundaries
Do not design parent/child pipelines on the assumption that all child report artifacts combine automatically in the parent. Current GitLab documentation notes that combined parent reports from child-pipeline artifacts are not generally supported. Coverage reports from child pipelines can still appear as MR diff annotations, but the underlying artifact relationship is separate. Preserve child pipeline/job IDs in the evidence packet.
8. Retention, parser timing, and artifact access
Report processing can happen after the job finishes. Coverage visualization, for example, is processed after pipeline completion. A blocking manual job can delay when coverage annotations appear. Artifact expiry may remove raw evidence later even though a reviewer remembers the UI summary. Set retention according to review/audit needs, and document which UI elements depend on retained artifacts versus persisted parsed metadata.
Current YAML also lets projects restrict artifact access. Treat report artifacts as data assets: classify them, avoid secrets, and set access intentionally.
9. Worked decision table
| Scenario | Choice | Why | Observable evidence |
|---|---|---|---|
| Fast unit suite on every MR | Blocking JUnit job + raw XML retained 7 days | Deterministic correctness signal; failures need immediate action. | Exit code, JUnit cases, raw checksum, MR/pipeline result. |
| Linter with many legacy findings | Advisory Code Quality report at first | Avoid freezing delivery while making new issues visible. | Stable fingerprints, MR report, documented baseline/owner. |
| Large test suite across 8 shards | Per-shard JUnit + coverage; verify shard inventory | Preserves failure ownership and reduces critical path. | Eight expected shard IDs, reports, aggregate coverage context. |
| Web performance regression check on Free | Raw/synthetic metric artifact + explicit threshold script | Mandatory path cannot rely on Premium Browser Performance widget. | Tool exit code + raw metrics + threshold rationale. |
| Regulated report retention | Raw report artifact with explicit restricted access/expiry | UI alone is insufficient audit evidence. | Checksum, access policy, expiry, pipeline/job/SHA. |
10. Anti-patterns
| Anti-pattern | Why it is weak | Safer alternative |
|---|---|---|
| Make every finding blocking on day one | Noisy tools create alert fatigue and bypass pressure. | Baseline, assign owners, then tighten policy with measured false-positive rate. |
Use || true everywhere so reports upload |
Destroys job-status semantics. | Let reports upload on failure and preserve the real exit status. |
| Trust coverage percentage without raw config | Denominator/exclusions can change. | Review raw report/config and changed-line coverage. |
| Delete raw reports immediately after parsing | Removes independent verification. | Retain for a bounded review/audit period. |
| Assume child reports aggregate in parent | Current behavior differs by report type. | Design explicit collection/inspection based on documented semantics. |
Knowledge check
When is an advisory check preferable to a blocking check?
When the signal still needs contextual judgment, has meaningful false positives, or is being introduced against a legacy baseline. The advisory status should still have an owner and review expectation.
Why should a sharded test suite keep shard identity in report evidence?
It makes missing/empty shards visible and gives failures clear ownership; aggregate UI alone can hide partition gaps.
What is the main weakness of a coverage-only gate?
Coverage can be gamed or misunderstood because it measures execution, not assertion quality or correctness. The denominator/configuration and changed-line context matter.
Why is a raw artifact useful even when a Code Quality MR report exists?
It preserves the exact parser input, tool output, and checksum for later diagnosis/audit, independent of UI presentation.
Can parent pipelines always combine report artifacts from child pipelines?
No. Current GitLab documentation explicitly notes this is not generally supported; design per report type.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. JUnit/unit-test reports, coverage percentage extraction, Cobertura/JaCoCo coverage visualization, Code Quality report import, and Accessibility reports are available on Free/Premium/Ultimate. Browser Performance reports are Premium/Ultimate. Code Quality's built-in CodeClimate-based template was deprecated in GitLab 17.3 and is planned for removal in GitLab 19.0; current guidance is to integrate a supported tool's report directly. Current Code Quality documentation recommends direct integration from supported tools; the CodeClimate-based built-in template is deprecated and planned for removal in GitLab 19.0.
- Unit test reports — official reference.
- Unit test report examples — official reference.
- Code coverage — official reference.
- Coverage reporting — official reference.
- Coverage visualization — official reference.
- Cobertura coverage visualization — official reference.
- Code Quality — official reference.
- Accessibility testing — official reference.
- Browser performance testing — official reference.
- Artifacts reports types — official reference.
- Job artifacts — official reference.
- CI/CD YAML syntax reference — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.