Test Reports, Coverage, JUnit, Code Quality, Browser Performance, Accessibility, and Pipeline Feedback: Diagnostics, Failure Modes, Security, and Performance
Diagnose false-green jobs, malformed or missing reports, misleading coverage, wrong-SHA evidence, oversized reports, and report-access risks from preserved pipeline/job/report evidence.
Learning objectives
- Diagnose a green job whose JUnit report still contains failures without hiding the original cause.
- Distinguish malformed/missing parser input from tool failure and UI timing.
- Detect misleading coverage caused by regex, exclusions, denominator changes, or wrong-SHA evidence.
- Control large-report cost with schema-aware splitting rather than dropping evidence.
- Protect report artifacts from leaking sensitive paths, URLs, traces, or other data.
1. Evidence-first diagnostic sequence
Start with the original pipeline and job IDs. Do not rerun first.
Capture CI_PIPELINE_SOURCE, ref,
CI_COMMIT_SHA, compiled configuration, report
declarations, tool command/version, exit code, runner context, raw
report checksum, artifact metadata, and the exact GitLab UI/parser
symptom. Then isolate the layer.
| Layer | Question | Evidence |
|---|---|---|
| Configuration | Did the compiled job declare the correct report type/path? |
Merged YAML, rules,
artifacts:reports declaration.
|
| Execution | Did the tool actually fail/pass as policy intended? | Original trace and exit status; no blind retry. |
| Report generation | Was the file created, complete, valid, and for this SHA? | Path, checksum, schema validation, embedded metadata if available. |
| Upload/retention | Did Runner upload the report and raw artifact? | Job artifact metadata and expiry/access settings. |
| GitLab parsing | Did the declared parser accept the schema? | Tests/report UI presence, parser errors/warnings where exposed. |
| Feedback/policy | Is the widget advisory or blocking by design? | Job status, allow-failure/gate settings, governance documentation. |
2. Failure mode: green job, failed JUnit test
This is the canonical chapter failure because it proves the dual-channel model. The intentionally broken job preserves a JUnit file that contains one failed case but neutralizes the test runner's exit code:
broken_test:
image: python:3.12-alpine
script:
- python tools/run_checks.py || true
artifacts:
paths: [reports/]
reports:
junit: reports/junit.xml
GitLab can parse and display the failed case while the shell returns success. Do not “repair” this by editing the XML. Preserve the pipeline/job/report checksum, remove the exit-code suppression, and rely on report upload behavior to keep the evidence available even when the job fails.
3. Failure mode: malformed or silently missing report
A job can succeed yet produce a file that GitLab cannot parse, or produce no file at the declared path. Diagnose locally before changing GitLab settings:
set -eu
ls -l reports/junit.xml reports/cobertura.xml
sha256sum reports/junit.xml reports/cobertura.xml
python - <<'PY2'
import xml.etree.ElementTree as ET
for f in ('reports/junit.xml','reports/cobertura.xml'):
try:
ET.parse(f)
print('valid', f)
except Exception as exc:
print('invalid', f, type(exc).__name__, str(exc))
raise
PY2
For JUnit, confirm filenames end in .xml, paths point
to files/patterns rather than an unsupported directory-only
declaration, and the report stays within documented size limits. For
Code Quality, validate JSON shape and required fields; a UTF-8
byte-order mark can also prevent parsing.
4. Failure mode: coverage percentage is technically correct but misleading
A coverage regex can match the wrong number, a tool upgrade can change output format, or project configuration can exclude difficult code. Preserve the raw tool output and report. Test the regex against the actual version, then inspect the Cobertura/JaCoCo denominator and changed-line annotations.
grep -n 'TOTAL' reports/tool-output.txt
# Compare the exact regex behavior outside GitLab using your language/tool of choice.
python - <<'PY2'
import re
text=open('reports/tool-output.txt',encoding='utf-8').read()
pat=re.compile(r'TOTAL .* ([0-9]+(?:\.[0-9]+)?)%$' , re.M)
m=pat.search(text)
print('coverage_match=', m.group(1) if m else None)
PY2
Do not raise a threshold just because the percentage looks low, and do not lower it because a pipeline blocks. First understand what the metric includes and whether the test suite meaningfully exercises critical behavior.
5. Failure mode: valid report from the wrong SHA
The file can be perfectly valid and still be wrong evidence. Common
causes include downloading “latest successful” artifacts by branch,
reusing a workspace without provenance, or manually copying an old
report. Compare the current CI_COMMIT_SHA with any
source metadata in the report/evidence packet and with the producer
job/pipeline ID. Prefer exact job IDs or same-pipeline
needs relationships over moving branch aliases.
If the report format cannot embed the SHA, create a sibling evidence file containing SHA, pipeline ID, job ID, tool version, and report checksum. Verification should fail if any binding differs.
6. Failure mode: huge report slows or fails ingestion
Do not solve report-size pressure by truncating failures silently. Use the report type's supported sharding/aggregation. JUnit documents 30 MB per file and 100 MB total per job. Cobertura visualization documents a 10 MiB file limit and a 100-source-node limit. Split by test shard/module while keeping unique identities, or reduce redundant data at the tool-export stage.
| Symptom | Likely layer | Correction |
|---|---|---|
| Upload rejected | Artifact/report size or instance limit | Measure file size; split according to supported report semantics. |
| Coverage lines missing | Cobertura path/source-node constraints |
Inspect <sources>, relative paths, and
report limits.
|
| Slow review UI | Very large finding/test set | Partition evidence and prioritize actionable summaries without discarding raw data. |
| One shard overwrites another | Artifact naming/path design | Use unique shard filenames/job artifacts. |
7. Security and privacy: quality reports can leak data
Reports may contain stack traces, absolute paths, test names, URLs, snippets, screenshots, or serialized values. Never print secrets to make a parser work. Use synthetic data in disposable labs, redact at the tool source where necessary, and restrict artifact access when reports should not be broadly downloadable.
A protected variable or production credential is not justified merely because an accessibility or browser test needs a URL. Review apps/test targets should use narrowly scoped, disposable identities and non-production data.
8. Queue and feedback performance
Structured feedback has two latency components: tool execution and GitLab post-processing. Optimize them separately. Parallelize independent tests only within runner capacity (Chapter 17), retain unique shard reports, and avoid generating a huge report that outweighs the runtime saved by sharding. For MR feedback, place fast deterministic checks early and move slower advisory analysis behind clear expectations.
9. Interpret one broken case end to end
| Observed fact | Interpretation |
|---|---|
| Job status: passed | Shell/process returned zero. |
| JUnit: 3 tests, 1 failed | Report parser accepted the file and saw a failure. |
| Raw XML checksum matches current job artifact | The UI is not reading a random stale file. |
Trace contains
python tools/run_checks.py || true
|
The failure was suppressed at the execution-policy layer. |
| Correct fix | Remove suppression, keep report upload, rerun only the affected job/pipeline scope after preserving evidence. |
Knowledge check
A JUnit widget shows failure but the job is green. Which layer should you inspect first?
Tool/shell exit semantics. JUnit does not set job status, so look for swallowed non-zero exits or a wrapper that always returns zero.
What should you do before rerunning a malformed-report job?
Preserve the original pipeline/job IDs, trace, report file/checksum, compiled report path/type, and parser symptom so the first cause is not erased.
Why can a valid Cobertura file still be wrong evidence?
It may belong to another source SHA/pipeline or use incorrect path/denominator assumptions. Syntax validity is not provenance.
How should you handle a JUnit report that exceeds documented limits?
Split evidence by supported files/shards while preserving test ownership and completeness; do not truncate failures silently.
Why is artifact access part of report security?
Raw reports can contain sensitive paths, stack traces, URLs, or data even when the MR widget shows only a summary.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. JUnit/unit-test reports, coverage percentage extraction, Cobertura/JaCoCo coverage visualization, Code Quality report import, and Accessibility reports are available on Free/Premium/Ultimate. Browser Performance reports are Premium/Ultimate. Code Quality's built-in CodeClimate-based template was deprecated in GitLab 17.3 and is planned for removal in GitLab 19.0; current guidance is to integrate a supported tool's report directly. The diagnostic lesson intentionally distinguishes report upload from job success and parser success. Current GitLab docs state that report artifacts are uploaded regardless of job success/failure, while JUnit contents do not affect job status.
- Unit test reports — official reference.
- Unit test report examples — official reference.
- Code coverage — official reference.
- Coverage reporting — official reference.
- Coverage visualization — official reference.
- Cobertura coverage visualization — official reference.
- Code Quality — official reference.
- Accessibility testing — official reference.
- Browser performance testing — official reference.
- Artifacts reports types — official reference.
- Job artifacts — official reference.
- CI/CD YAML syntax reference — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.