Continuous Integration Patterns for Build, Test, Lint, and Coverage: Diagnostics, Failure Modes, and Production Practices
Diagnose CI failures from exact revision through runner and toolchain to test, coverage, cache and required-check evidence without hiding the original cause.
Learning objectives
- Diagnose CI from run/attempt and exact SHA before changing tests or caches.
-
Detect stale/wrong checkout, swallowed exit codes and misleading
continue-on-error. - Separate cache/build-state problems from test/toolchain problems.
- Preserve failed-run reports and artifacts before rerun or cleanup.
- Apply the smallest correction that restores the intended evidence contract.
1. Evidence-first diagnostic sequence
-
Preserve identity: run ID, attempt, event, ref,
GITHUB_SHA, workflow revision and first-failure logs. -
Confirm checkout: compare
git rev-parse HEADwith the SHA the workflow claims to test. -
Confirm evaluated configuration: job graph,
ifconditions, permissions, matrix values and cache key. - Confirm runner/toolchain: runner OS/image metadata, Python/tool versions and dependency manifest.
- Inspect the first failing command: exit code, diagnostics, test/JUnit and coverage report.
- Inspect retained services: cache-hit semantics, artifact names/digests and branch required-check state.
- Apply the least destructive repair and rerun only the smallest equivalent scope that proves it.
2. Failure mode: CI builds a different SHA than the event SHA
This failure often appears when a workflow checks out a mutable branch name after GitHub created the run. The run claims one event SHA while build steps validate whatever commit the branch points to later.
# Broken for revision-bound CI: mutable ref can advance.
- uses: actions/checkout@FULL_SHA
with:
ref: main
# Safer for ordinary event-bound CI: use the event-selected ref/SHA.
- uses: actions/checkout@FULL_SHA
- run: test "$(git rev-parse HEAD)" = "$GITHUB_SHA"
Preserve the original run log showing both values. Do not rerun
first: if main moved again, the rerun can produce a
third observation and obscure the original mismatch.
3. Failure mode: a shell pipeline swallows the test exit code
A classic example pipes test output to tee. Depending
on shell options, the pipeline can report tee success
while the test command failed. The check turns green even though the
log contains failures.
# Broken if pipeline status is not handled:
pytest -q | tee evidence/pytest.txt
# Explicitly preserve pytest's status in Bash:
set +e
pytest -q 2>&1 | tee evidence/pytest.txt
rc=${PIPESTATUS[0]}
set -e
exit "$rc"
The diagnostic clue is contradictory evidence: failed assertions in the report but a successful step/job. Fix exit-code propagation rather than weakening branch policy.
4. Failure mode: continue-on-error becomes permanent failure masking
# Broken policy: job may remain green despite test failure.
- name: Tests
continue-on-error: true
run: pytest
If temporary tolerance is needed to upload reports, capture the raw
steps.ID.outcome, upload evidence with
if: always(), then explicitly fail a final gate when
the raw outcome is not success. The purpose is preservation, not
normalization of failure.
5. Failure mode: coverage is generated after the upload step
Artifacts snapshot files that exist when the upload step runs. If
coverage.xml is generated afterward, the artifact can
be green and valid while simply not containing the report reviewers
expect. Inspect artifact contents and step timestamps rather than
assuming “upload succeeded” means “correct files uploaded.”
# Correct order
- run: coverage xml -o evidence/coverage.xml
- uses: actions/upload-artifact@FULL_SHA
with:
path: evidence/coverage.xml
6. Failure mode: cached build outputs become undeclared source
Caching package-manager downloads is usually safer than caching generated application outputs. If a build cache key omits compiler flags, source generators, environment variables or tool versions, stale output can be restored and accepted as if newly built. A clean run that fails while a warm run passes is strong evidence that correctness depends on retained state.
Repair by narrowing what is cached, strengthening keys with every correctness-relevant input, or using a build system with validated content-addressed caching. Preserve the mismatching output hashes before deleting caches.
7. Failure mode: CI depends on developer-machine state
| Symptom | Likely hidden state | Proof |
|---|---|---|
| Works locally, import missing in CI | undeclared globally installed package | fresh venv/container reproduces CI failure |
| Generated file exists locally only | uncommitted/generated workspace state |
git status --short and clean clone differ
|
| Tests pass in one locale/timezone | implicit environment dependency | record locale/TZ; rerun under controlled values |
| Warm runner/self-hosted passes, hosted fails | persistent filesystem/tool installation | fresh hosted runner lacks undeclared state |
8. Intentionally broken example: stale checkout plus hidden failure
jobs:
tests:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@FULL_SHA
with:
ref: main # wrong: mutable ref
- run: pytest -q | tee pytest.txt # wrong: exit status may be hidden
- uses: actions/upload-artifact@FULL_SHA
with:
path: coverage.xml # wrong: never generated
Do not fix all three lines at once. First preserve
GITHUB_SHA, checked-out HEAD, the apparently green test
step and the artifact warning/missing file. Then repair revision
identity; then exit-code propagation; then report generation order.
The staged repair proves each causal layer independently.
9. Failure mode: CI passes but merge remains blocked
- Confirm the required check name exactly matches the current Actions job/check name.
- Confirm the check belongs to the latest commit SHA, not an earlier successful revision.
- If a specific GitHub App is required as the check source, confirm Actions produced the expected source.
-
If merge queue is enabled, confirm
merge_groupcreates the required check for the queue SHA. - If trigger-level path filtering suppresses the workflow, branch policy may wait for a check that was never created.
Do not troubleshoot by deleting the failed run, clearing every cache, granting write-all, disabling branch protection, or blindly rerunning. Those actions destroy or bypass evidence and can leave the causal defect intact.
10. Production practices that make failure cheap to understand
- Use stable descriptive job/check names and deterministic step names.
- Record run ID/attempt/SHA, runner image and tool versions in a small evidence manifest.
- Upload JUnit/coverage/lint reports on failure with bounded retention and no secrets.
- Keep cache-hit/miss visible but make clean installation/build valid on a miss.
- Separate untrusted PR validation from publishing/deployment permissions.
- Measure flaky-test rate, queue time and critical path before adding retries or parallelism.
11. Lesson summary
CI diagnosis begins with identity, not with rerun. Preserve the first run and exact SHA, then follow the job graph to runner/toolchain, command exit status, cache/artifact state and merge-policy check. Correct the smallest broken layer and prove the repair with a new, clearly identified run.
Knowledge check
A test log shows a failure, but the job is green after piping
through tee. What should you inspect first?
The shell pipeline exit status. The final command may have hidden the test process failure.
Why is checking out main dangerous in an
event-bound CI run?
The branch can move after the run is created, causing the job to
validate a different revision than GITHUB_SHA.
A warm build passes and a clean build fails. What hypothesis becomes likely?
Cached/generated or persistent state is part of correctness but is missing from declared inputs or cache keys.
Why can an artifact upload succeed while the expected coverage file is absent?
Upload success proves the action stored the selected existing paths; it does not prove a later-generated or mispathed file was included.
Why should you preserve the first failing run before rerun?
A rerun can encounter different branch, cache, dependency, service or timing state and erase the clearest evidence of the original condition.
Official references and version notes
- Workflow syntax — Current workflow/job/step, permissions, needs and condition syntax.
- Dependency caching — Current cache matching, cache-hit semantics, security restrictions and eviction behavior.
- Workflow commands — Native annotations, outputs, summaries and runner communication.
- Status checks — How GitHub Actions jobs become checks and how conclusions affect merge policy.
- Troubleshoot required checks — Latest-SHA requirements, skipped checks and merge-queue considerations.
- actions/setup-python v7.0.0 — Immutable action revision used in the labs.
- actions/cache v6.1.0 — Immutable cache action revision used in the labs.
- actions/upload-artifact v7.0.1 — Immutable artifact action revision used in the labs.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.