Chapter 24Lesson 04~175 minutes

Continuous Integration Patterns for Build, Test, Lint, and Coverage: Diagnostics, Failure Modes, and Production Practices

Diagnose CI failures from exact revision through runner and toolchain to test, coverage, cache and required-check evidence without hiding the original cause.

DiagnosticsExit codesSHA driftCache safetyFirst failure

Learning objectives

  • Diagnose CI from run/attempt and exact SHA before changing tests or caches.
  • Detect stale/wrong checkout, swallowed exit codes and misleading continue-on-error.
  • Separate cache/build-state problems from test/toolchain problems.
  • Preserve failed-run reports and artifacts before rerun or cleanup.
  • Apply the smallest correction that restores the intended evidence contract.

1. Evidence-first diagnostic sequence

  1. Preserve identity: run ID, attempt, event, ref, GITHUB_SHA, workflow revision and first-failure logs.
  2. Confirm checkout: compare git rev-parse HEAD with the SHA the workflow claims to test.
  3. Confirm evaluated configuration: job graph, if conditions, permissions, matrix values and cache key.
  4. Confirm runner/toolchain: runner OS/image metadata, Python/tool versions and dependency manifest.
  5. Inspect the first failing command: exit code, diagnostics, test/JUnit and coverage report.
  6. Inspect retained services: cache-hit semantics, artifact names/digests and branch required-check state.
  7. Apply the least destructive repair and rerun only the smallest equivalent scope that proves it.

2. Failure mode: CI builds a different SHA than the event SHA

This failure often appears when a workflow checks out a mutable branch name after GitHub created the run. The run claims one event SHA while build steps validate whatever commit the branch points to later.

# Broken for revision-bound CI: mutable ref can advance.
- uses: actions/checkout@FULL_SHA
  with:
    ref: main

# Safer for ordinary event-bound CI: use the event-selected ref/SHA.
- uses: actions/checkout@FULL_SHA
- run: test "$(git rev-parse HEAD)" = "$GITHUB_SHA"

Preserve the original run log showing both values. Do not rerun first: if main moved again, the rerun can produce a third observation and obscure the original mismatch.

3. Failure mode: a shell pipeline swallows the test exit code

A classic example pipes test output to tee. Depending on shell options, the pipeline can report tee success while the test command failed. The check turns green even though the log contains failures.

# Broken if pipeline status is not handled:
pytest -q | tee evidence/pytest.txt

# Explicitly preserve pytest's status in Bash:
set +e
pytest -q 2>&1 | tee evidence/pytest.txt
rc=${PIPESTATUS[0]}
set -e
exit "$rc"

The diagnostic clue is contradictory evidence: failed assertions in the report but a successful step/job. Fix exit-code propagation rather than weakening branch policy.

4. Failure mode: continue-on-error becomes permanent failure masking

# Broken policy: job may remain green despite test failure.
- name: Tests
  continue-on-error: true
  run: pytest

If temporary tolerance is needed to upload reports, capture the raw steps.ID.outcome, upload evidence with if: always(), then explicitly fail a final gate when the raw outcome is not success. The purpose is preservation, not normalization of failure.

5. Failure mode: coverage is generated after the upload step

Artifacts snapshot files that exist when the upload step runs. If coverage.xml is generated afterward, the artifact can be green and valid while simply not containing the report reviewers expect. Inspect artifact contents and step timestamps rather than assuming “upload succeeded” means “correct files uploaded.”

# Correct order
- run: coverage xml -o evidence/coverage.xml
- uses: actions/upload-artifact@FULL_SHA
  with:
    path: evidence/coverage.xml

6. Failure mode: cached build outputs become undeclared source

Caching package-manager downloads is usually safer than caching generated application outputs. If a build cache key omits compiler flags, source generators, environment variables or tool versions, stale output can be restored and accepted as if newly built. A clean run that fails while a warm run passes is strong evidence that correctness depends on retained state.

Repair by narrowing what is cached, strengthening keys with every correctness-relevant input, or using a build system with validated content-addressed caching. Preserve the mismatching output hashes before deleting caches.

7. Failure mode: CI depends on developer-machine state

Symptom Likely hidden state Proof
Works locally, import missing in CI undeclared globally installed package fresh venv/container reproduces CI failure
Generated file exists locally only uncommitted/generated workspace state git status --short and clean clone differ
Tests pass in one locale/timezone implicit environment dependency record locale/TZ; rerun under controlled values
Warm runner/self-hosted passes, hosted fails persistent filesystem/tool installation fresh hosted runner lacks undeclared state

8. Intentionally broken example: stale checkout plus hidden failure

jobs:
  tests:
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@FULL_SHA
        with:
          ref: main            # wrong: mutable ref
      - run: pytest -q | tee pytest.txt  # wrong: exit status may be hidden
      - uses: actions/upload-artifact@FULL_SHA
        with:
          path: coverage.xml   # wrong: never generated

Do not fix all three lines at once. First preserve GITHUB_SHA, checked-out HEAD, the apparently green test step and the artifact warning/missing file. Then repair revision identity; then exit-code propagation; then report generation order. The staged repair proves each causal layer independently.

9. Failure mode: CI passes but merge remains blocked

  • Confirm the required check name exactly matches the current Actions job/check name.
  • Confirm the check belongs to the latest commit SHA, not an earlier successful revision.
  • If a specific GitHub App is required as the check source, confirm Actions produced the expected source.
  • If merge queue is enabled, confirm merge_group creates the required check for the queue SHA.
  • If trigger-level path filtering suppresses the workflow, branch policy may wait for a check that was never created.

Do not troubleshoot by deleting the failed run, clearing every cache, granting write-all, disabling branch protection, or blindly rerunning. Those actions destroy or bypass evidence and can leave the causal defect intact.

10. Production practices that make failure cheap to understand

  • Use stable descriptive job/check names and deterministic step names.
  • Record run ID/attempt/SHA, runner image and tool versions in a small evidence manifest.
  • Upload JUnit/coverage/lint reports on failure with bounded retention and no secrets.
  • Keep cache-hit/miss visible but make clean installation/build valid on a miss.
  • Separate untrusted PR validation from publishing/deployment permissions.
  • Measure flaky-test rate, queue time and critical path before adding retries or parallelism.

11. Lesson summary

CI diagnosis begins with identity, not with rerun. Preserve the first run and exact SHA, then follow the job graph to runner/toolchain, command exit status, cache/artifact state and merge-policy check. Correct the smallest broken layer and prove the repair with a new, clearly identified run.

Next lesson

Checkpoint Lab — Continuous Integration Patterns for Build, Test, Lint, and Coverage

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

A test log shows a failure, but the job is green after piping through tee. What should you inspect first?

Why is checking out main dangerous in an event-bound CI run?

A warm build passes and a clean build fails. What hypothesis becomes likely?

Why can an artifact upload succeed while the expected coverage file is absent?

Why should you preserve the first failing run before rerun?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.