Chapter 24Lesson 02~190 minutes

Continuous Integration Patterns for Build, Test, Lint, and Coverage: Guided Hands-On Workflow

Implement a small evidence-rich Python CI pipeline whose cache can accelerate work but never becomes a correctness dependency.

Python 3.13RuffpytestCoverageArtifacts

Learning objectives

  • Create a tiny Python project with explicit CI-tool versions and a reviewable dependency input.
  • Use an Actions cache to accelerate pip downloads without bypassing installation or validation.
  • Preserve lint, test and coverage evidence even when the stage fails.
  • Use continue-on-error only as an evidence-collection technique and explicitly reassert failure.
  • Interpret how one failed stage changes downstream and aggregate job state.

1. Preflight: disposable repository, no credentials, no production resources

Use a throwaway repository such as gha-ci-evidence-lab. The workflow needs no cloud account, package registry, database, secret, environment or self-hosted runner. Standard GitHub-hosted Ubuntu is sufficient. The only repository permission requested is contents: read for checkout.

Current assumptions verified 2026-09-10: ubuntu-24.04; Python 3.13; checkout v7.0.1; setup-python v7.0.0; cache v6.1.0; upload-artifact v7.0.1. Every external action reference below is a full commit SHA.

2. Build the tiny project and pin the CI tools

gha-ci-evidence-lab/
├── .github/workflows/ci.yml
├── requirements-dev.txt
├── pyproject.toml
├── src/calc.py
└── tests/test_calc.py
# requirements-dev.txt — assumptions verified 2026-09-10
ruff==0.16.6
pytest==9.1.1
coverage==7.15.4

These top-level tools are exactly pinned for a repeatable lesson. A production project should use its ecosystem’s full lock/hashing strategy where stronger transitive reproducibility is required; this lesson deliberately keeps dependency mechanics small enough to inspect.

3. Minimal application and tests

# src/calc.py
def add(a: int, b: int) -> int:
    return a + b

def clamp(value: int, low: int, high: int) -> int:
    if low > high:
        raise ValueError("low must not exceed high")
    return max(low, min(value, high))
# tests/test_calc.py
from src.calc import add, clamp

def test_add():
    assert add(2, 3) == 5

def test_clamp_middle():
    assert clamp(5, 0, 10) == 5

def test_clamp_low():
    assert clamp(-2, 0, 10) == 0

def test_invalid_bounds():
    import pytest
    with pytest.raises(ValueError):
        clamp(1, 5, 2)
[tool.ruff]
target-version = "py313"
line-length = 100

[tool.coverage.run]
source = ["src"]
branch = true

[tool.coverage.report]
show_missing = true

4. The staged workflow: preserve claims, then aggregate them

The workflow below is intentionally explicit. dependencies establishes resolver evidence and warms the pip download cache. lint and build then run independently after that prerequisite. tests depends on the build hypothesis. Finally, CI required evaluates every upstream result even after failures.

name: CI evidence lab
on:
  pull_request:
  workflow_dispatch:

permissions:
  contents: read

env:
  PYTHON_VERSION: '3.13'

jobs:
  dependencies:
    name: Dependencies
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
      - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
        with:
          python-version: ${{ env.PYTHON_VERSION }}
      - id: cache
        uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
        with:
          path: ~/.cache/pip
          key: ${{ runner.os }}-py313-pip-${{ hashFiles('requirements-dev.txt') }}
      - name: Resolve and verify
        run: |
          mkdir -p evidence
          test "$(git rev-parse HEAD)" = "$GITHUB_SHA"
          python -m pip install --disable-pip-version-check -r requirements-dev.txt
          python -m pip check
          {
            echo "run_id=$GITHUB_RUN_ID"
            echo "attempt=$GITHUB_RUN_ATTEMPT"
            echo "sha=$GITHUB_SHA"
            echo "cache_hit=${{ steps.cache.outputs.cache-hit }}"
            python --version
            python -m pip --version
            python -m pip freeze
          } > evidence/dependencies.txt
      - uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
        with:
          name: dependencies-${{ github.run_id }}-${{ github.run_attempt }}
          path: evidence/dependencies.txt
          retention-days: 7

  lint:
    name: Lint
    needs: dependencies
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
      - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
        with:
          python-version: ${{ env.PYTHON_VERSION }}
      - uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
        with:
          path: ~/.cache/pip
          key: ${{ runner.os }}-py313-pip-${{ hashFiles('requirements-dev.txt') }}
      - run: python -m pip install --disable-pip-version-check -r requirements-dev.txt
      - id: lint
        name: Run Ruff but preserve its evidence
        continue-on-error: true
        shell: bash
        run: |
          mkdir -p evidence
          set +e
          ruff check src tests --output-format=concise 2>&1 | tee evidence/ruff.txt
          rc=${PIPESTATUS[0]}
          set -e
          echo "exit_code=$rc" >> "$GITHUB_OUTPUT"
          if [ "$rc" -ne 0 ]; then
            echo "::error title=Lint failed::Ruff exited $rc; inspect lint artifact."
          fi
          exit "$rc"
      - if: always()
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
        with:
          name: lint-${{ github.run_id }}-${{ github.run_attempt }}
          path: evidence/ruff.txt
          if-no-files-found: warn
          retention-days: 7
      - if: always()
        name: Reassert lint outcome
        run: test "${{ steps.lint.outcome }}" = "success"

  build:
    name: Build
    needs: dependencies
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
      - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
        with:
          python-version: ${{ env.PYTHON_VERSION }}
      - name: Compile and record
        run: |
          mkdir -p evidence
          test "$(git rev-parse HEAD)" = "$GITHUB_SHA"
          python -m compileall -q src
          find src -type f -name '*.py' -print0 | sort -z | xargs -0 sha256sum > evidence/source-sha256.txt
      - uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
        with:
          name: build-${{ github.run_id }}-${{ github.run_attempt }}
          path: evidence/source-sha256.txt
          retention-days: 7

  tests:
    name: Tests + coverage
    needs: [dependencies, build]
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
      - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
        with:
          python-version: ${{ env.PYTHON_VERSION }}
      - uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
        with:
          path: ~/.cache/pip
          key: ${{ runner.os }}-py313-pip-${{ hashFiles('requirements-dev.txt') }}
      - run: python -m pip install --disable-pip-version-check -r requirements-dev.txt
      - id: pytest
        continue-on-error: true
        shell: bash
        run: |
          mkdir -p evidence
          set +e
          coverage run -m pytest -q --junitxml=evidence/junit.xml
          rc=$?
          set -e
          coverage xml -o evidence/coverage.xml
          coverage report -m > evidence/coverage.txt
          exit "$rc"
      - id: coverage_gate
        continue-on-error: true
        run: coverage report --fail-under=90
      - if: always()
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
        with:
          name: tests-coverage-${{ github.run_id }}-${{ github.run_attempt }}
          path: evidence/
          if-no-files-found: warn
          retention-days: 7
      - if: always()
        name: Reassert test and coverage outcomes
        run: |
          test "${{ steps.pytest.outcome }}" = "success"
          test "${{ steps.coverage_gate.outcome }}" = "success"

  required:
    name: CI required
    if: always()
    needs: [dependencies, lint, build, tests]
    runs-on: ubuntu-24.04
    steps:
      - name: Aggregate required claims
        run: |
          echo "dependencies=${{ needs.dependencies.result }}"
          echo "lint=${{ needs.lint.result }}"
          echo "build=${{ needs.build.result }}"
          echo "tests=${{ needs.tests.result }}"
          test "${{ needs.dependencies.result }}" = "success"
          test "${{ needs.lint.result }}" = "success"
          test "${{ needs.build.result }}" = "success"
          test "${{ needs.tests.result }}" = "success"

5. Why the workflow uses continue-on-error without hiding failure

The lint and test commands are allowed to finish with an unsuccessful step outcome so later evidence-upload steps still execute. GitHub distinguishes the raw step outcome from the post-continue-on-error conclusion. The final “Reassert” step inspects the raw outcome and fails the job if the check actually failed.

This pattern is materially different from adding continue-on-error: true and doing nothing else. The latter can turn a real quality failure into a green job. Here the temporary tolerance exists only to preserve reports before the job conclusion is restored to failure.

6. Run A: establish a clean baseline

  1. Commit the files and push them to the disposable repository.
  2. Run the workflow manually or open a pull request.
  3. Record GITHUB_RUN_ID, GITHUB_RUN_ATTEMPT, GITHUB_SHA and the checked-out SHA from the logs.
  4. Confirm Dependencies, Lint, Build, Tests + coverage and CI required are all successful.
  5. Inspect the dependency, lint, build and test/coverage artifacts; record their names and digests shown by GitHub.
  6. Run the same revision again and compare the cache signal. A warm exact key may report true, but correctness must be unchanged even if it reports a miss.

7. Run B: intentionally fail a test and preserve first-failure evidence

Change only the test expectation below. Do not “fix” the application to make the deliberately wrong assertion pass.

def test_add():
    assert add(2, 3) == 6  # deliberate lesson failure

Predict before pushing: Dependencies, Lint and Build should remain green; Tests + coverage should become red; its artifact should still contain JUnit and coverage files; CI required should become red. After the run, verify each prediction and record the failing run/attempt before any rerun.

Then repair the assertion to == 5 in a new commit. That new run is a different source revision; it does not erase the diagnostic value of Run B.

8. Read the graph rather than guessing from color

Observation Interpretation
Lint red; Tests green The workflow deliberately allows tests to run independently of lint after dependency/build prerequisites. More evidence was collected before the aggregate gate failed.
Tests red; test artifact exists Artifact storage and job conclusion are independent states. This is expected evidence preservation.
CI required red At least one required upstream claim was not successful. Inspect printed needs.*.result values.
Cache miss but Dependencies green Performance state changed; correctness did not.
Checkout SHA differs from GITHUB_SHA Stop. The workflow is validating a different revision than its event identity claims.

9. Challenge: choose the correct layer

A teammate proposes adding continue-on-error: true to the entire Tests + coverage job because flaky tests block merges. Choose the correct response: trigger change, runner change, cache change, test/reliability repair, or governance bypass? Defend your answer with the state/evidence you would inspect first.

A strong answer preserves the failing test evidence, quantifies flakiness, fixes or quarantines the specific nondeterministic test under an explicit policy, and keeps the aggregate required check honest.

10. Cleanup and rollback

  • Delete the disposable repository when finished, or remove only the lesson workflow/branches you created.
  • Do not delete Run B before you have recorded the first-failure evidence required by the lab.
  • Artifacts in this lesson use seven-day retention; repository retention policy can further constrain actual availability.
  • No external service, credential, environment or package was created, so cleanup is bounded to the disposable repository.

11. Lesson summary

The guided workflow makes CI evidence explicit: dependency state, static checks, buildability, behavioral tests and coverage remain independently visible; the cache is optional acceleration; artifacts preserve failed-stage diagnostics; and one stable aggregate job becomes the merge-policy interface.

Next lesson

Continuous Integration Patterns for Build, Test, Lint, and Coverage: Configuration, Design Patterns, and Trade-Offs

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why does lint run even though it eventually fails the job?

Why does the pip installer still run after an exact cache hit?

What should happen to CI required when Tests + coverage fails?

Why include run attempt in artifact names?

The test artifact uploaded successfully, but the Tests job is red. Is that inconsistent?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.