Checkpoint Lab — Continuous Integration Patterns for Build, Test, Lint, and Coverage
Prove a reproducible CI operating model by preserving separate lint and test failures, their evidence artifacts, and one stable aggregate merge decision.
Learning objectives
- Operate the Chapter 24 CI workflow as a controlled experiment across clean, lint-failing and test-failing revisions.
- Predict job/check/artifact state before each run and verify predictions independently.
- Preserve first-failure evidence and map every merge check to its evidence source.
- Demonstrate that cache state can change performance without changing correctness.
- Produce a concise CI contract suitable for later release automation in Chapter 25.
1. Checkpoint scenario and success criteria
You will use the disposable Python repository from Lesson 2. Run A is clean. Run B introduces a lint-only defect. Run C restores lint and introduces a test-only defect. Run D repairs the test. The objective is not merely to end green; it is to prove that each failed revision produced the expected check graph and retained its own evidence.
The final deliverable is a check-to-evidence map showing source SHA,
toolchain/lock input, cache signal, lint/build/test/coverage
results, artifact identities and the stable
CI required conclusion for each run.
2. Verified assumptions and preflight (2026-09-10)
| Item | Lab baseline |
|---|---|
| Repository | disposable GitHub repository you control; public is sufficient for free path |
| Runner | ubuntu-24.04 GitHub-hosted |
| Python | 3.13 via setup-python v7.0.0 full SHA |
| CI tools | Ruff 0.16.6; pytest 9.1.1; coverage.py 7.15.4 |
| Cache | actions/cache v6.1.0 full SHA; pip download cache only |
| Evidence storage | actions/upload-artifact v7.0.1 full SHA; seven-day lesson retention |
| Permissions |
contents: read; no secrets, OIDC, release,
package or deployment permission
|
| Governance |
optional branch protection/ruleset requiring
CI required; simulate by treating it as the
merge decision if settings unavailable
|
Disposable-resource guard: do not use a production repository or make organization-wide policy changes for this checkpoint. If branch protection is unavailable or undesirable, record the intended required-check rule without enabling it.
3. Read-only preflight record
git status --short
git remote -v
git rev-parse HEAD
sha256sum requirements-dev.txt
python --version || true
In GitHub, inspect the workflow file at the exact revision that will run. Confirm every external action is pinned to the full SHA used in Lesson 2 and that no secret or write permission is present.
4. Write predictions before changing source
| Prediction | Expected state |
|---|---|
| Run A |
all stage jobs success; CI required success;
evidence artifacts exist
|
| Run B lint defect |
Dependencies/Build/Tests can still succeed; Lint fails; lint
artifact survives; CI required fails
|
| Run C test defect |
Dependencies/Lint/Build succeed; Tests + coverage fails;
JUnit/coverage evidence survives;
CI required fails
|
| Run D repaired | all required jobs succeed on a new SHA; prior failed runs remain available as historical evidence |
| Warm cache | a later same-lock run may report exact hit; a miss must still install and validate correctly |
5. Run A — establish the green contract
- Commit the known-good project and workflow.
-
Trigger CI and record run ID, attempt,
GITHUB_SHA, workflow name and job graph. - Verify each checkout HEAD equals
GITHUB_SHA. - Download or inspect the dependency, lint, build and tests/coverage artifacts. Record artifact names and GitHub-reported digests.
- Record Ruff/pytest/coverage/Python versions and the dependency-input hash.
-
If branch protection is enabled, confirm the exact required check
name is
CI requiredand it is satisfied on this SHA.
6. Run B — inject a lint-only failure
Add an unused import to src/calc.py without changing
behavior:
import os # deliberate checkpoint lint failure
def add(a: int, b: int) -> int:
return a + b
Commit and push. Do not rerun Run A. Verify that Ruff reports the
unused import, Lint ends in failure after its evidence upload, and
CI required fails. Because the test code and behavior
remain valid, Tests + coverage should still be able to pass under
this graph. Record the Run B SHA and artifact identity before
repair.
7. Run C — repair lint, then inject a test failure
Remove the unused import. Then change only test_add to
expect 6 instead of 5. Commit both changes together so Run C has a
clean lint state and an intentionally wrong behavioral expectation.
def test_add():
assert add(2, 3) == 6 # deliberate checkpoint test failure
Verify Lint and Build are successful, Tests + coverage fails, the
JUnit file identifies test_add, coverage files are
still uploaded, and CI required fails because
needs.tests.result is not success.
8. Run D — least-destructive repair
Restore the assertion to == 5 and commit. Do not
delete, overwrite or relabel the prior runs. Run D should be a
distinct SHA with a successful aggregate check. Compare its
dependency-input hash and cache signal with Runs B/C: cache state
may differ, but the logical dependency set should not.
9. Build the evidence packet
| Evidence field | Record for each relevant run |
|---|---|
| Run identity |
run ID, attempt, event, ref, exact GITHUB_SHA
|
| Workflow/action manifest | workflow path/revision; full SHAs for checkout/setup/cache/upload |
| Runner/toolchain |
ubuntu-24.04, image metadata from Set up job,
Python/Ruff/pytest/coverage versions
|
| Dependency evidence | requirements hash, cache-hit signal, pip resolver/check result, freeze manifest |
| Lint/build | Ruff diagnostics/exit outcome; compile result; source checksum manifest |
| Test/coverage | test count/failing name, JUnit XML, coverage XML/text, threshold outcome |
| Artifacts | name, producing run+attempt+SHA, digest, retention assumption |
| Governance |
CI required result and branch/ruleset
expectation or simulated rule
|
| Limitations | top-level dependency pins are not a full transitive hash lock; hosted image evolves; coverage is not test quality |
10. Check-to-evidence map
| Check | Authoritative question | Evidence | Must fail when |
|---|---|---|---|
| Dependencies | Can declared CI dependencies be established consistently? | requirements hash, pip output/check, freeze, cache signal | resolver/install/check fails |
| Lint | Does source meet static rules? | Ruff version, diagnostic artifact, raw step outcome | Ruff returns non-zero |
| Build | Does exact source compile/build under selected toolchain? | HEAD==GITHUB_SHA, compile result, source digests | compile/build fails or SHA mismatch |
| Tests + coverage | Does behavior pass and coverage policy hold? | JUnit, coverage XML/text, raw test and threshold outcomes | tests fail or threshold fails |
| CI required | Are all required claims successful on this run? | needs.*.result values |
any required upstream job is not success |
11. Optional governance proof
In a disposable public repository, branch protection can require
status checks on GitHub Free. If you enable it, require only the
deliberate stable check CI required. Verify a failing
Run B or C blocks merge and Run D satisfies the rule on the latest
SHA. Then remove the disposable rule during cleanup if you created
it solely for the lab.
If you do not enable branch protection, document the exact intended rule and use the aggregate job result as a faithful simulation. The learning goal is the check-policy contract, not changing repository administration.
12. Cleanup and rollback
- Keep the failed run IDs long enough to complete the evidence packet.
- Remove the deliberate lint/test defects; retain only the known-good source in the final branch.
- Remove any branch-protection/ruleset change created solely for this disposable lab.
- Allow seven-day lesson artifacts to expire or delete only the exact disposable artifacts after recording their identities.
- Delete the disposable repository when no longer needed. No cloud, package, registry, environment or secret cleanup is required.
13. Production CI contract
- Identity: every check is tied to the exact tested source and workflow/action revisions.
- Determinism: toolchain and dependency inputs are explicit; cache miss does not change correctness.
- Evidence: lint/test/coverage/build reports survive failures and retain run/attempt/SHA provenance.
- Policy: one stable aggregate check exposes a reviewable merge decision without erasing internal detail.
- Security: validation jobs use least privilege and do not require production secrets or trusted persistent runners.
- Recovery: failed revisions are preserved; repairs create new evidence instead of rewriting history.
14. Bridge to Chapter 25
Chapter 24 ends with a trustworthy answer to “is this exact revision acceptable to integrate?” Chapter 25 adds controlled publication side effects: tags, GitHub Releases, packages and container registries. The key transition is to promote the exact verified build output rather than silently rebuilding it during release.
15. Checkpoint summary
You have operated CI as an evidence system rather than a script runner: clean, lint-failing and test-failing revisions produced distinct, preserved records; caching remained optional; the final aggregate check represented policy; and the repaired run did not erase first-failure evidence.
Knowledge check
Why are Run B and Run C separate commits instead of two edits before one run?
Separate revisions isolate the lint and test hypotheses, making each failure causal and preserving an unambiguous check/evidence record.
Why should Run B’s lint artifact exist even though Lint is red?
The workflow intentionally uploads evidence before reasserting the raw lint failure. Storage success and check success are independent.
If Run D has a cache miss but still passes, what did the lab prove?
The cache is a performance optimization rather than a correctness prerequisite.
A teammate wants branch protection to require every internal job name. What maintenance risk should you explain?
Internal job renames/refactors become governance migrations. A stable aggregate check can keep policy interface stable while internals evolve.
Run C is red and someone immediately clicks rerun without recording its SHA/artifacts. What recovery principle was violated?
Preserve first-failure identity and evidence before rerun, because reruns can encounter different cache, dependency, runner or timing state.
Official references and version notes
- Workflow syntax — Current workflow/job/step, permissions, needs and condition syntax.
- Dependency caching — Current cache matching, cache-hit semantics, security restrictions and eviction behavior.
- Workflow commands — Native annotations, outputs, summaries and runner communication.
- Status checks — How GitHub Actions jobs become checks and how conclusions affect merge policy.
- Troubleshoot required checks — Latest-SHA requirements, skipped checks and merge-queue considerations.
- actions/setup-python v7.0.0 — Immutable action revision used in the labs.
- actions/cache v6.1.0 — Immutable cache action revision used in the labs.
- actions/upload-artifact v7.0.1 — Immutable artifact action revision used in the labs.
- Managing branch protection — Optional disposable governance path for requiring the aggregate check.
- Store and share workflow artifacts — Artifact upload/download and digest-validation concepts used for report retention.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.