Checkpoint Lab — Jobs, Steps, Matrices, Containers, Service Containers, and Dependency Graphs
This checkpoint combines the previous two Actions chapters with execution topology. You will bind every downstream job to one exact source SHA, prove that runner-local files do not cross job boundaries, execute a controlled matrix and service dependency, inject exactly one matrix-cell failure, and explain the final graph before optimizing it.
Learning objectives
- Build a two-stage validation graph whose downstream jobs independently reproduce the exact source SHA.
- Expand one bounded matrix with controlled parallelism and preserve complete evidence by disabling fail-fast.
- Add one Redis service dependency with explicit health/readiness and synthetic data only.
- Inject exactly one matrix-cell failure, preserve the failed run, and distinguish cell failure from graph/service failure.
- Analyze job graph, timing, result aggregation, and write a production operating policy before optimizing.
Checkpoint assumptions: GitHub.com, GitHub Free, a disposable public personal repository, write access, GitHub CLI, local Git, and standard Ubuntu GitHub-hosted runners. No secrets, private registries, self-hosted runners, production services, cloud accounts, or paid features are required. Redis runs only as a disposable service container using synthetic PING traffic.
1. Scenario and predictions before execution
You are qualifying a small release branch workflow. The
prepare job records one immutable source SHA and
deliberately creates one runner-local file. A matrix job
independently checks out that SHA, a service job proves Redis
readiness, and a summary job records prerequisite results without
hiding failures.
Write these predictions before running anything:
-
P1 — isolation:
prepare-only.txtexists inpreparebut not in any matrix child. -
P2 — clean run: with
inject_failure=false, all three matrix cells and the Redis job succeed; the workflow concludes success. -
P3 — injected run: with
inject_failure=true, onlyregression / bfails. Becausefail-fast:false, the other cells still complete. -
P4 — aggregation: the summary job runs in both
cases because it uses
always(); in the failing run it reports the matrix result as failure but does not erase that failure from the workflow.
2. Preflight and fresh repository
gh --version
gh auth status --active --hostname github.com
OWNER="$(gh api -H "X-GitHub-Api-Version: 2026-03-10" user --jq .login)"
REPO="$OWNER/atlas-c15-checkpoint"
gh repo view "$REPO" --json nameWithOwner >/dev/null 2>&1 && {
echo "Checkpoint repository already exists; use a fresh disposable name." >&2
exit 1
} || true
gh repo create "$REPO" --public --clone --add-readme
cd atlas-c15-checkpoint
DEFAULT_BRANCH="$(gh repo view "$REPO" --json defaultBranchRef --jq '.defaultBranchRef.name')"
gh workflow list -R "$REPO" --all --json id,name,path,state
If policy prevents Actions or public repositories, do not weaken policy. Use another personal disposable account/repository that satisfies the course’s free-path assumptions.
3. Create the checkpoint workflow
name: Chapter 15 checkpoint
on:
workflow_dispatch:
inputs:
inject_failure:
description: Fail only the regression/b matrix cell
required: true
type: boolean
default: false
permissions:
contents: read
jobs:
prepare:
runs-on: ubuntu-latest
outputs:
source_sha: ${{ steps.identity.outputs.sha }}
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
- id: identity
run: |
printf 'sha=%s\n' "$GITHUB_SHA" >> "$GITHUB_OUTPUT"
printf 'prepare-runner-only\n' > prepare-only.txt
test -f prepare-only.txt
printf 'prepared_sha=%s\n' "$GITHUB_SHA"
matrix_tests:
needs: prepare
runs-on: ubuntu-latest
strategy:
fail-fast: false
max-parallel: 2
matrix:
lane: [smoke, regression]
config: [a, b]
exclude:
- lane: regression
config: a
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
with:
ref: ${{ needs.prepare.outputs.source_sha }}
- name: Prove reproducible boundary
env:
EXPECTED_SHA: ${{ needs.prepare.outputs.source_sha }}
run: |
test ! -e prepare-only.txt
test "$(git rev-parse HEAD)" = "$EXPECTED_SHA"
- name: Deterministic lane validation
env:
LANE: ${{ matrix.lane }}
CONFIG: ${{ matrix.config }}
run: |
printf 'validated lane=%s config=%s\n' "$LANE" "$CONFIG"
- name: Inject one matrix-only failure
if: ${{ inputs.inject_failure && matrix.lane == 'regression' && matrix.config == 'b' }}
run: |
echo "intentional matrix-only failure: regression/b" >&2
exit 17
redis_check:
needs: prepare
runs-on: ubuntu-latest
services:
redis:
image: redis:7-alpine
ports:
- 6379:6379
options: >-
--health-cmd "redis-cli ping"
--health-interval 5s
--health-timeout 3s
--health-retries 10
steps:
- name: Verify service readiness from runner host
run: |
python - <<'PY2'
import socket
with socket.create_connection(('127.0.0.1', 6379), timeout=5) as s:
s.sendall(b'*1\r\n$4\r\nPING\r\n')
reply = s.recv(64)
print('redis_reply=', reply.decode().strip())
assert reply.startswith(b'+PONG')
PY2
summary:
if: ${{ always() }}
needs: [prepare, matrix_tests, redis_check]
runs-on: ubuntu-latest
steps:
- env:
PREPARE_RESULT: ${{ needs.prepare.result }}
MATRIX_RESULT: ${{ needs.matrix_tests.result }}
REDIS_RESULT: ${{ needs.redis_check.result }}
SOURCE_SHA: ${{ needs.prepare.outputs.source_sha }}
run: |
printf 'prepare=%s matrix=%s redis=%s sha=%s\n' \
"$PREPARE_RESULT" "$MATRIX_RESULT" "$REDIS_RESULT" "$SOURCE_SHA"
The matrix has exactly three cells after exclusion: smoke/a,
smoke/b, regression/b. The job does not use
continue-on-error; the injected failure must remain a
real workflow failure.
4. Commit the workflow and record its source identity
mkdir -p .github/workflows
# Save the YAML above as .github/workflows/ch15-checkpoint.yml
git add .github/workflows/ch15-checkpoint.yml
git commit -m "ci: add Chapter 15 checkpoint"
git push origin "$DEFAULT_BRANCH"
CHECKPOINT_SHA="$(git rev-parse HEAD)"
printf 'checkpoint_sha=%s\n' "$CHECKPOINT_SHA"
gh workflow view ch15-checkpoint.yml -R "$REPO" --yaml
Before dispatch, the only change is repository content and hosted workflow discovery. No runner has started and no Redis container exists yet.
5. Run the clean matrix and verify P1/P2
printf '%s\n' '{"inject_failure":false}' | gh workflow run ch15-checkpoint.yml -R "$REPO" --ref "$DEFAULT_BRANCH" --json
RUN_OK="$(gh run list -R "$REPO" --workflow ch15-checkpoint.yml --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$RUN_OK" -R "$REPO" --exit-status
gh run view "$RUN_OK" -R "$REPO" --json headSha,conclusion,jobs,url --jq '{headSha,conclusion,jobs:[.jobs[]|{name,conclusion,startedAt,completedAt,databaseId}]}'
gh run view "$RUN_OK" -R "$REPO" --log
Verify headSha == CHECKPOINT_SHA, three matrix child
jobs are visible, each child proves prepare-only.txt is
absent, Redis returns PONG, and summary reports
success/success/success.
6. Analyze graph/timing before changing anything
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/runs/$RUN_OK/jobs?filter=latest&per_page=100" --jq '.jobs[] | {id,name,conclusion,runner_name,started_at,completed_at}'
Look for three things:
dependency delay (matrix/Redis begin after
prepare), matrix overlap (up to two matrix cells
can overlap because of max-parallel:2), and
independent service work (Redis can run
concurrently with matrix work after prepare). Do not optimize yet;
first understand the measured graph.
7. Inject exactly one matrix-cell failure and verify P3/P4
The workflow definition does not change. Only the typed input changes, so run-to-run differences are attributable to controlled input rather than a new commit.
printf '%s\n' '{"inject_failure":true}' | gh workflow run ch15-checkpoint.yml -R "$REPO" --ref "$DEFAULT_BRANCH" --json
RUN_FAIL="$(gh run list -R "$REPO" --workflow ch15-checkpoint.yml --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
# Expected failure: preserve it rather than hiding the conclusion.
gh run watch "$RUN_FAIL" -R "$REPO" || true
gh run view "$RUN_FAIL" -R "$REPO" --json headSha,status,conclusion,jobs,url --jq '{headSha,status,conclusion,jobs:[.jobs[]|{name,conclusion,startedAt,completedAt,databaseId}]}'
gh run view "$RUN_FAIL" -R "$REPO" --log-failed
Expected evidence: only regression/b fails with exit 17; smoke/a and
smoke/b still complete because fail-fast:false; Redis
remains independent and succeeds; summary executes and reports
matrix=failure; the overall workflow conclusion remains
failure.
8. Compare clean and failing runs as evidence sets
| Evidence | Clean run | Injected run | Interpretation |
|---|---|---|---|
| Source SHA | Same CHECKPOINT_SHA |
Same CHECKPOINT_SHA |
Source/workflow revision held constant. |
| Matrix children | 3 success | 2 success + regression/b failure | Input targets one cell only. |
| Redis | success | success | Service topology is not the cause. |
| Summary | success, reports matrix success | success, reports matrix failure | Observer runs without masking upstream conclusion. |
| Workflow conclusion | success | failure | Required failure semantics preserved. |
This is result aggregation done correctly: the summary job can finish successfully as an evidence reporter while the workflow still fails because a required matrix job failed.
9. Optimize only after the evidence tells you where
In this tiny lab, optimization is intentionally unnecessary. In a
real repository, use observed timestamps and queue pressure to
decide whether to reduce matrix cells, lower/raise
max-parallel, split slow optional lanes, cache
dependencies, or use an artifact instead of rebuilding. Never
collapse a job boundary merely to make the graph look faster if that
removes needed isolation or release evidence.
10. Write the merge/release CI topology policy
Use this as the minimum operating policy for a busy protected branch:
- Every job declares why it is isolated and what state it consumes/produces.
- Source-consuming jobs bind to the exact event/source SHA; generated files cross jobs only through explicit artifact/package contracts.
-
Every matrix has a reviewed expected cell count,
fail-fastpolicy, and parallelism policy. -
No required validation lane uses undocumented
continue-on-error. - Service containers use synthetic data, explicit health readiness, and only required ports.
- Container images come from approved sources; release-critical images are pinned immutably where practical.
-
Summary/reporting jobs may use
always()but never convert upstream required failures into a green release decision. - Self-hosted runner registration, private network access, and production credentials require separate Chapter 16+/20 governance.
11. Verification checklist and cleanup/rollback
- Clean and failing runs both point at the exact checkpoint SHA.
- Three matrix children appear in each run—no accidental Cartesian growth.
- Runner-local file absence proves the cross-job filesystem boundary.
- Redis health and protocol PING both succeed.
- Only the intended regression/b cell fails in the injected run.
- Summary reports upstream results without hiding the workflow failure.
- No secret, package, deployment, self-hosted runner, branch-policy bypass, or force push was created.
gh workflow disable ch15-checkpoint.yml -R "$REPO"
gh workflow list -R "$REPO" --all --json id,name,path,state
gh repo archive "$REPO" --yes
gh repo view "$REPO" --json nameWithOwner,isArchived,url
Rollback: If you want to rerun the checkpoint, create a fresh disposable repository. Archiving preserves the run history and source for review while preventing accidental continued use; permanent repository deletion is not required.
12. What Chapter 15 adds to the production GitHub operating model
Chapters 13–14 established execution identity, permissions, activation, contexts, and data contracts. Chapter 15 adds topology governance: every job boundary is explicit, every matrix expansion is bounded, every service has a scoped lifecycle/readiness contract, and every downstream result can be traced to one source SHA and dependency graph.
Chapter 16 now moves one layer lower: which runners actually execute these jobs, how hosted and self-hosted runner labels/groups route work, how fleets scale, and why runner trust is a security boundary of its own.
13. Checkpoint summary
You built a reproducible two-stage validation graph, proved job isolation, executed a bounded matrix, attached and verified a healthy service container, preserved complete coverage during one deliberate matrix failure, inspected runner/timing records, separated summary reporting from release truth, disabled the workflow, and archived the disposable repository. That is a complete topology control loop.
Knowledge check
Why can the clean and injected runs be compared causally?
They use the same workflow/source SHA; only the typed
inject_failure input changes.
Why does prepare-only.txt intentionally
disappear?
Each matrix child runs on a fresh runner. The negative check proves isolation rather than indicating data loss.
The summary job succeeds while the workflow fails. Is that contradictory?
No. The summary is an observer using always(); the
required matrix job still has failure result, so the workflow
retains the failed evidence.
If regression/b fails and smoke/a is cancelled, which policy is likely different from the checkpoint?
strategy.fail-fast may be true/default, allowing
one non-tolerated matrix failure to cancel siblings.
A developer suggests registering a persistent self-hosted runner so jobs can share files. Why reject that shortcut here?
It replaces an explicit artifact/reproducibility contract with shared mutable runner state and creates a new trust/credential/network boundary. Chapter 16 governs self-hosted runners separately.
What is the bridge to Chapter 16?
Chapter 15 defines job/container/service topology; Chapter 16 explains how runner selection, labels/groups, scaling, and runner security determine where that topology actually executes.
Further reading — current official GitHub sources
- GitHub Docs — Workflow syntax
- GitHub Docs — Using jobs in a workflow
- GitHub Docs — Running variations of jobs
- GitHub Docs — GitHub-hosted runners
- GitHub Docs — Running jobs in a container
- GitHub Docs — Docker service containers
- GitHub Docs — Store and share workflow data
- GitHub CLI — gh run view
- GitHub REST — Workflow jobs
- GitHub Docs — Actions limits
- GitHub Docs — Choosing the runner for a job
- GitHub REST — Workflow jobs
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.