Jobs, Steps, Matrices, Containers, Service Containers, and Dependency Graphs: Diagnostics, Failure Modes, Security, and Performance
Topology failures often look like application failures: “file not found,” “connection refused,” or a missing test result. The correct repair starts by identifying which runner, workspace, matrix cell, service network, dependency edge, and result policy produced the symptom before adding retries or privileges.
Learning objectives
- Diagnose a cross-job missing-file failure without assuming a runner or GitHub outage.
- Detect accidental matrix growth from generated job evidence before it becomes a concurrency/cost incident.
- Separate service startup, health readiness, address/port selection, and application protocol failures.
-
Recognize when
fail-fastremoves useful diagnostic coverage and whencontinue-on-errormasks a required failure. - Apply a preserve-evidence-first correction sequence without force pushes, privilege broadening, secret logging, or runner registration.
Diagnostic safety: All broken examples belong in
disposable repositories. Do not troubleshoot by adding
write-all, printing secrets/github
contexts, disabling organization policy, force-updating refs,
registering a personal machine as a runner, or pointing tests at
production databases.
1. Diagnostic sequence: preserve → scope → inspect → minimally correct → verify
- Preserve evidence: run ID/attempt, event, head SHA, job list, matrix names, logs, and workflow file at that SHA.
-
Identify scope: repository, workflow, job/matrix
cell, runner, container/service topology, and direct
needsedges. - Inspect: job conclusions/timestamps, exact workflow YAML, service health/addressing, artifact/output contracts, and permissions.
- Correct the smallest violated contract: transfer/reproduce the file, bound the matrix, add health readiness, fix the address, or change failure policy deliberately.
- Verify with a new run: bind the repair to a new SHA while retaining the failed run as historical evidence.
REPO="OWNER/REPOSITORY"
RUN_ID="123456789"
gh run view "$RUN_ID" -R "$REPO" --json event,headBranch,headSha,status,conclusion,jobs,url
gh run view "$RUN_ID" -R "$REPO" --log-failed
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/runs/$RUN_ID/jobs?filter=latest&per_page=100" --jq '.jobs[] | {id,name,status,conclusion,runner_name,started_at,completed_at}'
2. Intentionally broken example: needs does not carry
files
Suppose the build job creates a file and the test job assumes
needs: build mounted the previous workspace:
jobs:
build:
runs-on: ubuntu-latest
steps:
- run: mkdir -p dist && printf 'binary\n' > dist/app.bin
test:
needs: build
runs-on: ubuntu-latest
steps:
- run: |
ls -la
test -f dist/app.bin
Expected failure: the test job runs on a fresh
runner; dist/app.bin is absent and
test -f exits non-zero. The original cause is topology,
not “flaky disk.”
Repair A: if test only needs repository source, check out the same exact SHA and rebuild deterministically. Repair B: if test must consume the exact generated bytes, upload/download a workflow artifact using reviewed, pinned action identities. Do not solve the problem with a shared self-hosted filesystem unless that shared state is an intentional, governed architecture.
3. Matrix explosion: count before you run
A matrix can multiply silently during maintenance. A three-dimensional matrix with 4 operating systems, 5 runtime versions, and 3 database modes creates 60 base jobs. Add more dimensions and you can approach GitHub’s current 256-job matrix cap quickly.
# Evidence after a run: count generated jobs with a stable name prefix.
gh run view "$RUN_ID" -R "$REPO" --json jobs --jq '[.jobs[] | select(.name | startswith("test ("))] | length'
Correction choices: remove unsupported combinations with
exclude, move exceptional lanes to explicit
scheduled/manual workflows, reduce dimensions, or cap
max-parallel. Do not assume free public minutes means
unlimited concurrency or external dependency capacity.
4. Service failure: separate lifecycle, readiness, and application protocol
| Symptom | Likely layer | Evidence | Least-destructive correction |
|---|---|---|---|
| Connection refused immediately | Service not ready / wrong mapped port | Service logs/health + job address/port | Add health check; correct port mapping; bounded readiness probe. |
Name redis does not resolve in host job |
Wrong network model |
Job has no container:; service has mapped host
port
|
Use 127.0.0.1:<host-port>. |
127.0.0.1 fails from containerized job |
Wrong network model | Job uses container: |
Use service-label hostname on shared Docker network. |
| TCP connects but test fails | Application/protocol/data layer | Client response and service logs | Fix fixture/protocol/config; do not add network retries blindly. |
A health check should express service readiness. It is not a blanket guarantee that schema migration, fixture loading, or external dependencies are ready; those require their own explicit preparation step.
5. Container networking and filesystem assumptions
Container jobs run their ordinary steps inside the declared image
and default to sh. Container actions may run as sibling
containers with shared mounts/network. A command that relies on a
host-only path, a Bash-only feature, or an uninstalled utility can
therefore fail even though the same command worked on the raw Ubuntu
runner.
Diagnose the image and shell first. Do not install a large toolchain into the image at runtime until you decide whether the job container is actually the right abstraction.
6. fail-fast can hide coverage;
continue-on-error can hide failure
If one matrix cell fails early and default
fail-fast: true cancels the rest, the run may lack
evidence about whether the defect affects one platform or every
platform. That can be acceptable for fast presubmit feedback but
weak for a release qualification matrix.
The opposite mistake is more dangerous:
continue-on-error: true on a required compatibility
lane can make the workflow appear successful even when the lane
fails. Tolerated failures must be deliberate, visible, and
non-authoritative for release policy.
7. Security-sensitive boundaries in topology debugging
- Container registry credentials: never print them; use least privilege and short lifetime where supported.
- Service data: use synthetic fixtures; do not copy production databases into public-repository service jobs.
- Self-hosted runners: do not register one merely to share files or reach an internal service; Chapter 16 treats runner trust explicitly.
- Ref/history repair: do not force-push to make a run “match” expected state. Create a new commit/run and preserve the failing SHA.
-
Permission repair: filesystem/network mistakes do
not justify broadening
GITHUB_TOKEN.
8. Performance and billing only where topology causes them
Every matrix cell is a job with startup and execution cost. Every job-container/service image may require pulling layers. Private repositories consume plan-based runner minutes/storage; public standard runners are currently free, but concurrency still has account limits and external services still have quotas. Optimize after observing job timestamps and bottlenecks, not by collapsing isolation blindly.
Typical safe optimizations are: reduce redundant matrix cells, cap
max-parallel, use a smaller approved image, reuse
dependency caches appropriately, split slow optional lanes from
blocking presubmit, or move a stable generated file through an
artifact instead of rebuilding it repeatedly. Each optimization must
preserve the evidence release policy needs.
9. Verification checklist after a repair
- New run binds to the intended new commit SHA.
- Generated job count matches the reviewed matrix count.
- Every required matrix cell reaches an explicit conclusion instead of accidental cancellation.
- Cross-job files use an explicit artifact/package or reproducible rebuild/checkout boundary.
- Service health and client protocol checks both succeed using the correct network address.
- No permission, secret, branch-policy, runner-registration, or history-rewrite workaround was introduced.
10. Lesson summary
Topology failures become tractable when you classify them by runner/workspace boundary, dependency graph, matrix generation, container user-space, service network/readiness, and result policy. Preserve the failed run, correct only the violated contract, and verify the repair in a new run. Do not hide execution-model mistakes with retries, broad permissions, persistent runners, or tolerated failures.
Knowledge check
A downstream job fails with “dist/app.bin: No such file.” The upstream job created it successfully. First diagnosis?
The jobs have separate runner filesystems. Verify whether the file should be reproduced or explicitly transferred as an artifact.
A host-runner job cannot resolve hostname redis,
but the service is healthy. What likely assumption is
wrong?
The job is using container-to-container addressing while running directly on the host. Use the mapped localhost port.
One matrix cell fails and three others are cancelled. What setting should you inspect before calling the cancellations independent failures?
strategy.fail-fast, which defaults to true.
Why is making a failing compatibility job
continue-on-error: true a governance
decision?
Because it changes whether that failed evidence can fail the workflow/release signal; it can mask a real compatibility defect.
Why should the fixed run use a new commit instead of rewriting the old SHA?
The old run is evidence tied to the original workflow/source. A new commit preserves auditability and lets the repair be independently verified.
Further reading — current official GitHub sources
- GitHub Docs — Workflow syntax
- GitHub Docs — Using jobs in a workflow
- GitHub Docs — Running variations of jobs
- GitHub Docs — GitHub-hosted runners
- GitHub Docs — Running jobs in a container
- GitHub Docs — Docker service containers
- GitHub Docs — Store and share workflow data
- GitHub CLI — gh run view
- GitHub REST — Workflow jobs
- GitHub Docs — Actions limits
- GitHub Docs — Secure use reference
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.