Chapter 15Lesson 04~175 minutes

Jobs, Steps, Matrices, Containers, Service Containers, and Dependency Graphs: Diagnostics, Failure Modes, Security, and Performance

Topology failures often look like application failures: “file not found,” “connection refused,” or a missing test result. The correct repair starts by identifying which runner, workspace, matrix cell, service network, dependency edge, and result policy produced the symptom before adding retries or privileges.

DiagnosticsMatrix failuresService readinessCross-job state

Learning objectives

  • Diagnose a cross-job missing-file failure without assuming a runner or GitHub outage.
  • Detect accidental matrix growth from generated job evidence before it becomes a concurrency/cost incident.
  • Separate service startup, health readiness, address/port selection, and application protocol failures.
  • Recognize when fail-fast removes useful diagnostic coverage and when continue-on-error masks a required failure.
  • Apply a preserve-evidence-first correction sequence without force pushes, privilege broadening, secret logging, or runner registration.

Diagnostic safety: All broken examples belong in disposable repositories. Do not troubleshoot by adding write-all, printing secrets/github contexts, disabling organization policy, force-updating refs, registering a personal machine as a runner, or pointing tests at production databases.

1. Diagnostic sequence: preserve → scope → inspect → minimally correct → verify

  1. Preserve evidence: run ID/attempt, event, head SHA, job list, matrix names, logs, and workflow file at that SHA.
  2. Identify scope: repository, workflow, job/matrix cell, runner, container/service topology, and direct needs edges.
  3. Inspect: job conclusions/timestamps, exact workflow YAML, service health/addressing, artifact/output contracts, and permissions.
  4. Correct the smallest violated contract: transfer/reproduce the file, bound the matrix, add health readiness, fix the address, or change failure policy deliberately.
  5. Verify with a new run: bind the repair to a new SHA while retaining the failed run as historical evidence.
REPO="OWNER/REPOSITORY"
RUN_ID="123456789"

gh run view "$RUN_ID" -R "$REPO"   --json event,headBranch,headSha,status,conclusion,jobs,url

gh run view "$RUN_ID" -R "$REPO" --log-failed

gh api -H "X-GitHub-Api-Version: 2026-03-10"   "repos/$REPO/actions/runs/$RUN_ID/jobs?filter=latest&per_page=100"   --jq '.jobs[] | {id,name,status,conclusion,runner_name,started_at,completed_at}'

2. Intentionally broken example: needs does not carry files

Suppose the build job creates a file and the test job assumes needs: build mounted the previous workspace:

jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - run: mkdir -p dist && printf 'binary\n' > dist/app.bin

  test:
    needs: build
    runs-on: ubuntu-latest
    steps:
      - run: |
          ls -la
          test -f dist/app.bin

Expected failure: the test job runs on a fresh runner; dist/app.bin is absent and test -f exits non-zero. The original cause is topology, not “flaky disk.”

Repair A: if test only needs repository source, check out the same exact SHA and rebuild deterministically. Repair B: if test must consume the exact generated bytes, upload/download a workflow artifact using reviewed, pinned action identities. Do not solve the problem with a shared self-hosted filesystem unless that shared state is an intentional, governed architecture.

3. Matrix explosion: count before you run

A matrix can multiply silently during maintenance. A three-dimensional matrix with 4 operating systems, 5 runtime versions, and 3 database modes creates 60 base jobs. Add more dimensions and you can approach GitHub’s current 256-job matrix cap quickly.

# Evidence after a run: count generated jobs with a stable name prefix.
gh run view "$RUN_ID" -R "$REPO" --json jobs   --jq '[.jobs[] | select(.name | startswith("test ("))] | length'

Correction choices: remove unsupported combinations with exclude, move exceptional lanes to explicit scheduled/manual workflows, reduce dimensions, or cap max-parallel. Do not assume free public minutes means unlimited concurrency or external dependency capacity.

4. Service failure: separate lifecycle, readiness, and application protocol

Symptom Likely layer Evidence Least-destructive correction
Connection refused immediately Service not ready / wrong mapped port Service logs/health + job address/port Add health check; correct port mapping; bounded readiness probe.
Name redis does not resolve in host job Wrong network model Job has no container:; service has mapped host port Use 127.0.0.1:<host-port>.
127.0.0.1 fails from containerized job Wrong network model Job uses container: Use service-label hostname on shared Docker network.
TCP connects but test fails Application/protocol/data layer Client response and service logs Fix fixture/protocol/config; do not add network retries blindly.

A health check should express service readiness. It is not a blanket guarantee that schema migration, fixture loading, or external dependencies are ready; those require their own explicit preparation step.

5. Container networking and filesystem assumptions

Container jobs run their ordinary steps inside the declared image and default to sh. Container actions may run as sibling containers with shared mounts/network. A command that relies on a host-only path, a Bash-only feature, or an uninstalled utility can therefore fail even though the same command worked on the raw Ubuntu runner.

Diagnose the image and shell first. Do not install a large toolchain into the image at runtime until you decide whether the job container is actually the right abstraction.

6. fail-fast can hide coverage; continue-on-error can hide failure

If one matrix cell fails early and default fail-fast: true cancels the rest, the run may lack evidence about whether the defect affects one platform or every platform. That can be acceptable for fast presubmit feedback but weak for a release qualification matrix.

The opposite mistake is more dangerous: continue-on-error: true on a required compatibility lane can make the workflow appear successful even when the lane fails. Tolerated failures must be deliberate, visible, and non-authoritative for release policy.

7. Security-sensitive boundaries in topology debugging

  • Container registry credentials: never print them; use least privilege and short lifetime where supported.
  • Service data: use synthetic fixtures; do not copy production databases into public-repository service jobs.
  • Self-hosted runners: do not register one merely to share files or reach an internal service; Chapter 16 treats runner trust explicitly.
  • Ref/history repair: do not force-push to make a run “match” expected state. Create a new commit/run and preserve the failing SHA.
  • Permission repair: filesystem/network mistakes do not justify broadening GITHUB_TOKEN.

8. Performance and billing only where topology causes them

Every matrix cell is a job with startup and execution cost. Every job-container/service image may require pulling layers. Private repositories consume plan-based runner minutes/storage; public standard runners are currently free, but concurrency still has account limits and external services still have quotas. Optimize after observing job timestamps and bottlenecks, not by collapsing isolation blindly.

Typical safe optimizations are: reduce redundant matrix cells, cap max-parallel, use a smaller approved image, reuse dependency caches appropriately, split slow optional lanes from blocking presubmit, or move a stable generated file through an artifact instead of rebuilding it repeatedly. Each optimization must preserve the evidence release policy needs.

9. Verification checklist after a repair

  • New run binds to the intended new commit SHA.
  • Generated job count matches the reviewed matrix count.
  • Every required matrix cell reaches an explicit conclusion instead of accidental cancellation.
  • Cross-job files use an explicit artifact/package or reproducible rebuild/checkout boundary.
  • Service health and client protocol checks both succeed using the correct network address.
  • No permission, secret, branch-policy, runner-registration, or history-rewrite workaround was introduced.

10. Lesson summary

Topology failures become tractable when you classify them by runner/workspace boundary, dependency graph, matrix generation, container user-space, service network/readiness, and result policy. Preserve the failed run, correct only the violated contract, and verify the repair in a new run. Do not hide execution-model mistakes with retries, broad permissions, persistent runners, or tolerated failures.

Knowledge check

A downstream job fails with “dist/app.bin: No such file.” The upstream job created it successfully. First diagnosis?

A host-runner job cannot resolve hostname redis, but the service is healthy. What likely assumption is wrong?

One matrix cell fails and three others are cancelled. What setting should you inspect before calling the cancellations independent failures?

Why is making a failing compatibility job continue-on-error: true a governance decision?

Why should the fixed run use a new commit instead of rewriting the old SHA?

Next lesson

Next: Checkpoint Lab — Jobs, Steps, Matrices, Containers, Service Containers, and Dependency Graphs

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.