Jobs, Steps, Matrices, Containers, Service Containers, and Dependency Graphs: Configuration, Design Choices, and Tradeoffs
Splitting work into more jobs or matrix cells is not automatically more scalable. Every boundary creates scheduling, state-transfer, permission, image, networking, and observability consequences. This lesson turns those consequences into explicit design choices rather than YAML habits.
Learning objectives
- Choose one job or multiple jobs based on isolation, state-transfer, permission, and observability requirements.
- Choose a matrix only when combinations share one execution shape and keep dimensions bounded.
- Compare host-runner dependency installation with a job container for reproducibility, startup cost, compatibility, and supply-chain control.
- Compare service containers with external test services by lifecycle, data isolation, credentials, fidelity, latency, and cost.
- Apply an explicit decision table covering maintainability, security, governance, reliability, compatibility, and cost.
Availability and deployment boundary: The mandatory design path remains GitHub.com + public GitHub Free repositories + standard GitHub-hosted Ubuntu runners. Private-repository minutes/storage/concurrency depend on plan. Container jobs and service containers require Linux. GitHub Enterprise Server and self-hosted runner topologies differ; this chapter does not ask learners to register a runner.
1. One job versus several isolated jobs
A single job is attractive when steps use the same workspace, installed tools, and permission context and the work is short enough that isolation adds little value. Multiple jobs become valuable when you need independent scheduling, different runners/containers, clearer policy boundaries, parallelism, or a release gate that should depend on separately visible evidence.
The cost of splitting is real: each GitHub-hosted job starts a fresh runner, so source must be checked out again and generated files must be transferred or reproduced. A job boundary should therefore express an architectural boundary, not merely make the YAML look organized.
| Question | Prefer one job | Prefer multiple jobs |
|---|---|---|
| Filesystem state | Several steps intentionally share generated local state. | State should be isolated/reproduced/transferred explicitly. |
| Runner/toolchain | Same runner and toolchain. | Different OS, architecture, container, permissions, or service topology. |
| Parallelism | Sequential work is intrinsic. | Independent work can safely run in parallel. |
| Evidence | One result is enough. | Separate check/result ownership matters to policy. |
2. Matrix dimensions versus explicit jobs
A matrix is best when the steps are structurally the same and only bounded parameters vary. Explicit jobs are clearer when lanes have different purpose, permissions, service dependencies, or failure policy. Forcing unrelated jobs into a clever matrix saves lines of YAML but hides intent.
Review the Cartesian product in code review. A matrix that begins as 2 × 2 can become 4 × 5 × 3 after “one more platform,” creating 60 jobs. GitHub’s 256-job matrix ceiling prevents unbounded expansion, but a workflow can become costly and slow long before the hard limit.
3. Complete coverage versus fast failure
fail-fast: true is a latency/resource optimization:
once a non-tolerated cell fails, queued/in-progress siblings may be
cancelled. fail-fast: false is an evidence policy:
every cell is allowed to finish, which is valuable when
platform-specific failures matter. Neither choice is universally
correct.
continue-on-error should be rarer. It changes whether a
failure is treated as fatal. Every tolerated cell needs an owner,
reason, and expiration condition. Otherwise “experimental” becomes a
permanent hole in release evidence.
4. Job container versus installing dependencies on the runner
| Dimension | Job container | Host runner + install/setup |
|---|---|---|
| Reproducibility | Strong user-space image boundary; pin digest for stronger identity. | Depends on runner image + setup commands/package repositories. |
| Startup | Image pull/start overhead. | Setup/install overhead; preinstalled tools may be fast. |
| OS support | Linux runner required for job container. | Linux/Windows/macOS according to runner. |
| Shell/filesystem |
Container-specific; default run shell is sh.
|
Runner image defaults and paths. |
| Supply chain | Container registry/image provenance must be governed. | Package managers/setup actions are dependencies to govern. |
| Debugging | Closer to containerized production if image is representative. | Closer to developer/host runner environment. |
A job container is not a security sandbox that makes untrusted code safe by itself. The underlying runner, mounted workspace, token permissions, and service/network access still matter.
5. Service container versus external test service
A service container is excellent for disposable databases/caches that can start per job, contain synthetic data, and be destroyed with the job. An external service is appropriate when tests require behavior that cannot be reproduced locally—managed identity, proprietary API, region semantics, or shared integration infrastructure—but then credentials, tenancy, cleanup, network latency, rate limits, and billing become part of the test design.
| Concern | Service container | External test service |
|---|---|---|
| Lifecycle | Job-scoped, runner-managed. | Independent; requires explicit ownership/cleanup. |
| Secrets | Often none for local synthetic fixture. | Usually credentials/secrets/federation. |
| Fidelity | Image-level approximation. | Potentially closer to managed production behavior. |
| Isolation | Per-job by default. | May be shared; collision/data leakage risk. |
| Latency/cost | Local network + image pull. | Network latency and provider cost/quotas. |
6. Outputs versus artifacts versus reproducible checkout
Do not use outputs as a file-transfer mechanism. Outputs are strings and are ideal for identifiers, digests, version numbers, or decisions. Workflow artifacts are designed to move files between jobs and preserve selected outputs after a run. Reproducible checkout is preferable when downstream work only needs repository source at an exact SHA.
# Documentation shape only: pin artifact actions to reviewed full commit SHAs in production.
- name: Upload generated files
uses: actions/upload-artifact@<PINNED_FULL_COMMIT_SHA>
with:
name: build-output
path: dist/
# In a dependent job:
- name: Download generated files
uses: actions/download-artifact@<PINNED_FULL_COMMIT_SHA>
with:
name: build-output
path: dist/
Artifacts have retention, storage, integrity, and sensitive-content implications covered deeply in Chapter 17. The design rule here is narrower: generated files do not teleport between jobs.
7. Plan, runner, registry, and deployment constraints
- Public GitHub.com: standard GitHub-hosted runners are currently free and unlimited, but concurrency and other Actions limits still apply.
- Private repositories: hosted-runner minutes/storage and concurrency depend on plan and may incur billing.
- Containers/services: GitHub-hosted jobs must use Ubuntu; self-hosted jobs require Linux plus Docker.
- Private container images: registry authentication creates a credential boundary; scope credentials minimally and never print them.
- GitHub Enterprise Server: use the runner model supported by the deployed GHES release; do not assume GitHub.com-hosted runner availability.
8. Worked decision table: choose the topology
| Scenario | Maintainability | Security/governance | Reliability | Compatibility | Cost | Choice |
|---|---|---|---|---|---|---|
| Small lint + unit test share same source/tools | Low complexity in one job | One permission context is acceptable | No cross-job transfer | Same Ubuntu toolchain | One runner startup | One job, multiple steps |
| Same test suite across 3 supported modes | One definition beats copy/paste | Bound dimensions/review additions | fail-fast policy explicit |
Same runner/toolchain | 3 cells; cap parallelism | Small matrix |
| Build once, consume generated binary in two checks | Clear producer/consumers | Artifact access/retention explicit | Consumers use identical bytes | Runner-independent if binary supports it | Artifact storage + runner jobs | Separate jobs + artifact |
| Integration test needs ephemeral Redis | Simple job-local fixture | Synthetic data/no shared credential | Health check before test | Linux/Docker required | Image pull + one runner | Service container |
| Test proprietary managed cloud behavior | More external setup | Credential/tenant/audit policy required | Provider availability affects test | Closest fidelity | Potential provider cost | Optional external integration lane |
The best design is the smallest topology that makes the required isolation and evidence explicit. Complexity should buy a concrete property.
9. Production policy examples
- Every cross-job generated file has an explicit artifact/package contract or is deterministically recreated.
- Every matrix declares an owner, dimensions, expected cell count, failure policy, and parallelism cap when resource pressure matters.
- Container image sources are approved; production-critical images are pinned by immutable digest where feasible.
- Service containers contain synthetic/non-sensitive data and expose only required ports.
- Tolerated failures are labeled experimental with an owner and removal criterion.
- Summary jobs report prerequisite results but do not transform failed required validation into a successful release signal.
10. Lesson summary
Job boundaries, matrix dimensions, containers, and services are architectural choices. Use one job for genuine shared execution state, multiple jobs for isolation/parallel evidence, matrices for bounded variations of one shape, job containers for controlled Linux user-space, and service containers for disposable job-local dependencies. Move metadata through outputs and files through artifacts or reproducible source/package boundaries.
Knowledge check
When is an explicit job clearer than a matrix cell?
When the lane has a different purpose, permissions, runner/service topology, or failure policy rather than merely different parameter values.
Why can a job container improve reproducibility without eliminating supply-chain risk?
It pins the user-space environment more tightly, but the image itself is a dependency whose registry, tag/digest, provenance, and vulnerabilities still require governance.
A generated binary must be tested on two fresh jobs. Output or artifact?
Artifact (or a package system), because it is a file. An output is for small string metadata such as an ID or digest.
What is the risk of routinely setting
continue-on-error: true?
Real validation failures can stop affecting workflow/release status, creating apparent green evidence without real confidence.
Why cap matrix parallelism even on a free public repository?
Resource pressure still matters: concurrency limits, external service capacity, rate limits, logs, and feedback quality can degrade even when standard runner minutes are free.
Further reading — current official GitHub sources
- GitHub Docs — Workflow syntax
- GitHub Docs — Using jobs in a workflow
- GitHub Docs — Running variations of jobs
- GitHub Docs — GitHub-hosted runners
- GitHub Docs — Running jobs in a container
- GitHub Docs — Docker service containers
- GitHub Docs — Store and share workflow data
- GitHub CLI — gh run view
- GitHub REST — Workflow jobs
- GitHub Docs — Actions limits
- GitHub Docs — Secure use reference
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.