Jobs, Steps, Matrices, Containers, Service Containers, and Dependency Graphs: Concepts, Architecture, and Mental Model
A workflow can contain correct commands and still produce unreliable CI if its execution graph is misunderstood. Files may vanish between jobs, a matrix may multiply work far beyond expectation, a service may be started but not ready, or one “allowed failure” may quietly turn evidence into false confidence. This lesson treats topology itself as a production control.
Learning objectives
- Explain exactly what state is shared between steps in one job and what is isolated between jobs.
-
Model
needsas a dependency graph with explicit success, failure, skip, cancel, and downstream-condition semantics. -
Predict matrix expansion,
include/exclude,fail-fast,continue-on-error, andmax-parallelbehavior before a run starts. - Distinguish a job container, a container action, and a service container by execution responsibility and network/filesystem boundary.
- Choose the correct way to move small metadata, source state, and generated files across job boundaries.
Availability: The mandatory path uses GitHub.com,
GitHub Free, a disposable public personal repository, and standard
ubuntu-latest GitHub-hosted runners. Standard hosted
runners are currently free and unlimited for public repositories.
Container jobs and service containers require Linux runners; on
GitHub-hosted runners that means an Ubuntu runner. Self-hosted
runners are conceptual only and are deferred to Chapter 16.
1. The problem: the graph is part of the test
Chapter 13 established what a workflow run is and Chapter 14 established when it activates and how small values move through outputs. Chapter 15 adds a new source of correctness risk: execution topology. A “build then test” pipeline is not one long shell session. It is a graph of jobs, and each job may execute on a different fresh runner, expand into multiple matrix children, enter a job container, or depend on one or more service containers.
If a team cannot draw that graph and label its state boundaries, it cannot reliably explain why a file exists, which source revision was tested, why a dependent job skipped, how many tests actually ran, or which network address a database should use. Those are change-control questions, not cosmetic YAML details.
2. Mental model: one workflow run, multiple isolated execution islands
flowchart LR
E["Workflow run Event + SHA"] --> B["build job Runner A"]
B -->|needs + output SHA| M1["test fast Runner B"]
B -->|needs + output SHA| M2["test full Runner C"]
B --> C["container probe Ubuntu runner + job container"]
B --> S["service probe Runner D + Redis service"]
M1 --> A["aggregate"]
M2 --> A
C --> A
S --> A
Each job is scheduled independently. The dependency arrows carry ordering/result metadata, not a shared filesystem. Matrix cells become separate jobs. A job container changes where that job’s steps execute; a service container is a sibling dependency scoped to one job.
Every box has its own runner lifecycle. On GitHub-hosted runners, each job receives a fresh runner instance. That is the first invariant to preserve when you reason about files, installed packages, background processes, and caches.
3. Job isolation versus step continuity
Jobs are the scheduling and isolation unit. They
run in parallel by default unless a dependency or another
concurrency control delays them. Steps are ordered
operations inside one job. All steps in a GitHub-hosted job run on
the same runner and can see changes in that job’s filesystem,
including GITHUB_WORKSPACE.
Do not overstate that continuity: each run: step starts
a new shell process. A shell-local export FOO=bar does
not automatically persist to the next step. Persist environment
values with GitHub’s workflow commands such as
GITHUB_ENV, persist small calculated values with
GITHUB_OUTPUT, and persist files across
jobs with an artifact/package/cache mechanism or recreate
them deterministically.
jobs:
same_job:
runs-on: ubuntu-latest
steps:
- run: printf 'evidence\n' > evidence.txt
- run: test -f evidence.txt # same runner/workspace: succeeds
different_job:
runs-on: ubuntu-latest
steps:
- run: test ! -e evidence.txt # fresh runner: no implicit file transfer
4. needs is an ordering/result edge, not a disk mount
jobs.<job_id>.needs says that one job depends on
one or more prerequisite jobs. By default the downstream job runs
only when those prerequisites complete successfully. If a required
job fails or is skipped, dependent jobs in that chain are skipped
unless their own condition deliberately permits execution—for
example an evidence-summary job using
if: ${{ always() }}.
jobs:
build:
runs-on: ubuntu-latest
test:
needs: build
runs-on: ubuntu-latest
evidence:
if: ${{ always() }}
needs: [build, test]
runs-on: ubuntu-latest
steps:
- env:
BUILD_RESULT: ${{ needs.build.result }}
TEST_RESULT: ${{ needs.test.result }}
run: printf 'build=%s test=%s\n' "$BUILD_RESULT" "$TEST_RESULT"
The needs context exposes result/output metadata from
directly named prerequisite jobs. It does not copy the prerequisite
runner’s workspace. Treat the graph edge as a contract: ordering +
declared outputs + result state.
5. A matrix is a job generator
A matrix strategy takes one job definition and expands it into job
instances. A two-value os dimension and three-value
runtime dimension create six jobs before exclusions or
additions. GitHub currently caps one matrix at 256 generated jobs
per workflow run, so dimension multiplication must be reviewed like
resource allocation.
strategy:
fail-fast: false
max-parallel: 2
matrix:
mode: [fast, thorough]
feature: [off, on]
exclude:
- mode: thorough
feature: off
include:
- mode: compatibility
feature: on
The base Cartesian product produces four combinations,
exclude removes one, and include adds a
deliberately named extra combination: four final jobs.
max-parallel: 2 limits simultaneous matrix execution
even if more runners are available.
6. Failure policy changes the evidence you collect
| Control | Scope | Current behavior | Operational question |
|---|---|---|---|
strategy.fail-fast |
Whole matrix | Defaults to true; a non-tolerated failure can cancel queued/in-progress siblings. | Do you prefer quick feedback or complete diagnostic coverage? |
continue-on-error |
One job/cell | Allows that job failure not to fail the workflow as normal. | Is this cell genuinely experimental, and who owns removing the exception? |
max-parallel |
Matrix | Caps simultaneous matrix jobs. | What runner/concurrency/service capacity is safe? |
needs...result |
Dependency consumer | Reports prerequisite result to downstream logic. | Does aggregation preserve a failure or accidentally mask it? |
For diagnostics, fail-fast: false is often useful
because every small matrix cell can finish and produce evidence. For
expensive or obviously redundant matrices, fail-fast can reduce
waste. The policy should be explicit—not an unnoticed default.
7. Job container, container action, and service container are different objects
| Mechanism | What it does | Lifetime / scope | Typical purpose |
|---|---|---|---|
Job container (jobs.<id>.container)
|
Runs ordinary job steps inside one specified container. | One job. | Pin a job’s user-space toolchain/runtime. |
| Container action | Packages one reusable action as a containerized step. | One action invocation. | Encapsulate a reusable operation. |
Service container (jobs.<id>.services)
|
Runs a sibling service managed by the runner. | One job. | Database/cache/message-broker dependency for integration tests. |
Container jobs and service containers require Linux. On a
GitHub-hosted runner use Ubuntu. When a job itself runs in a
container, ordinary run: steps default to
sh rather than bash unless you override
the shell.
8. Service networking depends on where the job runs
GitHub creates and tears down the service containers for the job. If
the job runs inside a job container, the job and
services share a Docker user-defined network and the service label
becomes a hostname such as redis; explicit host-port
mapping is not required for container-to-container access. If the
job runs directly on the runner host, map the
service port and connect through localhost or
127.0.0.1.
jobs:
host_job:
runs-on: ubuntu-latest
services:
redis:
image: redis:7-alpine
ports:
- 6379:6379
options: >-
--health-cmd "redis-cli ping"
--health-interval 5s
--health-timeout 3s
--health-retries 10
steps:
- run: echo "host job reaches Redis at 127.0.0.1:6379"
The service label is still useful metadata, but
redis:6379 is not the correct host-runner address
merely because it would be correct from a containerized job.
9. “Container started” is not the same as “service ready”
Databases and caches may accept Docker start before they are ready for client work. A service health check converts that race into an explicit readiness contract. For an external service, use the provider’s own readiness semantics and bounded retries; do not hide persistent connection failures with an infinite retry loop.
In the live lab, Redis uses its own
redis-cli ping Docker health check, and the test still
performs an application-side PING. Those are different observations:
the runner sees service health, while the test proves the network
path and protocol response.
10. Choose a state-transfer mechanism by data type
| Need | Correct boundary | Why |
|---|---|---|
| Small string / ID / SHA / digest | Step → job output → needs |
Explicit, lightweight metadata contract. |
| Repository source | Checkout the exact same SHA in each job | Reproducible and avoids pretending runner disks are shared. |
| Generated files needed by later jobs | Workflow artifact | Designed for file transfer/storage between jobs. |
| Dependency acceleration | Cache | Optimization only; never make correctness depend on cache presence. |
| Deployable release package | Package/release/artifact system | Needs version, integrity, permissions, retention, and provenance policy. |
11. Inspect the graph before changing it
Read-only evidence should answer how many jobs GitHub actually generated, which matrix names ran, how long they took, and which conclusion each produced.
REPO="OWNER/REPOSITORY"
RUN_ID="123456789"
gh run view "$RUN_ID" -R "$REPO" --json event,headBranch,headSha,status,conclusion,jobs,url --jq '{event,headBranch,headSha,status,conclusion,jobs:[.jobs[]|{name,conclusion,startedAt,completedAt,databaseId}]}'
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/runs/$RUN_ID/jobs?filter=latest&per_page=100" --jq '.jobs[] | {id,name,status,conclusion,runner_name,started_at,completed_at}'
These are hosted job records. They prove what GitHub scheduled and observed; they do not prove that an external system behaved correctly unless the workflow recorded that external evidence.
12. Why this matters in DevOps
Topology controls four production properties at once: correctness (which state crosses boundaries), reliability (what happens after failure), resource usage (how much matrix concurrency and container work is created), and evidence (whether release policy can reconstruct exactly what ran). A fast pipeline with ambiguous boundaries is not mature CI.
13. Lesson summary
A job is an isolated scheduling unit; steps share its
runner/workspace but not one persistent shell process.
needs creates ordering/result/output edges, not
filesystem sharing. Matrices generate jobs and therefore multiply
concurrency and evidence. Job containers change the execution
user-space; service containers provide job-scoped dependencies with
topology-specific networking. Cross-job data must move through
explicit outputs, reproducible source checkout, artifacts, caches,
or package systems according to what the data actually is.
Knowledge check
Step 1 creates report.txt and step 2 in the same
job reads it. Should that work on a GitHub-hosted
runner?
Yes. Steps in one job use the same runner and workspace. The caveat is that shell-process-local variables do not automatically persist between separate run steps.
Build job creates dist/app.bin. Test job has
needs: build. Does the file appear
automatically?
No. needs supplies ordering/result/declared
outputs, not the previous runner filesystem. Upload/download an
artifact or deterministically reproduce the required state.
A 4 × 5 × 3 matrix has how many base combinations?
60. Review the product before adding include rows; GitHub currently allows at most 256 matrix-generated jobs per workflow run.
A host-runner job maps Redis 6379:6379. Should the
client use redis:6379 or
127.0.0.1:6379?
Use localhost/127.0.0.1 with the mapped host port. The service-label hostname is the direct pattern when the job itself runs in a container on the shared Docker network.
Why might a team set fail-fast: false for a small
diagnostic matrix?
So one failing cell does not cancel siblings before they produce useful coverage. The tradeoff is more runner time and potentially slower feedback.
Further reading — current official GitHub sources
- GitHub Docs — Workflow syntax
- GitHub Docs — Using jobs in a workflow
- GitHub Docs — Running variations of jobs
- GitHub Docs — GitHub-hosted runners
- GitHub Docs — Running jobs in a container
- GitHub Docs — Docker service containers
- GitHub Docs — Store and share workflow data
- GitHub CLI — gh run view
- GitHub REST — Workflow jobs
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.