Chapter 15Lesson 03~150 minutes

Jobs, Steps, Matrices, Containers, Service Containers, and Dependency Graphs: Configuration, Design Choices, and Tradeoffs

Splitting work into more jobs or matrix cells is not automatically more scalable. Every boundary creates scheduling, state-transfer, permission, image, networking, and observability consequences. This lesson turns those consequences into explicit design choices rather than YAML habits.

Design tradeoffsCI topologyResource governanceCompatibility

Learning objectives

  • Choose one job or multiple jobs based on isolation, state-transfer, permission, and observability requirements.
  • Choose a matrix only when combinations share one execution shape and keep dimensions bounded.
  • Compare host-runner dependency installation with a job container for reproducibility, startup cost, compatibility, and supply-chain control.
  • Compare service containers with external test services by lifecycle, data isolation, credentials, fidelity, latency, and cost.
  • Apply an explicit decision table covering maintainability, security, governance, reliability, compatibility, and cost.

Availability and deployment boundary: The mandatory design path remains GitHub.com + public GitHub Free repositories + standard GitHub-hosted Ubuntu runners. Private-repository minutes/storage/concurrency depend on plan. Container jobs and service containers require Linux. GitHub Enterprise Server and self-hosted runner topologies differ; this chapter does not ask learners to register a runner.

1. One job versus several isolated jobs

A single job is attractive when steps use the same workspace, installed tools, and permission context and the work is short enough that isolation adds little value. Multiple jobs become valuable when you need independent scheduling, different runners/containers, clearer policy boundaries, parallelism, or a release gate that should depend on separately visible evidence.

The cost of splitting is real: each GitHub-hosted job starts a fresh runner, so source must be checked out again and generated files must be transferred or reproduced. A job boundary should therefore express an architectural boundary, not merely make the YAML look organized.

Question Prefer one job Prefer multiple jobs
Filesystem state Several steps intentionally share generated local state. State should be isolated/reproduced/transferred explicitly.
Runner/toolchain Same runner and toolchain. Different OS, architecture, container, permissions, or service topology.
Parallelism Sequential work is intrinsic. Independent work can safely run in parallel.
Evidence One result is enough. Separate check/result ownership matters to policy.

2. Matrix dimensions versus explicit jobs

A matrix is best when the steps are structurally the same and only bounded parameters vary. Explicit jobs are clearer when lanes have different purpose, permissions, service dependencies, or failure policy. Forcing unrelated jobs into a clever matrix saves lines of YAML but hides intent.

Review the Cartesian product in code review. A matrix that begins as 2 × 2 can become 4 × 5 × 3 after “one more platform,” creating 60 jobs. GitHub’s 256-job matrix ceiling prevents unbounded expansion, but a workflow can become costly and slow long before the hard limit.

3. Complete coverage versus fast failure

fail-fast: true is a latency/resource optimization: once a non-tolerated cell fails, queued/in-progress siblings may be cancelled. fail-fast: false is an evidence policy: every cell is allowed to finish, which is valuable when platform-specific failures matter. Neither choice is universally correct.

continue-on-error should be rarer. It changes whether a failure is treated as fatal. Every tolerated cell needs an owner, reason, and expiration condition. Otherwise “experimental” becomes a permanent hole in release evidence.

4. Job container versus installing dependencies on the runner

Dimension Job container Host runner + install/setup
Reproducibility Strong user-space image boundary; pin digest for stronger identity. Depends on runner image + setup commands/package repositories.
Startup Image pull/start overhead. Setup/install overhead; preinstalled tools may be fast.
OS support Linux runner required for job container. Linux/Windows/macOS according to runner.
Shell/filesystem Container-specific; default run shell is sh. Runner image defaults and paths.
Supply chain Container registry/image provenance must be governed. Package managers/setup actions are dependencies to govern.
Debugging Closer to containerized production if image is representative. Closer to developer/host runner environment.

A job container is not a security sandbox that makes untrusted code safe by itself. The underlying runner, mounted workspace, token permissions, and service/network access still matter.

5. Service container versus external test service

A service container is excellent for disposable databases/caches that can start per job, contain synthetic data, and be destroyed with the job. An external service is appropriate when tests require behavior that cannot be reproduced locally—managed identity, proprietary API, region semantics, or shared integration infrastructure—but then credentials, tenancy, cleanup, network latency, rate limits, and billing become part of the test design.

Concern Service container External test service
Lifecycle Job-scoped, runner-managed. Independent; requires explicit ownership/cleanup.
Secrets Often none for local synthetic fixture. Usually credentials/secrets/federation.
Fidelity Image-level approximation. Potentially closer to managed production behavior.
Isolation Per-job by default. May be shared; collision/data leakage risk.
Latency/cost Local network + image pull. Network latency and provider cost/quotas.

6. Outputs versus artifacts versus reproducible checkout

Do not use outputs as a file-transfer mechanism. Outputs are strings and are ideal for identifiers, digests, version numbers, or decisions. Workflow artifacts are designed to move files between jobs and preserve selected outputs after a run. Reproducible checkout is preferable when downstream work only needs repository source at an exact SHA.

# Documentation shape only: pin artifact actions to reviewed full commit SHAs in production.
- name: Upload generated files
  uses: actions/upload-artifact@<PINNED_FULL_COMMIT_SHA>
  with:
    name: build-output
    path: dist/

# In a dependent job:
- name: Download generated files
  uses: actions/download-artifact@<PINNED_FULL_COMMIT_SHA>
  with:
    name: build-output
    path: dist/

Artifacts have retention, storage, integrity, and sensitive-content implications covered deeply in Chapter 17. The design rule here is narrower: generated files do not teleport between jobs.

7. Plan, runner, registry, and deployment constraints

  • Public GitHub.com: standard GitHub-hosted runners are currently free and unlimited, but concurrency and other Actions limits still apply.
  • Private repositories: hosted-runner minutes/storage and concurrency depend on plan and may incur billing.
  • Containers/services: GitHub-hosted jobs must use Ubuntu; self-hosted jobs require Linux plus Docker.
  • Private container images: registry authentication creates a credential boundary; scope credentials minimally and never print them.
  • GitHub Enterprise Server: use the runner model supported by the deployed GHES release; do not assume GitHub.com-hosted runner availability.

8. Worked decision table: choose the topology

Scenario Maintainability Security/governance Reliability Compatibility Cost Choice
Small lint + unit test share same source/tools Low complexity in one job One permission context is acceptable No cross-job transfer Same Ubuntu toolchain One runner startup One job, multiple steps
Same test suite across 3 supported modes One definition beats copy/paste Bound dimensions/review additions fail-fast policy explicit Same runner/toolchain 3 cells; cap parallelism Small matrix
Build once, consume generated binary in two checks Clear producer/consumers Artifact access/retention explicit Consumers use identical bytes Runner-independent if binary supports it Artifact storage + runner jobs Separate jobs + artifact
Integration test needs ephemeral Redis Simple job-local fixture Synthetic data/no shared credential Health check before test Linux/Docker required Image pull + one runner Service container
Test proprietary managed cloud behavior More external setup Credential/tenant/audit policy required Provider availability affects test Closest fidelity Potential provider cost Optional external integration lane

The best design is the smallest topology that makes the required isolation and evidence explicit. Complexity should buy a concrete property.

9. Production policy examples

  • Every cross-job generated file has an explicit artifact/package contract or is deterministically recreated.
  • Every matrix declares an owner, dimensions, expected cell count, failure policy, and parallelism cap when resource pressure matters.
  • Container image sources are approved; production-critical images are pinned by immutable digest where feasible.
  • Service containers contain synthetic/non-sensitive data and expose only required ports.
  • Tolerated failures are labeled experimental with an owner and removal criterion.
  • Summary jobs report prerequisite results but do not transform failed required validation into a successful release signal.

10. Lesson summary

Job boundaries, matrix dimensions, containers, and services are architectural choices. Use one job for genuine shared execution state, multiple jobs for isolation/parallel evidence, matrices for bounded variations of one shape, job containers for controlled Linux user-space, and service containers for disposable job-local dependencies. Move metadata through outputs and files through artifacts or reproducible source/package boundaries.

Knowledge check

When is an explicit job clearer than a matrix cell?

Why can a job container improve reproducibility without eliminating supply-chain risk?

A generated binary must be tested on two fresh jobs. Output or artifact?

What is the risk of routinely setting continue-on-error: true?

Why cap matrix parallelism even on a free public repository?

Next lesson

Next: Jobs, Steps, Matrices, Containers, Service Containers, and Dependency Graphs: Diagnostics, Failure Modes, Security, and Performance

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.