Matrix Strategies, Dynamic Matrices, Fail-Fast, and Parallel Test Design: Core Concepts and Mental Model
Chapter 11 made job ordering and mutual exclusion explicit. Chapter 12 asks a different question: when one logical test job must be repeated across operating systems, runtimes, feature modes or other compatibility dimensions, how do we expand that work without losing control of cost, interpretation or evidence?
Learning objectives
- Explain a matrix as a job-definition expansion mechanism rather than a loop running inside one runner.
-
Predict Cartesian combinations and reason precisely about
includeandexclude. - Distinguish matrix definition/generation from per-cell runner, toolchain, logs, artifacts and conclusions.
-
Explain the interaction of
fail-fast, per-jobcontinue-on-errorandmax-parallel. - Define an aggregation policy that treats success, failure, cancelled and skipped cells as different evidence states.
1. The practical problem: one test definition, many support claims
A project may claim support for two operating systems and two runtime versions. Copying four jobs works initially, but the jobs drift: one gets a different setup action, another loses a test flag, and a third silently uses a newer runtime. A matrix lets one job definition expand into multiple generated jobs so the repeated structure stays consistent.
The important word is generated. The four cells are not four iterations inside one runner. Each cell becomes its own job with its own runner assignment, filesystem, logs, status, timing and possibly artifacts. Compatibility evidence therefore exists per cell.
2. Causal model: definition → expansion → policy → evidence
GitHub first evaluates the matrix definition. Static arrays or
JSON-derived values produce combinations.
exclude removes matching combinations;
include can enrich original combinations or add new
ones. Each remaining combination becomes a job instance. Strategy
policy then controls parallelism and failure propagation. Downstream
work must interpret the resulting cell conclusions and evidence.
flowchart TD A[Static axes or generated JSON] --> B[Matrix expansion] B --> C[Exclude matches] C --> D[Apply include objects] D --> E1[Cell 0: OS/runtime] D --> E2[Cell 1: OS/runtime] D --> E3[Cell N: experimental] E1 --> F[Runner + pinned toolchain] E2 --> F2[Runner + pinned toolchain] E3 --> F3[Runner + pinned toolchain] F --> G[Cell conclusion + evidence] F2 --> G F3 --> G G --> H[Aggregation / support decision]
The matrix answers which compatibility cells exist? The runner and setup actions answer what executed each cell? The strategy answers how much may execute at once and what happens when a cell fails? The aggregator answers what do the collected results mean for support?
3. Name the state before changing it
| State | Example | Owner | Evidence |
|---|---|---|---|
| Matrix source |
os, python arrays or planner JSON
|
Workflow/job output | workflow revision + planner output |
| Expanded cell | {os: ubuntu-24.04, python: 3.13} |
GitHub Actions run graph | generated job name / matrix context |
| Runner/toolchain | Ubuntu 24.04 + Python 3.13 | Runner + setup action | runner/image metadata + python --version |
| Strategy policy | fail-fast, max-parallel |
Workflow revision | evaluated strategy context |
| Tolerance |
experimental: true → job
continue-on-error
|
Workflow revision | matrix value + job conclusion |
| Cell evidence | test report / manifest | Runner then artifact service | artifact ID/digest + logs |
| Support decision | mandatory cells pass; experimental may fail | Project policy | aggregation report with source run/SHA |
4. Cartesian expansion: count before you run
Two axes multiply. Three operating systems × four runtimes × two
database modes creates 24 jobs before any
include entries are added. GitHub currently caps matrix
expansion at 256 generated jobs per workflow run, but a matrix well
below that limit can still waste minutes and obscure failures.
jobs:
compatibility:
strategy:
matrix:
os: [ubuntu-24.04, windows-2025]
python: ['3.12', '3.13']
runs-on: ${{ matrix.os }}
steps:
- run: echo "cell=${{ matrix.os }} python=${{ matrix.python }}"
This produces four jobs. The variable order influences job creation order, but runner availability and scheduling mean creation order is not a completion-order guarantee.
5. exclude removes matches;
include enriches or adds
An exclude object can be a partial match: every
generated combination matching all listed key/value pairs is
removed. GitHub processes include after
exclude. An include object is merged into original
combinations when that can happen without overwriting original
matrix values; otherwise it becomes an additional combination.
strategy:
matrix:
os: [ubuntu-24.04, windows-2025]
python: ['3.12', '3.13']
exclude:
- os: windows-2025
python: '3.12'
include:
- os: ubuntu-24.04
python: '3.13'
experimental: true
- os: ubuntu-24.04
python: '3.13'
mode: future-api
experimental: true
Do not infer the final cell count from indentation. Expand it on paper or with a small local script first, especially when multiple include objects can create additional combinations.
6. Failure policy belongs to the support model
fail-fast: true means a non-tolerated matrix failure
can cancel other queued/in-progress cells. That is valuable when the
remaining evidence has little value after the first mandatory
failure. continue-on-error is different: it is
evaluated for one generated job and can mark an experimental cell as
tolerated so its failure does not trigger fail-fast for the rest.
strategy:
fail-fast: true
matrix:
version: ['3.12', '3.13']
experimental: [false]
include:
- version: '3.15-dev'
experimental: true
continue-on-error: ${{ matrix.experimental }}
A tolerated experimental failure, a cancelled cell, or a skipped downstream job can coexist with an overall result that does not communicate the nuance you need. Preserve per-cell conclusions and evidence.
7. max-parallel is a capacity bound, not a lock
Without max-parallel, GitHub tries to run as many cells
as runner availability allows. Setting
max-parallel: 2 means at most two generated jobs from
that matrix execute simultaneously. It does not serialize other
workflows, reserve a deployment target, or replace Chapter 11
concurrency groups.
strategy:
fail-fast: false
max-parallel: 2
matrix:
os: [ubuntu-24.04, windows-2025]
python: ['3.12', '3.13']
8. Read-only inspection before editing a matrix
Before changing dimensions, capture the current matrix source, recent run IDs/attempts, the generated job list, runner labels, tool versions and any cells that are skipped/cancelled. If a dynamic matrix is involved, preserve the exact JSON output that produced it. Otherwise you may “fix” a cell that did not exist in the failing run.
# Optional read-only inspection in an authorized repository
gh run list --limit 10
gh run view RUN_ID --json databaseId,attempt,headSha,event,status,conclusion,jobs
Knowledge check
A 2×3×4 matrix has how many Cartesian cells before exclude/include?
24. Multiply the axis lengths before considering removals or added include combinations.
Does one matrix cell share a filesystem with another cell?
No. Each generated matrix cell is a distinct job and receives its own runner execution boundary.
What does fail-fast: true do after a non-tolerated
cell fails?
It can cancel queued and in-progress jobs in that matrix. It does not roll back evidence or external effects already produced.
Why is max-parallel: 2 not a deployment
lock?
It limits simultaneous jobs in one matrix. It does not coordinate unrelated jobs or workflow runs that target the same external system.
A workflow is green but an experimental cell failed. What evidence should a support decision use?
The cell-level matrix values, conclusion/tolerance policy and retained test evidence—not the top-level color alone.
Official references and version notes
- GitHub Docs — Running variations of jobs in a workflow — matrix expansion, contexts, include/exclude, dynamic outputs, failure handling and max-parallel.
- GitHub Docs — workflow syntax: strategy.matrix — current matrix limit and strategy semantics.
- GitHub Docs — expressions: fromJSON — converting job-output JSON into arrays/objects used by a later matrix.
- GitHub Docs — strategy context — fail-fast, job-index, job-total and max-parallel for the current generated job.
- GitHub Docs — matrix strategy with reusable workflows — matrix-driven reusable workflow calls and output caveats.
- actions/upload-artifact v7.0.1 and actions/download-artifact v8.0.1 — action releases pinned by full commit SHA in the checkpoint.
Version-sensitive behavior was rechecked against current
GitHub-maintained documentation on 2026-09-09. A
matrix can generate at most
256 jobs per workflow run.
strategy.fail-fast defaults to true;
continue-on-error is evaluated per generated job; and
max-parallel limits simultaneous matrix jobs but does
not redefine the matrix or create a repository-wide concurrency
lock. Current mandatory examples target GitHub.com with versioned
runner labels ubuntu-24.04 and
windows-2025. Official actions are pinned to full
immutable commit SHAs: checkout v7.0.1, setup-python v7.0.0,
upload-artifact v7.0.1 and download-artifact v8.0.1. GitHub
Enterprise Server users must verify artifact-action
backend/major-version compatibility for their appliance before
copying these examples.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.