Matrix Strategies, Dynamic Matrices, Fail-Fast, and Parallel Test Design: Guided Hands-On Workflow
This lesson turns the mental model into a disposable compatibility workflow. You will begin with a tiny static matrix, add include/exclude, generate one axis from a prior job, and then observe how fail-fast, tolerated experimental cells and max-parallel change execution without changing what each cell means.
Learning objectives
- Create a small static OS/runtime matrix and predict its generated jobs before dispatch.
-
Use
excludeandincludedeliberately and verify the resulting cell list. -
Generate a runtime axis in a prior job and consume it with
fromJSON. -
Compare
fail-fast: trueandfalse, plus a tolerated experimental cell. -
Measure the effect of
max-parallelwhile preserving run/SHA/runner/tool evidence.
1. Disposable scenario and preflight
Create a throwaway repository named gha-matrix-lab. The
mandatory path uses no secrets, no cloud service, no package
publication and no external mutation. Use GitHub.com Actions,
ubuntu-24.04 and windows-2025. The setup
action is pinned to the verified v7.0.0 commit.
- Repository may be deleted after the chapter.
-
Set workflow
permissions: {}; the lab needs no GitHub API write access. - Record the workflow source SHA before each experiment.
- Predict the number of generated cells before dispatching.
2. Start with the smallest useful static matrix
name: matrix-lab
on:
workflow_dispatch:
permissions: {}
jobs:
test:
name: ${{ matrix.os }} / py-${{ matrix.python }}
runs-on: ${{ matrix.os }}
strategy:
matrix:
os: [ubuntu-24.04, windows-2025]
python: ['3.12', '3.13']
steps:
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: ${{ matrix.python }}
- shell: bash
run: |
python --version
printf 'job-index=%s job-total=%s\n' '${{ strategy.job-index }}' '${{ strategy.job-total }}'
Before running, predict four jobs. After dispatch, verify four generated job names and record the exact Python version each setup action installed. A job name that includes matrix values is evidence-friendly because the compatibility cell is visible without opening YAML.
3. Add exclude and include without
guessing
Remove one combination because the lab does not claim it, then add metadata to one original cell. Keep the change small so you can explain the new expansion.
strategy:
matrix:
os: [ubuntu-24.04, windows-2025]
python: ['3.12', '3.13']
exclude:
- os: windows-2025
python: '3.12'
include:
- os: ubuntu-24.04
python: '3.13'
experimental: true
Now predict three cells. The include object enriches the matching Ubuntu/Python 3.13 original combination; it does not automatically create a duplicate cell. If you need a distinct experimental scenario, make it structurally distinct with another property or use an include-only matrix.
4. Generate one axis as typed JSON
The planner job owns policy about which runtimes are tested. It
writes a JSON array to GITHUB_OUTPUT. The test job
receives that exact string through needs and converts
it back to an array with fromJSON. This is dataflow,
not text substitution.
jobs:
plan:
runs-on: ubuntu-24.04
outputs:
python_versions: ${{ steps.matrix.outputs.python_versions }}
steps:
- id: matrix
shell: bash
run: echo 'python_versions=["3.12","3.13"]' >> "$GITHUB_OUTPUT"
test:
needs: plan
runs-on: ${{ matrix.os }}
strategy:
matrix:
os: [ubuntu-24.04, windows-2025]
python: ${{ fromJSON(needs.plan.outputs.python_versions) }}
If the later matrix is wrong, the planner output is first-failure evidence. Do not rerun after editing the planner without first saving the exact JSON that created the failing graph.
5. Make experimental tolerance explicit
For a cell that is informative but not part of the mandatory support
contract, put the policy in matrix data and map it to job-level
continue-on-error. Do not hide a supported runtime
behind this flag.
continue-on-error: ${{ matrix.experimental }}
strategy:
fail-fast: true
matrix:
include:
- os: ubuntu-24.04
python: '3.13'
experimental: false
- os: windows-2025
python: '3.13'
experimental: false
- os: ubuntu-24.04
python: '3.13'
mode: future-behavior
experimental: true
Intentionally fail only the experimental case with a condition such
as if: matrix.experimental followed by
exit 1. Verify that the tolerated cell is visibly
different from a supported failure and that it does not cause
fail-fast cancellation of mandatory cells.
6. Compare fail-fast true and false
Run two controlled attempts with the same matrix and same
intentional mandatory failure. In attempt A use
fail-fast: true; in attempt B use false.
Preserve both run IDs and attempts. With true, other
queued/in-progress cells may be cancelled after the mandatory
failure. With false, remaining cells are allowed to finish, which
produces fuller compatibility evidence but consumes more runner
time.
Do not compare only elapsed time. Compare the set of completed, failed and cancelled cells, because that is the evidence trade-off.
7. Bound parallelism and observe queue behavior
strategy:
fail-fast: false
max-parallel: 2
matrix:
os: [ubuntu-24.04, windows-2025]
python: ['3.12', '3.13']
Predict that no more than two cells from this matrix execute simultaneously. Runner availability can reduce actual parallelism further. Record job start/end timestamps instead of assuming the requested maximum was achieved.
8. Evidence packet for the guided lab
| Evidence | Capture | Why |
|---|---|---|
| Source identity | run ID, attempt, exact github.sha |
separates reruns/revisions |
| Matrix source | static YAML or planner JSON | proves which cells could exist |
| Generated jobs | name + strategy.job-index/job-total |
proves expansion |
| Runner/tool | OS label, ImageOS/ImageVersion where available, Python version | proves execution environment |
| Failure policy | fail-fast, experimental flag, max-parallel | explains cancellation/tolerance |
| Conclusion set | success/failure/cancelled per cell | prevents top-level-color reasoning |
9. Small design challenge
A library officially supports Ubuntu and Windows on Python
3.12/3.13, while a future Python build is advisory only. CI minutes
are constrained, but release-night evidence must include every
mandatory cell even after one failure. Choose the matrix dimensions,
experimental representation, fail-fast value and
max-parallel. Justify each choice by the support
contract, not by copied YAML.
Knowledge check
Why must fromJSON be used when a planner output
contains a JSON array for a matrix axis?
Job outputs are strings at the boundary. fromJSON converts the JSON text back into an array/object that matrix evaluation can consume as structured data.
After excluding one of four Cartesian cells, how many cells remain before include adds a new combination?
Three.
What should distinguish an experimental failure from a mandatory failure?
An explicit matrix policy value mapped to job continue-on-error, plus retained cell evidence—not an operator mentally ignoring a red job.
What additional evidence is created by
fail-fast: false?
Other cells can complete instead of being cancelled, giving a fuller compatibility picture at the cost of more runner time.
If max-parallel: 2 is configured but only one
suitable runner is available, how many cells execute at
once?
At most one. max-parallel is an upper bound; runner capacity can be lower.
Official references and version notes
- GitHub Docs — Running variations of jobs in a workflow — matrix expansion, contexts, include/exclude, dynamic outputs, failure handling and max-parallel.
- GitHub Docs — workflow syntax: strategy.matrix — current matrix limit and strategy semantics.
- GitHub Docs — expressions: fromJSON — converting job-output JSON into arrays/objects used by a later matrix.
- GitHub Docs — strategy context — fail-fast, job-index, job-total and max-parallel for the current generated job.
- GitHub Docs — matrix strategy with reusable workflows — matrix-driven reusable workflow calls and output caveats.
- actions/upload-artifact v7.0.1 and actions/download-artifact v8.0.1 — action releases pinned by full commit SHA in the checkpoint.
Version-sensitive behavior was rechecked against current
GitHub-maintained documentation on 2026-09-09. A
matrix can generate at most
256 jobs per workflow run.
strategy.fail-fast defaults to true;
continue-on-error is evaluated per generated job; and
max-parallel limits simultaneous matrix jobs but does
not redefine the matrix or create a repository-wide concurrency
lock. Current mandatory examples target GitHub.com with versioned
runner labels ubuntu-24.04 and
windows-2025. Official actions are pinned to full
immutable commit SHAs: checkout v7.0.1, setup-python v7.0.0,
upload-artifact v7.0.1 and download-artifact v8.0.1. GitHub
Enterprise Server users must verify artifact-action
backend/major-version compatibility for their appliance before
copying these examples.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.