Matrix Strategies, Dynamic Matrices, Fail-Fast, and Parallel Test Design: Diagnostics, Failure Modes, and Production Practices
Matrix failures become confusing when teams reason from the final workflow badge instead of the generated cells. This lesson diagnoses explosion, cancellation, hidden mandatory failures, malformed dynamic JSON and mutable toolchains by preserving the original run and tracing each cell back to its matrix source.
Learning objectives
- Diagnose matrix explosion before it reaches runner scheduling.
- Distinguish failed, cancelled and skipped cells instead of treating non-success as one state.
-
Detect a mandatory failure hidden behind incorrect
continue-on-error. - Repair malformed dynamic JSON while preserving the first planner output and run identity.
- Separate toolchain/image drift from application incompatibility and rerun the smallest equivalent scope.
1. Evidence-first diagnostic sequence
Preserve run ID, attempt, exact SHA and the first failed/cancelled job list before editing YAML or rerunning. Then confirm the matrix source and expanded graph, evaluated tolerance/fail-fast/max-parallel policy, runner/image/toolchain, failing step, per-cell outputs/artifacts, and only then downstream aggregation.
| Layer | Question | Evidence |
|---|---|---|
| Matrix source | What axes/JSON produced the graph? | workflow revision + planner output |
| Expansion | Which cells actually existed? | job names / matrix values / job-total |
| Strategy | Could a failure cancel siblings? | fail-fast, continue-on-error, max-parallel |
| Runner/tool | Did environment drift? | runner label/image + setup action SHA + tool version |
| Test | What assertion failed? | first-failure log/test report |
| Artifact | Did each completed cell publish evidence? | artifact name/ID/digest |
| Aggregation | Did policy classify all states correctly? | support report + source run/SHA |
2. Failure mode: matrix explosion
A new feature axis with eight values can turn an 8-cell
matrix into 64 cells. Add another six-value axis and the requested
384 cells exceed GitHub’s current 256-job matrix limit. The fix is
not “increase the limit.” Decide which combinations represent
independent risk and move exhaustive combinatorics to a lower-cost
local test layer if appropriate.
Before change: 2 OS × 2 runtimes × 2 modes = 8 cells
After change: 2 × 2 × 2 × 8 features = 64 cells
Add 6 database versions: 384 cells -> exceeds current 256 matrix-job limit
3. Failure mode: cancelled or skipped is counted as passed
With fail-fast enabled, a mandatory failure can cancel siblings. A
simplistic aggregator that counts only explicit
failure values and treats everything else as acceptable
can publish a false support verdict.
Cancelled and skipped cells are missing evidence,
not successful tests.
A support decision should require an observed successful result for every mandatory cell. Failure, cancellation, skip, missing artifact or malformed evidence must remain distinguishable.
4. Failure mode: a required cell is hidden behind continue-on-error
# BROKEN POLICY: all Windows cells are tolerated even though Windows is supported.
continue-on-error: ${{ matrix.os == 'windows-2025' }}
The workflow may continue despite a supported Windows failure.
Repair the policy by carrying an explicit
experimental field derived from the support contract
and mapping only that field to continue-on-error.
Preserve the original run because it proves the policy bug existed.
5. Failure mode: malformed dynamic JSON
# Broken planner output (single quotes are not valid JSON strings)
echo "versions=['3.12','3.13']" >> "$GITHUB_OUTPUT"
# Valid JSON array
echo 'versions=["3.12","3.13"]' >> "$GITHUB_OUTPUT"
If fromJSON cannot parse the planner output, the matrix
cannot be evaluated correctly. Capture the exact output string.
Validate generator output locally with a real JSON parser before
dispatching a large run.
6. Failure mode: mutable toolchains invalidate compatibility evidence
A versioned runner label still receives image updates, and an unpinned setup action tag can move. Record action commit SHAs and installed tool versions. If Python 3.13 passed last week but fails today, compare the exact setup action SHA, runner ImageVersion and resolved Python patch release before concluding the application regressed.
7. Intentionally broken example and repair
jobs:
plan:
runs-on: ubuntu-24.04
outputs:
versions: ${{ steps.out.outputs.versions }}
steps:
- id: out
run: echo "versions=['3.12','3.13']" >> "$GITHUB_OUTPUT" # invalid JSON
test:
needs: plan
strategy:
fail-fast: true
matrix:
python: ${{ fromJSON(needs.plan.outputs.versions) }}
runs-on: ubuntu-24.04
steps:
- run: python --version
Interpretation: this is a matrix-generation failure, not a Python test failure and not runner capacity. Preserve the planner step log/output, repair only the JSON serialization, and dispatch a new run. Do not change runner labels, permissions and application code simultaneously.
8. Smallest equivalent rerun
If a single mandatory cell failed from a deterministic test bug, use the smallest rerun scope GitHub supports while preserving the original run/attempt and first-failure evidence. If you edit matrix-generation YAML, that is a new workflow revision and therefore a new experiment—not a faithful rerun of the old graph.
Knowledge check
A mandatory cell is cancelled by fail-fast. Is that compatible configuration proven supported?
No. Cancellation means the test did not complete; it is missing evidence, not a pass.
What causal layer is responsible when fromJSON fails on planner output?
Matrix/dataflow evaluation—the generator/output JSON boundary—not the test runtime.
Why is continue-on-error based on OS name dangerous?
It can accidentally tolerate failures for an officially supported platform. Tolerance should come from an explicit support-policy field such as experimental.
What should you compare before blaming an app when a previously passing cell starts failing?
Exact source SHA, runner/image version, pinned action SHA, resolved toolchain version and first-failure test evidence.
Why is editing YAML and then pressing rerun not the same experiment?
The workflow revision and possibly the generated graph changed; it is a new configuration and must be treated as a new run/evidence set.
Official references and version notes
- GitHub Docs — Running variations of jobs in a workflow — matrix expansion, contexts, include/exclude, dynamic outputs, failure handling and max-parallel.
- GitHub Docs — workflow syntax: strategy.matrix — current matrix limit and strategy semantics.
- GitHub Docs — expressions: fromJSON — converting job-output JSON into arrays/objects used by a later matrix.
- GitHub Docs — strategy context — fail-fast, job-index, job-total and max-parallel for the current generated job.
- GitHub Docs — matrix strategy with reusable workflows — matrix-driven reusable workflow calls and output caveats.
- actions/upload-artifact v7.0.1 and actions/download-artifact v8.0.1 — action releases pinned by full commit SHA in the checkpoint.
Version-sensitive behavior was rechecked against current
GitHub-maintained documentation on 2026-09-09. A
matrix can generate at most
256 jobs per workflow run.
strategy.fail-fast defaults to true;
continue-on-error is evaluated per generated job; and
max-parallel limits simultaneous matrix jobs but does
not redefine the matrix or create a repository-wide concurrency
lock. Current mandatory examples target GitHub.com with versioned
runner labels ubuntu-24.04 and
windows-2025. Official actions are pinned to full
immutable commit SHAs: checkout v7.0.1, setup-python v7.0.0,
upload-artifact v7.0.1 and download-artifact v8.0.1. GitHub
Enterprise Server users must verify artifact-action
backend/major-version compatibility for their appliance before
copying these examples.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.