Chapter 12Lesson 04~180 minutes

Matrix Strategies, Dynamic Matrices, Fail-Fast, and Parallel Test Design: Diagnostics, Failure Modes, and Production Practices

Matrix failures become confusing when teams reason from the final workflow badge instead of the generated cells. This lesson diagnoses explosion, cancellation, hidden mandatory failures, malformed dynamic JSON and mutable toolchains by preserving the original run and tracing each cell back to its matrix source.

Matrix explosionCancelled cellsMalformed JSONToolchain driftFailure diagnosis

Learning objectives

  • Diagnose matrix explosion before it reaches runner scheduling.
  • Distinguish failed, cancelled and skipped cells instead of treating non-success as one state.
  • Detect a mandatory failure hidden behind incorrect continue-on-error.
  • Repair malformed dynamic JSON while preserving the first planner output and run identity.
  • Separate toolchain/image drift from application incompatibility and rerun the smallest equivalent scope.

1. Evidence-first diagnostic sequence

Preserve run ID, attempt, exact SHA and the first failed/cancelled job list before editing YAML or rerunning. Then confirm the matrix source and expanded graph, evaluated tolerance/fail-fast/max-parallel policy, runner/image/toolchain, failing step, per-cell outputs/artifacts, and only then downstream aggregation.

Layer Question Evidence
Matrix source What axes/JSON produced the graph? workflow revision + planner output
Expansion Which cells actually existed? job names / matrix values / job-total
Strategy Could a failure cancel siblings? fail-fast, continue-on-error, max-parallel
Runner/tool Did environment drift? runner label/image + setup action SHA + tool version
Test What assertion failed? first-failure log/test report
Artifact Did each completed cell publish evidence? artifact name/ID/digest
Aggregation Did policy classify all states correctly? support report + source run/SHA

2. Failure mode: matrix explosion

A new feature axis with eight values can turn an 8-cell matrix into 64 cells. Add another six-value axis and the requested 384 cells exceed GitHub’s current 256-job matrix limit. The fix is not “increase the limit.” Decide which combinations represent independent risk and move exhaustive combinatorics to a lower-cost local test layer if appropriate.

Before change: 2 OS × 2 runtimes × 2 modes = 8 cells
After change: 2 × 2 × 2 × 8 features = 64 cells
Add 6 database versions: 384 cells -> exceeds current 256 matrix-job limit

3. Failure mode: cancelled or skipped is counted as passed

With fail-fast enabled, a mandatory failure can cancel siblings. A simplistic aggregator that counts only explicit failure values and treats everything else as acceptable can publish a false support verdict. Cancelled and skipped cells are missing evidence, not successful tests.

Never normalize unknown evidence into success

A support decision should require an observed successful result for every mandatory cell. Failure, cancellation, skip, missing artifact or malformed evidence must remain distinguishable.

4. Failure mode: a required cell is hidden behind continue-on-error

# BROKEN POLICY: all Windows cells are tolerated even though Windows is supported.
continue-on-error: ${{ matrix.os == 'windows-2025' }}

The workflow may continue despite a supported Windows failure. Repair the policy by carrying an explicit experimental field derived from the support contract and mapping only that field to continue-on-error. Preserve the original run because it proves the policy bug existed.

5. Failure mode: malformed dynamic JSON

# Broken planner output (single quotes are not valid JSON strings)
echo "versions=['3.12','3.13']" >> "$GITHUB_OUTPUT"

# Valid JSON array
echo 'versions=["3.12","3.13"]' >> "$GITHUB_OUTPUT"

If fromJSON cannot parse the planner output, the matrix cannot be evaluated correctly. Capture the exact output string. Validate generator output locally with a real JSON parser before dispatching a large run.

6. Failure mode: mutable toolchains invalidate compatibility evidence

A versioned runner label still receives image updates, and an unpinned setup action tag can move. Record action commit SHAs and installed tool versions. If Python 3.13 passed last week but fails today, compare the exact setup action SHA, runner ImageVersion and resolved Python patch release before concluding the application regressed.

7. Intentionally broken example and repair

jobs:
  plan:
    runs-on: ubuntu-24.04
    outputs:
      versions: ${{ steps.out.outputs.versions }}
    steps:
      - id: out
        run: echo "versions=['3.12','3.13']" >> "$GITHUB_OUTPUT" # invalid JSON
  test:
    needs: plan
    strategy:
      fail-fast: true
      matrix:
        python: ${{ fromJSON(needs.plan.outputs.versions) }}
    runs-on: ubuntu-24.04
    steps:
      - run: python --version

Interpretation: this is a matrix-generation failure, not a Python test failure and not runner capacity. Preserve the planner step log/output, repair only the JSON serialization, and dispatch a new run. Do not change runner labels, permissions and application code simultaneously.

8. Smallest equivalent rerun

If a single mandatory cell failed from a deterministic test bug, use the smallest rerun scope GitHub supports while preserving the original run/attempt and first-failure evidence. If you edit matrix-generation YAML, that is a new workflow revision and therefore a new experiment—not a faithful rerun of the old graph.

Knowledge check

A mandatory cell is cancelled by fail-fast. Is that compatible configuration proven supported?

What causal layer is responsible when fromJSON fails on planner output?

Why is continue-on-error based on OS name dangerous?

What should you compare before blaming an app when a previously passing cell starts failing?

Why is editing YAML and then pressing rerun not the same experiment?

Checkpoint next

Produce an auditable support verdict

Lesson 5 deliberately breaks one supported and one experimental cell, retains every per-cell artifact, and aggregates the evidence without hiding either failure.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against current GitHub-maintained documentation on 2026-09-09. A matrix can generate at most 256 jobs per workflow run. strategy.fail-fast defaults to true; continue-on-error is evaluated per generated job; and max-parallel limits simultaneous matrix jobs but does not redefine the matrix or create a repository-wide concurrency lock. Current mandatory examples target GitHub.com with versioned runner labels ubuntu-24.04 and windows-2025. Official actions are pinned to full immutable commit SHAs: checkout v7.0.1, setup-python v7.0.0, upload-artifact v7.0.1 and download-artifact v8.0.1. GitHub Enterprise Server users must verify artifact-action backend/major-version compatibility for their appliance before copying these examples.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.