Chapter 21Lesson 03~140 minutes

Parallel Stages, Matrix Builds, Fail-Fast Behavior, Test Sharding, and High-Throughput Pipeline Design: Configuration, Design Choices, and Tradeoffs

Choose deliberately among parallel stages, matrix cells, static or duration-balanced sharding, fail-fast or collect-all behavior, and branch counts that match available capacity instead of blindly maximizing fan-out.

TradeoffsMatrix axesExcludesCapacityFan-inAuditability

Learning objectives

  • Choose parallel branches or a matrix based on the structure of the work.
  • Compare static deterministic sharding with duration-balanced assignment.
  • Select fail-fast or collect-all behavior based on diagnostic value and side-effect risk.
  • Relate branch count to actual executor/agent capacity and queue cost.
  • Design report/artifact fan-in without shared mutable workspaces.

1. Parallel versus matrix is a modeling decision

Use parallel branches when each branch has a different semantic purpose. Use matrix when the same workflow repeats across a small, explicit set of dimensions. The matrix cell identity becomes part of your build evidence, which is valuable when axes genuinely describe supported environments.

Choice Strength Cost/risk Good fit
Named parallel stages Clear semantic branch names Duplication if many branches differ only by values Unit / lint / integration
Declarative matrix Explicit Cartesian dimensions and excludes Cardinality can explode quickly Runtime × mode × platform
Sequential Simple resource/evidence model Longer wall time for independent work Small suites or scarce capacity

2. Static sharding versus duration-balanced sharding

Static modulo sharding is easy to reproduce: stable sorted test IDs map to shard index by modulo. It is a strong default for teaching and small suites. Its weakness is duration skew.

Duration-balanced sharding uses historical timing data to distribute long tests across shards. It can reduce the slowest shard, but now history/version of the timing dataset becomes part of the reproducibility contract. If the algorithm changes, record it as a build input.

Avoid opaque “smart split” tooling unless you can export the exact assignment manifest. Throughput that cannot explain which tests ran where is poor evidence.

3. Fail fast or collect all failures?

Policy Benefit Cost Prefer when
Collect all Maximum diagnostic coverage Consumes full capacity after first failure Independent test suites where all failures help triage
Fail fast Returns scarce capacity earlier Sibling evidence may be incomplete Redundant work or obvious blocker makes remaining cells low-value

Neither policy repairs external side effects. If a cell publishes something irreversible, separate that promotion from test fan-out and apply the Chapter 12 idempotency/guard rules.

4. More branches versus available capacity

With two eligible executors, moving from two to four branches does not double simultaneous work. It can add more checkout, workspace, container and reporting overhead while execution still happens in two waves. With dynamic agents, branch count may also amplify infrastructure demand, so keep cloud caps/quotas from Chapters 19–20.

Use three quantities:

  • Demand: runnable branch/cell count.
  • Capacity: eligible executors that can safely run the workload.
  • Critical path: the dependency path and slowest waves that determine wall time.

5. Use excludes to model unsupported combinations, not to hide cost after the fact

matrix {
  axes {
    axis { name 'RUNTIME'; values '21', '25' }
    axis { name 'MODE'; values 'normal', 'legacy' }
  }
  excludes {
    exclude {
      axis { name 'RUNTIME'; values '25' }
      axis { name 'MODE'; values 'legacy' }
    }
  }
  stages { /* ... */ }
}

The example says the excluded combination is unsupported. That is different from using an exclusion only because your platform cannot afford the test. Capacity decisions belong in the test strategy and documented support model.

6. Fan-in design: publish where evidence exists

Parallel workspaces are intentionally isolated. Publish JUnit in each branch/cell or stash/transfer uniquely named report files before leaving that workspace. For larger artifacts, archive/publish directly from the producing branch. Do not copy all cells into one shared filesystem directory and hope filenames do not collide.

Include cell identity in filenames and metadata: ${MODE}-${SHARD}-${GIT_COMMIT}. The overall Jenkins run can aggregate results while preserving the producer identity.

7. Concurrency multiplies trust exposure too

Running four branches with credentials can create four simultaneous secret-bearing processes/workspaces. Do not solve throughput by moving untrusted code onto privileged shared agents or by broadening credential scope. Prefer credential-free test branches, dedicated trust pools and short-lived bindings only where a cell truly needs them.

8. Worked scenario: 24 minutes of test work, two executors

Suppose four shards take 9, 7, 5 and 3 minutes. Total work is 24 minutes and the longest shard is 9. With two executors, an ideal lower bound is at least max(9, 24/2) = 12 minutes, before checkout/setup/fan-in. Depending on scheduling order, a naive assignment may take longer.

If historical balancing can produce two waves closer to 12 minutes without losing deterministic assignment evidence, that is useful. Creating eight smaller branches on the same two executors may only increase overhead.

9. Decision table

Question If yes Observable prerequisite
Are branches semantically different? Prefer named parallel stages Independent commands and outputs
Is work a small Cartesian product? Consider matrix Reviewed axes/cardinality/excludes
Are test durations stable and skewed? Consider duration balancing Versioned timing history + assignment manifest
Does first failure make siblings low-value? Consider fail-fast No unsafe side-effect rollback assumption
Do branches exceed useful capacity? Reduce/group work or add bounded capacity Queue/utilization measurements
Next lesson

Diagnostics, Failure Modes, Security, and Performance

Diagnose matrix explosion, shared-workspace races, shard coverage defects, queue contention and incomplete fail-fast evidence.

Knowledge check

Answer before revealing the explanation.

1. When is matrix clearer than hand-written parallel branches?

2. When can static modulo sharding be inefficient?

3. Why might collect-all failures be preferable to fail-fast?

4. What is the safe response when desired branch count exceeds capacity?

5. Do matrix excludes happen after all cells start?

Official references and version notes

Version note — 2026-09-17: executable examples assume Jenkins 2.568.3 LTS, Java 21, Pipeline 608.v67378e9d3db_1, Declarative 2.2293.v6e7193cec599, and JUnit 1425.v9c7318dca_96d. The mandatory sharding algorithm uses only POSIX shell/Python-style reasoning and Jenkins Pipeline features; no commercial or test-splitting plugin is required. Re-check current plugin/core compatibility and security advisories before production adoption.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.