Chapter 21Lesson 01~125 minutes

Parallel Stages, Matrix Builds, Fail-Fast Behavior, Test Sharding, and High-Throughput Pipeline Design: Concepts, Architecture, and Mental Model

Treat concurrency as a scheduling and evidence problem, not a syntax trick: define independent work, bound branch and matrix cardinality, understand queue/executor limits, preserve per-branch evidence, and make fail-fast behavior explicit before increasing throughput.

Parallel stagesMatrixQueue capacityFail-fastShardingCritical path

Learning objectives

  • Explain why Pipeline fan-out is constrained by queue, label, agent and executor capacity.
  • Calculate parallel/matrix cardinality before execution and identify the critical path.
  • Distinguish branch identity, workspace state, test-shard membership and retained report/artifact evidence.
  • Explain failFast as failure propagation, not rollback or transactionality.
  • Design deterministic shards and unique fan-in evidence before scaling concurrency.

1. The problem: syntax can create more runnable work than your platform can execute

Chapters 17–20 gave you explicit execution capacity: nodes, executors, Docker-capable agents and dynamic Kubernetes pods. Chapter 21 asks what happens when one Jenkins build deliberately creates many independently runnable branches. The answer is not “everything runs at once.” Jenkins must still queue each eligible branch or matrix cell until a matching executor exists.

Parallelism improves throughput only when work is genuinely independent, there is enough useful capacity, and fan-out/fan-in overhead is smaller than the saved serial time. Otherwise, a Pipeline can become slower, harder to diagnose and more expensive while producing less trustworthy evidence.

2. Mental model: dimension set → runnable cells → capacity → evidence fan-in

Start with a test/build dimension set. Jenkins expands that design into named parallel branches or matrix cells. Each branch requests an eligible agent/executor and receives its own execution context/workspace. The branch creates reports/artifacts identified by its branch/cell dimensions. Failure propagation may allow all cells to finish or request fail-fast cancellation. Finally, Jenkins aggregates retained reports/artifacts and computes the overall build result.

Mental model: dimension set → runnable cells → capacity → evidence fan-in
flowchart TD
  A[Dimension set / branch plan] --> B[Parallel branches or matrix cells]
  B --> C[Queue + label eligibility]
  C --> D[Free executor / workspace]
  D --> E[Independent test or build work]
  E --> F[Per-branch result + reports/artifacts]
  F --> G[Fail-fast or collect-all semantics]
  G --> H[Fan-in + overall build result]

The causal point is important: branch creation is not executor allocation, executor allocation is not step success, and step success is not report ingestion. Preserve those states separately.

3. State to define before adding concurrency

State Question Evidence
Matrix axes/cardinality How many cells can exist? Axis values, excludes, computed cell count
Branch identity Which logical work unit is this? Stage/cell name and axis environment variables
Queue/executor Did it wait, and where did it run? Queue reason, NODE_NAME, EXECUTOR_NUMBER
Workspace Which filesystem state belongs to the cell? WORKSPACE, unique subdirectories
Shard assignment Exactly which tests belong here? Versioned test manifest + shard manifest
Result Did the branch fail, abort or succeed? Branch logs, stage result, interruption evidence
Fan-in evidence What survived workspaces? JUnit results, archived artifacts/digests
Critical path What determined wall time? Queue + branch start/end + slowest dependency chain

4. Parallel versus matrix

Parallel stages are best when branches are semantically different: unit tests, linting and an integration suite. Matrix is best when work is the Cartesian product of explicit dimensions such as runtime × mode. Both ultimately create concurrently runnable Pipeline branches; matrix adds a declarative model for axes, cells, exclusions and per-cell directives.

stage('Verification') {
  parallel {
    stage('Unit') { steps { sh './run-unit.sh' } }
    stage('Lint') { steps { sh './run-lint.sh' } }
  }
}
stage('Matrix') {
  matrix {
    axes {
      axis { name 'MODE'; values 'normal', 'strict' }
      axis { name 'SHARD'; values '0', '1' }
    }
    stages {
      stage('Test') { steps { sh './run-shard.sh "$MODE" "$SHARD" 2' } }
    }
  }
}

5. Calculate cardinality before Jenkins calculates it for you

Before exclusions, a matrix has the product of axis sizes. Two modes × two shards creates four cells. Add four operating systems and five JDKs and the same design becomes 2 × 2 × 4 × 5 = 80 cells. On a two-executor lab agent, that means many queue waves even though the Jenkinsfile is syntactically valid.

Static excludes remove invalid combinations before execution. Runtime when conditions may skip cells after the static cell set is defined. Neither is a substitute for reviewing the intended coverage and platform cost.

6. Branch count is demand; executors are capacity

If four cells become runnable and only two eligible executors exist, at most two run at once. The others wait. Increasing executors can increase concurrency, but Chapter 17's resource rule still applies: executor count is not a CPU-core formula, and overcommitting one host can make every branch slower.

A useful lower-bound intuition is:

ideal_wall_time >= max(longest_single_branch,
                       total_branch_work / usable_parallel_capacity)
actual_wall_time = queue + setup + execution + contention + fan_in

The exact schedule depends on branch order and duration, so measure rather than promising a theoretical speedup.

7. Deterministic sharding means repeatable membership, not equal counts only

A shard is a deterministic subset of a larger test inventory. A simple teaching algorithm sorts stable test IDs and assigns item i to shard i mod N. Record the inventory, N, shard index and resulting IDs. Do not use an unstable runtime hash unless its seed and behavior are fixed and documented.

sort tests.txt > tests.sorted
awk -v shard=0 -v count=2 '((NR-1) % count)==shard {print}' tests.sorted > shard-0.txt
awk -v shard=1 -v count=2 '((NR-1) % count)==shard {print}' tests.sorted > shard-1.txt
cat shard-0.txt shard-1.txt | sort > union.txt
cmp tests.sorted union.txt

The final comparison proves coverage of the inventory, but you should also check that no ID appears twice.

8. Fail-fast is cancellation policy, not rollback

For Declarative parallel or matrix stages, failFast true requests that sibling branches be aborted when one fails. Pipeline-wide parallelsAlwaysFailFast() applies the same policy to subsequent parallel stages. This saves capacity when additional results have little value.

It does not undo a command that already published a package, changed a database or called an external service. Keep irreversible side effects outside retry/fail-fast experiments unless they have explicit idempotency and compensation controls.

9. Fan-in without evidence collisions

Each branch/cell should create uniquely named evidence such as results/normal-shard-0.xml and manifest-normal-shard-0.txt. Publish JUnit while the report still exists in the cell workspace. Archive or publish artifacts that must survive workspace cleanup. Never make parallel branches race to write one common filename.

10. Read-only inspection first

Before changing a Jenkinsfile, record the controller/core/plugin baseline, eligible agent labels and executor count, current queue depth, previous run duration, source SHA and existing test inventory. During a run, print non-secret execution context:

printf 'build=%s node=%s executor=%s workspace=%s\n' \
  "$BUILD_NUMBER" "$NODE_NAME" "$EXECUTOR_NUMBER" "$WORKSPACE"
git rev-parse HEAD

Do not print credentials or environment dumps. The goal is attribution, not maximal logging.

Next lesson

Guided Hands-On Workflow and Core Operations

Build a two-executor lab, create deterministic shards, run parallel and matrix variants, publish JUnit evidence, and measure queue versus execution time.

Knowledge check

Answer before revealing the explanation.

1. Does four parallel branches guarantee a four-times faster Pipeline?

2. How do you calculate the maximum number of Declarative matrix cells before excludes?

3. What makes a test shard reproducible?

4. What does fail-fast change?

5. Why must branch outputs have unique names?

Official references and version notes

Version note — 2026-09-17: executable examples assume Jenkins 2.568.3 LTS, Java 21, Pipeline 608.v67378e9d3db_1, Declarative 2.2293.v6e7193cec599, and JUnit 1425.v9c7318dca_96d. The mandatory sharding algorithm uses only POSIX shell/Python-style reasoning and Jenkins Pipeline features; no commercial or test-splitting plugin is required. Re-check current plugin/core compatibility and security advisories before production adoption.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.