Parallel Stages, Matrix Builds, Fail-Fast Behavior, Test Sharding, and High-Throughput Pipeline Design: Diagnostics, Failure Modes, Security, and Performance
Diagnose high-throughput Pipeline failures from preserved queue, branch, workspace, shard and report evidence. Separate syntax/model errors, queue capacity, workspace collisions, shard math, test failures and external side effects before changing the design.
Learning objectives
- Diagnose matrix explosion before adding executors or agents.
- Identify and repair shared mutable workspace collisions.
- Prove shard completeness and uniqueness from manifests.
- Separate queue/capacity slowdown from test or Pipeline failures.
- Preserve first-failure evidence when fail-fast interrupts sibling work.
1. Evidence-first diagnostic sequence
- Preserve job/build/queue IDs, source SHA and the first failure/interruption logs.
- Confirm Jenkins 2.568.3/Java/plugin baseline and exact Jenkinsfile revision.
- Calculate intended axis cardinality and list branch/cell names.
- Inspect queue reasons, eligible labels, executor count and actual node allocation.
- Inspect each branch workspace and output paths for collisions.
- Validate shard manifests: full union, zero duplicates, expected count.
- Inspect per-cell Pipeline/test result and whether fail-fast interrupted siblings.
- Inspect JUnit/artifact ingestion separately from step execution.
- Apply the least destructive correction and rerun only the smallest safe scope.
2. Failure mode: matrix explosion
A matrix with 4 operating systems × 5 runtimes × 4 databases × 3 modes creates 240 cells before excludes. On two executors that is at least 120 scheduling waves even if every cell takes the same time. The Jenkinsfile may validate perfectly while the operational design is unreasonable.
Evidence: axis list, computed cell count, queue depth, executor utilization, per-cell setup cost. Repair: test only supported combinations, use justified excludes, move broad compatibility to an appropriately bounded periodic suite, or add reviewed capacity. Do not set hundreds of executors on one node.
3. Failure mode: shared mutable workspace collision
This intentionally broken pattern makes independent branches write the same path:
parallel(
a: { node('ch21') { ws('/tmp/ch21-shared') { sh 'echo A > result.txt' } } },
b: { node('ch21') { ws('/tmp/ch21-shared') { sh 'echo B > result.txt' } } }
)
The last writer wins, and either branch may read evidence created by the other. Preserve the file timestamp/content that demonstrates the race. Repair by using Jenkins-managed independent workspaces (or unique per-branch directories) and branch-qualified output names. Do not add sleeps to “fix” a race.
5. Failure mode: parallel Pipeline is slower than sequential
Symptoms include long queue waits, repeated SCM checkout, cold container/tool startup, CPU/memory contention on one host, or a fan-in bottleneck. Preserve timestamps by cell and compare:
- time waiting for an executor;
- checkout/setup duration;
- actual test duration;
- slowest branch duration;
- report/artifact publication time.
If two executors are saturated and four branches add only setup overhead, reduce branches. If CPU is already saturated, adding executors may make every test slower.
7. Security and trust failure modes
Concurrency can magnify mistakes: one broadly scoped credential can be exposed to many simultaneous branches; untrusted code can fan out across privileged agents; report filenames can accidentally archive secret-bearing files. Keep credential scope minimal, use trusted agent pools, avoid environment dumps and archive only explicit safe paths.
8. Intentionally broken lab: 8 branches on 2 executors
Create eight synthetic 3-second branches on the two-executor agent and record start/end/executor for each. Expected evidence: no more than two branches execute simultaneously, the remaining branches wait, and total elapsed time approaches at least four waves plus overhead.
def branches = [:]
(0..7).each { n ->
def id = n
branches["b${id}"] = {
node('ch21') {
withEnv(["BRANCH_ID=${id}"]) {
sh '''
echo "b${BRANCH_ID} start=$(date +%s) executor=${EXECUTOR_NUMBER}"
sleep 3
echo "b${BRANCH_ID} end=$(date +%s)"
'''
}
}
}
}
parallel branches
Interpret the queue/executor evidence; do not “repair” the experiment by changing the agent to eight executors. A production fix must be justified by workload and host capacity.
9. Causal layer map
| Symptom | Likely first layer | Evidence |
|---|---|---|
| Cell never starts | Queue/label/capacity | Queue reason, labels, executors |
| Wrong tests run | Shard logic/source input | Inventory + shard manifest |
| Report missing | Workspace/path/report ingestion | File glob, workspace, JUnit log |
| Cell aborted after sibling failure | Fail-fast policy | First failure + interruption trace |
| All cells slow | Capacity/resource contention | Queue, host load, setup times |
| External duplicate effect | Side-effect/idempotency boundary | Provider state + request identity |
Knowledge check
Answer before revealing the explanation.
1. A 240-cell matrix waits for hours on two executors. What failed first?
The concurrency design. Cardinality exploded far beyond capacity. Preserve the queue evidence, then reduce dimensions/exclude invalid combinations or move only justified coverage to additional bounded capacity.
2. Two branches write to /tmp/shared/output.xml and one report disappears. What layer is at fault?
Workspace/output isolation. The branches are racing on shared mutable state. Use separate Jenkins workspaces/directories and unique report names rather than a common path.
3. How can you prove sharding has neither gaps nor duplicates?
Generate a manifest of every test ID and every shard assignment, then verify that the union equals the full manifest and intersections are empty before executing expensive tests.
4. Why can fail-fast hide useful evidence?
Sibling branches may be interrupted before they reach their own failure or report publication. Use collect-all mode when full diagnostic coverage matters, or publish partial evidence in interruption-safe post handling.
5. If parallel execution is slower than sequential execution, what should you inspect first?
Compare executor capacity, queue time, branch setup/checkout costs, shared-resource contention, slowest-branch duration and fan-in overhead before adding more concurrency.
Official references and version notes
-
Jenkins LTS changelog
— chapter baseline
Jenkins 2.568.3 LTS, released 2026-09-02 and tested with Java 21 and 25; labs use Java 21 for Jenkins components. - Jenkins Java support policy — verify controller/agent JVM support before reproducing the lab on another LTS line.
-
Declarative Pipeline: parallel
— nested parallel stages,
failFast trueandparallelsAlwaysFailFast(). - Declarative Pipeline: matrix — axes, static cell generation, excludes, per-cell directives and fail-fast behavior.
-
Pipeline: Declarative plugin
— version
2.2293.v6e7193cec599, requiring Jenkins 2.504.3. -
Pipeline plugin
— version
608.v67378e9d3db_1. -
JUnit plugin
— version
1425.v9c7318dca_96d, requiring Jenkins 2.541.3. - JUnit Pipeline step — JUnit XML publication semantics and report glob behavior.
Version note — 2026-09-17: executable examples
assume Jenkins 2.568.3 LTS, Java 21, Pipeline
608.v67378e9d3db_1, Declarative
2.2293.v6e7193cec599, and JUnit
1425.v9c7318dca_96d. The mandatory sharding algorithm
uses only POSIX shell/Python-style reasoning and Jenkins Pipeline
features; no commercial or test-splitting plugin is required.
Re-check current plugin/core compatibility and security advisories
before production adoption.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.