Chapter 21Lesson 04~155 minutes

Parallel Stages, Matrix Builds, Fail-Fast Behavior, Test Sharding, and High-Throughput Pipeline Design: Diagnostics, Failure Modes, Security, and Performance

Diagnose high-throughput Pipeline failures from preserved queue, branch, workspace, shard and report evidence. Separate syntax/model errors, queue capacity, workspace collisions, shard math, test failures and external side effects before changing the design.

DiagnosticsMatrix explosionWorkspace isolationShard integrityQueue contentionEvidence

Learning objectives

  • Diagnose matrix explosion before adding executors or agents.
  • Identify and repair shared mutable workspace collisions.
  • Prove shard completeness and uniqueness from manifests.
  • Separate queue/capacity slowdown from test or Pipeline failures.
  • Preserve first-failure evidence when fail-fast interrupts sibling work.

1. Evidence-first diagnostic sequence

  1. Preserve job/build/queue IDs, source SHA and the first failure/interruption logs.
  2. Confirm Jenkins 2.568.3/Java/plugin baseline and exact Jenkinsfile revision.
  3. Calculate intended axis cardinality and list branch/cell names.
  4. Inspect queue reasons, eligible labels, executor count and actual node allocation.
  5. Inspect each branch workspace and output paths for collisions.
  6. Validate shard manifests: full union, zero duplicates, expected count.
  7. Inspect per-cell Pipeline/test result and whether fail-fast interrupted siblings.
  8. Inspect JUnit/artifact ingestion separately from step execution.
  9. Apply the least destructive correction and rerun only the smallest safe scope.

2. Failure mode: matrix explosion

A matrix with 4 operating systems × 5 runtimes × 4 databases × 3 modes creates 240 cells before excludes. On two executors that is at least 120 scheduling waves even if every cell takes the same time. The Jenkinsfile may validate perfectly while the operational design is unreasonable.

Evidence: axis list, computed cell count, queue depth, executor utilization, per-cell setup cost. Repair: test only supported combinations, use justified excludes, move broad compatibility to an appropriately bounded periodic suite, or add reviewed capacity. Do not set hundreds of executors on one node.

3. Failure mode: shared mutable workspace collision

This intentionally broken pattern makes independent branches write the same path:

parallel(
  a: { node('ch21') { ws('/tmp/ch21-shared') { sh 'echo A > result.txt' } } },
  b: { node('ch21') { ws('/tmp/ch21-shared') { sh 'echo B > result.txt' } } }
)

The last writer wins, and either branch may read evidence created by the other. Preserve the file timestamp/content that demonstrates the race. Repair by using Jenkins-managed independent workspaces (or unique per-branch directories) and branch-qualified output names. Do not add sleeps to “fix” a race.

4. Failure mode: shard duplication or omission

A common defect is mixing zero-based and one-based shard conventions. For example, generating shard indexes 1..N but testing (NR-1) % N == shard silently leaves shard zero unexecuted.

set -eu
sort expected.ids > expected.sorted
cat shard-*.ids | sort > actual.sorted
printf 'duplicates:\n'
uniq -d actual.sorted
printf 'missing-or-extra:\n'
comm -3 expected.sorted actual.sorted

Make this manifest check a gate before expensive execution. The repair is deterministic math and one documented index convention, not rerunning until coverage “looks right.”

5. Failure mode: parallel Pipeline is slower than sequential

Symptoms include long queue waits, repeated SCM checkout, cold container/tool startup, CPU/memory contention on one host, or a fan-in bottleneck. Preserve timestamps by cell and compare:

  • time waiting for an executor;
  • checkout/setup duration;
  • actual test duration;
  • slowest branch duration;
  • report/artifact publication time.

If two executors are saturated and four branches add only setup overhead, reduce branches. If CPU is already saturated, adding executors may make every test slower.

6. Failure mode: fail-fast hides diagnostic coverage

With failFast true, a first failing cell can interrupt siblings before they publish complete reports. This is correct policy if early cancellation is the goal, but it is wrong if the incident requires all compatibility failures.

Preserve the cell that failed first, interrupted cell names and any partial reports. Repair by choosing collect-all for diagnostic runs, or move critical evidence publication into post { always { ... } } where practical. Do not convert genuine failures to green merely to keep the fan-out running.

7. Security and trust failure modes

Concurrency can magnify mistakes: one broadly scoped credential can be exposed to many simultaneous branches; untrusted code can fan out across privileged agents; report filenames can accidentally archive secret-bearing files. Keep credential scope minimal, use trusted agent pools, avoid environment dumps and archive only explicit safe paths.

Do not troubleshoot concurrency by weakening controls. Keep TLS, CSRF, authorization and Script Security enabled. Do not use broad admin credentials, permissive SSH host-key checks, unbounded executors/agents or delete-and-recreate fixes.

8. Intentionally broken lab: 8 branches on 2 executors

Create eight synthetic 3-second branches on the two-executor agent and record start/end/executor for each. Expected evidence: no more than two branches execute simultaneously, the remaining branches wait, and total elapsed time approaches at least four waves plus overhead.

def branches = [:]
(0..7).each { n ->
  def id = n
  branches["b${id}"] = {
    node('ch21') {
      withEnv(["BRANCH_ID=${id}"]) {
        sh '''
          echo "b${BRANCH_ID} start=$(date +%s) executor=${EXECUTOR_NUMBER}"
          sleep 3
          echo "b${BRANCH_ID} end=$(date +%s)"
        '''
      }
    }
  }
}
parallel branches

Interpret the queue/executor evidence; do not “repair” the experiment by changing the agent to eight executors. A production fix must be justified by workload and host capacity.

9. Causal layer map

Symptom Likely first layer Evidence
Cell never starts Queue/label/capacity Queue reason, labels, executors
Wrong tests run Shard logic/source input Inventory + shard manifest
Report missing Workspace/path/report ingestion File glob, workspace, JUnit log
Cell aborted after sibling failure Fail-fast policy First failure + interruption trace
All cells slow Capacity/resource contention Queue, host load, setup times
External duplicate effect Side-effect/idempotency boundary Provider state + request identity
Next lesson

Checkpoint Lab

Predict and measure a bounded four-cell matrix under two executors, then tune it without sacrificing shard or report evidence.

Knowledge check

Answer before revealing the explanation.

1. A 240-cell matrix waits for hours on two executors. What failed first?

2. Two branches write to /tmp/shared/output.xml and one report disappears. What layer is at fault?

3. How can you prove sharding has neither gaps nor duplicates?

4. Why can fail-fast hide useful evidence?

5. If parallel execution is slower than sequential execution, what should you inspect first?

Official references and version notes

Version note — 2026-09-17: executable examples assume Jenkins 2.568.3 LTS, Java 21, Pipeline 608.v67378e9d3db_1, Declarative 2.2293.v6e7193cec599, and JUnit 1425.v9c7318dca_96d. The mandatory sharding algorithm uses only POSIX shell/Python-style reasoning and Jenkins Pipeline features; no commercial or test-splitting plugin is required. Re-check current plugin/core compatibility and security advisories before production adoption.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.