Chapter 21Lesson 02~195 minutes

Parallel Stages, Matrix Builds, Fail-Fast Behavior, Test Sharding, and High-Throughput Pipeline Design: Guided Hands-On Workflow and Core Operations

Build a disposable two-executor lab, run synthetic tests sequentially and in bounded parallel/matrix forms, partition test IDs deterministically, publish JUnit evidence per branch, and measure how queue pressure changes wall-clock time.

Hands-onTwo executorsDeclarativeJUnitDeterministic shardsMeasurements

Learning objectives

  • Prepare a disposable two-executor Jenkins agent and synthetic test inventory.
  • Compare sequential, parallel and bounded matrix execution using the same work.
  • Create deterministic shard manifests and validate coverage before tests run.
  • Publish cell-local JUnit results without depending on a shared final workspace.
  • Compare fail-fast on/off using a controlled synthetic failure.

1. Disposable lab preflight

Use Jenkins 2.568.3 LTS on Java 21 with Pipeline/Declarative and JUnit versions listed in the references. Keep the built-in node at zero executors. Create or reuse one disposable trusted POSIX agent named ch21-twoexec with label ch21 and exactly two executors.

The job uses no credentials, no external services and no deployment side effects. If your agent does not have POSIX shell tools, use an equivalent disposable Linux container/VM agent rather than moving the workload onto the controller.

printf 'node=%s executor=%s workspace=%s\n' "$NODE_NAME" "$EXECUTOR_NUMBER" "$WORKSPACE"
java -version 2>&1 | head -1 || true
git --version

2. Create a synthetic source repository

Create a small repository containing eight stable test IDs and one script. The tests are synthetic sleeps; their only purpose is to make scheduling visible.

set -eu
mkdir -p ch21-lab && cd ch21-lab
git init
git config user.name 'DevOps Academy Lab'
git config user.email 'lab@example.invalid'
cat > tests.txt <<'EOF'
test-01|2
test-02|4
test-03|1
test-04|3
test-05|2
test-06|5
test-07|1
test-08|4
EOF
cat > run-shard.sh <<'SH'
#!/bin/sh
set -eu
MODE="$1"; SHARD="$2"; COUNT="$3"
mkdir -p results manifests
sort tests.txt > tests.sorted
awk -F'|' -v s="$SHARD" -v n="$COUNT" '((NR-1)%n)==s {print}' tests.sorted > "manifests/${MODE}-shard-${SHARD}.txt"
TOTAL=0; CASES=""
while IFS='|' read -r name sec; do
  [ -n "$name" ] || continue
  start=$(date +%s)
  sleep "$sec"
  end=$(date +%s)
  dur=$((end-start)); TOTAL=$((TOTAL+dur))
  CASES="$CASES"
done < "manifests/${MODE}-shard-${SHARD}.txt"
printf '<testsuite name="%s-shard-%s" tests="%s" failures="0" time="%s">%s</testsuite>\n' \
  "$MODE" "$SHARD" "$(wc -l < "manifests/${MODE}-shard-${SHARD}.txt")" "$TOTAL" "$CASES" \
  > "results/${MODE}-shard-${SHARD}.xml"
SH
chmod +x run-shard.sh
git add . && git commit -m 'ch21 synthetic test workload'
git rev-parse HEAD

Use this exact commit SHA in the Jenkins job. If you create the repository another way, record the resulting revision and do not mix results from different revisions.

3. Prove shard integrity before parallel execution

For two shards, generate the manifests without sleeping, then prove the union equals the source inventory and no test appears more than once.

cut -d'|' -f1 tests.txt | sort > expected.ids
for s in 0 1; do
  awk -F'|' -v shard="$s" -v count=2 '((NR-1)%count)==shard {print $1}' tests.txt | sort > "shard-$s.ids"
done
cat shard-0.ids shard-1.ids | sort > union.ids
cmp expected.ids union.ids
[ "$(cat shard-0.ids shard-1.ids | sort | uniq -d | wc -l)" -eq 0 ]

This validation is cheap and should fail before expensive test execution if shard math is wrong.

4. Establish a sequential baseline

First run the two shards sequentially on one allocated executor. This gives a baseline for setup and total work.

pipeline {
  agent { label 'ch21' }
  stages {
    stage('Sequential baseline') {
      steps {
        sh 'date +%s > baseline-start.txt'
        sh './run-shard.sh normal 0 2'
        sh './run-shard.sh normal 1 2'
        sh 'date +%s > baseline-end.txt'
      }
    }
  }
  post {
    always {
      junit 'results/*.xml'
      archiveArtifacts artifacts: 'manifests/*.txt,baseline-*.txt', fingerprint: true
    }
  }
}

Record wall-clock duration, node/executor, source SHA and the two JUnit suites. This is the comparison point, not a performance promise.

5. Run two named parallel shards

Now let each shard request its own executor. Because the lab has two executors, both can run concurrently if nothing else occupies the agent.

pipeline {
  agent none
  stages {
    stage('Parallel shards') {
      parallel {
        stage('Shard 0') {
          agent { label 'ch21' }
          steps { sh './run-shard.sh normal 0 2' }
          post {
            always {
              junit 'results/normal-shard-0.xml'
              archiveArtifacts artifacts: 'manifests/normal-shard-0.txt', fingerprint: true
            }
          }
        }
        stage('Shard 1') {
          agent { label 'ch21' }
          steps { sh './run-shard.sh normal 1 2' }
          post {
            always {
              junit 'results/normal-shard-1.xml'
              archiveArtifacts artifacts: 'manifests/normal-shard-1.txt', fingerprint: true
            }
          }
        }
      }
    }
  }
}

In a real SCM-backed job, each agent normally checks out the exact same source revision before running. Keep branch workspaces independent; do not redirect both branches to a common custom workspace.

6. Expand to a bounded four-cell matrix

Use two dimensions: MODE={normal,strict} and SHARD={0,1}. Cardinality is four cells, but capacity remains two executors, so expect approximately two waves.

pipeline {
  agent none
  stages {
    stage('Bounded matrix') {
      failFast false
      matrix {
        axes {
          axis { name 'MODE'; values 'normal', 'strict' }
          axis { name 'SHARD'; values '0', '1' }
        }
        agent { label 'ch21' }
        stages {
          stage('Synthetic test') {
            steps {
              sh '''
                set -eu
                printf 'cell=%s/%s node=%s executor=%s workspace=%s start=%s\\n' \
                  "$MODE" "$SHARD" "$NODE_NAME" "$EXECUTOR_NUMBER" "$WORKSPACE" "$(date +%s)"
                ./run-shard.sh "$MODE" "$SHARD" 2
                printf 'cell=%s/%s end=%s\\n' "$MODE" "$SHARD" "$(date +%s)"
              '''
            }
          }
        }
        post {
          always {
            junit "results/${MODE}-shard-${SHARD}.xml"
            archiveArtifacts artifacts: "manifests/${MODE}-shard-${SHARD}.txt", fingerprint: true
          }
        }
      }
    }
  }
}

Do not interpret the matrix graph alone as proof of simultaneous execution. Use timestamps, queue observations and executor IDs.

7. Compare collect-all and fail-fast with a synthetic failure

Create a guarded parameter such as FAIL_CELL whose default is none. In one disposable run, fail exactly strict/0 before its synthetic test. First run with failFast false and observe that other cells finish and publish reports. Then change only the stage policy to failFast true and compare interruption/evidence.

script {
  if (params.FAIL_CELL == "${MODE}/${SHARD}") {
    error "synthetic failure for ${MODE}/${SHARD}"
  }
}

Use no real publication/deployment in this experiment. Preserve both run numbers and source SHA so the only intended difference is fail-fast policy.

8. Measure queue and wall-clock effects

For each run, capture the build duration, per-cell start/end timestamps, node/executor, and whether a cell was aborted. The four-cell matrix on two executors cannot have four simultaneous cells. If branch startup and checkout dominate, the matrix can be slower than a simpler two-shard design even though its graph looks more parallel.

Run Cells Executors Fail-fast Expected observation
Sequential 2 sequential shards 1 occupied N/A Little queueing; sum of shard work
Parallel 2 branches 2 Off Both can execute together
Matrix 4 cells 2 Off At least two capacity waves
Matrix failure 4 cells 2 On Some siblings may be interrupted; less complete evidence

9. Challenge: identify the correct layer

Your Jenkinsfile creates eight cells, but only two run and six sit queued. No branch is failing. Should you fix Declarative syntax, JUnit, credentials, or capacity? The first layer is queue/capacity: verify label eligibility and executor availability. Only after measuring should you decide whether to reduce matrix cardinality or add reviewed capacity.

10. Cleanup

Remove the disposable job/repository only after exporting the evidence packet you want to keep. Do not delete first-failure runs merely to make the build history green. Reset the lab agent executor count to its prior value if you changed it for this chapter.

Next lesson

Configuration, Design Choices, and Tradeoffs

Choose parallel/matrix/sharding/fail-fast strategies based on coverage, capacity, auditability and critical-path evidence.

Knowledge check

Answer before revealing the explanation.

1. Why does the lab use exactly two executors?

2. Why partition a sorted test manifest instead of using a language runtime hash?

3. Where should JUnit publication happen when each matrix cell has its own workspace?

4. What should be measured besides total wall-clock time?

5. Should a synthetic fail-fast experiment include a real deployment?

Official references and version notes

Version note — 2026-09-17: executable examples assume Jenkins 2.568.3 LTS, Java 21, Pipeline 608.v67378e9d3db_1, Declarative 2.2293.v6e7193cec599, and JUnit 1425.v9c7318dca_96d. The mandatory sharding algorithm uses only POSIX shell/Python-style reasoning and Jenkins Pipeline features; no commercial or test-splitting plugin is required. Re-check current plugin/core compatibility and security advisories before production adoption.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.