Parallel Stages, Matrix Builds, Fail-Fast Behavior, Test Sharding, and High-Throughput Pipeline Design: Guided Hands-On Workflow and Core Operations
Build a disposable two-executor lab, run synthetic tests sequentially and in bounded parallel/matrix forms, partition test IDs deterministically, publish JUnit evidence per branch, and measure how queue pressure changes wall-clock time.
Learning objectives
- Prepare a disposable two-executor Jenkins agent and synthetic test inventory.
- Compare sequential, parallel and bounded matrix execution using the same work.
- Create deterministic shard manifests and validate coverage before tests run.
- Publish cell-local JUnit results without depending on a shared final workspace.
- Compare fail-fast on/off using a controlled synthetic failure.
1. Disposable lab preflight
Use Jenkins 2.568.3 LTS on Java 21 with
Pipeline/Declarative and JUnit versions listed in the references.
Keep the built-in node at zero executors. Create or reuse one
disposable trusted POSIX agent named ch21-twoexec with
label ch21 and exactly two executors.
The job uses no credentials, no external services and no deployment side effects. If your agent does not have POSIX shell tools, use an equivalent disposable Linux container/VM agent rather than moving the workload onto the controller.
printf 'node=%s executor=%s workspace=%s\n' "$NODE_NAME" "$EXECUTOR_NUMBER" "$WORKSPACE"
java -version 2>&1 | head -1 || true
git --version
2. Create a synthetic source repository
Create a small repository containing eight stable test IDs and one script. The tests are synthetic sleeps; their only purpose is to make scheduling visible.
set -eu
mkdir -p ch21-lab && cd ch21-lab
git init
git config user.name 'DevOps Academy Lab'
git config user.email 'lab@example.invalid'
cat > tests.txt <<'EOF'
test-01|2
test-02|4
test-03|1
test-04|3
test-05|2
test-06|5
test-07|1
test-08|4
EOF
cat > run-shard.sh <<'SH'
#!/bin/sh
set -eu
MODE="$1"; SHARD="$2"; COUNT="$3"
mkdir -p results manifests
sort tests.txt > tests.sorted
awk -F'|' -v s="$SHARD" -v n="$COUNT" '((NR-1)%n)==s {print}' tests.sorted > "manifests/${MODE}-shard-${SHARD}.txt"
TOTAL=0; CASES=""
while IFS='|' read -r name sec; do
[ -n "$name" ] || continue
start=$(date +%s)
sleep "$sec"
end=$(date +%s)
dur=$((end-start)); TOTAL=$((TOTAL+dur))
CASES="$CASES"
done < "manifests/${MODE}-shard-${SHARD}.txt"
printf '<testsuite name="%s-shard-%s" tests="%s" failures="0" time="%s">%s</testsuite>\n' \
"$MODE" "$SHARD" "$(wc -l < "manifests/${MODE}-shard-${SHARD}.txt")" "$TOTAL" "$CASES" \
> "results/${MODE}-shard-${SHARD}.xml"
SH
chmod +x run-shard.sh
git add . && git commit -m 'ch21 synthetic test workload'
git rev-parse HEAD
Use this exact commit SHA in the Jenkins job. If you create the repository another way, record the resulting revision and do not mix results from different revisions.
3. Prove shard integrity before parallel execution
For two shards, generate the manifests without sleeping, then prove the union equals the source inventory and no test appears more than once.
cut -d'|' -f1 tests.txt | sort > expected.ids
for s in 0 1; do
awk -F'|' -v shard="$s" -v count=2 '((NR-1)%count)==shard {print $1}' tests.txt | sort > "shard-$s.ids"
done
cat shard-0.ids shard-1.ids | sort > union.ids
cmp expected.ids union.ids
[ "$(cat shard-0.ids shard-1.ids | sort | uniq -d | wc -l)" -eq 0 ]
This validation is cheap and should fail before expensive test execution if shard math is wrong.
4. Establish a sequential baseline
First run the two shards sequentially on one allocated executor. This gives a baseline for setup and total work.
pipeline {
agent { label 'ch21' }
stages {
stage('Sequential baseline') {
steps {
sh 'date +%s > baseline-start.txt'
sh './run-shard.sh normal 0 2'
sh './run-shard.sh normal 1 2'
sh 'date +%s > baseline-end.txt'
}
}
}
post {
always {
junit 'results/*.xml'
archiveArtifacts artifacts: 'manifests/*.txt,baseline-*.txt', fingerprint: true
}
}
}
Record wall-clock duration, node/executor, source SHA and the two JUnit suites. This is the comparison point, not a performance promise.
5. Run two named parallel shards
Now let each shard request its own executor. Because the lab has two executors, both can run concurrently if nothing else occupies the agent.
pipeline {
agent none
stages {
stage('Parallel shards') {
parallel {
stage('Shard 0') {
agent { label 'ch21' }
steps { sh './run-shard.sh normal 0 2' }
post {
always {
junit 'results/normal-shard-0.xml'
archiveArtifacts artifacts: 'manifests/normal-shard-0.txt', fingerprint: true
}
}
}
stage('Shard 1') {
agent { label 'ch21' }
steps { sh './run-shard.sh normal 1 2' }
post {
always {
junit 'results/normal-shard-1.xml'
archiveArtifacts artifacts: 'manifests/normal-shard-1.txt', fingerprint: true
}
}
}
}
}
}
}
In a real SCM-backed job, each agent normally checks out the exact same source revision before running. Keep branch workspaces independent; do not redirect both branches to a common custom workspace.
6. Expand to a bounded four-cell matrix
Use two dimensions: MODE={normal,strict} and
SHARD={0,1}. Cardinality is four cells, but capacity
remains two executors, so expect approximately two waves.
pipeline {
agent none
stages {
stage('Bounded matrix') {
failFast false
matrix {
axes {
axis { name 'MODE'; values 'normal', 'strict' }
axis { name 'SHARD'; values '0', '1' }
}
agent { label 'ch21' }
stages {
stage('Synthetic test') {
steps {
sh '''
set -eu
printf 'cell=%s/%s node=%s executor=%s workspace=%s start=%s\\n' \
"$MODE" "$SHARD" "$NODE_NAME" "$EXECUTOR_NUMBER" "$WORKSPACE" "$(date +%s)"
./run-shard.sh "$MODE" "$SHARD" 2
printf 'cell=%s/%s end=%s\\n' "$MODE" "$SHARD" "$(date +%s)"
'''
}
}
}
post {
always {
junit "results/${MODE}-shard-${SHARD}.xml"
archiveArtifacts artifacts: "manifests/${MODE}-shard-${SHARD}.txt", fingerprint: true
}
}
}
}
}
}
Do not interpret the matrix graph alone as proof of simultaneous execution. Use timestamps, queue observations and executor IDs.
7. Compare collect-all and fail-fast with a synthetic failure
Create a guarded parameter such as FAIL_CELL whose
default is none. In one disposable run, fail exactly
strict/0 before its synthetic test. First run with
failFast false and observe that other cells finish and
publish reports. Then change only the stage policy to
failFast true and compare interruption/evidence.
script {
if (params.FAIL_CELL == "${MODE}/${SHARD}") {
error "synthetic failure for ${MODE}/${SHARD}"
}
}
Use no real publication/deployment in this experiment. Preserve both run numbers and source SHA so the only intended difference is fail-fast policy.
8. Measure queue and wall-clock effects
For each run, capture the build duration, per-cell start/end timestamps, node/executor, and whether a cell was aborted. The four-cell matrix on two executors cannot have four simultaneous cells. If branch startup and checkout dominate, the matrix can be slower than a simpler two-shard design even though its graph looks more parallel.
| Run | Cells | Executors | Fail-fast | Expected observation |
|---|---|---|---|---|
| Sequential | 2 sequential shards | 1 occupied | N/A | Little queueing; sum of shard work |
| Parallel | 2 branches | 2 | Off | Both can execute together |
| Matrix | 4 cells | 2 | Off | At least two capacity waves |
| Matrix failure | 4 cells | 2 | On | Some siblings may be interrupted; less complete evidence |
9. Challenge: identify the correct layer
Your Jenkinsfile creates eight cells, but only two run and six sit queued. No branch is failing. Should you fix Declarative syntax, JUnit, credentials, or capacity? The first layer is queue/capacity: verify label eligibility and executor availability. Only after measuring should you decide whether to reduce matrix cardinality or add reviewed capacity.
10. Cleanup
Remove the disposable job/repository only after exporting the evidence packet you want to keep. Do not delete first-failure runs merely to make the build history green. Reset the lab agent executor count to its prior value if you changed it for this chapter.
Knowledge check
Answer before revealing the explanation.
1. Why does the lab use exactly two executors?
A fixed small capacity makes queue waves observable. Learners can see that branch cardinality and executor capacity are separate quantities instead of assuming parallel syntax creates infinite compute.
2. Why partition a sorted test manifest instead of using a language runtime hash?
Many runtime hash functions are intentionally randomized between processes. Sorting stable test IDs and assigning by index modulo shard count gives repeatable membership without an extra plugin.
3. Where should JUnit publication happen when each matrix cell has its own workspace?
Publish each cell's XML while that cell still has its workspace, typically in a cell-level post block. Jenkins can aggregate multiple JUnit publications into the build record.
4. What should be measured besides total wall-clock time?
Record branch start/end times, NODE_NAME, EXECUTOR_NUMBER, workspace, queue observations, per-cell duration/result and report identity. Wall time alone cannot explain why a design is slow.
5. Should a synthetic fail-fast experiment include a real deployment?
No. Use a bounded fake failure or read-only test. Fail-fast cancellation is not a rollback mechanism and should not be demonstrated against irreversible production side effects.
Official references and version notes
-
Jenkins LTS changelog
— chapter baseline
Jenkins 2.568.3 LTS, released 2026-09-02 and tested with Java 21 and 25; labs use Java 21 for Jenkins components. - Jenkins Java support policy — verify controller/agent JVM support before reproducing the lab on another LTS line.
-
Declarative Pipeline: parallel
— nested parallel stages,
failFast trueandparallelsAlwaysFailFast(). - Declarative Pipeline: matrix — axes, static cell generation, excludes, per-cell directives and fail-fast behavior.
-
Pipeline: Declarative plugin
— version
2.2293.v6e7193cec599, requiring Jenkins 2.504.3. -
Pipeline plugin
— version
608.v67378e9d3db_1. -
JUnit plugin
— version
1425.v9c7318dca_96d, requiring Jenkins 2.541.3. - JUnit Pipeline step — JUnit XML publication semantics and report glob behavior.
Version note — 2026-09-17: executable examples
assume Jenkins 2.568.3 LTS, Java 21, Pipeline
608.v67378e9d3db_1, Declarative
2.2293.v6e7193cec599, and JUnit
1425.v9c7318dca_96d. The mandatory sharding algorithm
uses only POSIX shell/Python-style reasoning and Jenkins Pipeline
features; no commercial or test-splitting plugin is required.
Re-check current plugin/core compatibility and security advisories
before production adoption.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.