Checkpoint Lab — Parallel Stages, Matrix Builds, Fail-Fast Behavior, Test Sharding, and High-Throughput Pipeline Design
Design and run a bounded test matrix under a fixed two-executor budget, predict its queue waves and critical path, compare predictions with measurements, inject one controlled failure, and tune without losing per-shard evidence.
Learning objectives
- Design a four-cell matrix under a fixed two-executor budget and predict queue waves.
- Prove deterministic shard membership before execution.
- Capture per-cell execution, JUnit and manifest evidence.
- Compare collect-all and fail-fast behavior with a synthetic failure.
- Tune the design based on measured critical path without dropping coverage.
1. Scenario and fixed constraints
You own a synthetic verification Pipeline. Capacity is fixed: one
trusted ch21 agent with two executors; the built-in
node has zero executors. The source revision and eight-test
inventory from Lesson 2 are fixed. You may change matrix/shard
design, but you may not add executors, agents, credentials, cloud
resources or production side effects.
Required matrix: MODE={normal,strict} ×
SHARD={0,1} = four cells. Each cell publishes one
uniquely named JUnit XML and one shard manifest.
2. Version/tool assumptions
| Component | Checkpoint baseline |
|---|---|
| Jenkins | 2.568.3 LTS |
| Java for Jenkins components | 21 |
| Pipeline | 608.v67378e9d3db_1 |
| Pipeline: Declarative | 2.2293.v6e7193cec599 |
| JUnit | 1425.v9c7318dca_96d |
| Agent |
Disposable POSIX agent, label ch21, exactly 2
executors
|
| External services | None |
3. Preflight and prediction
Before running, capture job full name, source SHA, test inventory digest, agent name/labels, two-executor setting and current queue state. Then write your predictions:
- Four matrix cells will be created but at most two will execute simultaneously.
- At least two cells will wait for capacity unless their timing happens not to overlap.
- Every cell will produce a unique shard manifest and JUnit suite.
- With collect-all, one synthetic failing cell will not stop the others; with fail-fast, siblings may be interrupted.
Also estimate the lower bound from total synthetic work and the slowest cell. Keep the estimate with the evidence packet; do not edit it after seeing results.
4. Gate shard integrity
set -eu
cut -d'|' -f1 tests.txt | sort > expected.ids
for s in 0 1; do
awk -F'|' -v shard="$s" -v count=2 '((NR-1)%count)==shard {print $1}' tests.txt | sort > "shard-$s.ids"
done
cat shard-0.ids shard-1.ids | sort > union.ids
cmp expected.ids union.ids
[ "$(cat shard-0.ids shard-1.ids | sort | uniq -d | wc -l)" -eq 0 ]
sha256sum expected.ids shard-0.ids shard-1.ids
Fail the checkpoint here if coverage is incomplete or duplicated. High throughput is irrelevant if it changes what is tested.
5. Exact bounded matrix
pipeline {
agent none
parameters {
choice(name: 'FAIL_CELL', choices: ['none','strict/0'], description: 'Synthetic lab failure only')
}
stages {
stage('Matrix') {
// Run A: failFast false. Run B: change this one line to true and commit it.
failFast false
matrix {
axes {
axis { name 'MODE'; values 'normal', 'strict' }
axis { name 'SHARD'; values '0', '1' }
}
agent { label 'ch21' }
stages {
stage('Verify') {
steps {
sh '''
set -eu
mkdir -p timing
printf 'cell=%s/%s node=%s executor=%s workspace=%s source=%s start=%s\\n' \
"$MODE" "$SHARD" "$NODE_NAME" "$EXECUTOR_NUMBER" "$WORKSPACE" \
"$(git rev-parse HEAD)" "$(date +%s)" | tee "timing/${MODE}-${SHARD}-start.txt"
'''
script {
if (params.FAIL_CELL == "${MODE}/${SHARD}") {
error "synthetic checkpoint failure for ${MODE}/${SHARD}"
}
}
sh './run-shard.sh "$MODE" "$SHARD" 2'
sh '''
printf 'cell=%s/%s end=%s\\n' "$MODE" "$SHARD" "$(date +%s)" \
| tee "timing/${MODE}-${SHARD}-end.txt"
'''
}
}
}
post {
always {
junit testResults: "results/${MODE}-shard-${SHARD}.xml", allowEmptyResults: true
archiveArtifacts artifacts: "manifests/${MODE}-shard-${SHARD}.txt,timing/${MODE}-${SHARD}-*.txt", allowEmptyArchive: true, fingerprint: true
}
}
}
}
}
}
Why allowEmptyResults is true only in this
failure-injection checkpoint:
a deliberately failed cell may stop before creating XML. The missing
report is itself expected evidence for that controlled failure. In
normal verification jobs, prefer failure on missing reports so
misconfiguration is visible.
6. Execute three controlled runs
-
Run A:
FAIL_CELL=none,failFast false. Measure healthy four-cell scheduling. -
Run B:
FAIL_CELL=strict/0, stillfailFast false. Prove other cells can finish and publish evidence. -
Run C: commit only
failFast true, keep the same synthetic failure. Record which sibling cells are interrupted.
Keep every run number/source revision. Do not overwrite Run A/B evidence with Run C artifacts.
7. Compare prediction with measurements
Build a table from the archived timing files and console logs:
| Cell | Queued/start | Executor | End/result | JUnit? | Manifest? |
|---|---|---|---|---|---|
| normal/0 | record | record | record | yes/no | yes/no |
| normal/1 | record | record | record | yes/no | yes/no |
| strict/0 | record | record | record | yes/no | yes/no |
| strict/1 | record | record | record | yes/no | yes/no |
Explain which two cells overlapped, which waited, and what branch or wave determined the critical path. If the observed order differs from your prediction, record why instead of changing the original prediction.
8. Tune under the same two-executor budget
Now compare the four-cell matrix with a simpler two-shard parallel
design where MODE becomes a test parameter inside each
shard rather than a separate cell. Coverage remains the same, but
branch count drops from four to two. Re-run and compare
setup/queue/wall time.
A valid improvement is lower critical-path or queue/setup overhead with identical test/shard evidence. If four cells are clearer for production support-matrix attribution, keeping them can also be justified—the decision must cite measured evidence, not aesthetics.
9. Required evidence packet
| Evidence | Capture |
|---|---|
| Controller baseline | Jenkins, Java, Pipeline/Declarative/JUnit versions |
| Source identity | Job full name, build number/URL, cause, exact Git SHA |
| Capacity | Agent/label, executor count, queue observations |
| Matrix design | Axes, cardinality, failFast policy, exclusions if any |
| Shard identity | Inventory + per-shard manifests and SHA-256 |
| Execution | Per-cell node/executor/workspace/start/end/result |
| Reports | JUnit suites/results and intentionally missing/aborted report evidence |
| Artifacts | Archived timing/manifests and fingerprints/digests |
| Performance | Predicted vs measured critical path/wall time |
| Assumptions | Disposable two-executor lab; no credentials/external side effects |
10. Cleanup and rollback
Restore any temporary executor-setting change, remove only the
disposable ch21 job/agent/repository when you no longer
need them, and retain the checkpoint build records/evidence until
review is complete. Do not delete failed runs just because they are
red.
11. What Chapter 21 adds to a production Jenkins operating model
You can now describe high-throughput Pipeline behavior in auditable terms: exact matrix cardinality, branch identity, deterministic shard membership, capacity/queue constraints, failure propagation policy, per-branch report/artifact identity, and measured critical path. Throughput is no longer “more parallel is better”; it is a bounded scheduling design with reproducible coverage.
Chapter 22 builds on this by creating Pipelines per branch and pull request. That adds another multiplicative dimension—SCM branch discovery—so trust, indexing and Jenkinsfile provenance must be explicit before high-concurrency branch builds are allowed to consume credentials or capacity.
Knowledge check
Answer before revealing the explanation.
1. What should you predict before starting the checkpoint?
Predict matrix cardinality, maximum simultaneous cells under two executors, which cells will queue, an approximate lower bound/critical path, and which evidence files each cell should publish.
2. What result proves the shard algorithm is deterministic?
The same source revision, axis values, shard count and test manifest produce the same per-shard membership manifest across reruns.
3. Why publish JUnit from every cell instead of only at the end?
Each cell owns a separate workspace that may disappear or be reused. Publishing while the report exists avoids losing evidence and lets Jenkins aggregate results into the run record.
4. What is a valid tuning outcome under a fixed executor budget?
A smaller or better-balanced cell/shard design that reduces queue/setup overhead or the slowest shard while keeping coverage and evidence intact. More branches alone is not a success metric.
5. What does Chapter 21 add to the operating model?
It makes throughput reviewable: axis cardinality, branch identity, deterministic shard membership, queue/executor time, fail-fast semantics, per-branch evidence and measured critical path are explicit production design inputs.
Official references and version notes
-
Jenkins LTS changelog
— chapter baseline
Jenkins 2.568.3 LTS, released 2026-09-02 and tested with Java 21 and 25; labs use Java 21 for Jenkins components. - Jenkins Java support policy — verify controller/agent JVM support before reproducing the lab on another LTS line.
-
Declarative Pipeline: parallel
— nested parallel stages,
failFast trueandparallelsAlwaysFailFast(). - Declarative Pipeline: matrix — axes, static cell generation, excludes, per-cell directives and fail-fast behavior.
-
Pipeline: Declarative plugin
— version
2.2293.v6e7193cec599, requiring Jenkins 2.504.3. -
Pipeline plugin
— version
608.v67378e9d3db_1. -
JUnit plugin
— version
1425.v9c7318dca_96d, requiring Jenkins 2.541.3. - JUnit Pipeline step — JUnit XML publication semantics and report glob behavior.
Version note — 2026-09-17: executable examples
assume Jenkins 2.568.3 LTS, Java 21, Pipeline
608.v67378e9d3db_1, Declarative
2.2293.v6e7193cec599, and JUnit
1425.v9c7318dca_96d. The mandatory sharding algorithm
uses only POSIX shell/Python-style reasoning and Jenkins Pipeline
features; no commercial or test-splitting plugin is required.
Re-check current plugin/core compatibility and security advisories
before production adoption.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.