Chapter 21Lesson 05~210 minutes

Checkpoint Lab — Parallel Stages, Matrix Builds, Fail-Fast Behavior, Test Sharding, and High-Throughput Pipeline Design

Design and run a bounded test matrix under a fixed two-executor budget, predict its queue waves and critical path, compare predictions with measurements, inject one controlled failure, and tune without losing per-shard evidence.

Checkpoint labFixed budgetPredictionJUnitCritical pathTuning

Learning objectives

  • Design a four-cell matrix under a fixed two-executor budget and predict queue waves.
  • Prove deterministic shard membership before execution.
  • Capture per-cell execution, JUnit and manifest evidence.
  • Compare collect-all and fail-fast behavior with a synthetic failure.
  • Tune the design based on measured critical path without dropping coverage.

1. Scenario and fixed constraints

You own a synthetic verification Pipeline. Capacity is fixed: one trusted ch21 agent with two executors; the built-in node has zero executors. The source revision and eight-test inventory from Lesson 2 are fixed. You may change matrix/shard design, but you may not add executors, agents, credentials, cloud resources or production side effects.

Required matrix: MODE={normal,strict} × SHARD={0,1} = four cells. Each cell publishes one uniquely named JUnit XML and one shard manifest.

2. Version/tool assumptions

Component Checkpoint baseline
Jenkins 2.568.3 LTS
Java for Jenkins components 21
Pipeline 608.v67378e9d3db_1
Pipeline: Declarative 2.2293.v6e7193cec599
JUnit 1425.v9c7318dca_96d
Agent Disposable POSIX agent, label ch21, exactly 2 executors
External services None

3. Preflight and prediction

Before running, capture job full name, source SHA, test inventory digest, agent name/labels, two-executor setting and current queue state. Then write your predictions:

  1. Four matrix cells will be created but at most two will execute simultaneously.
  2. At least two cells will wait for capacity unless their timing happens not to overlap.
  3. Every cell will produce a unique shard manifest and JUnit suite.
  4. With collect-all, one synthetic failing cell will not stop the others; with fail-fast, siblings may be interrupted.

Also estimate the lower bound from total synthetic work and the slowest cell. Keep the estimate with the evidence packet; do not edit it after seeing results.

4. Gate shard integrity

set -eu
cut -d'|' -f1 tests.txt | sort > expected.ids
for s in 0 1; do
  awk -F'|' -v shard="$s" -v count=2 '((NR-1)%count)==shard {print $1}' tests.txt | sort > "shard-$s.ids"
done
cat shard-0.ids shard-1.ids | sort > union.ids
cmp expected.ids union.ids
[ "$(cat shard-0.ids shard-1.ids | sort | uniq -d | wc -l)" -eq 0 ]
sha256sum expected.ids shard-0.ids shard-1.ids

Fail the checkpoint here if coverage is incomplete or duplicated. High throughput is irrelevant if it changes what is tested.

5. Exact bounded matrix

pipeline {
  agent none
  parameters {
    choice(name: 'FAIL_CELL', choices: ['none','strict/0'], description: 'Synthetic lab failure only')
  }
  stages {
    stage('Matrix') {
      // Run A: failFast false. Run B: change this one line to true and commit it.
      failFast false
      matrix {
        axes {
          axis { name 'MODE'; values 'normal', 'strict' }
          axis { name 'SHARD'; values '0', '1' }
        }
        agent { label 'ch21' }
        stages {
          stage('Verify') {
            steps {
              sh '''
                set -eu
                mkdir -p timing
                printf 'cell=%s/%s node=%s executor=%s workspace=%s source=%s start=%s\\n' \
                  "$MODE" "$SHARD" "$NODE_NAME" "$EXECUTOR_NUMBER" "$WORKSPACE" \
                  "$(git rev-parse HEAD)" "$(date +%s)" | tee "timing/${MODE}-${SHARD}-start.txt"
              '''
              script {
                if (params.FAIL_CELL == "${MODE}/${SHARD}") {
                  error "synthetic checkpoint failure for ${MODE}/${SHARD}"
                }
              }
              sh './run-shard.sh "$MODE" "$SHARD" 2'
              sh '''
                printf 'cell=%s/%s end=%s\\n' "$MODE" "$SHARD" "$(date +%s)" \
                  | tee "timing/${MODE}-${SHARD}-end.txt"
              '''
            }
          }
        }
        post {
          always {
            junit testResults: "results/${MODE}-shard-${SHARD}.xml", allowEmptyResults: true
            archiveArtifacts artifacts: "manifests/${MODE}-shard-${SHARD}.txt,timing/${MODE}-${SHARD}-*.txt", allowEmptyArchive: true, fingerprint: true
          }
        }
      }
    }
  }
}

Why allowEmptyResults is true only in this failure-injection checkpoint: a deliberately failed cell may stop before creating XML. The missing report is itself expected evidence for that controlled failure. In normal verification jobs, prefer failure on missing reports so misconfiguration is visible.

6. Execute three controlled runs

  1. Run A: FAIL_CELL=none, failFast false. Measure healthy four-cell scheduling.
  2. Run B: FAIL_CELL=strict/0, still failFast false. Prove other cells can finish and publish evidence.
  3. Run C: commit only failFast true, keep the same synthetic failure. Record which sibling cells are interrupted.

Keep every run number/source revision. Do not overwrite Run A/B evidence with Run C artifacts.

7. Compare prediction with measurements

Build a table from the archived timing files and console logs:

Cell Queued/start Executor End/result JUnit? Manifest?
normal/0 record record record yes/no yes/no
normal/1 record record record yes/no yes/no
strict/0 record record record yes/no yes/no
strict/1 record record record yes/no yes/no

Explain which two cells overlapped, which waited, and what branch or wave determined the critical path. If the observed order differs from your prediction, record why instead of changing the original prediction.

8. Tune under the same two-executor budget

Now compare the four-cell matrix with a simpler two-shard parallel design where MODE becomes a test parameter inside each shard rather than a separate cell. Coverage remains the same, but branch count drops from four to two. Re-run and compare setup/queue/wall time.

A valid improvement is lower critical-path or queue/setup overhead with identical test/shard evidence. If four cells are clearer for production support-matrix attribution, keeping them can also be justified—the decision must cite measured evidence, not aesthetics.

9. Required evidence packet

Evidence Capture
Controller baseline Jenkins, Java, Pipeline/Declarative/JUnit versions
Source identity Job full name, build number/URL, cause, exact Git SHA
Capacity Agent/label, executor count, queue observations
Matrix design Axes, cardinality, failFast policy, exclusions if any
Shard identity Inventory + per-shard manifests and SHA-256
Execution Per-cell node/executor/workspace/start/end/result
Reports JUnit suites/results and intentionally missing/aborted report evidence
Artifacts Archived timing/manifests and fingerprints/digests
Performance Predicted vs measured critical path/wall time
Assumptions Disposable two-executor lab; no credentials/external side effects

10. Cleanup and rollback

Restore any temporary executor-setting change, remove only the disposable ch21 job/agent/repository when you no longer need them, and retain the checkpoint build records/evidence until review is complete. Do not delete failed runs just because they are red.

11. What Chapter 21 adds to a production Jenkins operating model

You can now describe high-throughput Pipeline behavior in auditable terms: exact matrix cardinality, branch identity, deterministic shard membership, capacity/queue constraints, failure propagation policy, per-branch report/artifact identity, and measured critical path. Throughput is no longer “more parallel is better”; it is a bounded scheduling design with reproducible coverage.

Chapter 22 builds on this by creating Pipelines per branch and pull request. That adds another multiplicative dimension—SCM branch discovery—so trust, indexing and Jenkinsfile provenance must be explicit before high-concurrency branch builds are allowed to consume credentials or capacity.

Next chapter

Multibranch Pipelines, Branch Sources, Pull Request Discovery, Jenkinsfile Trust, and Branch Indexing

Apply the same capacity and evidence discipline when Jenkins automatically discovers branch and pull-request Pipelines.

Knowledge check

Answer before revealing the explanation.

1. What should you predict before starting the checkpoint?

2. What result proves the shard algorithm is deterministic?

3. Why publish JUnit from every cell instead of only at the end?

4. What is a valid tuning outcome under a fixed executor budget?

5. What does Chapter 21 add to the operating model?

Official references and version notes

Version note — 2026-09-17: executable examples assume Jenkins 2.568.3 LTS, Java 21, Pipeline 608.v67378e9d3db_1, Declarative 2.2293.v6e7193cec599, and JUnit 1425.v9c7318dca_96d. The mandatory sharding algorithm uses only POSIX shell/Python-style reasoning and Jenkins Pipeline features; no commercial or test-splitting plugin is required. Re-check current plugin/core compatibility and security advisories before production adoption.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.