Checkpoint Lab — Pipeline Steps, Error Handling, retry, timeout, catchError, input, waitUntil, and Resilient Flow Control
Build and verify an auditable resilient workflow with a flaky read-only step, bounded waiting and approval, a guarded simulated side effect, an intentional timeout/failure, preserved first-failure evidence, and idempotent rerun behavior.
Learning objectives
- Predict state transitions for retry, timeout/input, build result, target mutation, and rerun before execution.
- Run a safe flaky read-only operation under bounded retry and prove attempt count.
- Preserve an intentional first failure/timeout without hiding it or repeating the simulated side effect.
- Record approver identity and verify no agent is held during the human wait.
- Produce an evidence packet that correlates Jenkins run state, source revision, attempt history, and synthetic target state.
1. Scenario and acceptance target
You operate a synthetic maintenance workflow. It must perform a flaky read-only precheck, wait for readiness, request approval, create one local target marker, run an optional verifier, and retain evidence. You will deliberately break the readiness path once, then repair it. The side effect must never duplicate.
2. Record current assumptions
- Jenkins 2.568.3 LTS; Java 21 or 25.
- Pipeline aggregator 608.v67378e9d3db_1.
- Pipeline: Basic Steps 1098.v808b_fd7f8cf4.
- Pipeline: Input Step 560.v56198a_642157.
-
Disposable
lab-linuxagent; built-in controller node not used for routine commands. - Synthetic repository and local target only; no real credential or external provider.
3. Predict before running
| Event | Prediction | Evidence to collect |
|---|---|---|
| Flaky precheck | Attempt 1 fails; attempt 2 succeeds | Console + attempt counter |
| Broken readiness wait | Outer timeout interrupts run before approval/mutation | Timeout message + no target marker |
| Approval in repaired run | No lab agent reserved while waiting | Executor view + stage timestamps |
| Guarded mutation | Exactly one marker for operation ID | Target file + checksum |
| Optional verifier | Run becomes UNSTABLE, evidence still archives | Stage/build result + artifacts |
| Second execution with same synthetic operation ID | No duplicate mutation | already_applied evidence |
4. Create the synthetic repository
mkdir jenkins-ch12-checkpoint && cd jenkins-ch12-checkpoint
git init
git config user.name "Jenkins Chapter 12 Checkpoint"
git config user.email "jenkins-ch12-checkpoint@example.invalid"
mkdir -p scripts
cat > scripts/precheck.sh <<'EOF'
#!/usr/bin/env sh
set -eu
mkdir -p .lab-state
f=.lab-state/precheck-count
n=0
[ -f "$f" ] && n=$(cat "$f")
n=$((n + 1))
printf '%s\n' "$n" > "$f"
printf 'precheck_attempt=%s\n' "$n"
[ "$n" -ge 2 ]
EOF
chmod +x scripts/precheck.sh
printf 'chapter12-checkpoint\n' > README.txt
git add . && git commit -m 'ch12 checkpoint baseline'
git rev-parse HEAD
5. Jenkinsfile with explicit boundaries
pipeline {
agent none
options { timestamps() }
parameters {
string(name: 'OPERATION_ID', defaultValue: 'maintenance-demo-001', description: 'Synthetic immutable operation id')
}
stages {
stage('Baseline') {
agent { label 'lab-linux' }
steps {
checkout scm
sh '''
set -eu
rm -rf .lab-state evidence
mkdir -p evidence synthetic-target
{
printf 'job=%s\n' "$JOB_NAME"
printf 'build=%s\n' "$BUILD_NUMBER"
printf 'build_url=%s\n' "$BUILD_URL"
printf 'node=%s\n' "$NODE_NAME"
printf 'workspace=%s\n' "$WORKSPACE"
printf 'source_sha=%s\n' "$(git rev-parse HEAD)"
} > evidence/context.txt
'''
}
}
stage('Precheck') {
agent { label 'lab-linux' }
steps {
retry(3) { sh './scripts/precheck.sh' }
sh 'cp .lab-state/precheck-count evidence/precheck-count.txt'
}
}
stage('Readiness') {
agent { label 'lab-linux' }
steps {
timeout(time: 12, unit: 'SECONDS') {
waitUntil(initialRecurrencePeriod: 500, quiet: true) {
return fileExists('synthetic-ready.flag')
}
}
}
}
stage('Approval') {
steps {
script {
timeout(time: 10, unit: 'MINUTES') {
env.LAB_APPROVER = input(
message: "Apply ${params.OPERATION_ID} to the synthetic local target?",
ok: 'Apply',
submitter: 'lab-approver',
submitterParameter: 'APPROVER'
)
}
}
}
}
stage('Apply once') {
agent { label 'lab-linux' }
steps {
sh '''
set -eu
case "$OPERATION_ID" in
maintenance-demo-[0-9][0-9][0-9]) ;;
*) printf 'invalid operation id\n' >&2; exit 2 ;;
esac
marker="synthetic-target/${OPERATION_ID}.done"
if [ -f "$marker" ]; then
printf 'already_applied=%s\n' "$OPERATION_ID" > evidence/apply.txt
else
printf 'operation=%s\nbuild=%s\n' "$OPERATION_ID" "$BUILD_NUMBER" > "$marker"
printf 'applied=%s\n' "$OPERATION_ID" > evidence/apply.txt
fi
sha256sum "$marker" > evidence/target.sha256
'''
writeFile file: 'evidence/approver.txt', text: "${env.LAB_APPROVER}\n"
}
}
stage('Optional verification') {
agent { label 'lab-linux' }
steps {
catchError(buildResult: 'UNSTABLE', stageResult: 'UNSTABLE', catchInterruptions: false,
message: 'Optional checkpoint verifier failed') {
sh 'test -f deliberately-absent-optional-report.txt'
}
}
}
}
post {
always {
node('lab-linux') {
sh 'find evidence -maxdepth 1 -type f -print | sort > evidence/manifest.txt || true'
archiveArtifacts artifacts: 'evidence/*,synthetic-target/*', allowEmptyArchive: true, fingerprint: true
}
}
}
}
The post block deliberately reacquires an agent to
retain whatever evidence exists. In a production design, choose an
evidence path that is reliable even if the original agent is gone;
Chapter 14 will go deeper into artifacts and stash semantics.
6. Broken run: let readiness time out
Do not create synthetic-ready.flag. Run the baseline
revision. Preserve the build URL, source SHA, precheck attempts,
timeout console text, final result, and proof that
synthetic-target/${OPERATION_ID}.done does not exist.
This is the intentionally broken example.
Do not add catchError around the readiness timeout
merely to force continuation. The checkpoint requires the
interruption to stop the mutation.
7. Repair the readiness precondition
Commit a new revision that creates the readiness flag from a bounded
synthetic producer before waitUntil:
sh '(sleep 3; touch synthetic-ready.flag) >/dev/null 2>&1 &'
timeout(time: 12, unit: 'SECONDS') {
waitUntil(initialRecurrencePeriod: 500, quiet: true) {
return fileExists('synthetic-ready.flag')
}
}
Run the repaired revision, wait at the Approval stage, and verify the agent is not occupied by that stage. Approve using the lab-only approver identity, then allow Apply once and Optional verification to run.
8. Prove idempotent repeated execution
Run again with the same OPERATION_ID. The synthetic
target persists outside the cleaned workspace state for this
exercise. Expected evidence:
already_applied=maintenance-demo-001, and the existing
marker checksum remains unchanged. Do not create a second marker.
9. Required evidence packet
-
baseline.txt: Jenkins core/Java and relevant Pipeline plugin versions. -
predictions.md: predictions made before the runs. - Broken and repaired source/Jenkinsfile SHAs.
- Job full name, build numbers/URLs, causes, queue IDs where observed.
- Node/label/executor/workspace evidence for agent-backed stages.
precheck-count.txtproving retry count.- Broken-run timeout/interruption excerpt and result.
- Approver identity and evidence that Approval did not hold the lab executor.
-
apply.txt, target marker, andtarget.sha256. - Optional verification UNSTABLE stage/build evidence.
- Second-run
already_appliedevidence. -
limitations.md: local marker is only a simulation of durable target-side idempotency.
10. Verification checklist
| Check | Pass condition |
|---|---|
| Retry safe | Only read-only precheck is retried; count is bounded and archived |
| Timeout truthful | Broken run stops before approval/mutation and remains visibly interrupted/failed |
| Approval efficient | No lab executor is held while waiting |
| Side effect guarded | One deterministic target marker per operation ID |
| Error downgrade explicit | Optional verifier yields UNSTABLE, not green |
| Rerun safe | Same operation ID reports already applied and does not duplicate target state |
| Evidence preserved | Broken and repaired runs remain distinguishable by build URL and source SHA |
11. Cleanup and rollback
Download the complete evidence packet. Delete only the Chapter 12 job and synthetic repository/target created for this checkpoint. Do not remove shared Pipeline plugins, shared agents, or unrelated credentials/jobs. If a lab approver account was created solely for this chapter, remove it only after confirming it is not reused elsewhere.
12. What Chapter 12 adds to the operating model
You can now distinguish resilience from concealment: transient read-only work can retry safely, waits have deadlines, approvals are auditable and executor-efficient, failures remain visible, interruption semantics are preserved, and side effects require idempotency or target-state verification before repetition.
Chapter 13 continues from here into Groovy CPS internals,
serialization, @NonCPS, Pipeline state across restarts,
and the programming pitfalls that become visible once resilient
flows suspend and resume repeatedly.
Knowledge check
What must happen in the intentionally broken run?
The readiness timeout must stop the flow before approval and before the synthetic target marker is created; the failure evidence must remain preserved.
What proves the precheck retry is bounded and successful?
Archived attempt count plus console evidence showing the first failure and later success within the configured retry limit.
Why does the checkpoint use a deterministic
OPERATION_ID?
It provides an identity that can be checked against target state so repeated execution can detect an already-applied operation instead of duplicating it.
Why is the optional verifier allowed to produce UNSTABLE?
It is explicitly classified as non-blocking; the downgrade is intentional and visible rather than pretending the check succeeded.
What is the bridge to Chapter 13?
These control steps suspend/resume Pipeline execution, so the next chapter examines which Groovy/CPS state survives those suspension points and restarts safely.
Official references and version notes
- Jenkins LTS changelog — current LTS and tested Java configurations.
-
Pipeline: Basic Steps
— current
retry,timeout,catchError,waitUntil,sleep, and related step contracts. -
Pipeline: Input Step reference
—
input, submitter restrictions, identifiers, and captured approver identity. - Pipeline: Basic Steps plugin — current plugin baseline and dependencies.
- Pipeline: Input Step plugin — current plugin baseline and security history.
- Jenkins Pipeline handbook — durable Pipeline execution and Jenkinsfile concepts.
- Pipeline Best Practices — controller/agent boundaries and safe Pipeline design.
Rechecked on 2026-09-16. Examples assume Jenkins 2.568.3 LTS, tested with Java 21 and 25, Pipeline aggregator 608.v67378e9d3db_1, Pipeline: Basic Steps 1098.v808b_fd7f8cf4, and Pipeline: Input Step 560.v56198a_642157. The mandatory path is local/disposable, uses synthetic state and fake identities, performs no production deployment/publication, and does not require commercial services. Always record the versions actually installed on your controller because plugins release independently from Jenkins core.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.