Chapter 12Lesson 05~170 minutes

Checkpoint Lab — Pipeline Steps, Error Handling, retry, timeout, catchError, input, waitUntil, and Resilient Flow Control

Build and verify an auditable resilient workflow with a flaky read-only step, bounded waiting and approval, a guarded simulated side effect, an intentional timeout/failure, preserved first-failure evidence, and idempotent rerun behavior.

Checkpoint labResilienceEvidence packetIdempotencyApprovalRecovery

Learning objectives

  • Predict state transitions for retry, timeout/input, build result, target mutation, and rerun before execution.
  • Run a safe flaky read-only operation under bounded retry and prove attempt count.
  • Preserve an intentional first failure/timeout without hiding it or repeating the simulated side effect.
  • Record approver identity and verify no agent is held during the human wait.
  • Produce an evidence packet that correlates Jenkins run state, source revision, attempt history, and synthetic target state.

1. Scenario and acceptance target

You operate a synthetic maintenance workflow. It must perform a flaky read-only precheck, wait for readiness, request approval, create one local target marker, run an optional verifier, and retain evidence. You will deliberately break the readiness path once, then repair it. The side effect must never duplicate.

Success means truthful evidence, not green at any cost. The broken run must remain visibly failed/aborted with its first-failure evidence. The repaired run must prove exactly one target mutation for its deterministic operation ID.

2. Record current assumptions

  • Jenkins 2.568.3 LTS; Java 21 or 25.
  • Pipeline aggregator 608.v67378e9d3db_1.
  • Pipeline: Basic Steps 1098.v808b_fd7f8cf4.
  • Pipeline: Input Step 560.v56198a_642157.
  • Disposable lab-linux agent; built-in controller node not used for routine commands.
  • Synthetic repository and local target only; no real credential or external provider.

3. Predict before running

Event Prediction Evidence to collect
Flaky precheck Attempt 1 fails; attempt 2 succeeds Console + attempt counter
Broken readiness wait Outer timeout interrupts run before approval/mutation Timeout message + no target marker
Approval in repaired run No lab agent reserved while waiting Executor view + stage timestamps
Guarded mutation Exactly one marker for operation ID Target file + checksum
Optional verifier Run becomes UNSTABLE, evidence still archives Stage/build result + artifacts
Second execution with same synthetic operation ID No duplicate mutation already_applied evidence

4. Create the synthetic repository

mkdir jenkins-ch12-checkpoint && cd jenkins-ch12-checkpoint
git init
git config user.name "Jenkins Chapter 12 Checkpoint"
git config user.email "jenkins-ch12-checkpoint@example.invalid"
mkdir -p scripts
cat > scripts/precheck.sh <<'EOF'
#!/usr/bin/env sh
set -eu
mkdir -p .lab-state
f=.lab-state/precheck-count
n=0
[ -f "$f" ] && n=$(cat "$f")
n=$((n + 1))
printf '%s\n' "$n" > "$f"
printf 'precheck_attempt=%s\n' "$n"
[ "$n" -ge 2 ]
EOF
chmod +x scripts/precheck.sh
printf 'chapter12-checkpoint\n' > README.txt
git add . && git commit -m 'ch12 checkpoint baseline'
git rev-parse HEAD

5. Jenkinsfile with explicit boundaries

pipeline {
  agent none
  options { timestamps() }
  parameters {
    string(name: 'OPERATION_ID', defaultValue: 'maintenance-demo-001', description: 'Synthetic immutable operation id')
  }
  stages {
    stage('Baseline') {
      agent { label 'lab-linux' }
      steps {
        checkout scm
        sh '''
          set -eu
          rm -rf .lab-state evidence
          mkdir -p evidence synthetic-target
          {
            printf 'job=%s\n' "$JOB_NAME"
            printf 'build=%s\n' "$BUILD_NUMBER"
            printf 'build_url=%s\n' "$BUILD_URL"
            printf 'node=%s\n' "$NODE_NAME"
            printf 'workspace=%s\n' "$WORKSPACE"
            printf 'source_sha=%s\n' "$(git rev-parse HEAD)"
          } > evidence/context.txt
        '''
      }
    }

    stage('Precheck') {
      agent { label 'lab-linux' }
      steps {
        retry(3) { sh './scripts/precheck.sh' }
        sh 'cp .lab-state/precheck-count evidence/precheck-count.txt'
      }
    }

    stage('Readiness') {
      agent { label 'lab-linux' }
      steps {
        timeout(time: 12, unit: 'SECONDS') {
          waitUntil(initialRecurrencePeriod: 500, quiet: true) {
            return fileExists('synthetic-ready.flag')
          }
        }
      }
    }

    stage('Approval') {
      steps {
        script {
          timeout(time: 10, unit: 'MINUTES') {
            env.LAB_APPROVER = input(
              message: "Apply ${params.OPERATION_ID} to the synthetic local target?",
              ok: 'Apply',
              submitter: 'lab-approver',
              submitterParameter: 'APPROVER'
            )
          }
        }
      }
    }

    stage('Apply once') {
      agent { label 'lab-linux' }
      steps {
        sh '''
          set -eu
          case "$OPERATION_ID" in
            maintenance-demo-[0-9][0-9][0-9]) ;;
            *) printf 'invalid operation id\n' >&2; exit 2 ;;
          esac
          marker="synthetic-target/${OPERATION_ID}.done"
          if [ -f "$marker" ]; then
            printf 'already_applied=%s\n' "$OPERATION_ID" > evidence/apply.txt
          else
            printf 'operation=%s\nbuild=%s\n' "$OPERATION_ID" "$BUILD_NUMBER" > "$marker"
            printf 'applied=%s\n' "$OPERATION_ID" > evidence/apply.txt
          fi
          sha256sum "$marker" > evidence/target.sha256
        '''
        writeFile file: 'evidence/approver.txt', text: "${env.LAB_APPROVER}\n"
      }
    }

    stage('Optional verification') {
      agent { label 'lab-linux' }
      steps {
        catchError(buildResult: 'UNSTABLE', stageResult: 'UNSTABLE', catchInterruptions: false,
                   message: 'Optional checkpoint verifier failed') {
          sh 'test -f deliberately-absent-optional-report.txt'
        }
      }
    }
  }
  post {
    always {
      node('lab-linux') {
        sh 'find evidence -maxdepth 1 -type f -print | sort > evidence/manifest.txt || true'
        archiveArtifacts artifacts: 'evidence/*,synthetic-target/*', allowEmptyArchive: true, fingerprint: true
      }
    }
  }
}

The post block deliberately reacquires an agent to retain whatever evidence exists. In a production design, choose an evidence path that is reliable even if the original agent is gone; Chapter 14 will go deeper into artifacts and stash semantics.

6. Broken run: let readiness time out

Do not create synthetic-ready.flag. Run the baseline revision. Preserve the build URL, source SHA, precheck attempts, timeout console text, final result, and proof that synthetic-target/${OPERATION_ID}.done does not exist. This is the intentionally broken example.

Do not add catchError around the readiness timeout merely to force continuation. The checkpoint requires the interruption to stop the mutation.

7. Repair the readiness precondition

Commit a new revision that creates the readiness flag from a bounded synthetic producer before waitUntil:

sh '(sleep 3; touch synthetic-ready.flag) >/dev/null 2>&1 &'
timeout(time: 12, unit: 'SECONDS') {
  waitUntil(initialRecurrencePeriod: 500, quiet: true) {
    return fileExists('synthetic-ready.flag')
  }
}

Run the repaired revision, wait at the Approval stage, and verify the agent is not occupied by that stage. Approve using the lab-only approver identity, then allow Apply once and Optional verification to run.

8. Prove idempotent repeated execution

Run again with the same OPERATION_ID. The synthetic target persists outside the cleaned workspace state for this exercise. Expected evidence: already_applied=maintenance-demo-001, and the existing marker checksum remains unchanged. Do not create a second marker.

Interpretation: this is a local simulation of target-side idempotency. A workspace file alone is not adequate production protection because workspaces can disappear or be reused.

9. Required evidence packet

  • baseline.txt: Jenkins core/Java and relevant Pipeline plugin versions.
  • predictions.md: predictions made before the runs.
  • Broken and repaired source/Jenkinsfile SHAs.
  • Job full name, build numbers/URLs, causes, queue IDs where observed.
  • Node/label/executor/workspace evidence for agent-backed stages.
  • precheck-count.txt proving retry count.
  • Broken-run timeout/interruption excerpt and result.
  • Approver identity and evidence that Approval did not hold the lab executor.
  • apply.txt, target marker, and target.sha256.
  • Optional verification UNSTABLE stage/build evidence.
  • Second-run already_applied evidence.
  • limitations.md: local marker is only a simulation of durable target-side idempotency.

10. Verification checklist

Check Pass condition
Retry safe Only read-only precheck is retried; count is bounded and archived
Timeout truthful Broken run stops before approval/mutation and remains visibly interrupted/failed
Approval efficient No lab executor is held while waiting
Side effect guarded One deterministic target marker per operation ID
Error downgrade explicit Optional verifier yields UNSTABLE, not green
Rerun safe Same operation ID reports already applied and does not duplicate target state
Evidence preserved Broken and repaired runs remain distinguishable by build URL and source SHA

11. Cleanup and rollback

Download the complete evidence packet. Delete only the Chapter 12 job and synthetic repository/target created for this checkpoint. Do not remove shared Pipeline plugins, shared agents, or unrelated credentials/jobs. If a lab approver account was created solely for this chapter, remove it only after confirming it is not reused elsewhere.

12. What Chapter 12 adds to the operating model

You can now distinguish resilience from concealment: transient read-only work can retry safely, waits have deadlines, approvals are auditable and executor-efficient, failures remain visible, interruption semantics are preserved, and side effects require idempotency or target-state verification before repetition.

Chapter 13 continues from here into Groovy CPS internals, serialization, @NonCPS, Pipeline state across restarts, and the programming pitfalls that become visible once resilient flows suspend and resume repeatedly.

Next lesson

Chapter 13 — Groovy CPS, Serialization, @NonCPS, Pipeline State, Restarts, and Common Pipeline Programming Pitfalls

Deepen the durability model by examining CPS transformation, serializable state, @NonCPS boundaries, restart behavior, and method-mismatch pitfalls.

Knowledge check

What must happen in the intentionally broken run?

What proves the precheck retry is bounded and successful?

Why does the checkpoint use a deterministic OPERATION_ID?

Why is the optional verifier allowed to produce UNSTABLE?

What is the bridge to Chapter 13?

Official references and version notes

Version and compatibility note

Rechecked on 2026-09-16. Examples assume Jenkins 2.568.3 LTS, tested with Java 21 and 25, Pipeline aggregator 608.v67378e9d3db_1, Pipeline: Basic Steps 1098.v808b_fd7f8cf4, and Pipeline: Input Step 560.v56198a_642157. The mandatory path is local/disposable, uses synthetic state and fake identities, performs no production deployment/publication, and does not require commercial services. Always record the versions actually installed on your controller because plugins release independently from Jenkins core.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.