Chapter 12Lesson 04~130 minutes

Pipeline Steps, Error Handling, retry, timeout, catchError, input, waitUntil, and Resilient Flow Control: Diagnostics, Failure Modes, Security, and Performance

Diagnose resilience failures without erasing evidence: blind retries, false-green error handling, executor-holding approvals, infinite polling, swallowed interruption semantics, and ambiguous external side effects.

DiagnosticsFailure evidenceAbort semanticsSecurityExecutor capacitySide effects

Learning objectives

  • Apply an evidence-first diagnostic sequence before editing retry or error-handling logic.
  • Diagnose and repair blind retry around a simulated external side effect.
  • Recognize catchError configurations that hide critical failure or convert interruptions unexpectedly.
  • Detect agent/executor waste caused by approval waits and unbounded polling.
  • Preserve abort/timeout semantics and external target evidence during repair.

1. Evidence-first diagnostic sequence

  1. Preserve job full name, build/queue IDs, URL, cause, exact source/Jenkinsfile SHA, result, and the first failing console excerpt.
  2. Confirm controller/core/Java and Pipeline plugin baseline.
  3. Confirm the failing stage/step and whether it ran inside a node/workspace.
  4. Inspect exception type: ordinary failure, timeout/interruption, agent loss, or user abort.
  5. Count retry attempts and identify exactly which body was repeated.
  6. Check whether catchError changed build/stage result or swallowed interruption semantics.
  7. Inspect approver identity and whether an agent was held during input.
  8. Inspect polling duration/recurrence and outer timeout.
  9. Verify external target state independently before retrying or rerunning.
  10. Apply the smallest correction and rerun only the safe scope.

2. Intentionally broken example: blind side-effect retry

retry(3) {
  sh './create-synthetic-release.sh'
}

Suppose attempt 1 creates synthetic-target/release-42.done and then exits nonzero because the response path is simulated to fail. Attempt 2 tries to create the same release again. The retry wrapper did exactly what it was told; the design is wrong.

# Synthetic first-attempt evidence
created=release-42
simulated_response=lost
exit=1

# Target-side evidence
synthetic-target/release-42.done exists
Do not “fix” this by increasing retry count. The original cause is uncertainty after a side effect. Preserve attempt 1 evidence, query the target, and make operation identity explicit.

3. Repair: preflight, deterministic operation ID, verify target

def releaseId = 'release-42'

stage('Preflight') {
  sh "./verify-target-accepts '${releaseId}'"
}

stage('Apply once') {
  sh "./apply-idempotent-synthetic-release '${releaseId}'"
}

stage('Verify target') {
  sh "./verify-synthetic-release '${releaseId}'"
}

The repair is not “never retry.” The read-only preflight and verification may have bounded retry if transient. The mutation uses a deterministic ID and target-state check so Jenkins can distinguish “already applied” from “not applied.”

4. Failure mode: catchError creates misleading green/continuation

catchError(buildResult: 'SUCCESS', stageResult: 'SUCCESS') {
  sh './critical-security-gate.sh'
}
sh './publish.sh'

This is structurally dangerous: a critical gate failure can be hidden and publication can continue. Build results cannot improve once worse, but explicitly keeping results successful when an error is caught can still make the stage look acceptable. Correct the policy: critical gate failure should abort, or at minimum remain failed and prevent the side-effect stage.

5. Failure mode: swallowing timeout or abort

catchError can catch Pipeline interruptions by default. If the intention is for timeout/manual abort to stop the flow, use catchInterruptions: false or explicit exception handling that rethrows interruption. Do not broadly catch every exception and continue.

catchError(
  buildResult: 'FAILURE',
  stageResult: 'FAILURE',
  catchInterruptions: false
) {
  timeout(time: 5, unit: 'MINUTES') {
    sh './bounded-required-check.sh'
  }
}

6. Failure mode: approval holds scarce executor

node('lab-linux') {
  input 'Approve?'
  sh './run-approved-step.sh'
}

While waiting, the node allocation can occupy an executor/workspace. Repair by persisting necessary evidence, leaving the node, waiting for input, then allocating an agent again for deterministic approved work.

7. Failure mode: unbounded polling

waitUntil {
  return sh(script: './is-ready.sh', returnStatus: true) == 0
}

Without a surrounding timeout, this can wait indefinitely. It may also call the target forever. Add a deadline and choose a recurrence appropriate for the service; for long waits, prefer event-driven integration.

8. Security-sensitive boundaries

  • Do not expose API tokens or approval secrets in URLs/logs.
  • Do not use broad administrator credentials just to make retry succeed.
  • Do not disable CSRF, authorization, TLS, or Script Security to bypass a failed step.
  • Do not run untrusted remediation code on the controller.
  • Do not let a human approval alone prove that the external target is still in the same state as when approval was requested.

9. Performance symptoms and evidence

Symptom Likely design issue Evidence
Many busy executors, little command activity Input/sleep/wait placed inside node Executor view + stage timestamps
Repeated target traffic Tight or duplicated polling/retry Request logs + attempt count
Long queue after transient outage Retry storm Queue depth + correlated build starts
Green build after required check failed Incorrect catch/downgrade policy Console failure + build/stage result
Duplicate release records Unsafe retry/rerun of mutation Target IDs + build attempts

10. Repair without hiding the cause

Never delete failed evidence just because the repaired run succeeds. Keep the original build URL, exact revision, first exception, attempt count, and target state. Commit the repair as a new source revision, run the smallest safe scope, and link the before/after evidence.

Next lesson

Checkpoint Lab — Pipeline Steps, Error Handling, retry, timeout, catchError, input, waitUntil, and Resilient Flow Control

Combine safe retry, bounded polling, approval, a guarded simulated side effect, and preserved failure evidence in one auditable checkpoint.

Knowledge check

Why is increasing retry count not a valid repair for an uncertain side effect?

What is wrong with holding input inside a long-lived node block?

How do you preserve timeout/manual-abort semantics through catchError?

What proves a false-green design?

What should remain after repair?

Official references and version notes

Version and compatibility note

Rechecked on 2026-09-16. Examples assume Jenkins 2.568.3 LTS, tested with Java 21 and 25, Pipeline aggregator 608.v67378e9d3db_1, Pipeline: Basic Steps 1098.v808b_fd7f8cf4, and Pipeline: Input Step 560.v56198a_642157. The mandatory path is local/disposable, uses synthetic state and fake identities, performs no production deployment/publication, and does not require commercial services. Always record the versions actually installed on your controller because plugins release independently from Jenkins core.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.