Chapter 12Lesson 01~120 minutes

Pipeline Steps, Error Handling, retry, timeout, catchError, input, waitUntil, and Resilient Flow Control: Concepts, Architecture, and Mental Model

Model Pipeline resilience explicitly: classify failure, bound waiting and retries, preserve interruption semantics, keep human approval auditable, and prevent recovery logic from repeating unsafe side effects.

Pipeline stepsretrytimeoutcatchErrorinputwaitUntilFailure semantics

Learning objectives

  • Explain the difference between step failure, timeout, manual abort, infrastructure failure, and an external side effect that may already have happened.
  • Describe how retry, timeout, catchError, input, and waitUntil transform Pipeline control flow and build/stage results.
  • Choose a safe retry boundary by separating read-only/idempotent work from irreversible external actions.
  • Explain why human approval and polling should usually avoid holding scarce agent executors.
  • Identify the evidence required to distinguish original failure from later recovery attempts.

1. Resilience is not “keep going at any cost”

Chapters 09–11 established how Jenkins persists Pipeline state, allocates agents, and executes Declarative or Scripted control flow. The next problem is what happens when work does not complete normally. Networks fail, agents disappear, services return temporary errors, operators abort a run, approvals take hours, and an external deployment may succeed even when Jenkins loses the response.

A resilient Pipeline does not convert every problem into a retry or a green build. It first classifies what happened, preserves the first failure, and decides whether another attempt is safe. A read-only status query is usually safe to repeat. Creating a release, charging a customer, rotating a key, publishing a package, or deploying a version may not be safe unless the operation has an idempotency key or independently verified target state.

Resilience rule: retry uncertainty only when repeating the enclosed operation is safe. If an external side effect may already have happened, verify the external system before deciding whether to compensate, continue, or stop.

2. Mental model: outcome → wrapper decision → durable state

Each Pipeline step produces a control-flow outcome. A normal return means success. An exception may mean command failure, timeout, manual abort, controller/agent interruption, or another plugin-specific failure. Wrappers such as retry, timeout, and catchError do not erase the underlying event; they decide what Pipeline should do next.

Resilient control-flow decision
flowchart TD
 A[Step attempt] --> B{Outcome}
 B -->|success| C[Record success evidence]
 B -->|ordinary failure| D{Retry safe?}
 B -->|timeout / abort| E[Interrupted state]
 D -->|yes| F[Bounded retry]
 D -->|no| G[Preserve failure]
 F --> A
 E --> H{Policy preserves interruption?}
 H -->|yes| G
 H -->|explicitly handled| I[Record handled interruption]
 G --> J[Build/stage result]
 I --> J
 C --> J
 J --> K{External side effect?}
 K -->|none / verified idempotent| L[Continue if policy allows]
 K -->|uncertain| M[Verify target / compensate / stop]

3. The state you must track

State Question Evidence
Step result / exception What failed first? Console excerpt, exception type, exit status
Build and stage result Did handling change visualization or final status? Run API/UI, stage result
Retry count How many attempts occurred? Attempt log + marker file/artifact
Timeout boundary What exactly was bounded? Jenkinsfile revision + timeout message
Input approver Who authorized continuation? Input record / captured submitter
Wait condition What fact ended the wait? Polled synthetic state + timestamps
Interrupted state Was it a user abort or timeout? FlowInterruptedException context / run result
External side effect Did the target change before Jenkins lost certainty? Target-side immutable ID or state query

4. retry: repeat a bounded body, not an unsafe transaction

The Basic Steps plugin retries its body when an exception occurs, up to the configured count. A user abort is not retried. Current Jenkins can also apply retry conditions for infrastructure-oriented cases such as agent loss.

retry(3) {
  sh './synthetic-read-only-check.sh'
}

This is appropriate when the script only observes synthetic state. It is not automatically appropriate around a deployment or publication. The wrapper cannot know whether an external system performed the action before the connection failed.

5. timeout: define a maximum wait boundary

timeout(time: 2, unit: 'MINUTES') {
  waitUntil(initialRecurrencePeriod: 1000) {
    return fileExists('synthetic-ready.flag')
  }
}

A timeout interrupts the nested body. The interruption is a first-class outcome, not just another false condition. Keep the boundary close to the wait you intend to bound so operators can tell what exceeded its budget.

6. catchError: continue while preserving a non-green result

catchError(
  buildResult: 'FAILURE',
  stageResult: 'FAILURE',
  catchInterruptions: false,
  message: 'Synthetic verification failed'
) {
  sh './verify-synthetic-output.sh'
}

catchError is useful when later evidence collection or cleanup must run even after a failure. It can set build and stage results separately. In this chapter we normally use catchInterruptions: false when an abort or timeout should remain an abort instead of being silently converted into an ordinary handled error.

7. input: human decision is durable Pipeline state

The Input Step pauses the Pipeline and can restrict who may proceed. An administrator can still respond, and users with appropriate cancellation permission may abort. If you need accountability, capture the approving submitter.

def approver = input(
  message: 'Proceed with the synthetic side effect?',
  ok: 'Proceed',
  submitter: 'lab-approver',
  submitterParameter: 'APPROVER'
)
echo "approved_by=${approver}"

Do not hold an agent solely while waiting for a person unless the workspace/executor truly must remain reserved. Prefer approval outside node, then allocate the agent for the approved work.

8. waitUntil: polling needs an outer stop condition

waitUntil repeats until its body returns true. Its recurrence period backs off, but there is no built-in maximum attempt count. An outer timeout prevents an infinite wait.

timeout(time: 90, unit: 'SECONDS') {
  waitUntil(initialRecurrencePeriod: 500, quiet: true) {
    return fileExists('synthetic-ready.flag')
  }
}

9. Safe recovery needs an idempotency story

For this course, irreversible behavior is simulated with a local marker keyed by an immutable operation ID. A repeated attempt first checks whether the marker already exists. Production systems should prefer native idempotency keys, immutable release IDs, deployment transaction IDs, or a target-state query rather than relying on a workspace file.

set -eu
mkdir -p synthetic-target
op_id="release-${BUILD_NUMBER}"
marker="synthetic-target/${op_id}.done"
if [ -f "$marker" ]; then
  printf 'already_applied=%s\n' "$op_id"
else
  printf 'applied_by_build=%s\n' "$BUILD_NUMBER" > "$marker"
fi

10. Inspect before changing control flow

  • Record Jenkins core/Java and Pipeline plugin versions.
  • Record job full name, build number/URL, cause, exact Jenkinsfile/source SHA, and current result.
  • Capture the first failing console section before adding retry/catch logic.
  • Record node/label/executor/workspace when an agent was involved.
  • If an external side effect exists, verify target state independently before rerunning.
Next lesson

Guided Hands-On Workflow and Core Operations

Build a disposable resilient workflow with safe retry, bounded polling, approval, explicit result changes, and a guarded simulated side effect.

Knowledge check

Why is a retry wrapper unsafe around an arbitrary deployment?

What does timeout do when its limit is reached?

When is catchError useful?

Why place a human input outside a node when possible?

Why should waitUntil usually be wrapped by timeout?

Official references and version notes

Version and compatibility note

Rechecked on 2026-09-16. Examples assume Jenkins 2.568.3 LTS, tested with Java 21 and 25, Pipeline aggregator 608.v67378e9d3db_1, Pipeline: Basic Steps 1098.v808b_fd7f8cf4, and Pipeline: Input Step 560.v56198a_642157. The mandatory path is local/disposable, uses synthetic state and fake identities, performs no production deployment/publication, and does not require commercial services. Always record the versions actually installed on your controller because plugins release independently from Jenkins core.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.