Pipeline Steps, Error Handling, retry, timeout, catchError, input, waitUntil, and Resilient Flow Control: Concepts, Architecture, and Mental Model
Model Pipeline resilience explicitly: classify failure, bound waiting and retries, preserve interruption semantics, keep human approval auditable, and prevent recovery logic from repeating unsafe side effects.
Learning objectives
- Explain the difference between step failure, timeout, manual abort, infrastructure failure, and an external side effect that may already have happened.
- Describe how retry, timeout, catchError, input, and waitUntil transform Pipeline control flow and build/stage results.
- Choose a safe retry boundary by separating read-only/idempotent work from irreversible external actions.
- Explain why human approval and polling should usually avoid holding scarce agent executors.
- Identify the evidence required to distinguish original failure from later recovery attempts.
1. Resilience is not “keep going at any cost”
Chapters 09–11 established how Jenkins persists Pipeline state, allocates agents, and executes Declarative or Scripted control flow. The next problem is what happens when work does not complete normally. Networks fail, agents disappear, services return temporary errors, operators abort a run, approvals take hours, and an external deployment may succeed even when Jenkins loses the response.
A resilient Pipeline does not convert every problem into a retry or a green build. It first classifies what happened, preserves the first failure, and decides whether another attempt is safe. A read-only status query is usually safe to repeat. Creating a release, charging a customer, rotating a key, publishing a package, or deploying a version may not be safe unless the operation has an idempotency key or independently verified target state.
2. Mental model: outcome → wrapper decision → durable state
Each Pipeline step produces a control-flow outcome. A normal return
means success. An exception may mean command failure, timeout,
manual abort, controller/agent interruption, or another
plugin-specific failure. Wrappers such as retry,
timeout, and catchError do not erase the
underlying event; they decide what Pipeline should do next.
flowchart TD
A[Step attempt] --> B{Outcome}
B -->|success| C[Record success evidence]
B -->|ordinary failure| D{Retry safe?}
B -->|timeout / abort| E[Interrupted state]
D -->|yes| F[Bounded retry]
D -->|no| G[Preserve failure]
F --> A
E --> H{Policy preserves interruption?}
H -->|yes| G
H -->|explicitly handled| I[Record handled interruption]
G --> J[Build/stage result]
I --> J
C --> J
J --> K{External side effect?}
K -->|none / verified idempotent| L[Continue if policy allows]
K -->|uncertain| M[Verify target / compensate / stop]
3. The state you must track
| State | Question | Evidence |
|---|---|---|
| Step result / exception | What failed first? | Console excerpt, exception type, exit status |
| Build and stage result | Did handling change visualization or final status? | Run API/UI, stage result |
| Retry count | How many attempts occurred? | Attempt log + marker file/artifact |
| Timeout boundary | What exactly was bounded? | Jenkinsfile revision + timeout message |
| Input approver | Who authorized continuation? | Input record / captured submitter |
| Wait condition | What fact ended the wait? | Polled synthetic state + timestamps |
| Interrupted state | Was it a user abort or timeout? | FlowInterruptedException context / run result |
| External side effect | Did the target change before Jenkins lost certainty? | Target-side immutable ID or state query |
4. retry: repeat a bounded body, not an unsafe
transaction
The Basic Steps plugin retries its body when an exception occurs, up to the configured count. A user abort is not retried. Current Jenkins can also apply retry conditions for infrastructure-oriented cases such as agent loss.
retry(3) {
sh './synthetic-read-only-check.sh'
}
This is appropriate when the script only observes synthetic state. It is not automatically appropriate around a deployment or publication. The wrapper cannot know whether an external system performed the action before the connection failed.
5. timeout: define a maximum wait boundary
timeout(time: 2, unit: 'MINUTES') {
waitUntil(initialRecurrencePeriod: 1000) {
return fileExists('synthetic-ready.flag')
}
}
A timeout interrupts the nested body. The interruption is a first-class outcome, not just another false condition. Keep the boundary close to the wait you intend to bound so operators can tell what exceeded its budget.
6. catchError: continue while preserving a non-green
result
catchError(
buildResult: 'FAILURE',
stageResult: 'FAILURE',
catchInterruptions: false,
message: 'Synthetic verification failed'
) {
sh './verify-synthetic-output.sh'
}
catchError is useful when later evidence collection or
cleanup must run even after a failure. It can set build and stage
results separately. In this chapter we normally use
catchInterruptions: false when an abort or timeout
should remain an abort instead of being silently converted into an
ordinary handled error.
7. input: human decision is durable Pipeline state
The Input Step pauses the Pipeline and can restrict who may proceed. An administrator can still respond, and users with appropriate cancellation permission may abort. If you need accountability, capture the approving submitter.
def approver = input(
message: 'Proceed with the synthetic side effect?',
ok: 'Proceed',
submitter: 'lab-approver',
submitterParameter: 'APPROVER'
)
echo "approved_by=${approver}"
Do not hold an agent solely while waiting for a person unless the
workspace/executor truly must remain reserved. Prefer approval
outside node, then allocate the agent for the approved
work.
8. waitUntil: polling needs an outer stop condition
waitUntil repeats until its body returns
true. Its recurrence period backs off, but there is no
built-in maximum attempt count. An outer
timeout prevents an infinite wait.
timeout(time: 90, unit: 'SECONDS') {
waitUntil(initialRecurrencePeriod: 500, quiet: true) {
return fileExists('synthetic-ready.flag')
}
}
9. Safe recovery needs an idempotency story
For this course, irreversible behavior is simulated with a local marker keyed by an immutable operation ID. A repeated attempt first checks whether the marker already exists. Production systems should prefer native idempotency keys, immutable release IDs, deployment transaction IDs, or a target-state query rather than relying on a workspace file.
set -eu
mkdir -p synthetic-target
op_id="release-${BUILD_NUMBER}"
marker="synthetic-target/${op_id}.done"
if [ -f "$marker" ]; then
printf 'already_applied=%s\n' "$op_id"
else
printf 'applied_by_build=%s\n' "$BUILD_NUMBER" > "$marker"
fi
10. Inspect before changing control flow
- Record Jenkins core/Java and Pipeline plugin versions.
- Record job full name, build number/URL, cause, exact Jenkinsfile/source SHA, and current result.
- Capture the first failing console section before adding retry/catch logic.
- Record node/label/executor/workspace when an agent was involved.
- If an external side effect exists, verify target state independently before rerunning.
Knowledge check
Why is a retry wrapper unsafe around an arbitrary deployment?
Because Jenkins may not know whether the external deployment already succeeded before the failure was observed; a repeated attempt can duplicate or corrupt the side effect.
What does timeout do when its limit is
reached?
It interrupts the nested body by throwing a Pipeline interruption exception, which normally aborts the run unless explicitly handled.
When is catchError useful?
When later cleanup/evidence/notification should continue while the build or stage still records the failure or unstable state.
Why place a human input outside a node when
possible?
So the Pipeline can wait without reserving a scarce executor and workspace.
Why should waitUntil usually be wrapped by
timeout?
Because waitUntil has no built-in overall
attempt/time limit and can otherwise wait indefinitely.
Official references and version notes
- Jenkins LTS changelog — current LTS and tested Java configurations.
-
Pipeline: Basic Steps
— current
retry,timeout,catchError,waitUntil,sleep, and related step contracts. -
Pipeline: Input Step reference
—
input, submitter restrictions, identifiers, and captured approver identity. - Pipeline: Basic Steps plugin — current plugin baseline and dependencies.
- Pipeline: Input Step plugin — current plugin baseline and security history.
- Jenkins Pipeline handbook — durable Pipeline execution and Jenkinsfile concepts.
- Pipeline Best Practices — controller/agent boundaries and safe Pipeline design.
Rechecked on 2026-09-16. Examples assume Jenkins 2.568.3 LTS, tested with Java 21 and 25, Pipeline aggregator 608.v67378e9d3db_1, Pipeline: Basic Steps 1098.v808b_fd7f8cf4, and Pipeline: Input Step 560.v56198a_642157. The mandatory path is local/disposable, uses synthetic state and fake identities, performs no production deployment/publication, and does not require commercial services. Always record the versions actually installed on your controller because plugins release independently from Jenkins core.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.