Pipeline Steps, Error Handling, retry, timeout, catchError, input, waitUntil, and Resilient Flow Control: Diagnostics, Failure Modes, Security, and Performance
Diagnose resilience failures without erasing evidence: blind retries, false-green error handling, executor-holding approvals, infinite polling, swallowed interruption semantics, and ambiguous external side effects.
Learning objectives
- Apply an evidence-first diagnostic sequence before editing retry or error-handling logic.
- Diagnose and repair blind retry around a simulated external side effect.
- Recognize catchError configurations that hide critical failure or convert interruptions unexpectedly.
- Detect agent/executor waste caused by approval waits and unbounded polling.
- Preserve abort/timeout semantics and external target evidence during repair.
1. Evidence-first diagnostic sequence
- Preserve job full name, build/queue IDs, URL, cause, exact source/Jenkinsfile SHA, result, and the first failing console excerpt.
- Confirm controller/core/Java and Pipeline plugin baseline.
- Confirm the failing stage/step and whether it ran inside a node/workspace.
- Inspect exception type: ordinary failure, timeout/interruption, agent loss, or user abort.
- Count retry attempts and identify exactly which body was repeated.
-
Check whether
catchErrorchanged build/stage result or swallowed interruption semantics. - Inspect approver identity and whether an agent was held during input.
- Inspect polling duration/recurrence and outer timeout.
- Verify external target state independently before retrying or rerunning.
- Apply the smallest correction and rerun only the safe scope.
2. Intentionally broken example: blind side-effect retry
retry(3) {
sh './create-synthetic-release.sh'
}
Suppose attempt 1 creates
synthetic-target/release-42.done and then exits nonzero
because the response path is simulated to fail. Attempt 2 tries to
create the same release again. The retry wrapper did exactly what it
was told; the design is wrong.
# Synthetic first-attempt evidence
created=release-42
simulated_response=lost
exit=1
# Target-side evidence
synthetic-target/release-42.done exists
3. Repair: preflight, deterministic operation ID, verify target
def releaseId = 'release-42'
stage('Preflight') {
sh "./verify-target-accepts '${releaseId}'"
}
stage('Apply once') {
sh "./apply-idempotent-synthetic-release '${releaseId}'"
}
stage('Verify target') {
sh "./verify-synthetic-release '${releaseId}'"
}
The repair is not “never retry.” The read-only preflight and verification may have bounded retry if transient. The mutation uses a deterministic ID and target-state check so Jenkins can distinguish “already applied” from “not applied.”
4. Failure mode: catchError creates misleading
green/continuation
catchError(buildResult: 'SUCCESS', stageResult: 'SUCCESS') {
sh './critical-security-gate.sh'
}
sh './publish.sh'
This is structurally dangerous: a critical gate failure can be hidden and publication can continue. Build results cannot improve once worse, but explicitly keeping results successful when an error is caught can still make the stage look acceptable. Correct the policy: critical gate failure should abort, or at minimum remain failed and prevent the side-effect stage.
5. Failure mode: swallowing timeout or abort
catchError can catch Pipeline interruptions by default.
If the intention is for timeout/manual abort to stop the flow, use
catchInterruptions: false or explicit exception
handling that rethrows interruption. Do not broadly catch every
exception and continue.
catchError(
buildResult: 'FAILURE',
stageResult: 'FAILURE',
catchInterruptions: false
) {
timeout(time: 5, unit: 'MINUTES') {
sh './bounded-required-check.sh'
}
}
6. Failure mode: approval holds scarce executor
node('lab-linux') {
input 'Approve?'
sh './run-approved-step.sh'
}
While waiting, the node allocation can occupy an executor/workspace. Repair by persisting necessary evidence, leaving the node, waiting for input, then allocating an agent again for deterministic approved work.
7. Failure mode: unbounded polling
waitUntil {
return sh(script: './is-ready.sh', returnStatus: true) == 0
}
Without a surrounding timeout, this can wait indefinitely. It may also call the target forever. Add a deadline and choose a recurrence appropriate for the service; for long waits, prefer event-driven integration.
8. Security-sensitive boundaries
- Do not expose API tokens or approval secrets in URLs/logs.
- Do not use broad administrator credentials just to make retry succeed.
- Do not disable CSRF, authorization, TLS, or Script Security to bypass a failed step.
- Do not run untrusted remediation code on the controller.
- Do not let a human approval alone prove that the external target is still in the same state as when approval was requested.
9. Performance symptoms and evidence
| Symptom | Likely design issue | Evidence |
|---|---|---|
| Many busy executors, little command activity | Input/sleep/wait placed inside node | Executor view + stage timestamps |
| Repeated target traffic | Tight or duplicated polling/retry | Request logs + attempt count |
| Long queue after transient outage | Retry storm | Queue depth + correlated build starts |
| Green build after required check failed | Incorrect catch/downgrade policy | Console failure + build/stage result |
| Duplicate release records | Unsafe retry/rerun of mutation | Target IDs + build attempts |
10. Repair without hiding the cause
Never delete failed evidence just because the repaired run succeeds. Keep the original build URL, exact revision, first exception, attempt count, and target state. Commit the repair as a new source revision, run the smallest safe scope, and link the before/after evidence.
Knowledge check
Why is increasing retry count not a valid repair for an uncertain side effect?
Because it increases the chance of duplication without resolving whether the first attempt already changed the target.
What is wrong with holding input inside a
long-lived node block?
It can reserve a scarce executor/workspace while no agent work is happening.
How do you preserve timeout/manual-abort semantics through
catchError?
Use a policy such as catchInterruptions: false when
interruptions should propagate, or explicitly rethrow
interruption exceptions.
What proves a false-green design?
Console evidence of a required check failure combined with a successful-looking stage/build that continued into protected side effects.
What should remain after repair?
The original failed run and first-failure/target evidence, plus the new revision/run that demonstrates the correction.
Official references and version notes
- Jenkins LTS changelog — current LTS and tested Java configurations.
-
Pipeline: Basic Steps
— current
retry,timeout,catchError,waitUntil,sleep, and related step contracts. -
Pipeline: Input Step reference
—
input, submitter restrictions, identifiers, and captured approver identity. - Pipeline: Basic Steps plugin — current plugin baseline and dependencies.
- Pipeline: Input Step plugin — current plugin baseline and security history.
- Jenkins Pipeline handbook — durable Pipeline execution and Jenkinsfile concepts.
- Pipeline Best Practices — controller/agent boundaries and safe Pipeline design.
Rechecked on 2026-09-16. Examples assume Jenkins 2.568.3 LTS, tested with Java 21 and 25, Pipeline aggregator 608.v67378e9d3db_1, Pipeline: Basic Steps 1098.v808b_fd7f8cf4, and Pipeline: Input Step 560.v56198a_642157. The mandatory path is local/disposable, uses synthetic state and fake identities, performs no production deployment/publication, and does not require commercial services. Always record the versions actually installed on your controller because plugins release independently from Jenkins core.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.