Pipeline Steps, Error Handling, retry, timeout, catchError, input, waitUntil, and Resilient Flow Control: Configuration, Design Choices, and Tradeoffs
Choose resilient control-flow patterns by comparing fail-fast behavior, bounded retries, result downgrades, timeout placement, executor occupancy, polling, approval, and rerun semantics against observable Jenkins and target state.
Learning objectives
- Choose retry versus fail-fast behavior using failure class and side-effect safety.
- Compare catchError with explicit try/catch/finally and preserve interruption semantics deliberately.
- Place timeout boundaries around queue allocation, agent work, polling, and approval intentionally.
- Decide when human input belongs outside an agent and when a reserved workspace is justified.
- Design rerun behavior around immutable inputs and target-side idempotency rather than workspace leftovers.
1. Retry versus fail fast
Retry is appropriate for transient failures when the body is safe to repeat and a later attempt can reasonably succeed. Fail fast when the input is invalid, a policy gate failed deterministically, a required artifact is missing, or the operation may already have caused an external side effect that Jenkins cannot verify.
| Failure | Default choice | Why |
|---|---|---|
| Read-only service query gets transient 503 | Bounded retry | Repeat is safe; transient recovery plausible |
| Invalid release version | Fail fast | Retry cannot repair bad input |
| Agent disappeared during a pure test step | Infrastructure-aware retry may fit | New agent can reproduce work |
| Package upload timed out after request was sent | Verify repository state first | Upload may already exist |
| Security gate found a real violation | Fail/block | Downgrading would misrepresent policy state |
2. catchError versus try/catch/finally
catchError is concise when the desired outcome is
“continue, but mark the stage/build non-green.” Plain Groovy
try/catch/finally is clearer when you need explicit
compensation, selective exception handling, or to rethrow after
collecting evidence.
try {
sh './required-check.sh'
} catch (err) {
writeFile file: 'evidence/required-check-failed.txt', text: "${err.class.name}\n"
throw err
} finally {
echo 'Evidence collection finished'
}
Do not catch broadly and forget to rethrow a critical failure. Also avoid swallowing user aborts or timeouts unless the behavior is intentionally documented.
3. Timeout placement changes what is bounded
| Boundary | What timeout includes | Tradeoff |
|---|---|---|
Around node |
Queue wait + allocated work | Caps total acquisition/execution but may fail during capacity shortage |
Inside node |
Only work after allocation | Queue can wait longer than command budget |
Around input |
Human approval window | Prevents indefinite pending approvals |
Around waitUntil |
Total polling duration | Provides a hard stop for otherwise unbounded polling |
| Around a whole deployment stage | Potentially multiple side effects | Simple but harder to diagnose/compensate safely |
4. Input inside versus outside an agent
Human approval usually belongs before agent allocation: the decision can take minutes or hours, while an executor is scarce capacity. Keep approval inside a node only when the very same allocated workspace/process must remain intact and you consciously accept the occupancy cost. Prefer persisting required evidence first, releasing the node, then reacquiring deterministic inputs after approval.
5. waitUntil versus event-driven integration
Polling is straightforward for a lab or a system with no callback
mechanism. It costs controller scheduling activity and can amplify
load if many runs poll aggressively. waitUntil backs
off its recurrence interval, but event-driven webhooks/callbacks or
an external orchestrator may scale better for long waits. Regardless
of strategy, define a deadline and retain the observed target state.
6. Rerun is a new attempt, not time travel
A rerun occurs in a changed world: credentials may rotate, targets may already contain the release, dependencies may disappear, and previous workspace files may be gone. Record immutable source/artifact IDs and inspect external state before repeating side effects. A replay of Pipeline logic does not automatically reproduce the external preconditions of the earlier run.
7. Worked design: maintenance simulation
Requirement: validate a target, wait for approval, execute one local synthetic mutation, then verify it. Choose:
| Decision | Choice | Observable justification |
|---|---|---|
| Target input | Allowlisted choice | Rejected value never reaches side-effect stage |
| Preflight health query | retry(3) |
Read-only and attempt count archived |
| Approval | input outside node + timeout |
No executor held; approver identity recorded |
| Mutation | No blind retry | Operation ID checked against target marker |
| Optional report | catchError to UNSTABLE |
Failure visible while later evidence still archives |
| Target readiness | Bounded waitUntil |
Polling has explicit deadline |
8. Cost, performance, and maintainability
Retries consume agent time and can amplify downstream load. Human
waits inside nodes waste executors. Tight polling increases
controller and service pressure. Overuse of
catchError creates ambiguous “successful” pipelines.
Optimize for truthful state first, then capacity: small bounded
retries, sparse polling, released agents during waits, and explicit
build/stage results.
Knowledge check
When should an invalid parameter be retried?
Normally never; invalid input is deterministic and should fail fast until the input changes.
Why can catchError be dangerous around a critical
policy gate?
It can let later stages run and may downgrade or obscure a failure that should have blocked the workflow.
What is the difference between timeout around and inside
node?
Around node includes queue wait; inside it starts
after allocation.
Why is rerunning a side effect different from rerunning a unit test?
External target state may already have changed, so the side effect needs target-state/idempotency verification before repetition.
What is the main scaling concern with many
waitUntil loops?
Repeated polling consumes controller/service activity; event-driven integration may be preferable for long waits.
Official references and version notes
- Jenkins LTS changelog — current LTS and tested Java configurations.
-
Pipeline: Basic Steps
— current
retry,timeout,catchError,waitUntil,sleep, and related step contracts. -
Pipeline: Input Step reference
—
input, submitter restrictions, identifiers, and captured approver identity. - Pipeline: Basic Steps plugin — current plugin baseline and dependencies.
- Pipeline: Input Step plugin — current plugin baseline and security history.
- Jenkins Pipeline handbook — durable Pipeline execution and Jenkinsfile concepts.
- Pipeline Best Practices — controller/agent boundaries and safe Pipeline design.
Rechecked on 2026-09-16. Examples assume Jenkins 2.568.3 LTS, tested with Java 21 and 25, Pipeline aggregator 608.v67378e9d3db_1, Pipeline: Basic Steps 1098.v808b_fd7f8cf4, and Pipeline: Input Step 560.v56198a_642157. The mandatory path is local/disposable, uses synthetic state and fake identities, performs no production deployment/publication, and does not require commercial services. Always record the versions actually installed on your controller because plugins release independently from Jenkins core.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.