Chapter 09Lesson 04~125 minutes

Pipeline Fundamentals, Jenkinsfile, Pipeline Engine, Durable Execution, Nodes, Workspaces, and Stages: Diagnostics, Failure Modes, Security, and Performance

Diagnose Pipeline incidents by preserving run identity and separating controller/CPS state, queue/node allocation, workspace/process state, plugin behavior, and external side effects before attempting the smallest repair.

DiagnosticsCPSAgent lossWorkspace lossRestartPerformance

Learning objectives

  • Apply an evidence-first diagnostic sequence to a stuck, failed, or interrupted Pipeline.
  • Diagnose queue/node/workspace failures separately from Pipeline Groovy or controller failures.
  • Recognize controller-heavy Groovy and inappropriate in-process computation as performance/reliability risks.
  • Explain why workspace persistence and external side-effect state cannot be inferred from Pipeline resume state.
  • Repair an intentionally broken cross-node Pipeline without hiding the original cause.

1. Preserve the first failure

Before retrying, restarting, deleting a workspace, moving the job to another agent, or editing the Jenkinsfile, record the build URL/number, cause, source/Jenkinsfile SHA, current stage/step, queue item if any, assigned node, workspace path, plugin/core baseline, console excerpt, artifacts/reports, and any independently observed external target state.

Why: a rerun creates a different run identity and may select a different node, workspace, source revision, credential state, or external-world precondition.

2. Evidence-first diagnostic ladder

  1. Definition/run: confirm exact job, build number, source/Jenkinsfile revision, cause.
  2. Controller/plugin: confirm Jenkins core, Java, Pipeline/Declarative/Groovy/Job/Durable Task versions and controller health.
  3. Queue/allocation: inspect label eligibility, offline causes, executor availability, queue reason.
  4. Agent/workspace: confirm Remoting connection, OS/toolchain, workspace existence/ownership.
  5. Pipeline/CPS/step: identify the exact stage and Pipeline step that stopped advancing.
  6. External process: inspect shell/tool process and durable-task log state without killing it first.
  7. Credentials/network/provider: inspect authorization/connection evidence with secret-safe logging.
  8. Artifacts/target: verify reports, artifact IDs/checksums, and external target state.
  9. Repair: change the smallest layer that evidence implicates, then rerun/resume the smallest safe scope.

3. Failure mode: “all Pipeline code runs on the agent”

Symptom: a Jenkinsfile performs large loops, parses huge files, or makes extensive network calls directly in Groovy and the controller becomes CPU/memory constrained although agents are mostly idle.

Cause: Pipeline Groovy/CPS interpretation and orchestration occur on the controller. A node allocation does not teleport arbitrary Groovy computation to the agent. External commands launched through steps run on agents.

Repair: move heavy computation into versioned scripts/tools executed with sh/bat/powershell on appropriate agents, returning compact results to Pipeline orchestration.

4. Failure mode: Pipeline state survives but workspace state does not

Symptom: after an agent is replaced, the Pipeline still exists but a later stage fails because it expects build/output.bin from an earlier workspace.

Cause: controller run/CPS persistence is separate from agent filesystem persistence.

Repair: re-create deterministic workspace state from source, or move the needed output through a deliberate mechanism such as stash/unstash, archived artifact, or external repository. Do not rely on “the file was there last time.”

5. Failure mode: stage appears stuck but is actually waiting in queue

A stage-level agent with label gpu-test may wait because no online node matches the label or all executors are busy. This is allocation state, not a Groovy deadlock.

Inspect the queue reason and node labels/executors. Preserve the queue item/build identity, then repair capacity/label configuration rather than editing unrelated Pipeline logic.

6. Failure mode: restart resumes orchestration but external side effect is ambiguous

Suppose a Pipeline called an external release API and the controller restarted after the HTTP request left Jenkins but before the step recorded a clean success response. On resume, Jenkins may know the step was in progress, but it cannot infer whether the provider accepted the request.

Correct diagnostic: query the external system using an idempotency key, release ID, artifact digest, or other immutable identifier before retrying. Blind retry is unsafe.

7. Intentionally broken example: cross-node workspace assumption

pipeline {
  agent none
  stages {
    stage('Build') {
      agent { label 'linux-builder' }
      steps { sh 'mkdir -p out && echo payload > out/app.txt' }
    }
    stage('Verify') {
      agent { label 'linux-test' }
      steps { sh 'cat out/app.txt' }
    }
  }
}

Observed failure: cat: out/app.txt: No such file or directory on the test node. The Pipeline is valid; the dataflow assumption is wrong.

Repair: preserve the failure, then explicitly transfer the run-scoped file:

stage('Build') {
  agent { label 'linux-builder' }
  steps {
    sh 'mkdir -p out && echo payload > out/app.txt'
    stash name: 'app', includes: 'out/app.txt'
  }
}
stage('Verify') {
  agent { label 'linux-test' }
  steps {
    unstash 'app'
    sh 'cat out/app.txt'
  }
}

The repair makes state transfer explicit instead of making both stages accidentally depend on one machine.

8. Performance: distinguish Pipeline persistence cost from executor cost

Pipeline persistence writes execution metadata to controller storage. Agent executors consume build capacity. A slow Pipeline can be limited by controller disk I/O, CPS/Groovy work, plugin behavior, queue capacity, workspace I/O, dependency downloads, or external services. Measure the layer before tuning it.

Changing the durability mode is a reliability tradeoff, not a universal performance fix. Adding executors can make controller/disk/network contention worse. Similarly, large stashes can shift heavy file transfer through Jenkins when an artifact repository would be more appropriate.

9. Security boundaries

  • Do not run untrusted builds on the controller.
  • Do not use Script Console as routine Pipeline debugging; it is administrative code execution.
  • Do not print complete environments or credentials to explain a step failure.
  • Do not disable Script Security, CSRF, TLS, or authorization to make a Pipeline work.
  • Do not “fix” agent connectivity with privileged host mounts or permissive SSH shortcuts.
  • Do not rerun a non-idempotent external operation until target state has been verified.

10. Minimal incident packet

Keep: job/build URL and number; cause; exact Jenkinsfile/source SHA; controller/core/Java/Pipeline plugin baseline; current stage/step; queue reason; node/label/executor/workspace; relevant console/controller/agent log excerpts; artifact/report identifiers; and external target verification. This is enough to let another engineer reconstruct the failure path without deleting the original evidence.

Next lesson

Checkpoint Lab — Pipeline Fundamentals

Apply the model end-to-end: run an SCM-backed Pipeline, predict restart behavior, restart the disposable controller during a safe wait, preserve evidence, and prove which state survived.

Knowledge check

A Pipeline is waiting for a label with no matching online node. Which layer should you inspect first?

Why can large Groovy loops hurt the controller?

What does the broken cross-node example prove?

After a restart, why is an ambiguous external API call dangerous to rerun?

Should you tune durability before measuring the actual bottleneck?

Official references and version notes

Version and compatibility note

Rechecked on 2026-09-16. Examples assume Jenkins 2.568.3 LTS, tested with Java 21 and 25. Current reference versions used for compatibility discussion are Pipeline aggregator 608.v67378e9d3db_1, Declarative Pipeline 2.2293.v6e7193cec599, Pipeline: Groovy 4380.v6eb_8378b_9647, Pipeline: Job 1600.v6f36ed83529d, Pipeline: Nodes and Processes 1479.v56e587f413a_7, and optional Pipeline: Stage View 2.41. Plugin releases move independently from Jenkins core; record the installed controller baseline before reproducing a lab. The mandatory exercises use only disposable local resources, no production credentials, and no controller-side untrusted builds.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.