Chapter 01Lesson 04~110 minutes

Jenkins and CI/CD Foundations, Controller-Agent Architecture, Jobs, Builds, and Automation Boundaries: Diagnostics, Failure Modes, Security, and Performance

Production Jenkins failures are easier to repair when you preserve the first-failure evidence and diagnose from scheduling outward. This lesson builds a failure taxonomy across controller/JVM/plugin, queue/capacity, agent/Remoting/workspace, Pipeline/shell, credentials/network, artifact/report, and external-system layers, then uses deliberately broken Chapter 01 scenarios to practice the least-destructive fix.

DiagnosticsQueueAgent failuresSecurityPerformance

Learning objectives

  • Apply an evidence-first diagnostic sequence before restart, rerun, workspace deletion, plugin upgrade, or other state-changing “fix.”
  • Distinguish queue/label/capacity failures from agent connectivity, Pipeline/Groovy, shell/tool, credential, artifact, and external-service failures.
  • Diagnose why builds on the built-in node are a security/performance problem even when they are green.
  • Explain why a green build, a mutable branch, or a surviving workspace are insufficient evidence for delivery correctness.
  • Capture a minimal incident packet with job/build/queue/source/agent identities, logs, versions, and external verification.

1. Preserve first-failure evidence before you “fix” Jenkins

Retries, restarts, plugin updates, reconnecting agents, deleting workspaces, or re-running a job can all change the evidence. Sometimes they also repeat external side effects. The first response to a Jenkins failure should therefore be observation: record exact job/build/queue identity, source revision, cause, controller/core/Java/plugin baseline, node/agent, timestamps, and the first causal error.

Do not begin an incident by clicking “Build Now” again. A rerun can succeed because an external dependency recovered, fail differently because the branch moved, or duplicate a publication/deployment side effect. Preserve the original record first.

2. Evidence-first diagnostic ladder

  1. Trigger/source/item: did the expected job exist, and what cause/source/Jenkinsfile revision was selected?
  2. Queue: was a queue item created, and what reason explains waiting?
  3. Node/label/executor: is there an eligible online node with free capacity?
  4. Agent/Remoting/workspace/toolchain: did the agent connect, create a workspace, and provide the expected Java/tools?
  5. Pipeline/Groovy/step: did Pipeline compile/evaluate, and which step first failed?
  6. Shell/tool/network: what process exit status, stderr, DNS/TLS/HTTP response, or tool log is causal?
  7. Credentials/authorization: was the correct identity available and authorized without leaking secret values?
  8. Reports/artifacts: were expected files produced and successfully ingested/archived?
  9. External system: did the registry/cloud/cluster/scanner/notification service reach the intended final state?

Stop at the first layer that contradicts the expected state. Do not rewrite shell commands to solve a queue eligibility problem or restart the controller to solve an application test failure.

3. Failure taxonomy: symptom, evidence, and owning layer

Symptom First evidence to inspect Likely owning layer
“Waiting for next available executor” Queue reason, eligible nodes, executor occupancy. Queue/capacity.
“There are no nodes with the label …” Job label expression, node labels, offline causes. Scheduling/configuration.
Agent repeatedly disconnects Controller/agent logs, Remoting/Java versions, network path. Agent/Remoting/network.
Pipeline fails before a shell starts Pipeline syntax/evaluation error, plugin stack trace, Jenkinsfile revision. Pipeline/Groovy/plugin.
Shell exits non-zero Exact command, exit code, stderr, tool version, workspace input. Build tool/script.
Archive step says no artifacts found Expected path, working directory, producer step, glob. Artifact/dataflow.
HTTP 403 from registry/cloud Credential ID/scope and provider authorization response. Identity/authorization.
Jenkins build green but service unhealthy Provider deployment record and health checks. External target/runtime.

4. Broken example A: the build is green—but it ran on the controller

Suppose a learner leaves the built-in node at one executor and writes agent any. If the separate agent is offline, Jenkins may choose the built-in node. The Pipeline can be completely green, yet the architecture violates the intended trust boundary.

pipeline {
  agent any
  stages {
    stage('Where am I?') {
      steps {
        sh 'printf "node=%s workspace=%s\\n" "$NODE_NAME" "$WORKSPACE"'
      }
    }
  }
}

Evidence: console output and build metadata show NODE_NAME=built-in (or the controller's built-in node name). The correct fix is not to add retries. Restore the intended scheduling model: create/repair the agent, set built-in executors to zero, and use an explicit label such as lab. Then create a new build and preserve both build records so the change is auditable.

5. Broken example B: the job never reaches the shell

Change the Chapter 01 Pipeline to agent { label 'lab-missing' }. Jenkins can accept the item configuration and create a build, but the run waits because no node satisfies the label expression.

Evidence: the queue/build page reports that no matching node is available. There is no shell exit code because no executor has been allocated. The least-destructive repair is to correct the label or intentionally add a node with that label—not to restart Jenkins, reinstall Pipeline plugins, or delete the workspace.

6. Broken example C: green step, missing durable evidence

Consider a Pipeline that creates out/report.txt and exits successfully but never calls archiveArtifacts. The file exists only in the workspace. If the agent is destroyed or the workspace is cleaned, the evidence disappears while the historical build can remain green.

pipeline {
  agent { label 'lab' }
  stages {
    stage('Produce') {
      steps {
        sh 'mkdir -p out && printf "ok\\n" > out/report.txt'
      }
    }
  }
}

The repair is not to rely on longer workspace retention. Decide whether the file is build evidence or a publishable package. If it is build evidence, archive it with producer build/source identity. If it is a release package, publish it to an appropriate external repository and retain the immutable package/digest identity.

7. Broken assumption D: “main” tells you what the old build executed

A branch is a pointer that can move after a build. If a failed build only records “branch=main” and someone pushes again, re-running or debugging from the current branch tip can reproduce different code. Before any retry, preserve the original build's SCM revision from the build/workspace metadata.

When reproducing a failure, check out the recorded SHA, not “whatever main is now.” If the Jenkinsfile itself came from SCM, preserve its revision too; a changed Jenkinsfile can alter scheduling, credentials, tools, and side effects independently of application source.

8. Security-sensitive shortcuts that turn incidents into compromises

Several tempting “fixes” should be treated as red flags: disabling CSRF or authorization, approving arbitrary Groovy/Script Console code, making an agent privileged, mounting the Docker socket, turning off SSH host-key checking, copying admin credentials into the job, or dumping environment variables to find a secret. These changes enlarge the trust boundary and can destroy the distinction between build code and controller administration.

Use the narrowest reversible change that addresses the causal layer. If the problem is a credential scope, fix only that credential binding/permission. If the problem is an agent image missing a tool, rebuild/version the agent image or configure the tool explicitly. If the problem is DNS/TLS, capture the network evidence and repair trust/network configuration rather than bypassing verification.

9. Performance: queue time and controller health are evidence, not guesses

A slow Jenkins system can be caused by queue capacity, agent provisioning, controller CPU/heap/GC, disk I/O, oversized logs/artifacts, plugin behavior, SCM/API latency, or build-tool work. Increasing executors can make things worse if agents are already resource constrained. Running builds on the controller can also turn build CPU/I/O pressure into UI/scheduling instability.

For Chapter 01, measure only simple signals: queue wait, executor occupancy, controller/agent CPU/memory from Docker, and build duration. Change one variable at a time. Later chapters introduce metrics, thread dumps, support bundles, and JVM tuning.

docker stats --no-stream \
  jenkins-ch01-controller jenkins-ch01-agent

10. Minimal incident packet before escalation

  • Controller: Jenkins core/LTS, Java version, relevant plugin versions, timestamp/timezone.
  • Item: full job path, configuration/Jenkinsfile source and revision.
  • Build: number, URL, cause, start time, result, first causal log excerpt.
  • Queue: queue ID/reason if waiting.
  • Source: exact commit SHA and any Shared Library revision.
  • Agent: node name, labels, executor, Java/Remoting/image/tool versions, workspace.
  • Outputs: expected report/artifact path, archived/published identity, digest.
  • External: endpoint/provider resource ID and health/status evidence, with secrets redacted.

Collect only what is needed. Support bundles, heap dumps, environment dumps, and configuration exports can contain sensitive operational data and should be reviewed before sharing.

11. Summary: fix the first broken boundary, not the loudest symptom

Jenkins troubleshooting is fastest when you reason in order: source/item → queue → eligibility/capacity → agent/workspace/toolchain → Pipeline/step → credential/network → artifact/report → external target. Preserve the original build and its immutable identities before rerunning. That discipline prevents both accidental evidence loss and duplicate side effects.

Next lesson

Checkpoint Lab

Integrate the chapter into one evidence-driven lab with two build causes, immutable source identity, retained artifacts, failure injection, and exact cleanup.

Knowledge check

A job is queued with “no nodes with label lab.” Should you inspect the shell exit code?

Why is a successful build on the built-in node still a problem?

A report file existed during the build but is gone tomorrow. What likely boundary was misunderstood?

Why can a blind retry be dangerous for a failed deploy job?

What is the first principle when a production Jenkins incident begins?

Official references and version notes

Version and compatibility note

Rechecked against primary Jenkins sources on 2026-09-13. The executable Chapter 01 baseline is 2.568.3 LTS on Java 21 (Java 25 is also supported by this LTS line), controller image jenkins/jenkins:2.568.3-lts-jdk21, and inbound-agent image jenkins/inbound-agent:3391.va_37fa_a_305d6d-2-jdk21. The current LTS line can change after this date, so future generation and later lab reuse must re-check the LTS changelog, Java support matrix, Docker tags, plugin minimum core versions, and security advisories.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.