Chapter 17Lesson 04~130 minutes

Nodes, Agents, Labels, Executors, Workspaces, Offline Causes, and Agent Capacity Design: Diagnostics, Failure Modes, Security, and Performance

Diagnose stuck or slow work evidence-first: distinguish label mismatch, offline state, executor saturation, host contention, workspace contamination, and trust-policy errors before changing capacity.

DiagnosticsQueue reasonsContentionWorkspace stateTrust boundariesRepair

Learning objectives

  • Separate label/eligibility failures from executor saturation and agent connection failures.
  • Preserve queue/build/node evidence before changing labels, executor counts or offline state.
  • Diagnose CPU/memory contention caused by overcommitted executors.
  • Recognize workspace residue and cross-trust scheduling as security failures, not only performance defects.
  • Apply the least destructive correction and rerun only the smallest safe scope.

1. Evidence-first diagnostic sequence

Evidence-first diagnostic sequence
flowchart TD
A[Preserve queue/build IDs + first reason] --> B[Confirm Jenkins/Java/plugin baseline]
B --> C[Confirm job/Jenkinsfile/source/cause]
C --> D[Evaluate label expression + eligible nodes]
D --> E[Check online/offline cause + executor occupancy]
E --> F[Inspect agent/Remoting/workspace/tools]
F --> G[Inspect host CPU/RAM/disk/network]
G --> H[Inspect credentials/artifacts/external systems]
H --> I[Apply smallest correction]
I --> J[Rerun smallest safe scope + compare evidence]

Do not begin with restart, node deletion or executor inflation. Queue reasons are often enough to identify the layer.

2. Failure: labels never match

Broken example: the Jenkinsfile requests linux && jdk25, but the only node has labels linux jdk21. The build remains queued even though the node is online and idle.

pipeline {
  agent { label 'linux && jdk25' }
  stages { stage('Evidence') { steps { sh 'java -version' } } }
}

Preserve the queue ID and reason, then compare the required label expression with assignedLabels. If the workload really requires JDK 25, provisioning/installing that capability is the fix; lying by adding jdk25 to a JDK 21 node only hides the defect. If the requirement was mistaken, repair the Jenkinsfile/source revision.

3. Failure: too many executors cause host contention

An eight-core host does not automatically need eight executors. Four memory-heavy compilers may exhaust RAM and cause swap or OOM failures. Preserve host load, memory, disk I/O, executor occupancy, per-build duration and failures. Reduce concurrency before rerunning the failed test; do not use blind retries to hide resource starvation.

Evidence Likely conclusion Least destructive response
Queue long, host mostly idle Potentially insufficient concurrency/workers Test one additional executor/worker
Queue long, host saturated, builds slower Worker resource bottleneck Add hosts or optimize workload; do not add local executors
Idle executors but job waiting Eligibility/usage/label restriction Repair scheduling configuration
Node offline with explicit monitor cause Node health threshold/connection issue Fix underlying host condition first

4. Failure: shared workspace leaks stale state

A test unexpectedly passes because an old generated file remains. Preserve WORKSPACE, build numbers, file timestamps/hashes and checkout source SHA. Repair by making the build create/validate its inputs and by applying an appropriate cleanup/isolation policy.

Do not delete every workspace on every incident before collecting evidence. If the workspace might contain a real secret, restrict access and handle it as an incident; ordinary cleanup does not prove the secret was not copied elsewhere.

5. Failure: privileged and untrusted jobs share an agent

A persistent node labeled both pr-untrusted and release-trusted is a dangerous design if both execute as the same OS identity or share filesystem/process visibility. A pull request build could plant files, inspect residual data or attack later privileged work.

Repair the topology: separate nodes/pools and OS/runtime trust boundaries, restrict protected-source jobs, and verify protected credentials are only exposed after the trust transition. Merely removing one label while leaving the same hostile process/host access may be insufficient.

6. Failure: misreading offline cause as queue starvation

If a node monitor takes an agent offline due to disk/temp/swap/clock/response thresholds, increasing executors is the opposite of a fix. Capture the node monitor/offline cause and host condition. Correct disk pressure, clock sync, agent connectivity or the relevant health problem; only then bring the node back and re-evaluate queue behavior.

7. Agent/Remoting and launcher evidence

When the node is disconnected rather than intentionally offline, inspect the agent connection log, Java version, network path and launcher-specific evidence. For SSH-launched agents, preserve host-key verification and least-privilege credentials. Do not disable host-key checks to make an agent connect.

For inbound agents, protect the agent secret and avoid copying a failing secret into tickets/logs. If compromise is suspected, recreate/rotate the node secret after evidence preservation.

8. Repair matrix

Failure Do not Prefer Verification
Label mismatch Add fake labels Repair requirement or real capability Queue becomes eligible on intended nodes
Saturation Unbounded executor increase Measure, add bounded capacity Queue falls without host/build regression
Workspace residue Delete evidence first Preserve hashes/times, then clean/isolate Fresh build no longer depends on residue
Trust mixing Rely on masking/label names Separate execution pools/identities Untrusted jobs are ineligible for privileged pool
Offline monitor Force online repeatedly Fix underlying host threshold issue Healthy monitor + stable connection

9. Intentionally broken diagnostic exercise

Configure agent-ci-1 with label linux-ci, one executor, and temporarily offline cause disk maintenance drill. Queue capacity-hold. Before fixing, capture:

  • queue item ID and why text;
  • node online/temporarily-offline state and cause;
  • executor count/busy count;
  • job/Jenkinsfile/source revision and requested label;
  • agent/host free disk evidence.

Repair only by bringing the disposable node online after confirming the cause is a simulation, then verify the same queued item can progress. Do not delete/recreate the job or change its label to hide the original cause.

Next lesson

Checkpoint Lab

Design two execution pools, produce controlled pressure, inject one routing/capacity defect, repair it, and justify utilization and trust choices with preserved evidence.

Knowledge check

Answer before revealing the explanation.

1. A queue item says no nodes match linux && arm64. What layer should you inspect first?

2. A node is online with four busy executors and jobs wait. Is the agent offline?

3. Why is increasing executors a poor first fix for slow builds on a memory-starved host?

4. Why can workspace reuse be a security as well as correctness problem?

5. What evidence must remain after repairing a temporarily-offline or label problem?

Official references and version notes

Assumption timestamp: 2026-09-17. Recheck Jenkins LTS/Java support, agent/remoting compatibility, launcher-plugin advisories, and current node/monitor behavior before reproducing the lab later.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.