Nodes, Agents, Labels, Executors, Workspaces, Offline Causes, and Agent Capacity Design: Diagnostics, Failure Modes, Security, and Performance
Diagnose stuck or slow work evidence-first: distinguish label mismatch, offline state, executor saturation, host contention, workspace contamination, and trust-policy errors before changing capacity.
Learning objectives
- Separate label/eligibility failures from executor saturation and agent connection failures.
- Preserve queue/build/node evidence before changing labels, executor counts or offline state.
- Diagnose CPU/memory contention caused by overcommitted executors.
- Recognize workspace residue and cross-trust scheduling as security failures, not only performance defects.
- Apply the least destructive correction and rerun only the smallest safe scope.
1. Evidence-first diagnostic sequence
flowchart TD A[Preserve queue/build IDs + first reason] --> B[Confirm Jenkins/Java/plugin baseline] B --> C[Confirm job/Jenkinsfile/source/cause] C --> D[Evaluate label expression + eligible nodes] D --> E[Check online/offline cause + executor occupancy] E --> F[Inspect agent/Remoting/workspace/tools] F --> G[Inspect host CPU/RAM/disk/network] G --> H[Inspect credentials/artifacts/external systems] H --> I[Apply smallest correction] I --> J[Rerun smallest safe scope + compare evidence]
Do not begin with restart, node deletion or executor inflation. Queue reasons are often enough to identify the layer.
2. Failure: labels never match
Broken example: the Jenkinsfile requests
linux && jdk25, but the only node has labels
linux jdk21. The build remains queued even though the
node is online and idle.
pipeline {
agent { label 'linux && jdk25' }
stages { stage('Evidence') { steps { sh 'java -version' } } }
}
Preserve the queue ID and reason, then compare the required label
expression with assignedLabels. If the workload really
requires JDK 25, provisioning/installing that capability is the fix;
lying by adding jdk25 to a JDK 21 node only hides the
defect. If the requirement was mistaken, repair the
Jenkinsfile/source revision.
3. Failure: too many executors cause host contention
An eight-core host does not automatically need eight executors. Four memory-heavy compilers may exhaust RAM and cause swap or OOM failures. Preserve host load, memory, disk I/O, executor occupancy, per-build duration and failures. Reduce concurrency before rerunning the failed test; do not use blind retries to hide resource starvation.
| Evidence | Likely conclusion | Least destructive response |
|---|---|---|
| Queue long, host mostly idle | Potentially insufficient concurrency/workers | Test one additional executor/worker |
| Queue long, host saturated, builds slower | Worker resource bottleneck | Add hosts or optimize workload; do not add local executors |
| Idle executors but job waiting | Eligibility/usage/label restriction | Repair scheduling configuration |
| Node offline with explicit monitor cause | Node health threshold/connection issue | Fix underlying host condition first |
4. Failure: shared workspace leaks stale state
A test unexpectedly passes because an old generated file remains.
Preserve WORKSPACE, build numbers, file
timestamps/hashes and checkout source SHA. Repair by making the
build create/validate its inputs and by applying an appropriate
cleanup/isolation policy.
Do not delete every workspace on every incident before collecting evidence. If the workspace might contain a real secret, restrict access and handle it as an incident; ordinary cleanup does not prove the secret was not copied elsewhere.
5. Failure: privileged and untrusted jobs share an agent
A persistent node labeled both pr-untrusted and
release-trusted is a dangerous design if both execute
as the same OS identity or share filesystem/process visibility. A
pull request build could plant files, inspect residual data or
attack later privileged work.
Repair the topology: separate nodes/pools and OS/runtime trust boundaries, restrict protected-source jobs, and verify protected credentials are only exposed after the trust transition. Merely removing one label while leaving the same hostile process/host access may be insufficient.
6. Failure: misreading offline cause as queue starvation
If a node monitor takes an agent offline due to disk/temp/swap/clock/response thresholds, increasing executors is the opposite of a fix. Capture the node monitor/offline cause and host condition. Correct disk pressure, clock sync, agent connectivity or the relevant health problem; only then bring the node back and re-evaluate queue behavior.
7. Agent/Remoting and launcher evidence
When the node is disconnected rather than intentionally offline, inspect the agent connection log, Java version, network path and launcher-specific evidence. For SSH-launched agents, preserve host-key verification and least-privilege credentials. Do not disable host-key checks to make an agent connect.
For inbound agents, protect the agent secret and avoid copying a failing secret into tickets/logs. If compromise is suspected, recreate/rotate the node secret after evidence preservation.
8. Repair matrix
| Failure | Do not | Prefer | Verification |
|---|---|---|---|
| Label mismatch | Add fake labels | Repair requirement or real capability | Queue becomes eligible on intended nodes |
| Saturation | Unbounded executor increase | Measure, add bounded capacity | Queue falls without host/build regression |
| Workspace residue | Delete evidence first | Preserve hashes/times, then clean/isolate | Fresh build no longer depends on residue |
| Trust mixing | Rely on masking/label names | Separate execution pools/identities | Untrusted jobs are ineligible for privileged pool |
| Offline monitor | Force online repeatedly | Fix underlying host threshold issue | Healthy monitor + stable connection |
9. Intentionally broken diagnostic exercise
Configure agent-ci-1 with label linux-ci,
one executor, and temporarily offline cause
disk maintenance drill. Queue
capacity-hold. Before fixing, capture:
- queue item ID and
whytext; - node online/temporarily-offline state and cause;
- executor count/busy count;
- job/Jenkinsfile/source revision and requested label;
- agent/host free disk evidence.
Repair only by bringing the disposable node online after confirming the cause is a simulation, then verify the same queued item can progress. Do not delete/recreate the job or change its label to hide the original cause.
Knowledge check
Answer before revealing the explanation.
1. A queue item says no nodes match linux && arm64. What layer should you inspect first?
Scheduling configuration: the label expression and current node labels. Adding executors cannot fix a zero-eligibility label mismatch.
2. A node is online with four busy executors and jobs wait. Is the agent offline?
No. It is online but saturated. Preserve executor/build assignments, queue age and host resource evidence before deciding whether to add capacity.
3. Why is increasing executors a poor first fix for slow builds on a memory-starved host?
It raises concurrency and often worsens memory/CPU/I/O contention. First prove the bottleneck and compare per-build duration, load and queue metrics.
4. Why can workspace reuse be a security as well as correctness problem?
Leftover source, generated files, tokens, caches or permissions can affect later builds. Across trust domains, residual state can disclose protected material or alter results.
5. What evidence must remain after repairing a temporarily-offline or label problem?
Original queue/build IDs, first queue reason/offline cause, node/label/executor state, the narrow configuration change, rerun identity and before/after observations.
Official references and version notes
-
Jenkins LTS changelog
and
Java support policy
— lab baseline
Jenkins 2.568.3 LTS, Java 21; this LTS line supports Java 21 and 25. - Managing Nodes — controller, node, agent, executor and node-monitor concepts.
- Using Jenkins agents — labels, executor counts, usage modes and distributed builds.
-
Controller Isolation
— do not run builds on the built-in node; set its executor count
to
0once agents exist. - Securing Builds — isolate builds and separate trust domains instead of treating every shared agent as equivalent.
-
Pipeline: Nodes and Processes step reference
—
node,ws, workspace context and supported label-expression operators. -
Pipeline: Nodes and Processes plugin
— baseline
1479.v56e587f413a_7, requires Jenkins 2.479.3 or newer. -
SSH Build Agents plugin
— optional launcher baseline
3.1097.v868116049892; use verified host keys and least-privilege agent accounts. -
Built-In Node Name and Label Migration
— terminology and
NODE_NAME/NODE_LABELSbehavior.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.