Nodes, Agents, Labels, Executors, Workspaces, Offline Causes, and Agent Capacity Design: Configuration, Design Choices, and Tradeoffs
Choose executor counts, agent lifecycle, label taxonomy, trust pools, and workspace isolation by connecting each decision to measurable queue behavior, host resources, source trust, credentials, and rollback.
Learning objectives
- Choose executor counts from measured host/workload behavior rather than arbitrary ratios.
- Compare static and ephemeral workers, including provisioning, cleanup, observability and failure tradeoffs.
- Design labels around capabilities and trust instead of individual machine names.
- Separate privileged and untrusted work even when both need the same operating system/toolchain.
- Define reversible workspace and capacity policies with observable acceptance criteria.
1. One executor versus many
One executor provides the simplest isolation model: only one Jenkins task consumes the node at a time. Multiple executors can improve utilization when jobs are lightweight or spend time waiting on external systems, but concurrency multiplies demand for CPU, RAM, disk bandwidth, temporary space, process limits, licenses and network connections.
| Design | Benefits | Risks | Evidence before/after |
|---|---|---|---|
| 1 executor | Predictable isolation, simple diagnosis | Queue may grow while host is underused | Queue age, host utilization, build duration |
| 2–N executors | Higher concurrency on capable hosts | Contention, noisy neighbors, memory pressure | Same metrics plus failure/GC/disk pressure |
| Ephemeral 1-per-worker | Strong cleanup/isolation, elastic fleet | Provision latency, image/plugin/cloud dependency | Provision time, image identity, queue-to-start latency |
2. Static versus ephemeral workers
Static agents are easy to inspect and cache aggressively, but persistent workspaces and manually installed tools drift. Ephemeral agents can start from reviewed images and disappear after work, reducing residue and easing horizontal scaling. They shift durability requirements outward: artifacts, reports and important logs must not depend on the worker surviving.
Chapter 20 will cover Kubernetes-based dynamic agents. Here, the design rule is provider-neutral: worker lifecycle is separate from the build record. If a worker disappears, Jenkins should still retain enough source/build/artifact evidence to diagnose what happened.
3. Label taxonomy: capability, platform and trust
Prefer labels that are stable properties administrators can verify:
linux, windows, arm64,
jdk21, high-memory, gpu,
release-trusted. Avoid making every job depend on
physical host names such as server42 unless affinity to
that exact machine is the real requirement.
// Good: expresses capabilities and trust
node('linux && jdk21 && !release-trusted') {
sh './gradlew test'
}
// Deliberate privileged pool
node('linux && release-trusted') {
echo 'Simulated release evidence only in this lab'
}
Labels are administrator-controlled scheduling claims. Periodically verify that nodes carrying high-trust labels still satisfy the intended hardening and access requirements.
4. Dedicated versus shared agents
A shared CI pool improves utilization and developer throughput. A dedicated signing/release/hardware pool improves isolation and makes privileged access easier to audit. The key is to separate conflicting trust levels, not to create a dedicated machine for every project by habit.
5. Workspace reuse versus clean isolation
Workspace reuse is fast and cache-friendly, but it can preserve untracked files, old dependencies and generated state. Clean isolation improves reproducibility at the cost of checkout/download time. Separate dependency caches from the source workspace when possible and make cache keys/version ownership explicit.
For high-trust or cross-tenant pools, stronger isolation is usually worth the cost. For a single trusted project, controlled reuse may be reasonable when the build explicitly recreates outputs and validates cache identity.
6. Node usage mode
A node configured “Use this node as much as possible” can accept suitable unlabelled work. “Only build jobs with label expressions matching this node” requires explicit scheduling intent. For scarce, privileged or specialized workers, explicit label targeting reduces accidental use.
Usage mode is controller configuration; it does not alter host resources. Changing it can immediately change queue eligibility, so record the previous setting and affected queued jobs before mutation.
7. Worked scenario: growing service with release signing
| Need | Recommended design | Prerequisite | Observable proof |
|---|---|---|---|
| High-volume unit tests |
Shared linux-ci pool; scale workers before
overcommitting one host
|
Repeatable tools/caches | Queue age and host utilization fall without longer builds |
| Release signing |
Dedicated release-trusted pool, explicit label
usage
|
Reviewed source/permissions and protected credential/device | Only authorized builds appear on that pool |
| ARM validation | arm64 capability pool |
Correct architecture/toolchain | uname -m, node labels, test report |
| Burst demand | Ephemeral workers or additional static nodes | Provisioning automation and externalized evidence | Provision/start latency, queue reduction, image identity |
8. Capacity changes need rollback criteria
Before raising executors from 1 to 2, state what success means—for example p95 queue wait decreases while median build duration changes less than 10%, memory remains above a safety threshold, disk latency stays healthy and failure rate does not increase. If those conditions fail, revert the executor count and investigate the real bottleneck.
Likewise, changing label taxonomy should have a compatibility plan: temporarily support old and new labels, update jobs, verify routing, then remove the deprecated label. Do not perform a fleet-wide rename without proving who still depends on it.
Knowledge check
Answer before revealing the explanation.
1. When can multiple executors on one agent be reasonable?
When the host has measured spare CPU/memory/I/O and workloads are sufficiently isolated and lightweight. Increase slowly from one and validate queue latency plus host saturation.
2. Why prefer capability labels such as linux, arm64 or signing over project-name labels everywhere?
Stable capability/trust labels express why a job needs a pool, reduce brittle coupling to individual machines and make replacements/ephemeral workers easier.
3. What is a major advantage of ephemeral workers?
They reduce persistent workspace/tool drift and can scale with demand. The tradeoff is provisioning latency, image/launcher dependencies and the need to externalize durable artifacts/evidence.
4. Why separate privileged release agents from untrusted pull-request agents?
A shared agent can expose processes, files, caches, credentials or host privileges across trust domains. Separate pools reduce the blast radius and make authorization intent auditable.
5. What should trigger an executor-count rollback?
Evidence such as CPU/memory pressure, disk contention, longer build duration, agent instability or increased failures after the change—not just a preference for lower numbers.
Official references and version notes
-
Jenkins LTS changelog
and
Java support policy
— lab baseline
Jenkins 2.568.3 LTS, Java 21; this LTS line supports Java 21 and 25. - Managing Nodes — controller, node, agent, executor and node-monitor concepts.
- Using Jenkins agents — labels, executor counts, usage modes and distributed builds.
-
Controller Isolation
— do not run builds on the built-in node; set its executor count
to
0once agents exist. - Securing Builds — isolate builds and separate trust domains instead of treating every shared agent as equivalent.
-
Pipeline: Nodes and Processes step reference
—
node,ws, workspace context and supported label-expression operators. -
Pipeline: Nodes and Processes plugin
— baseline
1479.v56e587f413a_7, requires Jenkins 2.479.3 or newer. -
SSH Build Agents plugin
— optional launcher baseline
3.1097.v868116049892; use verified host keys and least-privilege agent accounts. -
Built-In Node Name and Label Migration
— terminology and
NODE_NAME/NODE_LABELSbehavior.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.