Chapter 17Lesson 03~120 minutes

Nodes, Agents, Labels, Executors, Workspaces, Offline Causes, and Agent Capacity Design: Configuration, Design Choices, and Tradeoffs

Choose executor counts, agent lifecycle, label taxonomy, trust pools, and workspace isolation by connecting each decision to measurable queue behavior, host resources, source trust, credentials, and rollback.

Capacity designLabel taxonomyStatic vs ephemeralIsolationRetentionRollback

Learning objectives

  • Choose executor counts from measured host/workload behavior rather than arbitrary ratios.
  • Compare static and ephemeral workers, including provisioning, cleanup, observability and failure tradeoffs.
  • Design labels around capabilities and trust instead of individual machine names.
  • Separate privileged and untrusted work even when both need the same operating system/toolchain.
  • Define reversible workspace and capacity policies with observable acceptance criteria.

1. One executor versus many

One executor provides the simplest isolation model: only one Jenkins task consumes the node at a time. Multiple executors can improve utilization when jobs are lightweight or spend time waiting on external systems, but concurrency multiplies demand for CPU, RAM, disk bandwidth, temporary space, process limits, licenses and network connections.

Design Benefits Risks Evidence before/after
1 executor Predictable isolation, simple diagnosis Queue may grow while host is underused Queue age, host utilization, build duration
2–N executors Higher concurrency on capable hosts Contention, noisy neighbors, memory pressure Same metrics plus failure/GC/disk pressure
Ephemeral 1-per-worker Strong cleanup/isolation, elastic fleet Provision latency, image/plugin/cloud dependency Provision time, image identity, queue-to-start latency

2. Static versus ephemeral workers

Static agents are easy to inspect and cache aggressively, but persistent workspaces and manually installed tools drift. Ephemeral agents can start from reviewed images and disappear after work, reducing residue and easing horizontal scaling. They shift durability requirements outward: artifacts, reports and important logs must not depend on the worker surviving.

Chapter 20 will cover Kubernetes-based dynamic agents. Here, the design rule is provider-neutral: worker lifecycle is separate from the build record. If a worker disappears, Jenkins should still retain enough source/build/artifact evidence to diagnose what happened.

3. Label taxonomy: capability, platform and trust

Prefer labels that are stable properties administrators can verify: linux, windows, arm64, jdk21, high-memory, gpu, release-trusted. Avoid making every job depend on physical host names such as server42 unless affinity to that exact machine is the real requirement.

// Good: expresses capabilities and trust
node('linux && jdk21 && !release-trusted') {
  sh './gradlew test'
}

// Deliberate privileged pool
node('linux && release-trusted') {
  echo 'Simulated release evidence only in this lab'
}

Labels are administrator-controlled scheduling claims. Periodically verify that nodes carrying high-trust labels still satisfy the intended hardening and access requirements.

4. Dedicated versus shared agents

A shared CI pool improves utilization and developer throughput. A dedicated signing/release/hardware pool improves isolation and makes privileged access easier to audit. The key is to separate conflicting trust levels, not to create a dedicated machine for every project by habit.

Do not co-schedule incompatible trust domains. A build able to execute arbitrary source code on a persistent host may inspect same-user processes, caches or files. Masking and folder permissions do not erase the risk once the secret or privileged device reaches that agent.

5. Workspace reuse versus clean isolation

Workspace reuse is fast and cache-friendly, but it can preserve untracked files, old dependencies and generated state. Clean isolation improves reproducibility at the cost of checkout/download time. Separate dependency caches from the source workspace when possible and make cache keys/version ownership explicit.

For high-trust or cross-tenant pools, stronger isolation is usually worth the cost. For a single trusted project, controlled reuse may be reasonable when the build explicitly recreates outputs and validates cache identity.

6. Node usage mode

A node configured “Use this node as much as possible” can accept suitable unlabelled work. “Only build jobs with label expressions matching this node” requires explicit scheduling intent. For scarce, privileged or specialized workers, explicit label targeting reduces accidental use.

Usage mode is controller configuration; it does not alter host resources. Changing it can immediately change queue eligibility, so record the previous setting and affected queued jobs before mutation.

7. Worked scenario: growing service with release signing

Need Recommended design Prerequisite Observable proof
High-volume unit tests Shared linux-ci pool; scale workers before overcommitting one host Repeatable tools/caches Queue age and host utilization fall without longer builds
Release signing Dedicated release-trusted pool, explicit label usage Reviewed source/permissions and protected credential/device Only authorized builds appear on that pool
ARM validation arm64 capability pool Correct architecture/toolchain uname -m, node labels, test report
Burst demand Ephemeral workers or additional static nodes Provisioning automation and externalized evidence Provision/start latency, queue reduction, image identity

8. Capacity changes need rollback criteria

Before raising executors from 1 to 2, state what success means—for example p95 queue wait decreases while median build duration changes less than 10%, memory remains above a safety threshold, disk latency stays healthy and failure rate does not increase. If those conditions fail, revert the executor count and investigate the real bottleneck.

Likewise, changing label taxonomy should have a compatibility plan: temporarily support old and new labels, update jobs, verify routing, then remove the deprecated label. Do not perform a fleet-wide rename without proving who still depends on it.

Next lesson

Diagnostics, Failure Modes, Security, and Performance

Use queue IDs, offline causes, labels, executor occupancy, agent logs, workspace evidence and host metrics to repair realistic scheduling and isolation failures.

Knowledge check

Answer before revealing the explanation.

1. When can multiple executors on one agent be reasonable?

2. Why prefer capability labels such as linux, arm64 or signing over project-name labels everywhere?

3. What is a major advantage of ephemeral workers?

4. Why separate privileged release agents from untrusted pull-request agents?

5. What should trigger an executor-count rollback?

Official references and version notes

Assumption timestamp: 2026-09-17. Recheck Jenkins LTS/Java support, agent/remoting compatibility, launcher-plugin advisories, and current node/monitor behavior before reproducing the lab later.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.