Chapter 38Lesson 01~200 minutes

Performance Engineering: Executors, Queueing, Heap, Garbage Collection, Disk I/O, Workspaces, and Controller Load: Concepts, Architecture, and Mental Model

Performance engineering starts by decomposing end-to-end latency into queueing, controller orchestration, agent execution, storage traffic and external dependency time—then changing only the layer proven to be constrained.

performancequeueingexecutorsJVMdisk I/OPipeline CPS

Learning objectives

  • Decompose build elapsed time into queue wait, orchestration, agent execution, I/O and external-service latency.
  • Distinguish executor saturation from label/eligibility problems and controller overload.
  • Interpret controller CPU, heap, GC, threads and disk evidence without assuming one metric proves the root cause.
  • Explain why Pipeline/CPS structure, console volume, workspace traffic and artifact handling can load the controller.
  • Form one measurable bottleneck hypothesis before tuning.

1. The practical problem: “Jenkins is slow” is not a diagnosis

Two builds may both take ten minutes while needing completely different fixes. One might wait eight minutes in the queue because every eligible executor is busy. Another might start immediately but spend eight minutes downloading dependencies from a slow remote repository. A third might execute quickly on the agent while the controller struggles to persist a very large Pipeline graph and huge console output.

Performance engineering therefore begins with a boundary: measure where time and resources are consumed before changing capacity or configuration. Increasing heap, adding executors or creating more parallel branches without that evidence can move the bottleneck, increase contention, or reduce durability without improving user-visible latency.

2. Mental model: workload arrival to completed evidence

Mental model: workload arrival to completed evidence
flowchart TD
  A["Trigger / workload arrival"] --> B["Queue"]
  B --> C["Controller scheduling + Pipeline orchestration"]
  C --> D["Eligible agent + executor"]
  D --> E["Workspace + tool execution"]
  E --> F["Logs / reports / artifacts / storage I/O"]
  F --> G["External services"]
  G --> H["Completion + retained evidence"]

  B -. queue telemetry .-> I["Measurements"]
  C -. CPU / heap / GC / threads / CPS .-> I
  D -. utilization / labels / online state .-> I
  E -. CPU / memory / disk / workspace .-> I
  F -. bytes / latency / retention .-> I
  G -. request latency / errors .-> I

  I --> J["One bottleneck hypothesis"]
  J --> K["One controlled change"]
  K --> L["Same workload rerun"]
  

The model deliberately separates scheduling from execution. A job can spend most of its elapsed time waiting before it ever gets an executor. Likewise, an executor can be idle while a job waits because labels, node mode or workspace/tool constraints make that executor ineligible.

3. Decompose end-to-end latency

Component Question Evidence
Queue wait How long between queue admission and executor allocation? queue item ID, inQueueSince/buildableStartMilliseconds, queue reason, build start time.
Controller orchestration Is Pipeline scheduling/persistence/controller work delayed? controller CPU, heap/GC, thread dumps, Pipeline timing, disk I/O, logs.
Agent execution Is the actual build/test work constrained? agent CPU/memory/I/O, executor utilization, step duration, tool logs.
Workspace I/O Are checkout/cache/build files dominating time? workspace size, file count, disk throughput/latency, cleanup behavior.
Evidence I/O Are logs, test reports or artifacts expensive to retain/transfer? console bytes, report counts, archive sizes, storage latency.
External dependency Is Jenkins waiting on SCM, package registry, artifact repo, Sonar, IdP or API? request latency, external logs/metrics, timestamps around calls.
Recovery/retry Did reruns or retries add hidden elapsed time? build causes, retry counts, first-failure timestamps, external side-effect state.

4. Queue pressure is not the same as controller overload

A long queue means demand cannot currently advance to execution; it does not by itself identify why. The key evidence is the why reason and the eligible capacity for the requested label. If all eligible executors are busy, capacity may be the constraint. If executors are idle but labels do not match, adding more generic executors does not solve the scheduling contract.

Measure at least queue age/wait, requested label, eligible online nodes, busy/idle executors and workload arrival rate before changing executor counts.

5. Executors: concurrency is a resource claim

An executor is a concurrency slot, not additional CPU, RAM or I/O. Jenkins current guidance recommends keeping the controller's executor count at 0 so builds do not compete with orchestration. On agents, multiple executors can be effective for small, mostly waiting tasks, but CPU-heavy, memory-heavy or I/O-heavy jobs can become slower when too many share one host.

Observation Possible interpretation Do not conclude yet
Executors 100% busy + agents healthy Eligible pool may be saturated Do not add executors until host CPU/memory/I/O headroom is measured.
Executors idle + queue growing Eligibility/labels/node mode may be wrong Do not increase controller heap as a scheduling fix.
More executors increase per-build duration Resource contention on agent Do not equate throughput with individual latency.
Queue shrinks but controller CPU rises sharply Scheduling/Pipeline/controller load may have moved upstream Do not call the change a win without end-to-end/resource comparison.

6. Controller CPU, heap and garbage collection

Heap utilization is only one JVM signal. A healthy controller can use a high percentage of its heap and still perform well if allocation/collection behavior is stable. Conversely, repeated long GC pauses, high CPU, growing live-set size, thread contention or disk stalls can make the UI and Pipeline orchestration slow even when heap has not reached its configured maximum.

Increasing -Xmx can be appropriate when the workload genuinely needs more live memory and the host has capacity. It is not a universal response to a leak, unbounded plugin cache, giant Pipeline object graph or excessive job history. Always compare before/after GC pause behavior and host memory pressure.

7. Pipeline/CPS is controller work

Pipeline Groovy executes through the CPS engine in the controller process. Current plugin guidance explicitly recommends keeping Pipeline Groovy as high-level orchestration “glue” and running project computation in external tools on agents. Large in-memory Groovy collections, deeply dynamic Pipeline generation, excessive step counts and enormous parallel graphs increase controller-side work and persisted Pipeline state.

Durability adds another dimension. Jenkins Pipeline can write execution state frequently so it can survive unexpected restarts. Performance-optimized durability writes less frequently and can reduce I/O, but it trades away some crash survivability. It is a reliability decision—not a workaround for serialization bugs and not the first tuning lever for every job.

8. Disk I/O: JENKINS_HOME, workspaces and retained evidence are different

Controller storage handles configuration, build records, Pipeline state, logs and often archived artifacts. Agent workspaces handle checkout, compilation, caches and temporary files. Treating all disk usage as one number hides the owner of the traffic.

Traffic Primary location Performance risk Safer pattern
Pipeline state/build records Controller JENKINS_HOME Metadata latency and persistence cost Fast reliable controller storage; tune durability only with explicit reliability tradeoff.
Checkout/build outputs Agent workspace Large files, inode pressure, cleanup leaks Disposable/managed workspaces; bounded caches; monitor disk and I/O.
Console logs Controller/build record path Huge log volume and rendering/storage cost Summarize; write detailed tool output to artifacts when appropriate; retain intentionally.
Artifacts/reports Controller or external artifact manager/repository Network/disk bandwidth and retention growth Use external durable artifact systems when scale/semantics justify it; preserve immutable identity.
Dependency caches Agent/cache service Contention, stale data, large footprint Version/validate caches and bound retention rather than accumulating indefinitely.

9. Parallelism changes the critical path—but also demand

Parallel stages can reduce a Pipeline's critical path when independent work has sufficient eligible executors and underlying resources. They can also increase queue wait, controller step count, artifact fan-in, external API concurrency and agent contention. “More parallel” is therefore an experiment, not an axiom.

For every parallelization proposal, compare: total elapsed time, queue wait, maximum concurrency, controller CPU/heap/GC, agent utilization, disk/network traffic and external-service latency. If elapsed time improves by 10% while controller CPU doubles and queue wait for other teams triples, the platform outcome may be worse.

10. Read-only inspection before tuning

set -euo pipefail
JENKINS_URL='http://127.0.0.1:8080'
OUT='/tmp/jenkins-perf-inspect'
mkdir -p "$OUT"
TS="$(date -u +%Y%m%dT%H%M%SZ)"
# Add approved read-only lab authentication in your environment when required.
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,inQueueSince,why,task[name,url]]" > "$OUT/queue-$TS.json"
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,numExecutors,assignedLabels[name]]" > "$OUT/nodes-$TS.json"
# Optional when the pinned Prometheus plugin is installed and access is authorized:
curl -fsS "$JENKINS_URL/prometheus/" > "$OUT/metrics-$TS.txt"
printf 'captured=%s\n' "$TS" | tee "$OUT/manifest.txt"

Do not paste credentials into commands committed to source or lesson output. If metrics require authentication, use the disposable lab identity and keep authorization headers/tokens out of evidence artifacts.

11. Write a falsifiable bottleneck hypothesis

Bad:  Jenkins is slow; add 4 executors.
Better: During workload PERF-BASE-01, p95 queue wait is 180 s while the only
        eligible perf-linux executor is >95% busy; controller CPU/heap/GC and
        agent CPU/I/O remain below our lab saturation thresholds.
Prediction: Adding one executor on a second equally sized disposable agent will
            reduce queue wait without increasing controller critical-path time.
Test: Re-run the exact same workload and compare the same measurements.

A good hypothesis names the workload, layer, evidence, predicted effect and comparison method. It can be disproven.

12. Common wrong approaches

  • Add executors until the queue disappears: concurrency can outrun CPU, memory or disk and make builds slower.
  • Increase heap because memory is high: live-set growth or plugin behavior may remain; host swapping can worsen latency.
  • Turn on performance-optimized Pipeline durability everywhere: you changed restart/crash semantics, not just speed.
  • Delete workspaces/logs immediately: you may erase the evidence needed to find the actual bottleneck.
  • Run builds on the controller to avoid queueing: this trades queue delay for controller contention and violates controller-isolation guidance.
  • Optimize from one fast build: use a repeatable workload and multiple comparable runs.
Next

Build a controlled performance experiment

Lesson 2 creates a disposable workload, measures queue/JVM/disk/workspace evidence, changes executor/parallelism variables one at a time, and compares identical reruns.

Knowledge check

Answer before revealing the explanation.

1. Why can a growing queue coexist with idle executors?

2. Why is an executor not equivalent to a CPU core?

3. What is the main performance cost of Pipeline durability?

4. Why should project computation stay out of large Pipeline Groovy loops?

5. What makes a bottleneck hypothesis useful?

Official references and version notes

Performance behavior depends on controller scale, workload shape, storage, plugins and runtime versions. Re-check current primary documentation before applying any tuning change to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.