Performance Engineering: Executors, Queueing, Heap, Garbage Collection, Disk I/O, Workspaces, and Controller Load: Diagnostics, Failure Modes, Security, and Performance
Preserve the first performance failure, identify the constrained layer, and repair the smallest scope—without using heap, executors, retries or deletion as universal shortcuts.
Learning objectives
- Diagnose executor, controller/JVM, Pipeline/CPS, disk/workspace and external-latency failures separately.
- Interpret intentionally broken examples without hiding first-failure evidence.
- Use thread/heap/support diagnostics with appropriate sensitivity controls.
- Recognize when a tuning change masks rather than fixes a defect.
- Apply the least destructive correction and remeasure.
1. Evidence-first diagnostic sequence
- Preserve exact job/build/queue IDs, source SHA, node/workspace and the first slow/failing run.
- Confirm controller core, Java and plugin baseline.
- Inspect queue reason, labels, eligibility and executor occupancy.
- Inspect agent CPU/memory/I/O/workspace/tooling.
- Inspect controller CPU, heap/GC, threads, disk and Pipeline/CPS behavior.
- Inspect log/report/artifact volume and external service latency.
- State one bottleneck hypothesis and one competing hypothesis.
- Apply the least destructive correction to the identified layer.
- Rerun the same workload and compare the same measurements.
Do not restart, retry, delete workspaces or increase capacity before preserving the evidence that lets you distinguish the layers.
2. Failure map
| Failure mode | Evidence pattern | Wrong shortcut | Better correction |
|---|---|---|---|
| More executors on CPU-starved host | queue may shrink; per-build time/CPU saturation worsen | Add still more executors | Reduce concurrency or add correctly sized agent capacity. |
| Huge console logs | controller disk/write/rendering grows with log volume | Increase heap | Reduce producer verbosity; preserve detailed file artifact if needed. |
| Controller doing build work | built-in executor busy; controller CPU/I/O tied to build | Raise controller executors | Set controller executors to 0; move workload to agents. |
| Giant CPS object/graph | controller CPU/heap and Pipeline persistence grow; many FlowNodes | Tune GC only | Move computation/data to agent-side tools; simplify Pipeline. |
| Workspace leak fills disk | agent free space falls across runs; stale workspaces/caches grow | Delete “latest” workspace blindly | Identify exact workspace/cache owner; snapshot evidence; bounded cleanup. |
| Heap increase masks plugin leak | outage delayed but live set keeps trending upward | Keep increasing Xmx | Reproduce/investigate plugin/object retention; update/fix/rollback. |
| Parallelism increases queue time | single build critical path improves, global queue wait worsens | Add branches everywhere | Bound concurrency and optimize platform objective, not one build. |
| External service slow | step time dominated by remote request latency | Add Jenkins executors/retries | Fix dependency, timeout/cache policy; avoid amplifying load. |
3. Broken example 1: “fix queueing” by adding executors
Assume a 2-core disposable agent already runs two CPU-bound builds and is at sustained high CPU. An administrator raises executors from 2 to 4. Queue wait drops, but each build takes more than twice as long.
Before
eligible executors: 2
agent CPU: 92-100%
median queue wait: 95 s
median build: 80 s
After executors=4
agent CPU: 100%
median queue wait: 20 s
median build: 190 s
end-to-end batch: not improved
The first causal evidence is agent CPU contention. Repair by restoring the previous executor count or adding independent capacity, then repeat the same workload. Do not claim the queue graph alone improved performance.
4. Broken example 2: giant console output
This bounded lab example demonstrates the pattern without producing gigabytes:
set -euo pipefail
# Disposable lab only: enough output to observe growth, not a production load test.
for i in $(seq 1 5000); do
printf 'diagnostic-line=%05d status=synthetic\n' "$i"
done
Measure the console log size and controller write activity. The safer production pattern is to keep concise actionable console output and write large machine-readable diagnostics to a controlled artifact/file with explicit retention and access. Never put secrets in either location.
5. Broken example 3: controller-side Groovy data processing
Large collection processing inside script {} runs in
the Pipeline CPS/controller domain. The exact threshold depends on
workload, but the architecture is wrong when Pipeline becomes the
data-processing engine.
// Intentionally poor pattern — do not copy to production.
script {
def rows = (1..10000).collect { i -> [id: i, value: "item-${i}"] }
def selected = rows.findAll { it.id % 2 == 0 }
echo "selected=${selected.size()}"
}
Repair by putting computation in a normal program on the agent and returning only the small result Jenkins needs for orchestration:
sh 'python3 tools/select_items.py input.json > selected.json'
archiveArtifacts artifacts: 'selected.json', fingerprint: true
6. Broken example 4: workspace leak
A workspace grows by several gigabytes each day because generated caches and test data never expire. Before cleanup, record the exact node/workspace, build numbers, largest paths and whether any retained release/evidence depends on them.
set -euo pipefail
# Read-only inspection inside the exact disposable workspace.
pwd
du -x -d 2 -k . 2>/dev/null | sort -nr | head -n 30
find . -xdev -type f -printf '%T@ %s %p\n' | sort -nr | head -n 30
Then implement a bounded cache/workspace policy. Do not run recursive deletion against a path chosen by “latest”, a glob you have not verified, or a controller filesystem location.
7. Broken example 5: heap increase masks a leak
If the controller's live-set trend rises steadily after every similar job and never returns after GC, doubling heap may only postpone the failure. Capture repeated post-GC heap observations, plugin/version baseline and thread/object evidence on a disposable/test controller.
# Examples only when you have authorized local access to the disposable controller JVM.
jcmd <controller-pid> GC.heap_info
jcmd <controller-pid> Thread.print > /restricted/jenkins-thread-dump.txt
# A full heap dump is much more sensitive and expensive; do not create one casually.
8. Broken example 6: parallelism helps one build, hurts the platform
Suppose a test Pipeline changes from 4 serial shards to 16 parallel branches. Its own wall clock falls from 18 to 11 minutes, but the shared pool now sees p95 queue wait rise from 30 seconds to 7 minutes and other jobs miss their feedback SLO. The correct unit of optimization is the platform/service objective, not only one build.
Bound matrix/parallel cardinality, use dedicated pools where workload priority differs, and re-evaluate critical path versus queue/capacity cost.
9. Thread dumps, metrics and support evidence
Use controller thread dumps when you suspect lock/contention or long controller-side execution, not as a first reaction to every queue. Correlate the dump timestamp with queue/build IDs and metrics. A thread dump without a timeline can be misleading.
Likewise, a support bundle or JVM histogram is not a tuning prescription. It is evidence that must be interpreted with the workload, versions and first-failure sequence.
10. Security and reliability guardrails
- Never expose metrics/JMX/debug endpoints publicly to make collection easier.
- Never disable authorization, TLS, CSRF or Script Security for performance testing.
- Never run untrusted builds on the controller to “save agent startup time.”
- Never log credentials or copy secret-bearing environment data into diagnostics.
- Never use a production controller as an uncontrolled load generator.
- Never delete artifacts/workspaces by ambiguous ordering such as “latest.”
- Do not create full heap dumps on sensitive systems unless you have a secure handling plan.
11. Repair pattern: smallest safe scope
Preserve -> classify -> hypothesize -> change one layer -> rerun same workload -> compare
Example:
1. Preserve job perf-lab/bounded-workload #42, queue item 991, node perf-agent-01.
2. Evidence: queue low; controller GC normal; agent disk latency high; workspace 28 GiB.
3. Hypothesis: stale workspace/cache I/O dominates execution.
4. Change: rotate only the disposable cache directory with exact path guard.
5. Rerun same commit/job/tool baseline.
6. Compare build duration, disk latency, cache-hit behavior and workspace size.
Knowledge check
Answer before revealing the explanation.
1. Why can reducing queue wait still be a performance regression?
If the change increases per-build duration or controller/agent contention enough that end-to-end completion or platform SLOs worsen.
2. What is wrong with using a heap increase as the first response to a rising live set?
It can mask a leak or unbounded workload; first determine object/live-set behavior and causal plugin/Pipeline load.
3. Why are large Pipeline Groovy collections problematic?
They create controller-side CPS computation/state; project data processing belongs in agent-side tools.
4. What should be captured before cleaning a leaking workspace?
Exact node/workspace/build identity, largest paths/ages, dependency on retained evidence, and current disk/I/O state.
5. Why are thread/heap/support artifacts security-sensitive?
They may contain operational identifiers, paths, URLs or other sensitive runtime data and require restricted handling.
Official references and version notes
Performance behavior depends on controller scale, workload shape, storage, plugins and runtime versions. Re-check current primary documentation before applying any tuning change to a real controller.
- Jenkins — Scaling Jenkins
- Jenkins — Hardware Recommendations
- Jenkins — Architecting for Scale
- Jenkins — Scaling Pipelines / durability
- Jenkins — Pipeline Best Practices
- Jenkins — Using agents
- Jenkins — Managing nodes
- Jenkins — Java Support Policy
- Metrics plugin
- Prometheus metrics plugin
- Pipeline: Groovy plugin
- Pipeline: Supporting APIs plugin
- Jenkins LTS changelog
- Jenkins Security Advisories
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.