Chapter 38Lesson 04~220 minutes

Performance Engineering: Executors, Queueing, Heap, Garbage Collection, Disk I/O, Workspaces, and Controller Load: Diagnostics, Failure Modes, Security, and Performance

Preserve the first performance failure, identify the constrained layer, and repair the smallest scope—without using heap, executors, retries or deletion as universal shortcuts.

diagnosticsthread dumpGCworkspace leakCPSfirst-failure evidence

Learning objectives

  • Diagnose executor, controller/JVM, Pipeline/CPS, disk/workspace and external-latency failures separately.
  • Interpret intentionally broken examples without hiding first-failure evidence.
  • Use thread/heap/support diagnostics with appropriate sensitivity controls.
  • Recognize when a tuning change masks rather than fixes a defect.
  • Apply the least destructive correction and remeasure.

1. Evidence-first diagnostic sequence

  1. Preserve exact job/build/queue IDs, source SHA, node/workspace and the first slow/failing run.
  2. Confirm controller core, Java and plugin baseline.
  3. Inspect queue reason, labels, eligibility and executor occupancy.
  4. Inspect agent CPU/memory/I/O/workspace/tooling.
  5. Inspect controller CPU, heap/GC, threads, disk and Pipeline/CPS behavior.
  6. Inspect log/report/artifact volume and external service latency.
  7. State one bottleneck hypothesis and one competing hypothesis.
  8. Apply the least destructive correction to the identified layer.
  9. Rerun the same workload and compare the same measurements.

Do not restart, retry, delete workspaces or increase capacity before preserving the evidence that lets you distinguish the layers.

2. Failure map

Failure mode Evidence pattern Wrong shortcut Better correction
More executors on CPU-starved host queue may shrink; per-build time/CPU saturation worsen Add still more executors Reduce concurrency or add correctly sized agent capacity.
Huge console logs controller disk/write/rendering grows with log volume Increase heap Reduce producer verbosity; preserve detailed file artifact if needed.
Controller doing build work built-in executor busy; controller CPU/I/O tied to build Raise controller executors Set controller executors to 0; move workload to agents.
Giant CPS object/graph controller CPU/heap and Pipeline persistence grow; many FlowNodes Tune GC only Move computation/data to agent-side tools; simplify Pipeline.
Workspace leak fills disk agent free space falls across runs; stale workspaces/caches grow Delete “latest” workspace blindly Identify exact workspace/cache owner; snapshot evidence; bounded cleanup.
Heap increase masks plugin leak outage delayed but live set keeps trending upward Keep increasing Xmx Reproduce/investigate plugin/object retention; update/fix/rollback.
Parallelism increases queue time single build critical path improves, global queue wait worsens Add branches everywhere Bound concurrency and optimize platform objective, not one build.
External service slow step time dominated by remote request latency Add Jenkins executors/retries Fix dependency, timeout/cache policy; avoid amplifying load.

3. Broken example 1: “fix queueing” by adding executors

Assume a 2-core disposable agent already runs two CPU-bound builds and is at sustained high CPU. An administrator raises executors from 2 to 4. Queue wait drops, but each build takes more than twice as long.

Before
eligible executors: 2
agent CPU:         92-100%
median queue wait: 95 s
median build:      80 s

After executors=4
agent CPU:         100%
median queue wait: 20 s
median build:      190 s
end-to-end batch:  not improved

The first causal evidence is agent CPU contention. Repair by restoring the previous executor count or adding independent capacity, then repeat the same workload. Do not claim the queue graph alone improved performance.

4. Broken example 2: giant console output

This bounded lab example demonstrates the pattern without producing gigabytes:

set -euo pipefail
# Disposable lab only: enough output to observe growth, not a production load test.
for i in $(seq 1 5000); do
  printf 'diagnostic-line=%05d status=synthetic\n' "$i"
done

Measure the console log size and controller write activity. The safer production pattern is to keep concise actionable console output and write large machine-readable diagnostics to a controlled artifact/file with explicit retention and access. Never put secrets in either location.

5. Broken example 3: controller-side Groovy data processing

Large collection processing inside script {} runs in the Pipeline CPS/controller domain. The exact threshold depends on workload, but the architecture is wrong when Pipeline becomes the data-processing engine.

// Intentionally poor pattern — do not copy to production.
script {
  def rows = (1..10000).collect { i -> [id: i, value: "item-${i}"] }
  def selected = rows.findAll { it.id % 2 == 0 }
  echo "selected=${selected.size()}"
}

Repair by putting computation in a normal program on the agent and returning only the small result Jenkins needs for orchestration:

sh 'python3 tools/select_items.py input.json > selected.json'
archiveArtifacts artifacts: 'selected.json', fingerprint: true

6. Broken example 4: workspace leak

A workspace grows by several gigabytes each day because generated caches and test data never expire. Before cleanup, record the exact node/workspace, build numbers, largest paths and whether any retained release/evidence depends on them.

set -euo pipefail
# Read-only inspection inside the exact disposable workspace.
pwd
du -x -d 2 -k . 2>/dev/null | sort -nr | head -n 30
find . -xdev -type f -printf '%T@ %s %p\n' | sort -nr | head -n 30

Then implement a bounded cache/workspace policy. Do not run recursive deletion against a path chosen by “latest”, a glob you have not verified, or a controller filesystem location.

7. Broken example 5: heap increase masks a leak

If the controller's live-set trend rises steadily after every similar job and never returns after GC, doubling heap may only postpone the failure. Capture repeated post-GC heap observations, plugin/version baseline and thread/object evidence on a disposable/test controller.

# Examples only when you have authorized local access to the disposable controller JVM.
jcmd <controller-pid> GC.heap_info
jcmd <controller-pid> Thread.print > /restricted/jenkins-thread-dump.txt
# A full heap dump is much more sensitive and expensive; do not create one casually.

8. Broken example 6: parallelism helps one build, hurts the platform

Suppose a test Pipeline changes from 4 serial shards to 16 parallel branches. Its own wall clock falls from 18 to 11 minutes, but the shared pool now sees p95 queue wait rise from 30 seconds to 7 minutes and other jobs miss their feedback SLO. The correct unit of optimization is the platform/service objective, not only one build.

Bound matrix/parallel cardinality, use dedicated pools where workload priority differs, and re-evaluate critical path versus queue/capacity cost.

9. Thread dumps, metrics and support evidence

Use controller thread dumps when you suspect lock/contention or long controller-side execution, not as a first reaction to every queue. Correlate the dump timestamp with queue/build IDs and metrics. A thread dump without a timeline can be misleading.

Likewise, a support bundle or JVM histogram is not a tuning prescription. It is evidence that must be interpreted with the workload, versions and first-failure sequence.

10. Security and reliability guardrails

  • Never expose metrics/JMX/debug endpoints publicly to make collection easier.
  • Never disable authorization, TLS, CSRF or Script Security for performance testing.
  • Never run untrusted builds on the controller to “save agent startup time.”
  • Never log credentials or copy secret-bearing environment data into diagnostics.
  • Never use a production controller as an uncontrolled load generator.
  • Never delete artifacts/workspaces by ambiguous ordering such as “latest.”
  • Do not create full heap dumps on sensitive systems unless you have a secure handling plan.

11. Repair pattern: smallest safe scope

Preserve -> classify -> hypothesize -> change one layer -> rerun same workload -> compare

Example:
1. Preserve job perf-lab/bounded-workload #42, queue item 991, node perf-agent-01.
2. Evidence: queue low; controller GC normal; agent disk latency high; workspace 28 GiB.
3. Hypothesis: stale workspace/cache I/O dominates execution.
4. Change: rotate only the disposable cache directory with exact path guard.
5. Rerun same commit/job/tool baseline.
6. Compare build duration, disk latency, cache-hit behavior and workspace size.
Next

Prove a performance improvement

Lesson 5 is a checkpoint experiment: measure a repeatable workload, state a bottleneck hypothesis, change one variable, rerun and decide from evidence whether the platform actually improved.

Knowledge check

Answer before revealing the explanation.

1. Why can reducing queue wait still be a performance regression?

2. What is wrong with using a heap increase as the first response to a rising live set?

3. Why are large Pipeline Groovy collections problematic?

4. What should be captured before cleaning a leaking workspace?

5. Why are thread/heap/support artifacts security-sensitive?

Official references and version notes

Performance behavior depends on controller scale, workload shape, storage, plugins and runtime versions. Re-check current primary documentation before applying any tuning change to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.