Performance Engineering: Executors, Queueing, Heap, Garbage Collection, Disk I/O, Workspaces, and Controller Load: Guided Hands-On Workflow and Core Operations
Run a bounded local performance experiment: capture a baseline, generate repeatable queue pressure and I/O, vary one concurrency variable, and compare the same evidence before deciding whether the change helped.
Learning objectives
- Create a reproducible disposable performance lab without using production repositories or infrastructure.
- Measure queue wait, executor state, controller/JVM signals and workspace/log/artifact volume.
- Change executor count and Pipeline parallelism in controlled experiments.
- Separate throughput improvement from per-build latency and controller cost.
- Preserve an evidence packet and clean up only exact lab resources.
1. Scenario and safety boundary
You operate a disposable controller called
jenkins-perf-lab with the built-in node set to
0 executors. One disposable Linux agent is labeled
perf-linux. The job
perf-lab/bounded-workload performs synthetic CPU-light
work, bounded filesystem work, modest log output and a simulated
external wait. Nothing points at production systems.
2. Preflight and record the baseline
| Item | Lab requirement |
|---|---|
| Controller | Jenkins LTS 2.568.3; built-in node executors = 0. |
| Runtime | Java 21 for controller/agent JVMs. |
| Agent |
Disposable perf-agent-01, label
perf-linux, initially 1 executor.
|
| Metrics | Remote API required; Metrics/Prometheus optional but pinned if installed. |
| Source | Synthetic local Git repository or Pipeline job; record exact Jenkinsfile revision. |
| External service | Local mock/sleep only; no production API. |
| Evidence path |
/tmp/jenkins-perf-lab-evidence on the operator
host.
|
| Bounds | At most 4 queued builds, 2 executors, 2 parallel branches; no unbounded loops. |
set -euo pipefail
OUT='/tmp/jenkins-perf-lab-evidence'
mkdir -p "$OUT"
date -u +baseline_utc=%Y-%m-%dT%H:%M:%SZ | tee "$OUT/baseline.txt"
java -version 2>&1 | tee -a "$OUT/baseline.txt"
printf 'jenkins_lts=2.568.3\nagent=perf-agent-01\nlabel=perf-linux\nexecutors=1\n' | tee -a "$OUT/baseline.txt"
3. Use a bounded, parameterized synthetic Pipeline
The Pipeline below keeps the work intentionally small. The point is not to benchmark hardware; it is to learn measurement discipline.
pipeline {
agent { label 'perf-linux' }
parameters {
choice(name: 'MODE', choices: ['serial', 'parallel2'], description: 'bounded workload shape')
}
options { timestamps(); disableConcurrentBuilds(abortPrevious: false) }
stages {
stage('Identity') {
steps {
sh '''set -eu
printf 'job=%s\n' "$JOB_NAME"
printf 'build=%s\n' "$BUILD_NUMBER"
printf 'node=%s\n' "$NODE_NAME"
printf 'workspace=%s\n' "$WORKSPACE"
printf 'mode=%s\n' "$MODE"
date -u +started_utc=%Y-%m-%dT%H:%M:%SZ
'''
}
}
stage('Bounded work') {
steps {
script {
def task = {
sh '''set -eu
mkdir -p perf-data
dd if=/dev/zero of=perf-data/payload.bin bs=1M count=8 status=none
sha256sum perf-data/payload.bin > perf-data/payload.sha256
sleep 12
printf 'bounded-complete\n'
'''
}
if (params.MODE == 'parallel2') {
parallel a: task, b: task
} else {
task()
task()
}
}
}
}
stage('Evidence') {
steps {
sh '''set -eu
du -sk . | tee workspace-kib.txt
find perf-data -maxdepth 1 -type f -printf '%f %s bytes\n' | sort > files.txt
'''
archiveArtifacts artifacts: 'workspace-kib.txt,files.txt,perf-data/*.sha256', fingerprint: true
}
}
}
}
Because the two tasks create the same bounded file path, a production-quality parallel version would allocate branch-specific directories. For this training example, Jenkins gives parallel branches distinct workspace suffixes only in some allocation patterns, so the safer exercise is to put each branch inside an explicit subdirectory if your environment shares the same workspace. Do not rely on accidental filesystem separation.
4. Capture a repeatable measurement snapshot
set -euo pipefail
JENKINS_URL='http://127.0.0.1:8080'
OUT='/tmp/jenkins-perf-lab-evidence'
TS="$(date -u +%Y%m%dT%H%M%SZ)"
# Add approved read-only lab authentication through your environment when required.
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,inQueueSince,why,task[name,url]]" > "$OUT/queue-$TS.json"
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,numExecutors,assignedLabels[name]]" > "$OUT/nodes-$TS.json"
if curl -fsS "$JENKINS_URL/prometheus/" > "$OUT/prometheus-$TS.txt"; then
printf 'prometheus_snapshot=%s\n' "$TS" >> "$OUT/timeline.txt"
fi
For controller JVM evidence, prefer the mechanisms already approved
for your lab. If you have local process access and a full JDK,
commands such as jcmd <pid> VM.info,
jcmd <pid> GC.heap_info and a thread dump can be
useful. Treat their output as potentially sensitive and do not
publish it unreviewed.
5. Experiment A — baseline: one executor, serial workload
-
Confirm
perf-agent-01has exactly one executor. -
Trigger exactly four
MODE=serialbuilds within a short interval. - Capture queue snapshots every 5–10 seconds until all builds have started.
- Record build start/end times, node name, workspace, queue reason and controller metrics.
- After completion, record artifact sizes and workspace size.
Expected shape: one build runs while later builds wait for the single eligible executor. The exact queue wait is hardware/environment dependent; the lesson does not hard-code a “correct” result.
6. Normalize the evidence into one comparison table
| Run | Executors | Mode | Median queue wait | Median build duration | Controller CPU/heap/GC | Workspace/artifact bytes | Notes |
|---|---|---|---|---|---|---|---|
| A1–A4 | 1 | serial | measure | measure | measure | measure | baseline |
| B1–B4 | 2 | serial | measure | measure | measure | measure | executor experiment |
| C1–C4 | 1 | parallel2 | measure | measure | measure | measure | Pipeline-parallel experiment |
Do not compare a warm-cache run with a cold-cache run and then attribute the difference to executors. Record cache/tooling state or recreate the same state between experiments.
7. Experiment B — change only executor count
- Preserve Experiment A evidence.
- Change
perf-agent-01from 1 executor to 2. -
Keep the Pipeline in
serialmode and trigger the same four builds. - Capture the same queue/JVM/agent/workspace measurements.
- Compare queue wait, per-build duration and controller/agent resource use.
- Return the agent to 1 executor before Experiment C.
A successful result is not “two is faster.” A successful result is a defensible conclusion such as: “queue wait fell while per-build duration and resource headroom stayed acceptable,” or “queue wait fell but agent I/O contention doubled build duration, so total completion did not improve.”
8. Experiment C — change only Pipeline parallelism
With the agent back at one executor, run the same four builds with
MODE=parallel2. Depending on how the workload allocates
node/executor context, parallel branches may compete inside one
allocation or require additional capacity. Observe what Jenkins
actually schedules; do not assume the syntax guarantees hardware
concurrency.
Record Pipeline step/stage timing, controller CPU, number of active steps, queue wait, workspace traffic and total elapsed time. Parallelism that reduces a single build’s step duration but increases platform queueing for all builds may not be an overall improvement.
9. Inspect workspace, console and artifact volume
set -euo pipefail
# Run only inside the disposable agent workspace whose identity you already recorded.
pwd
du -sk .
find . -xdev -type f -printf '%s %p\n' | sort -nr | head -n 20
# Inspect Jenkins-retained artifact sizes through the build UI/API rather than deleting files by hand.
Do not delete a live workspace to “fix performance” before proving it is the cause. Workspace cleanup changes cache state and may erase first-failure evidence; it must be an intentional experiment with before/after measurements.
10. Choose one actual bottleneck and make one change
Use your measured data—not the tutorial—to choose the next change. Examples:
| Evidence pattern | Hypothesis | One valid experiment |
|---|---|---|
| High queue wait; eligible executor continuously busy; agent resources have headroom | Insufficient execution concurrency | Add one disposable agent or one bounded executor, then rerun identical workload. |
| Low queue wait; long steps; agent disk latency high | Agent storage bottleneck | Move only the lab workspace to faster disposable storage, rerun. |
| Controller GC pauses correlate with Pipeline scheduling delay | Controller JVM/live-set pressure | Reduce Pipeline/controller workload first; if memory sizing is justified, change one JVM memory setting and remeasure. |
| Huge console output and controller disk writes dominate | Logging/evidence volume | Reduce console verbosity while retaining detailed file artifact, rerun. |
| External mock call dominates step duration | External dependency latency | Optimize/cache/retry policy at dependency boundary; more Jenkins executors may amplify load instead. |
11. Verification checklist
- Built-in node remained at 0 executors.
- All workload and resource changes were limited to disposable lab resources.
- Each experiment used the same job/source/tool versions and bounded trigger count.
- Only one intended variable changed between comparable runs.
- Queue wait was measured separately from build duration.
- Controller/JVM and agent resource evidence was captured before claiming a bottleneck.
- Workspace/log/artifact volume was recorded rather than guessed.
- No credentials, heap dump or sensitive support data was included in public artifacts.
12. Cleanup and rollback
Restore the agent to its original executor count, preserve the non-sensitive comparison report, dispose of the lab controller/agent, then remove only the exact evidence directory.
set -euo pipefail
OUT='/tmp/jenkins-perf-lab-evidence'
case "$OUT" in
/tmp/jenkins-perf-lab-evidence) rm -rf -- "$OUT" ;;
*) echo 'refusing unexpected cleanup path' >&2; exit 70 ;;
esac
13. Challenge: choose the next measurement, not the next knob
After adding a second executor, queue wait drops from 150 seconds to 30 seconds but median build duration rises from 70 seconds to 170 seconds. Controller metrics are normal; agent disk latency rises sharply. What should you do?
Do not add a third executor. The new evidence indicates agent storage contention. Restore the baseline concurrency and test a storage/workspace hypothesis while keeping the workload constant.
Knowledge check
Answer before revealing the explanation.
1. Why trigger exactly four builds rather than as many as possible?
The lab needs bounded, reproducible queue pressure—not an uncontrolled load test or denial-of-service condition.
2. Why restore the executor count before testing Pipeline parallelism?
So the experiments change one variable at a time and remain comparable.
3. What if queue wait improves but build duration becomes much worse?
The change may have shifted the bottleneck into shared agent resources; compare end-to-end completion and resource evidence before accepting it.
4. Why record workspace and artifact size?
Filesystem and evidence traffic can dominate disk/network cost and explain performance changes that queue/JVM metrics alone cannot.
5. What must be true before increasing heap?
Evidence should show memory/live-set pressure that the host can support, and you should rule out unbounded workload/plugin behavior that more heap would only mask.
Official references and version notes
Performance behavior depends on controller scale, workload shape, storage, plugins and runtime versions. Re-check current primary documentation before applying any tuning change to a real controller.
- Jenkins — Scaling Jenkins
- Jenkins — Hardware Recommendations
- Jenkins — Architecting for Scale
- Jenkins — Scaling Pipelines / durability
- Jenkins — Pipeline Best Practices
- Jenkins — Using agents
- Jenkins — Managing nodes
- Jenkins — Java Support Policy
- Metrics plugin
- Prometheus metrics plugin
- Pipeline: Groovy plugin
- Pipeline: Supporting APIs plugin
- Jenkins LTS changelog
- Jenkins Security Advisories
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.