Chapter 38Lesson 02~230 minutes

Performance Engineering: Executors, Queueing, Heap, Garbage Collection, Disk I/O, Workspaces, and Controller Load: Guided Hands-On Workflow and Core Operations

Run a bounded local performance experiment: capture a baseline, generate repeatable queue pressure and I/O, vary one concurrency variable, and compare the same evidence before deciding whether the change helped.

hands-onbaselinesynthetic workloadmetricsexperimentscleanup

Learning objectives

  • Create a reproducible disposable performance lab without using production repositories or infrastructure.
  • Measure queue wait, executor state, controller/JVM signals and workspace/log/artifact volume.
  • Change executor count and Pipeline parallelism in controlled experiments.
  • Separate throughput improvement from per-build latency and controller cost.
  • Preserve an evidence packet and clean up only exact lab resources.

1. Scenario and safety boundary

You operate a disposable controller called jenkins-perf-lab with the built-in node set to 0 executors. One disposable Linux agent is labeled perf-linux. The job perf-lab/bounded-workload performs synthetic CPU-light work, bounded filesystem work, modest log output and a simulated external wait. Nothing points at production systems.

2. Preflight and record the baseline

Item Lab requirement
Controller Jenkins LTS 2.568.3; built-in node executors = 0.
Runtime Java 21 for controller/agent JVMs.
Agent Disposable perf-agent-01, label perf-linux, initially 1 executor.
Metrics Remote API required; Metrics/Prometheus optional but pinned if installed.
Source Synthetic local Git repository or Pipeline job; record exact Jenkinsfile revision.
External service Local mock/sleep only; no production API.
Evidence path /tmp/jenkins-perf-lab-evidence on the operator host.
Bounds At most 4 queued builds, 2 executors, 2 parallel branches; no unbounded loops.
set -euo pipefail
OUT='/tmp/jenkins-perf-lab-evidence'
mkdir -p "$OUT"
date -u +baseline_utc=%Y-%m-%dT%H:%M:%SZ | tee "$OUT/baseline.txt"
java -version 2>&1 | tee -a "$OUT/baseline.txt"
printf 'jenkins_lts=2.568.3\nagent=perf-agent-01\nlabel=perf-linux\nexecutors=1\n' | tee -a "$OUT/baseline.txt"

3. Use a bounded, parameterized synthetic Pipeline

The Pipeline below keeps the work intentionally small. The point is not to benchmark hardware; it is to learn measurement discipline.

pipeline {
  agent { label 'perf-linux' }
  parameters {
    choice(name: 'MODE', choices: ['serial', 'parallel2'], description: 'bounded workload shape')
  }
  options { timestamps(); disableConcurrentBuilds(abortPrevious: false) }
  stages {
    stage('Identity') {
      steps {
        sh '''set -eu
          printf 'job=%s\n' "$JOB_NAME"
          printf 'build=%s\n' "$BUILD_NUMBER"
          printf 'node=%s\n' "$NODE_NAME"
          printf 'workspace=%s\n' "$WORKSPACE"
          printf 'mode=%s\n' "$MODE"
          date -u +started_utc=%Y-%m-%dT%H:%M:%SZ
        '''
      }
    }
    stage('Bounded work') {
      steps {
        script {
          def task = {
            sh '''set -eu
              mkdir -p perf-data
              dd if=/dev/zero of=perf-data/payload.bin bs=1M count=8 status=none
              sha256sum perf-data/payload.bin > perf-data/payload.sha256
              sleep 12
              printf 'bounded-complete\n'
            '''
          }
          if (params.MODE == 'parallel2') {
            parallel a: task, b: task
          } else {
            task()
            task()
          }
        }
      }
    }
    stage('Evidence') {
      steps {
        sh '''set -eu
          du -sk . | tee workspace-kib.txt
          find perf-data -maxdepth 1 -type f -printf '%f %s bytes\n' | sort > files.txt
        '''
        archiveArtifacts artifacts: 'workspace-kib.txt,files.txt,perf-data/*.sha256', fingerprint: true
      }
    }
  }
}

Because the two tasks create the same bounded file path, a production-quality parallel version would allocate branch-specific directories. For this training example, Jenkins gives parallel branches distinct workspace suffixes only in some allocation patterns, so the safer exercise is to put each branch inside an explicit subdirectory if your environment shares the same workspace. Do not rely on accidental filesystem separation.

4. Capture a repeatable measurement snapshot

set -euo pipefail
JENKINS_URL='http://127.0.0.1:8080'
OUT='/tmp/jenkins-perf-lab-evidence'
TS="$(date -u +%Y%m%dT%H%M%SZ)"
# Add approved read-only lab authentication through your environment when required.
curl -fsS "$JENKINS_URL/queue/api/json?tree=items[id,inQueueSince,why,task[name,url]]" > "$OUT/queue-$TS.json"
curl -fsS "$JENKINS_URL/computer/api/json?tree=computer[displayName,offline,numExecutors,assignedLabels[name]]" > "$OUT/nodes-$TS.json"
if curl -fsS "$JENKINS_URL/prometheus/" > "$OUT/prometheus-$TS.txt"; then
  printf 'prometheus_snapshot=%s\n' "$TS" >> "$OUT/timeline.txt"
fi

For controller JVM evidence, prefer the mechanisms already approved for your lab. If you have local process access and a full JDK, commands such as jcmd <pid> VM.info, jcmd <pid> GC.heap_info and a thread dump can be useful. Treat their output as potentially sensitive and do not publish it unreviewed.

5. Experiment A — baseline: one executor, serial workload

  1. Confirm perf-agent-01 has exactly one executor.
  2. Trigger exactly four MODE=serial builds within a short interval.
  3. Capture queue snapshots every 5–10 seconds until all builds have started.
  4. Record build start/end times, node name, workspace, queue reason and controller metrics.
  5. After completion, record artifact sizes and workspace size.

Expected shape: one build runs while later builds wait for the single eligible executor. The exact queue wait is hardware/environment dependent; the lesson does not hard-code a “correct” result.

6. Normalize the evidence into one comparison table

Run Executors Mode Median queue wait Median build duration Controller CPU/heap/GC Workspace/artifact bytes Notes
A1–A4 1 serial measure measure measure measure baseline
B1–B4 2 serial measure measure measure measure executor experiment
C1–C4 1 parallel2 measure measure measure measure Pipeline-parallel experiment

Do not compare a warm-cache run with a cold-cache run and then attribute the difference to executors. Record cache/tooling state or recreate the same state between experiments.

7. Experiment B — change only executor count

  1. Preserve Experiment A evidence.
  2. Change perf-agent-01 from 1 executor to 2.
  3. Keep the Pipeline in serial mode and trigger the same four builds.
  4. Capture the same queue/JVM/agent/workspace measurements.
  5. Compare queue wait, per-build duration and controller/agent resource use.
  6. Return the agent to 1 executor before Experiment C.

A successful result is not “two is faster.” A successful result is a defensible conclusion such as: “queue wait fell while per-build duration and resource headroom stayed acceptable,” or “queue wait fell but agent I/O contention doubled build duration, so total completion did not improve.”

8. Experiment C — change only Pipeline parallelism

With the agent back at one executor, run the same four builds with MODE=parallel2. Depending on how the workload allocates node/executor context, parallel branches may compete inside one allocation or require additional capacity. Observe what Jenkins actually schedules; do not assume the syntax guarantees hardware concurrency.

Record Pipeline step/stage timing, controller CPU, number of active steps, queue wait, workspace traffic and total elapsed time. Parallelism that reduces a single build’s step duration but increases platform queueing for all builds may not be an overall improvement.

9. Inspect workspace, console and artifact volume

set -euo pipefail
# Run only inside the disposable agent workspace whose identity you already recorded.
pwd
du -sk .
find . -xdev -type f -printf '%s %p\n' | sort -nr | head -n 20
# Inspect Jenkins-retained artifact sizes through the build UI/API rather than deleting files by hand.

Do not delete a live workspace to “fix performance” before proving it is the cause. Workspace cleanup changes cache state and may erase first-failure evidence; it must be an intentional experiment with before/after measurements.

10. Choose one actual bottleneck and make one change

Use your measured data—not the tutorial—to choose the next change. Examples:

Evidence pattern Hypothesis One valid experiment
High queue wait; eligible executor continuously busy; agent resources have headroom Insufficient execution concurrency Add one disposable agent or one bounded executor, then rerun identical workload.
Low queue wait; long steps; agent disk latency high Agent storage bottleneck Move only the lab workspace to faster disposable storage, rerun.
Controller GC pauses correlate with Pipeline scheduling delay Controller JVM/live-set pressure Reduce Pipeline/controller workload first; if memory sizing is justified, change one JVM memory setting and remeasure.
Huge console output and controller disk writes dominate Logging/evidence volume Reduce console verbosity while retaining detailed file artifact, rerun.
External mock call dominates step duration External dependency latency Optimize/cache/retry policy at dependency boundary; more Jenkins executors may amplify load instead.

11. Verification checklist

  • Built-in node remained at 0 executors.
  • All workload and resource changes were limited to disposable lab resources.
  • Each experiment used the same job/source/tool versions and bounded trigger count.
  • Only one intended variable changed between comparable runs.
  • Queue wait was measured separately from build duration.
  • Controller/JVM and agent resource evidence was captured before claiming a bottleneck.
  • Workspace/log/artifact volume was recorded rather than guessed.
  • No credentials, heap dump or sensitive support data was included in public artifacts.

12. Cleanup and rollback

Restore the agent to its original executor count, preserve the non-sensitive comparison report, dispose of the lab controller/agent, then remove only the exact evidence directory.

set -euo pipefail
OUT='/tmp/jenkins-perf-lab-evidence'
case "$OUT" in
  /tmp/jenkins-perf-lab-evidence) rm -rf -- "$OUT" ;;
  *) echo 'refusing unexpected cleanup path' >&2; exit 70 ;;
esac

13. Challenge: choose the next measurement, not the next knob

After adding a second executor, queue wait drops from 150 seconds to 30 seconds but median build duration rises from 70 seconds to 170 seconds. Controller metrics are normal; agent disk latency rises sharply. What should you do?

Do not add a third executor. The new evidence indicates agent storage contention. Restore the baseline concurrency and test a storage/workspace hypothesis while keeping the workload constant.

Next

Choose tuning strategies deliberately

Lesson 3 compares more executors versus more agents, heap versus workload reduction, external artifact storage, Pipeline optimization, plugin changes and cleanup/retention tradeoffs.

Knowledge check

Answer before revealing the explanation.

1. Why trigger exactly four builds rather than as many as possible?

2. Why restore the executor count before testing Pipeline parallelism?

3. What if queue wait improves but build duration becomes much worse?

4. Why record workspace and artifact size?

5. What must be true before increasing heap?

Official references and version notes

Performance behavior depends on controller scale, workload shape, storage, plugins and runtime versions. Re-check current primary documentation before applying any tuning change to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.