Chapter 38Lesson 03~210 minutes

Performance Engineering: Executors, Queueing, Heap, Garbage Collection, Disk I/O, Workspaces, and Controller Load: Configuration, Design Choices, and Tradeoffs

Choose performance changes by ownership and evidence: capacity, JVM sizing, Pipeline structure, storage topology and retention each solve different bottlenecks and carry different reliability costs.

designcapacityheapartifact storagedurabilityretention

Learning objectives

  • Compare more executors with more agents using resource and isolation evidence.
  • Decide when controller heap sizing is appropriate versus reducing controller-side work.
  • Evaluate external artifact storage and workspace/cache strategies.
  • Use Pipeline structure/durability changes only for the bottleneck they actually affect.
  • Balance cleanup frequency with diagnostic and audit retention.

1. Performance changes are architecture changes

Every tuning knob changes a resource contract. Adding executors changes concurrency and contention. Adding an agent changes scheduling capacity, isolation and operational cost. Increasing heap changes GC behavior and host memory reserve. Moving artifacts changes network, storage and recovery boundaries. Changing Pipeline durability changes crash-recovery semantics.

Therefore the design question is never just “what is faster?” It is “which layer is constrained, which component owns that layer, what side effects will this change introduce, and what evidence proves the outcome?”

2. More executors versus more agents

Choice Helps when Tradeoffs / risks Evidence required
More executors on same agent Tasks are lightweight/wait-heavy and host has CPU/memory/I/O headroom Shared-resource contention; weaker isolation; one host failure affects more work host CPU/memory/I/O, per-build duration, queue wait, concurrency.
More agents Eligible pool truly lacks capacity and workloads can distribute Provisioning cost; image/tool drift; more Remoting connections; cache cold starts queue/eligibility, agent utilization, provisioning time, controller connection load.
Larger agents Individual workloads need more memory/CPU/I/O Vertical scaling cost and larger failure domain per-job resource profile and host saturation.
Ephemeral agents Bursty workload, isolation, reproducible images Startup latency, external scheduler/API dependence, lost local caches/logs queue wait, provisioning latency, image pull/cache metrics, failure recovery.

Current Jenkins guidance treats agents as the place for build execution and recommends zero controller executors. That boundary should not be sacrificed as a “performance optimization.”

3. Controller heap increase versus workload reduction

A larger heap can reduce collection frequency or prevent genuine out-of-memory conditions, but it can also lengthen some GC cycles, consume host memory reserved for filesystem cache/other processes, and mask unbounded retained objects. Before changing heap, establish the live-set trend, allocation/GC behavior, host memory/swap state and controller workload.

Evidence Prefer first
Large Pipeline Groovy collections / many CPS objects Refactor Pipeline: keep Groovy as glue and move computation to agent tools.
Huge job/build history / logs / views Review retention and presentation workload; do not increase heap blindly.
Plugin cache/object growth over time Investigate plugin behavior/version and reproduce on test controller.
Healthy object profile but heap genuinely undersized for stable workload Increase heap cautiously if host memory supports it, then remeasure GC/latency.
Host swapping / memory pressure Do not increase heap; reduce process/workload pressure or add host memory first.

4. Artifact manager/external storage versus controller disk

Jenkins archived artifacts are convenient build evidence, but very large binaries and long retention can make controller backup, restore and disk I/O expensive. External artifact repositories or artifact-manager integrations can move durable binary storage to systems built for that role. The move is not “free”: network latency, credential scope, repository availability and traceability become external dependencies.

Keep immutable artifact coordinates/digests tied to the Jenkins build. Do not “optimize” by leaving release bytes only in an agent workspace.

5. Pipeline structure optimization versus plugin change

When controller load is driven by huge dynamic Pipeline graphs, many tiny steps or controller-side Groovy computation, adding or swapping plugins may not solve the causal workload. Simplify the Pipeline first: run data processing in scripts/binaries on agents, group excessively granular shell calls where sensible, bound matrix/parallel cardinality, and preserve only necessary step-level orchestration.

Plugin upgrades may still deliver important performance fixes. Treat them as privileged controller dependency changes: read changelogs/security advisories, test representative Pipelines and define rollback. Never change multiple major plugins and Pipeline structure in the same performance experiment if you want causal evidence.

6. Durability: speed versus restart survivability

Jenkins documents three broad Pipeline durability levels. The performance-optimized mode writes execution state less frequently and can greatly reduce disk I/O; maximum survivability persists more aggressively and is the slowest. “Survivable nonatomic” sits between them.

Workload Typical durability direction Reasoning
Re-runnable CI build/test Performance-optimized may be acceptable after testing If an abrupt crash loses running state, rerunning may be low risk.
Production deployment / critical infrastructure mutation Favor stronger survivability Execution history and resume behavior matter; side effects may not be safely repeatable.
Audit-sensitive long-running workflow Favor stronger durability and explicit evidence Lost Pipeline state can undermine audit/recovery expectations.
Pipeline with serialization bug Fix the code Do not lower durability merely to avoid a NotSerializableException.

Change durability on a test controller/job and verify restart semantics as part of acceptance, not just throughput.

7. Cleanup frequency versus diagnostic retention

Frequent workspace/log cleanup can reduce disk pressure and cache staleness, but aggressive deletion can increase dependency download time and erase incident evidence. Retention should be policy-driven by data class:

Data Retention decision
Ephemeral workspace outputs Delete/recreate when safe; keep only needed caches/evidence.
Dependency/tool caches Bound size/age; validate cache keys; measure hit rate and restore cost.
Build console Retain long enough for diagnosis/audit; reduce needless verbosity at producer.
Test/quality reports Retain according to trend/audit needs; externalize raw bulk when appropriate.
Release artifacts Use durable repository retention/immutability policy, not workspace cleanup.
Thread/heap/support dumps Short, restricted diagnostic retention; sensitive review before sharing.

8. Worked scenario: a controller with high queue wait and rising GC

Suppose p95 queue wait is 4 minutes, every linux-small agent is 90–100% busy, controller CPU averages 75%, GC pauses spike during morning bursts, and Pipeline runs create hundreds of tiny steps. The temptation is to add executors and heap simultaneously.

A better sequence is:

  1. Split scheduling evidence from controller evidence.
  2. Test whether adding one disposable agent reduces queue wait without materially increasing controller load.
  3. Independently refactor one representative Pipeline to reduce tiny controller-orchestrated steps and compare controller CPU/GC.
  4. Only if the stable live set still requires more memory, test one heap change on the same workload.
  5. Accept changes separately so rollback remains possible.

9. Decision table

Symptom + evidence Likely owner Candidate change Rollback / guardrail
Queue wait high; eligible agents saturated; controller healthy Agent capacity Add agent capacity, not controller executor Scale back pool; compare queue and controller overhead.
Agent CPU/I/O saturated after executor increase Agent resource contention Reduce executors or use more/larger agents Restore original executor count.
Controller disk write latency high from Pipeline state Pipeline/controller storage Simplify Pipeline; test durability/storage change Restore durability setting; validate restart behavior.
Controller disk full from workspaces/artifacts Retention/storage ownership Move/clean exact data class; external repository where appropriate No broad delete; snapshot/verify before cleanup.
GC pause high + live set stable near heap limit JVM sizing Test modest heap increase with host headroom Restore JVM flags after test; compare GC/latency.
External repository latency dominates External dependency Fix/cache/repository/network boundary Do not add Jenkins concurrency that amplifies load.

10. Cost and developer experience

Performance improvements consume capacity, operational complexity or durability budget. A 30-second improvement that doubles cloud-agent cost or removes crucial recovery evidence may be a poor trade. Likewise, extremely conservative durability and retention can make feedback unnecessarily slow. Tie every platform choice to an explicit service objective and owner.

Next

Diagnose performance failures causally

Lesson 4 engineers common anti-patterns—too many executors, giant logs, controller-side work, giant CPS state, workspace leaks, heap masking and excessive parallelism—and repairs them without erasing first-failure evidence.

Knowledge check

Answer before revealing the explanation.

1. When is adding an executor safer than adding an agent?

2. Why can a larger heap make diagnosis harder?

3. What is the release-artifact rule when moving storage away from Jenkins?

4. When is performance-optimized Pipeline durability inappropriate as a first choice?

5. Why should cleanup policy distinguish caches, logs, workspaces and release artifacts?

Official references and version notes

Performance behavior depends on controller scale, workload shape, storage, plugins and runtime versions. Re-check current primary documentation before applying any tuning change to a real controller.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.