Performance Engineering: Executors, Queueing, Heap, Garbage Collection, Disk I/O, Workspaces, and Controller Load: Configuration, Design Choices, and Tradeoffs
Choose performance changes by ownership and evidence: capacity, JVM sizing, Pipeline structure, storage topology and retention each solve different bottlenecks and carry different reliability costs.
Learning objectives
- Compare more executors with more agents using resource and isolation evidence.
- Decide when controller heap sizing is appropriate versus reducing controller-side work.
- Evaluate external artifact storage and workspace/cache strategies.
- Use Pipeline structure/durability changes only for the bottleneck they actually affect.
- Balance cleanup frequency with diagnostic and audit retention.
1. Performance changes are architecture changes
Every tuning knob changes a resource contract. Adding executors changes concurrency and contention. Adding an agent changes scheduling capacity, isolation and operational cost. Increasing heap changes GC behavior and host memory reserve. Moving artifacts changes network, storage and recovery boundaries. Changing Pipeline durability changes crash-recovery semantics.
Therefore the design question is never just “what is faster?” It is “which layer is constrained, which component owns that layer, what side effects will this change introduce, and what evidence proves the outcome?”
2. More executors versus more agents
| Choice | Helps when | Tradeoffs / risks | Evidence required |
|---|---|---|---|
| More executors on same agent | Tasks are lightweight/wait-heavy and host has CPU/memory/I/O headroom | Shared-resource contention; weaker isolation; one host failure affects more work | host CPU/memory/I/O, per-build duration, queue wait, concurrency. |
| More agents | Eligible pool truly lacks capacity and workloads can distribute | Provisioning cost; image/tool drift; more Remoting connections; cache cold starts | queue/eligibility, agent utilization, provisioning time, controller connection load. |
| Larger agents | Individual workloads need more memory/CPU/I/O | Vertical scaling cost and larger failure domain | per-job resource profile and host saturation. |
| Ephemeral agents | Bursty workload, isolation, reproducible images | Startup latency, external scheduler/API dependence, lost local caches/logs | queue wait, provisioning latency, image pull/cache metrics, failure recovery. |
Current Jenkins guidance treats agents as the place for build execution and recommends zero controller executors. That boundary should not be sacrificed as a “performance optimization.”
3. Controller heap increase versus workload reduction
A larger heap can reduce collection frequency or prevent genuine out-of-memory conditions, but it can also lengthen some GC cycles, consume host memory reserved for filesystem cache/other processes, and mask unbounded retained objects. Before changing heap, establish the live-set trend, allocation/GC behavior, host memory/swap state and controller workload.
| Evidence | Prefer first |
|---|---|
| Large Pipeline Groovy collections / many CPS objects | Refactor Pipeline: keep Groovy as glue and move computation to agent tools. |
| Huge job/build history / logs / views | Review retention and presentation workload; do not increase heap blindly. |
| Plugin cache/object growth over time | Investigate plugin behavior/version and reproduce on test controller. |
| Healthy object profile but heap genuinely undersized for stable workload | Increase heap cautiously if host memory supports it, then remeasure GC/latency. |
| Host swapping / memory pressure | Do not increase heap; reduce process/workload pressure or add host memory first. |
4. Artifact manager/external storage versus controller disk
Jenkins archived artifacts are convenient build evidence, but very large binaries and long retention can make controller backup, restore and disk I/O expensive. External artifact repositories or artifact-manager integrations can move durable binary storage to systems built for that role. The move is not “free”: network latency, credential scope, repository availability and traceability become external dependencies.
Keep immutable artifact coordinates/digests tied to the Jenkins build. Do not “optimize” by leaving release bytes only in an agent workspace.
5. Pipeline structure optimization versus plugin change
When controller load is driven by huge dynamic Pipeline graphs, many tiny steps or controller-side Groovy computation, adding or swapping plugins may not solve the causal workload. Simplify the Pipeline first: run data processing in scripts/binaries on agents, group excessively granular shell calls where sensible, bound matrix/parallel cardinality, and preserve only necessary step-level orchestration.
Plugin upgrades may still deliver important performance fixes. Treat them as privileged controller dependency changes: read changelogs/security advisories, test representative Pipelines and define rollback. Never change multiple major plugins and Pipeline structure in the same performance experiment if you want causal evidence.
6. Durability: speed versus restart survivability
Jenkins documents three broad Pipeline durability levels. The performance-optimized mode writes execution state less frequently and can greatly reduce disk I/O; maximum survivability persists more aggressively and is the slowest. “Survivable nonatomic” sits between them.
| Workload | Typical durability direction | Reasoning |
|---|---|---|
| Re-runnable CI build/test | Performance-optimized may be acceptable after testing | If an abrupt crash loses running state, rerunning may be low risk. |
| Production deployment / critical infrastructure mutation | Favor stronger survivability | Execution history and resume behavior matter; side effects may not be safely repeatable. |
| Audit-sensitive long-running workflow | Favor stronger durability and explicit evidence | Lost Pipeline state can undermine audit/recovery expectations. |
| Pipeline with serialization bug | Fix the code |
Do not lower durability merely to avoid a
NotSerializableException.
|
Change durability on a test controller/job and verify restart semantics as part of acceptance, not just throughput.
7. Cleanup frequency versus diagnostic retention
Frequent workspace/log cleanup can reduce disk pressure and cache staleness, but aggressive deletion can increase dependency download time and erase incident evidence. Retention should be policy-driven by data class:
| Data | Retention decision |
|---|---|
| Ephemeral workspace outputs | Delete/recreate when safe; keep only needed caches/evidence. |
| Dependency/tool caches | Bound size/age; validate cache keys; measure hit rate and restore cost. |
| Build console | Retain long enough for diagnosis/audit; reduce needless verbosity at producer. |
| Test/quality reports | Retain according to trend/audit needs; externalize raw bulk when appropriate. |
| Release artifacts | Use durable repository retention/immutability policy, not workspace cleanup. |
| Thread/heap/support dumps | Short, restricted diagnostic retention; sensitive review before sharing. |
8. Worked scenario: a controller with high queue wait and rising GC
Suppose p95 queue wait is 4 minutes, every
linux-small agent is 90–100% busy, controller CPU
averages 75%, GC pauses spike during morning bursts, and Pipeline
runs create hundreds of tiny steps. The temptation is to add
executors and heap simultaneously.
A better sequence is:
- Split scheduling evidence from controller evidence.
- Test whether adding one disposable agent reduces queue wait without materially increasing controller load.
- Independently refactor one representative Pipeline to reduce tiny controller-orchestrated steps and compare controller CPU/GC.
- Only if the stable live set still requires more memory, test one heap change on the same workload.
- Accept changes separately so rollback remains possible.
9. Decision table
| Symptom + evidence | Likely owner | Candidate change | Rollback / guardrail |
|---|---|---|---|
| Queue wait high; eligible agents saturated; controller healthy | Agent capacity | Add agent capacity, not controller executor | Scale back pool; compare queue and controller overhead. |
| Agent CPU/I/O saturated after executor increase | Agent resource contention | Reduce executors or use more/larger agents | Restore original executor count. |
| Controller disk write latency high from Pipeline state | Pipeline/controller storage | Simplify Pipeline; test durability/storage change | Restore durability setting; validate restart behavior. |
| Controller disk full from workspaces/artifacts | Retention/storage ownership | Move/clean exact data class; external repository where appropriate | No broad delete; snapshot/verify before cleanup. |
| GC pause high + live set stable near heap limit | JVM sizing | Test modest heap increase with host headroom | Restore JVM flags after test; compare GC/latency. |
| External repository latency dominates | External dependency | Fix/cache/repository/network boundary | Do not add Jenkins concurrency that amplifies load. |
10. Cost and developer experience
Performance improvements consume capacity, operational complexity or durability budget. A 30-second improvement that doubles cloud-agent cost or removes crucial recovery evidence may be a poor trade. Likewise, extremely conservative durability and retention can make feedback unnecessarily slow. Tie every platform choice to an explicit service objective and owner.
Knowledge check
Answer before revealing the explanation.
1. When is adding an executor safer than adding an agent?
When tasks are light/wait-heavy and the existing agent has measured CPU/memory/I/O headroom; still verify per-build and queue effects.
2. Why can a larger heap make diagnosis harder?
It can postpone symptoms of leaks/unbounded retained objects and consume host memory without fixing the causal workload.
3. What is the release-artifact rule when moving storage away from Jenkins?
Preserve immutable artifact coordinates/digests tied to the exact Jenkins build; do not rely on workspace files.
4. When is performance-optimized Pipeline durability inappropriate as a first choice?
When crash survivability/audit of critical side effects matters, or when it is being used to hide serialization problems.
5. Why should cleanup policy distinguish caches, logs, workspaces and release artifacts?
They have different owners, recovery value, durability requirements and performance effects; one blanket deletion policy is unsafe.
Official references and version notes
Performance behavior depends on controller scale, workload shape, storage, plugins and runtime versions. Re-check current primary documentation before applying any tuning change to a real controller.
- Jenkins — Scaling Jenkins
- Jenkins — Hardware Recommendations
- Jenkins — Architecting for Scale
- Jenkins — Scaling Pipelines / durability
- Jenkins — Pipeline Best Practices
- Jenkins — Using agents
- Jenkins — Managing nodes
- Jenkins — Java Support Policy
- Metrics plugin
- Prometheus metrics plugin
- Pipeline: Groovy plugin
- Pipeline: Supporting APIs plugin
- Jenkins LTS changelog
- Jenkins Security Advisories
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.