Kubernetes Plugin, Pod Templates, Dynamic Kubernetes Agents, Volumes, Service Accounts, and Cluster Scaling: Diagnostics, Failure Modes, Security, and Performance
Diagnose Kubernetes-agent failures from preserved queue/build/pod evidence. Separate Jenkins provisioning, Kubernetes admission/scheduling, Remoting, workspace/tool execution, RBAC, images, storage, and cluster capacity before changing anything.
Learning objectives
- Use a layered diagnostic sequence from Jenkins queue state through Kubernetes admission/scheduling, Remoting, workspace/tool execution, and evidence publication.
- Diagnose excessive ServiceAccount/RBAC authority and dangerous pod security settings without normalizing them.
- Interpret Pending/Failed/Evicted pods and separate cluster-capacity failures from Jenkins label/capacity failures.
- Recover from pod/workspace loss without pretending ephemeral output is a retained release artifact.
- Bound retry and scaling behavior and preserve pod UID/events/logs before repair.
1. Evidence-first diagnostic sequence
- Preserve Jenkins job/build/queue IDs, source SHA, console output and first-failure timestamp.
- Record Jenkins core/Java/Kubernetes-plugin baseline and cloud/template identity.
- Confirm the intended cluster context, namespace and provisioner credential reference.
- Inspect whether a pod was requested, admitted and scheduled; retain pod name, UID and events.
- Confirm agent container Java/ImageID and Remoting connection state.
- Inspect ServiceAccount, RBAC, workspace volume, tool container and resource requests/limits.
- Inspect reports/artifacts/external publications separately from workspace success.
- Apply the smallest correction and retry only a safe scope.
2. Failure: a pod exists but stays Pending
kubectl -n ch20-lab get pod "$POD" -o wide
kubectl -n ch20-lab describe pod "$POD"
kubectl get nodes
kubectl describe nodes | sed -n '/Allocated resources:/,/Events:/p'
kubectl -n ch20-lab get resourcequota,limitrange
Interpret the scheduler event. Insufficient cpu or
Insufficient memory points to requests/capacity. An
untolerated taint, node selector, PVC binding, or quota produces
different evidence. Do not change Jenkins labels until the
Kubernetes reason supports that conclusion.
3. Failure: broad cluster-admin ServiceAccount
The dangerous state is not subtle: a build pod can read/modify cluster-wide resources or create workloads that mount secrets/host resources. Preserve the current binding as evidence, then replace it in the disposable lab with a dedicated namespaced identity or no token at all.
kubectl auth can-i --list --as=system:serviceaccount:ch20-lab:ch20-agent -n ch20-lab
kubectl auth can-i get secrets --as=system:serviceaccount:ch20-lab:ch20-agent -n ch20-lab
kubectl get clusterrolebindings -o wide | grep -F ch20-agent || true
4. Failure: privileged/host-mounted build pod
During review, reject pod templates that introduce broad host authority for ordinary CI. The following is an anti-pattern excerpt for recognition only, not a lab manifest:
# DO NOT USE for ordinary Jenkins build pods
securityContext:
privileged: true
volumes:
- name: host-root
hostPath:
path: /
Repair by removing host access and privileged mode, then use a purpose-built isolated worker pattern if a privileged build capability is genuinely required. Separate trusted release infrastructure from untrusted PR execution.
5. Failure: the pod died before outputs left emptyDir
If the only binary or report existed in the workspace, pod deletion can make it unrecoverable. Preserve Jenkins/pod failure evidence. Re-run from the exact source/tool inputs only if rebuilding is acceptable for that workflow, but do not claim the rebuilt binary is the same artifact that was previously tested or approved.
The prevention is simple: archive/publish immutable build evidence before a risky boundary, and promote the original retained artifact rather than reconstructing it later.
6. Failure: image tag drift
kubectl -n ch20-lab get pod "$POD" \
-o jsonpath='{range .status.containerStatuses[*]}{.name}{" requested="}{.image}{" actual="}{.imageID}{"\\n"}{end}'
If two builds requested the same mutable tag but Kubernetes reports different ImageIDs, execution inputs changed. Pin reviewed digests where practical or at minimum retain the actual ImageID/digest and treat tag changes as a controlled dependency update.
7. Failure: pod Running, Jenkins agent offline
This is usually not a Kubernetes scheduler problem. Inspect the agent container logs, Jenkins controller log, Jenkins URL, DNS/TLS route and WebSocket/TCP mode. Confirm the agent image runs a supported Java version. A Running pod only proves Kubernetes started containers; it does not prove Remoting authenticated and connected.
kubectl -n ch20-lab logs "$POD" -c jnlp --tail=200
kubectl -n ch20-lab get pod "$POD" -o jsonpath='{.status.containerStatuses[?(@.name=="jnlp")].state}'
8. Failure: stale/orphaned agent pods
The plugin has garbage collection for exceptional orphaned pods, but current documentation notes it is disabled by default because it adds Kubernetes API load. First identify why Jenkins lost ownership, preserve pod labels/UID/age/logs, then delete only confirmed lab-owned orphans. Do not make a broad namespace sweep your first troubleshooting action.
9. Failure: demand amplification and unbounded scaling
One flaky test can trigger job concurrency, pod provisioning, Pipeline retry and node autoscaling at the same time. Symptoms include many Pending pods, API throttling, node provisioning churn, startup latency and cost growth.
Bound each layer: Jenkins job concurrency, Kubernetes cloud/template caps, small retry counts, namespace ResourceQuota, realistic pod requests, and node-pool autoscaler maxima. Scaling limits are reliability controls as well as cost controls.
10. Failure: CI and production workloads share a cluster without isolation
Build code often executes arbitrary repository-controlled commands and downloads toolchains. Running it beside sensitive production workloads without namespace/RBAC/network/node isolation increases blast radius. Kubernetes namespaces alone are not a complete hard multi-tenancy boundary.
For higher-risk workloads, separate clusters or dedicated nodes/pools may be justified. At minimum apply least privilege, Pod Security controls, network policy, taints/affinity, resource quotas and explicit trust-based Jenkins routing.
11. Intentionally broken example: impossible request
In the disposable kind cluster, temporarily set an agent container
request to cpu: "8" on a one-node laptop cluster.
Preserve the resulting pod event showing why it cannot schedule.
Then repair only the request to a realistic value such as
100m.
kubectl -n ch20-lab get pods
kubectl -n ch20-lab describe pod "$POD" | sed -n '/Events:/,$p'
This failure is valuable because it produces an explicit Kubernetes scheduling reason. Do not “fix” it by removing all resource requests; accurate requests are what let the scheduler and autoscaler reason about capacity.
Knowledge check
Answer before revealing the explanation.
1. A pod is Pending with Insufficient cpu. Is this a Jenkins label bug?
Not necessarily. Jenkins already requested the pod; Kubernetes scheduling cannot satisfy its resource request. Inspect pod events, node allocatable resources, quotas, affinities/taints, and autoscaler state.
2. A build pod has cluster-admin. What is the correction?
Replace it with a dedicated namespaced ServiceAccount and the minimum Role/RoleBinding required—or no API token at all for ordinary build code. Do not rely on container isolation to contain cluster-admin credentials.
3. A pod was deleted and the only compiled binary was in emptyDir. Can Jenkins recover it?
Not from that workspace. The evidence/export design failed. Reproduce from source if appropriate, but do not claim the lost binary is the same release artifact; future runs must archive/publish before pod loss.
4. Why is hostPath a serious CI trust concern?
It can expose node filesystem state directly into a build pod and undermine workload isolation. Kubernetes RBAC guidance also treats workload and host access as privilege-escalation risk.
5. A tag-pinned build changes after an image update. Which layer failed?
Image identity/reproducibility failed. Record and use immutable digests or a controlled image-promotion policy rather than assuming a tag is stable.
Official references and version notes
-
Jenkins LTS changelog
— baseline
Jenkins 2.568.3 LTS, released 2026-09-02 and tested with Java 21 and 25; labs use Java 21 for Jenkins components. - Jenkins Java Support Policy — current Jenkins system components, including agents, require a supported JVM; this chapter uses Java 21.
-
Jenkins Kubernetes plugin
— reviewed version
4547.v52f3080db_8cd, requires Jenkins2.516.3, and has no current security advisory shown by the plugin health page at the chapter timestamp. -
Kubernetes plugin Pipeline steps
—
podTemplate,container, pod retention, workspace volumes, and related fields. -
kind v0.33.0
— local lab pin; its Kubernetes v1.37.0 node image is
kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5. - Kubernetes RBAC good practices and ServiceAccounts for Pods — least privilege, namespaced roles, dedicated service accounts, and token-automount guidance.
-
Kubernetes volumes
—
emptyDiris pod-lifetime storage; persistent volume types have a different lifecycle. - Resource management and Node autoscaling — scheduling uses requests; node autoscaling is a separate control loop from Jenkins agent provisioning.
-
Jenkins inbound-agent image
— reviewed agent tag
3391.va_37fa_a_305d6d-2-jdk21; record the actual architecture-specific digest/ImageID used by the pod.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.