Chapter 20Lesson 04~155 minutes

Kubernetes Plugin, Pod Templates, Dynamic Kubernetes Agents, Volumes, Service Accounts, and Cluster Scaling: Diagnostics, Failure Modes, Security, and Performance

Diagnose Kubernetes-agent failures from preserved queue/build/pod evidence. Separate Jenkins provisioning, Kubernetes admission/scheduling, Remoting, workspace/tool execution, RBAC, images, storage, and cluster capacity before changing anything.

DiagnosticsRBACPod securityPending podsImage driftEvidence

Learning objectives

  • Use a layered diagnostic sequence from Jenkins queue state through Kubernetes admission/scheduling, Remoting, workspace/tool execution, and evidence publication.
  • Diagnose excessive ServiceAccount/RBAC authority and dangerous pod security settings without normalizing them.
  • Interpret Pending/Failed/Evicted pods and separate cluster-capacity failures from Jenkins label/capacity failures.
  • Recover from pod/workspace loss without pretending ephemeral output is a retained release artifact.
  • Bound retry and scaling behavior and preserve pod UID/events/logs before repair.

1. Evidence-first diagnostic sequence

  1. Preserve Jenkins job/build/queue IDs, source SHA, console output and first-failure timestamp.
  2. Record Jenkins core/Java/Kubernetes-plugin baseline and cloud/template identity.
  3. Confirm the intended cluster context, namespace and provisioner credential reference.
  4. Inspect whether a pod was requested, admitted and scheduled; retain pod name, UID and events.
  5. Confirm agent container Java/ImageID and Remoting connection state.
  6. Inspect ServiceAccount, RBAC, workspace volume, tool container and resource requests/limits.
  7. Inspect reports/artifacts/external publications separately from workspace success.
  8. Apply the smallest correction and retry only a safe scope.

2. Failure: a pod exists but stays Pending

kubectl -n ch20-lab get pod "$POD" -o wide
kubectl -n ch20-lab describe pod "$POD"
kubectl get nodes
kubectl describe nodes | sed -n '/Allocated resources:/,/Events:/p'
kubectl -n ch20-lab get resourcequota,limitrange

Interpret the scheduler event. Insufficient cpu or Insufficient memory points to requests/capacity. An untolerated taint, node selector, PVC binding, or quota produces different evidence. Do not change Jenkins labels until the Kubernetes reason supports that conclusion.

3. Failure: broad cluster-admin ServiceAccount

The dangerous state is not subtle: a build pod can read/modify cluster-wide resources or create workloads that mount secrets/host resources. Preserve the current binding as evidence, then replace it in the disposable lab with a dedicated namespaced identity or no token at all.

kubectl auth can-i --list --as=system:serviceaccount:ch20-lab:ch20-agent -n ch20-lab
kubectl auth can-i get secrets --as=system:serviceaccount:ch20-lab:ch20-agent -n ch20-lab
kubectl get clusterrolebindings -o wide | grep -F ch20-agent || true
Do not “test” privilege by reading real Secrets. Authorization checks are sufficient. Kubernetes RBAC guidance explicitly warns that workload creation, Secrets access, privileged pods and host-level resources can become escalation paths.

4. Failure: privileged/host-mounted build pod

During review, reject pod templates that introduce broad host authority for ordinary CI. The following is an anti-pattern excerpt for recognition only, not a lab manifest:

# DO NOT USE for ordinary Jenkins build pods
securityContext:
  privileged: true
volumes:
- name: host-root
  hostPath:
    path: /

Repair by removing host access and privileged mode, then use a purpose-built isolated worker pattern if a privileged build capability is genuinely required. Separate trusted release infrastructure from untrusted PR execution.

5. Failure: the pod died before outputs left emptyDir

If the only binary or report existed in the workspace, pod deletion can make it unrecoverable. Preserve Jenkins/pod failure evidence. Re-run from the exact source/tool inputs only if rebuilding is acceptable for that workflow, but do not claim the rebuilt binary is the same artifact that was previously tested or approved.

The prevention is simple: archive/publish immutable build evidence before a risky boundary, and promote the original retained artifact rather than reconstructing it later.

6. Failure: image tag drift

kubectl -n ch20-lab get pod "$POD" \
  -o jsonpath='{range .status.containerStatuses[*]}{.name}{" requested="}{.image}{" actual="}{.imageID}{"\\n"}{end}'

If two builds requested the same mutable tag but Kubernetes reports different ImageIDs, execution inputs changed. Pin reviewed digests where practical or at minimum retain the actual ImageID/digest and treat tag changes as a controlled dependency update.

7. Failure: pod Running, Jenkins agent offline

This is usually not a Kubernetes scheduler problem. Inspect the agent container logs, Jenkins controller log, Jenkins URL, DNS/TLS route and WebSocket/TCP mode. Confirm the agent image runs a supported Java version. A Running pod only proves Kubernetes started containers; it does not prove Remoting authenticated and connected.

kubectl -n ch20-lab logs "$POD" -c jnlp --tail=200
kubectl -n ch20-lab get pod "$POD" -o jsonpath='{.status.containerStatuses[?(@.name=="jnlp")].state}'

8. Failure: stale/orphaned agent pods

The plugin has garbage collection for exceptional orphaned pods, but current documentation notes it is disabled by default because it adds Kubernetes API load. First identify why Jenkins lost ownership, preserve pod labels/UID/age/logs, then delete only confirmed lab-owned orphans. Do not make a broad namespace sweep your first troubleshooting action.

9. Failure: demand amplification and unbounded scaling

One flaky test can trigger job concurrency, pod provisioning, Pipeline retry and node autoscaling at the same time. Symptoms include many Pending pods, API throttling, node provisioning churn, startup latency and cost growth.

Bound each layer: Jenkins job concurrency, Kubernetes cloud/template caps, small retry counts, namespace ResourceQuota, realistic pod requests, and node-pool autoscaler maxima. Scaling limits are reliability controls as well as cost controls.

10. Failure: CI and production workloads share a cluster without isolation

Build code often executes arbitrary repository-controlled commands and downloads toolchains. Running it beside sensitive production workloads without namespace/RBAC/network/node isolation increases blast radius. Kubernetes namespaces alone are not a complete hard multi-tenancy boundary.

For higher-risk workloads, separate clusters or dedicated nodes/pools may be justified. At minimum apply least privilege, Pod Security controls, network policy, taints/affinity, resource quotas and explicit trust-based Jenkins routing.

11. Intentionally broken example: impossible request

In the disposable kind cluster, temporarily set an agent container request to cpu: "8" on a one-node laptop cluster. Preserve the resulting pod event showing why it cannot schedule. Then repair only the request to a realistic value such as 100m.

kubectl -n ch20-lab get pods
kubectl -n ch20-lab describe pod "$POD" | sed -n '/Events:/,$p'

This failure is valuable because it produces an explicit Kubernetes scheduling reason. Do not “fix” it by removing all resource requests; accurate requests are what let the scheduler and autoscaler reason about capacity.

Next lesson

Checkpoint Lab

Apply the complete model: bounded pod template, restrictive identity, pod-loss injection, retained evidence, safe retry, and verified cleanup.

Knowledge check

Answer before revealing the explanation.

1. A pod is Pending with Insufficient cpu. Is this a Jenkins label bug?

2. A build pod has cluster-admin. What is the correction?

3. A pod was deleted and the only compiled binary was in emptyDir. Can Jenkins recover it?

4. Why is hostPath a serious CI trust concern?

5. A tag-pinned build changes after an image update. Which layer failed?

Official references and version notes

  • Jenkins LTS changelog — baseline Jenkins 2.568.3 LTS, released 2026-09-02 and tested with Java 21 and 25; labs use Java 21 for Jenkins components.
  • Jenkins Java Support Policy — current Jenkins system components, including agents, require a supported JVM; this chapter uses Java 21.
  • Jenkins Kubernetes plugin — reviewed version 4547.v52f3080db_8cd, requires Jenkins 2.516.3, and has no current security advisory shown by the plugin health page at the chapter timestamp.
  • Kubernetes plugin Pipeline steps — podTemplate, container, pod retention, workspace volumes, and related fields.
  • kind v0.33.0 — local lab pin; its Kubernetes v1.37.0 node image is kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5.
  • Kubernetes RBAC good practices and ServiceAccounts for Pods — least privilege, namespaced roles, dedicated service accounts, and token-automount guidance.
  • Kubernetes volumes — emptyDir is pod-lifetime storage; persistent volume types have a different lifecycle.
  • Resource management and Node autoscaling — scheduling uses requests; node autoscaling is a separate control loop from Jenkins agent provisioning.
  • Jenkins inbound-agent image — reviewed agent tag 3391.va_37fa_a_305d6d-2-jdk21; record the actual architecture-specific digest/ImageID used by the pod.
Assumption timestamp: 2026-09-17. Recheck Jenkins LTS/Java, Kubernetes plugin version/dependencies/security status, kind/Kubernetes node digest, agent image identity, and Kubernetes API semantics before repeating later.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.