Checkpoint Lab — Kubernetes Plugin, Pod Templates, Dynamic Kubernetes Agents, Volumes, Service Accounts, and Cluster Scaling
Run a bounded dynamic-agent workflow on a local Kubernetes cluster, use restrictive resources and identity, inject pod loss, recover only from retained evidence, and verify that every build-owned pod and ephemeral volume is gone afterward.
Learning objectives
- Run one SCM-backed synthetic Pipeline on a dynamic Kubernetes agent with an explicit template, resources, images and ServiceAccount.
- Predict and verify Jenkins-build identity versus Kubernetes pod identity across pod loss.
- Prove workspace-only state disappears while archived evidence remains.
- Recover through a bounded Kubernetes-agent retry without repeating an irreversible side effect.
- Produce a reviewable evidence packet and prove complete pod/ephemeral-volume teardown.
1. Scenario and acceptance criteria
You operate a disposable Jenkins controller and a local kind
cluster. Job jenkins/ch20-k8s-checkpoint runs only
synthetic work. A dynamic agent pod records
source/build/pod/resource evidence, archives a pre-loss record,
waits for an intentional pod deletion, then recovers on a fresh pod
using a bounded infrastructure retry.
The lab passes only if the same Jenkins build remains attributable to exact source while the pod name/UID changes, workspace-only state from the first pod is gone, archived evidence remains, and the namespace contains no build-owned pods/PVCs after completion.
2. Pinned assumptions
| Component | Lab baseline | Verification |
|---|---|---|
| Jenkins | 2.568.3 LTS, Java 21 |
Manage Jenkins → System Information /
java -version
|
| Kubernetes plugin | 4547.v52f3080db_8cd | Plugin inventory |
| kind | v0.33.0 | kind version |
| Kubernetes node | v1.37.0 kind image digest a1ed56…580ae5 |
docker inspect / kind create command |
| Agent image |
jenkins/inbound-agent:3391.va_37fa_a_305d6d-2-jdk21
|
Pod imageID; record full runtime digest |
| Namespace | ch20-lab |
kubectl get ns |
3. Preflight and guardrails
set -eu
CTX=$(kubectl config current-context)
printf 'context=%s\\n' "$CTX"
case "$CTX" in kind-ch20) ;; *) echo 'Refusing: not the ch20 disposable cluster' >&2; exit 2;; esac
kubectl get nodes
kubectl -n ch20-lab get serviceaccount ch20-agent
kubectl -n ch20-lab auth can-i get pods --as=system:serviceaccount:ch20-lab:ch20-agent
The final command should return no. Also verify the
Jenkins built-in node has zero executors, the cloud is restricted to
the lab folder where supported, and its cap is small.
4. Predict before execution
Write these predictions into the build description or a checked-in lab note before starting:
-
The first attempt creates one pod with a unique UID and an
emptyDir-backed workspace. Deleting that pod destroys that workspace. - The archived pre-loss evidence is copied out of the workspace before deletion and remains attached to the Jenkins build.
-
A qualifying Kubernetes-agent retry may create a second pod with a
different name/UID while the Jenkins
BUILD_NUMBERremains unchanged. - No production/external side effect occurs, so repeating the synthetic attempt is safe.
5. Synthetic SCM content
Commit a tiny payload.txt,
ci/ch20-pod.yaml, and Jenkinsfile to a disposable
repository so GIT_COMMIT is available. The pod YAML:
apiVersion: v1
kind: Pod
metadata:
labels:
devops-academy: ch20
spec:
serviceAccountName: ch20-agent
automountServiceAccountToken: false
securityContext:
runAsNonRoot: true
runAsUser: 1000
runAsGroup: 1000
fsGroup: 1000
containers:
- name: jnlp
image: jenkins/inbound-agent:3391.va_37fa_a_305d6d-2-jdk21
resources:
requests:
cpu: 100m
memory: 192Mi
limits:
cpu: 500m
memory: 512Mi
Do not add privileged mode, hostPath, host networking or broad Kubernetes credentials.
6. Checkpoint Jenkinsfile
pipeline {
agent none
options {
timeout(time: 20, unit: 'MINUTES')
timestamps()
}
stages {
stage('Kubernetes attempt') {
steps {
script {
podTemplate(
cloud: 'ch20-kind',
namespace: 'ch20-lab',
serviceAccount: 'ch20-agent',
podRetention: never(),
yaml: readTrusted('ci/ch20-pod.yaml')
) {
retry(count: 2, conditions: [kubernetesAgent(), nonresumable()]) {
node(POD_LABEL) {
checkout scm
sh '''
set -eu
POD=$(hostname)
printf 'job=%s\\nbuild=%s\\nsource=%s\\nnode=%s\\npod=%s\\nworkspace=%s\\n' \
"$JOB_NAME" "$BUILD_NUMBER" "$GIT_COMMIT" "$NODE_NAME" "$POD" "$WORKSPACE" \
| tee "retained-${POD}.txt"
printf 'pod-local-only=%s\\n' "$POD" > workspace-only.txt
sha256sum payload.txt | tee -a "retained-${POD}.txt"
'''
archiveArtifacts artifacts: 'retained-*.txt', fingerprint: true
echo 'Preserve pod evidence now, then delete ONLY this lab pod from a separate terminal.'
sleep 90
sh '''
set -eu
test -f payload.txt
printf 'post-wait-pod=%s\\n' "$(hostname)" | tee post-wait.txt
'''
archiveArtifacts artifacts: 'post-wait.txt', fingerprint: true
}
}
}
}
}
}
}
post {
always {
echo "build=${env.BUILD_URL} result=${currentBuild.currentResult}"
}
}
}
The retry scope contains only source checkout, hashing, local files and artifact upload. It deliberately excludes deployment, signing, package publication, registry promotion or destructive external API calls.
7. Preserve evidence, then inject pod loss
While the Pipeline is in sleep 90, use a separate
terminal:
set -eu
kubectl -n ch20-lab get pods -o wide
POD=$(kubectl -n ch20-lab get pods -l devops-academy=ch20 \
--sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].metadata.name}')
test -n "$POD"
kubectl -n ch20-lab get pod "$POD" -o yaml > "/tmp/${POD}.yaml"
kubectl -n ch20-lab describe pod "$POD" > "/tmp/${POD}.describe.txt"
kubectl -n ch20-lab get pod "$POD" -o jsonpath='uid={.metadata.uid}{"\\n"}node={.spec.nodeName}{"\\n"}sa={.spec.serviceAccountName}{"\\n"}'
kubectl -n ch20-lab delete pod "$POD" --wait=false
Keep the console output showing the original pod name. The retry
should not be forced if the failure is a normal test error;
kubernetesAgent() is specifically for qualifying
agent/pod infrastructure loss.
8. Verify recovery independently
After Jenkins provisions another agent, record the new pod name/UID
and compare it with the first. Confirm the build number/source SHA
remain the same. Then inspect the Jenkins artifacts: at least one
retained-<pod>.txt from before failure should
exist even though the first workspace is gone.
kubectl -n ch20-lab get pods -o wide
kubectl -n ch20-lab get events --sort-by=.lastTimestamp | tail -60
If the first pod is already deleted, its saved YAML/describe files and Jenkins console/artifact record are the first-failure evidence. Do not recreate that evidence by guessing.
9. Verify teardown and volume lifetime
After the build finishes and plugin retention deletes the final agent pod:
kubectl -n ch20-lab get pods -l devops-academy=ch20
kubectl -n ch20-lab get pvc
Expected: no build-owned pod remains; this template created no PVC;
the emptyDir workspaces disappeared with their pods.
Jenkins artifacts and build metadata remain because they were
transferred out before teardown.
10. Evidence packet
| Evidence | Capture | Secret? |
|---|---|---|
| Controller baseline | Jenkins core, Java, Kubernetes plugin version | No |
| Source/build | Job full name, build number/URL, cause, Git SHA | No |
| Cloud/template |
ch20-kind, namespace, pod YAML hash/revision
|
No |
| First pod | Name, UID, node, imageID, resources, events/termination evidence | No |
| Replacement pod | Name, UID, node, imageID | No |
| Identity | Pod ServiceAccount and auth can-i denial |
No |
| Workspace |
Path and statement that emptyDir was pod-scoped
|
No |
| Retained outputs |
Archived retained-*.txt digests/fingerprints
|
No |
| Limits | CPU/memory requests/limits; cloud cap; retry count | No |
| Assumptions | Local-only kind cluster; no production side effects | No |
Do not include the Kubernetes provisioner token, Jenkins agent secret, kubeconfig private key material, or unrelated cluster Secrets.
11. Final cleanup and rollback
Remove the lab Jenkins cloud credential/configuration first so the controller no longer holds a path to the local cluster. Then:
set -eu
case "$(kubectl config current-context)" in kind-ch20) ;; *) exit 2;; esac
kubectl -n ch20-lab get all,serviceaccount,role,rolebinding,pvc
kubectl delete namespace ch20-lab
kind delete cluster --name ch20
rm -f /tmp/ch20-kind.yaml /tmp/ch20-token /tmp/ch20-*.txt /tmp/ch20-*.yaml
Verify the cluster is absent and the Jenkins job retains only the intended build records/artifacts.
12. What this chapter adds to the production model
You can now treat Kubernetes agents as attributable execution infrastructure rather than anonymous disposable pods. A production record can state which Jenkins build/source requested which template, which pod UID/images/identity/resources executed it, what evidence was exported, why the pod terminated, and whether cleanup completed.
Chapter 21 builds on this by increasing concurrency through parallel stages, matrices and test sharding. The same lesson becomes more important: throughput must remain bounded by explicit agent capacity, artifact identity and fail-fast semantics rather than uncontrolled fan-out.
Knowledge check
Answer before revealing the explanation.
1. What two state changes should be predicted before failure injection?
Predict that deleting the current agent pod removes its pod-scoped workspace, and that a qualifying Kubernetes-agent retry can create a fresh pod while the same Jenkins build retains already archived evidence.
2. Why archive evidence before the intentional pod deletion?
Because the workspace is ephemeral. Anything not exported before pod loss may be unrecoverable from that attempt.
3. How do you prove the replacement is really a new pod?
Compare pod name and UID from the first and recovered attempts. The Jenkins build number can remain the same while the Kubernetes execution identity changes.
4. Why must the checkpoint use bounded resources and retry counts?
Dynamic agents can otherwise amplify queue pressure into unbounded pod or node demand. Resource requests/limits, plugin/cloud caps, quotas, and bounded retries are independent safety controls.
5. What does Chapter 20 add to the Jenkins operating model?
It makes execution infrastructure itself attributable: cloud/template, pod UID, images, ServiceAccount/RBAC, resources, volumes, termination cause, and retained build evidence are all explicit and reviewable.
Official references and version notes
-
Jenkins LTS changelog
— baseline
Jenkins 2.568.3 LTS, released 2026-09-02 and tested with Java 21 and 25; labs use Java 21 for Jenkins components. - Jenkins Java Support Policy — current Jenkins system components, including agents, require a supported JVM; this chapter uses Java 21.
-
Jenkins Kubernetes plugin
— reviewed version
4547.v52f3080db_8cd, requires Jenkins2.516.3, and has no current security advisory shown by the plugin health page at the chapter timestamp. -
Kubernetes plugin Pipeline steps
—
podTemplate,container, pod retention, workspace volumes, and related fields. -
kind v0.33.0
— local lab pin; its Kubernetes v1.37.0 node image is
kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5. - Kubernetes RBAC good practices and ServiceAccounts for Pods — least privilege, namespaced roles, dedicated service accounts, and token-automount guidance.
-
Kubernetes volumes
—
emptyDiris pod-lifetime storage; persistent volume types have a different lifecycle. - Resource management and Node autoscaling — scheduling uses requests; node autoscaling is a separate control loop from Jenkins agent provisioning.
-
Jenkins inbound-agent image
— reviewed agent tag
3391.va_37fa_a_305d6d-2-jdk21; record the actual architecture-specific digest/ImageID used by the pod.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.