Chapter 20Lesson 05~215 minutes

Checkpoint Lab — Kubernetes Plugin, Pod Templates, Dynamic Kubernetes Agents, Volumes, Service Accounts, and Cluster Scaling

Run a bounded dynamic-agent workflow on a local Kubernetes cluster, use restrictive resources and identity, inject pod loss, recover only from retained evidence, and verify that every build-owned pod and ephemeral volume is gone afterward.

Checkpoint labDynamic agentFailure injectionArtifactsTeardownCapacity

Learning objectives

  • Run one SCM-backed synthetic Pipeline on a dynamic Kubernetes agent with an explicit template, resources, images and ServiceAccount.
  • Predict and verify Jenkins-build identity versus Kubernetes pod identity across pod loss.
  • Prove workspace-only state disappears while archived evidence remains.
  • Recover through a bounded Kubernetes-agent retry without repeating an irreversible side effect.
  • Produce a reviewable evidence packet and prove complete pod/ephemeral-volume teardown.

1. Scenario and acceptance criteria

You operate a disposable Jenkins controller and a local kind cluster. Job jenkins/ch20-k8s-checkpoint runs only synthetic work. A dynamic agent pod records source/build/pod/resource evidence, archives a pre-loss record, waits for an intentional pod deletion, then recovers on a fresh pod using a bounded infrastructure retry.

The lab passes only if the same Jenkins build remains attributable to exact source while the pod name/UID changes, workspace-only state from the first pod is gone, archived evidence remains, and the namespace contains no build-owned pods/PVCs after completion.

2. Pinned assumptions

Component Lab baseline Verification
Jenkins 2.568.3 LTS, Java 21 Manage Jenkins → System Information / java -version
Kubernetes plugin 4547.v52f3080db_8cd Plugin inventory
kind v0.33.0 kind version
Kubernetes node v1.37.0 kind image digest a1ed56…580ae5 docker inspect / kind create command
Agent image jenkins/inbound-agent:3391.va_37fa_a_305d6d-2-jdk21 Pod imageID; record full runtime digest
Namespace ch20-lab kubectl get ns

3. Preflight and guardrails

set -eu
CTX=$(kubectl config current-context)
printf 'context=%s\\n' "$CTX"
case "$CTX" in kind-ch20) ;; *) echo 'Refusing: not the ch20 disposable cluster' >&2; exit 2;; esac
kubectl get nodes
kubectl -n ch20-lab get serviceaccount ch20-agent
kubectl -n ch20-lab auth can-i get pods --as=system:serviceaccount:ch20-lab:ch20-agent

The final command should return no. Also verify the Jenkins built-in node has zero executors, the cloud is restricted to the lab folder where supported, and its cap is small.

4. Predict before execution

Write these predictions into the build description or a checked-in lab note before starting:

  1. The first attempt creates one pod with a unique UID and an emptyDir-backed workspace. Deleting that pod destroys that workspace.
  2. The archived pre-loss evidence is copied out of the workspace before deletion and remains attached to the Jenkins build.
  3. A qualifying Kubernetes-agent retry may create a second pod with a different name/UID while the Jenkins BUILD_NUMBER remains unchanged.
  4. No production/external side effect occurs, so repeating the synthetic attempt is safe.

5. Synthetic SCM content

Commit a tiny payload.txt, ci/ch20-pod.yaml, and Jenkinsfile to a disposable repository so GIT_COMMIT is available. The pod YAML:

apiVersion: v1
kind: Pod
metadata:
  labels:
    devops-academy: ch20
spec:
  serviceAccountName: ch20-agent
  automountServiceAccountToken: false
  securityContext:
    runAsNonRoot: true
    runAsUser: 1000
    runAsGroup: 1000
    fsGroup: 1000
  containers:
  - name: jnlp
    image: jenkins/inbound-agent:3391.va_37fa_a_305d6d-2-jdk21
    resources:
      requests:
        cpu: 100m
        memory: 192Mi
      limits:
        cpu: 500m
        memory: 512Mi

Do not add privileged mode, hostPath, host networking or broad Kubernetes credentials.

6. Checkpoint Jenkinsfile

pipeline {
  agent none
  options {
    timeout(time: 20, unit: 'MINUTES')
    timestamps()
  }
  stages {
    stage('Kubernetes attempt') {
      steps {
        script {
          podTemplate(
            cloud: 'ch20-kind',
            namespace: 'ch20-lab',
            serviceAccount: 'ch20-agent',
            podRetention: never(),
            yaml: readTrusted('ci/ch20-pod.yaml')
          ) {
            retry(count: 2, conditions: [kubernetesAgent(), nonresumable()]) {
              node(POD_LABEL) {
                checkout scm
                sh '''
                  set -eu
                  POD=$(hostname)
                  printf 'job=%s\\nbuild=%s\\nsource=%s\\nnode=%s\\npod=%s\\nworkspace=%s\\n' \
                    "$JOB_NAME" "$BUILD_NUMBER" "$GIT_COMMIT" "$NODE_NAME" "$POD" "$WORKSPACE" \
                    | tee "retained-${POD}.txt"
                  printf 'pod-local-only=%s\\n' "$POD" > workspace-only.txt
                  sha256sum payload.txt | tee -a "retained-${POD}.txt"
                '''
                archiveArtifacts artifacts: 'retained-*.txt', fingerprint: true
                echo 'Preserve pod evidence now, then delete ONLY this lab pod from a separate terminal.'
                sleep 90
                sh '''
                  set -eu
                  test -f payload.txt
                  printf 'post-wait-pod=%s\\n' "$(hostname)" | tee post-wait.txt
                '''
                archiveArtifacts artifacts: 'post-wait.txt', fingerprint: true
              }
            }
          }
        }
      }
    }
  }
  post {
    always {
      echo "build=${env.BUILD_URL} result=${currentBuild.currentResult}"
    }
  }
}

The retry scope contains only source checkout, hashing, local files and artifact upload. It deliberately excludes deployment, signing, package publication, registry promotion or destructive external API calls.

7. Preserve evidence, then inject pod loss

While the Pipeline is in sleep 90, use a separate terminal:

set -eu
kubectl -n ch20-lab get pods -o wide
POD=$(kubectl -n ch20-lab get pods -l devops-academy=ch20 \
  --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].metadata.name}')
test -n "$POD"
kubectl -n ch20-lab get pod "$POD" -o yaml > "/tmp/${POD}.yaml"
kubectl -n ch20-lab describe pod "$POD" > "/tmp/${POD}.describe.txt"
kubectl -n ch20-lab get pod "$POD" -o jsonpath='uid={.metadata.uid}{"\\n"}node={.spec.nodeName}{"\\n"}sa={.spec.serviceAccountName}{"\\n"}'
kubectl -n ch20-lab delete pod "$POD" --wait=false

Keep the console output showing the original pod name. The retry should not be forced if the failure is a normal test error; kubernetesAgent() is specifically for qualifying agent/pod infrastructure loss.

8. Verify recovery independently

After Jenkins provisions another agent, record the new pod name/UID and compare it with the first. Confirm the build number/source SHA remain the same. Then inspect the Jenkins artifacts: at least one retained-<pod>.txt from before failure should exist even though the first workspace is gone.

kubectl -n ch20-lab get pods -o wide
kubectl -n ch20-lab get events --sort-by=.lastTimestamp | tail -60

If the first pod is already deleted, its saved YAML/describe files and Jenkins console/artifact record are the first-failure evidence. Do not recreate that evidence by guessing.

9. Verify teardown and volume lifetime

After the build finishes and plugin retention deletes the final agent pod:

kubectl -n ch20-lab get pods -l devops-academy=ch20
kubectl -n ch20-lab get pvc

Expected: no build-owned pod remains; this template created no PVC; the emptyDir workspaces disappeared with their pods. Jenkins artifacts and build metadata remain because they were transferred out before teardown.

10. Evidence packet

Evidence Capture Secret?
Controller baseline Jenkins core, Java, Kubernetes plugin version No
Source/build Job full name, build number/URL, cause, Git SHA No
Cloud/template ch20-kind, namespace, pod YAML hash/revision No
First pod Name, UID, node, imageID, resources, events/termination evidence No
Replacement pod Name, UID, node, imageID No
Identity Pod ServiceAccount and auth can-i denial No
Workspace Path and statement that emptyDir was pod-scoped No
Retained outputs Archived retained-*.txt digests/fingerprints No
Limits CPU/memory requests/limits; cloud cap; retry count No
Assumptions Local-only kind cluster; no production side effects No

Do not include the Kubernetes provisioner token, Jenkins agent secret, kubeconfig private key material, or unrelated cluster Secrets.

11. Final cleanup and rollback

Remove the lab Jenkins cloud credential/configuration first so the controller no longer holds a path to the local cluster. Then:

set -eu
case "$(kubectl config current-context)" in kind-ch20) ;; *) exit 2;; esac
kubectl -n ch20-lab get all,serviceaccount,role,rolebinding,pvc
kubectl delete namespace ch20-lab
kind delete cluster --name ch20
rm -f /tmp/ch20-kind.yaml /tmp/ch20-token /tmp/ch20-*.txt /tmp/ch20-*.yaml

Verify the cluster is absent and the Jenkins job retains only the intended build records/artifacts.

12. What this chapter adds to the production model

You can now treat Kubernetes agents as attributable execution infrastructure rather than anonymous disposable pods. A production record can state which Jenkins build/source requested which template, which pod UID/images/identity/resources executed it, what evidence was exported, why the pod terminated, and whether cleanup completed.

Chapter 21 builds on this by increasing concurrency through parallel stages, matrices and test sharding. The same lesson becomes more important: throughput must remain bounded by explicit agent capacity, artifact identity and fail-fast semantics rather than uncontrolled fan-out.

Next chapter

Parallel Stages, Matrix Builds, Fail-Fast Behavior, Test Sharding, and High-Throughput Pipeline Design

Scale Pipeline concurrency without losing deterministic evidence, capacity controls, or failure attribution.

Knowledge check

Answer before revealing the explanation.

1. What two state changes should be predicted before failure injection?

2. Why archive evidence before the intentional pod deletion?

3. How do you prove the replacement is really a new pod?

4. Why must the checkpoint use bounded resources and retry counts?

5. What does Chapter 20 add to the Jenkins operating model?

Official references and version notes

  • Jenkins LTS changelog — baseline Jenkins 2.568.3 LTS, released 2026-09-02 and tested with Java 21 and 25; labs use Java 21 for Jenkins components.
  • Jenkins Java Support Policy — current Jenkins system components, including agents, require a supported JVM; this chapter uses Java 21.
  • Jenkins Kubernetes plugin — reviewed version 4547.v52f3080db_8cd, requires Jenkins 2.516.3, and has no current security advisory shown by the plugin health page at the chapter timestamp.
  • Kubernetes plugin Pipeline steps — podTemplate, container, pod retention, workspace volumes, and related fields.
  • kind v0.33.0 — local lab pin; its Kubernetes v1.37.0 node image is kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5.
  • Kubernetes RBAC good practices and ServiceAccounts for Pods — least privilege, namespaced roles, dedicated service accounts, and token-automount guidance.
  • Kubernetes volumes — emptyDir is pod-lifetime storage; persistent volume types have a different lifecycle.
  • Resource management and Node autoscaling — scheduling uses requests; node autoscaling is a separate control loop from Jenkins agent provisioning.
  • Jenkins inbound-agent image — reviewed agent tag 3391.va_37fa_a_305d6d-2-jdk21; record the actual architecture-specific digest/ImageID used by the pod.
Assumption timestamp: 2026-09-17. Recheck Jenkins LTS/Java, Kubernetes plugin version/dependencies/security status, kind/Kubernetes node digest, agent image identity, and Kubernetes API semantics before repeating later.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.