Chapter 26Lesson 05~230 minutes

Checkpoint Lab — Continuous Delivery to AWS, Azure, Google Cloud, and Kubernetes

Prove a provider-neutral deployment contract with an exact artifact digest, environment boundary, disposable Kubernetes rollout, health check, and evidence-preserving rollback.

CheckpointExact digestDeployment evidenceRollbackContract

Learning objectives

  • Define one provider-neutral deployment contract whose inputs/outputs can be implemented by AWS, Azure, Google Cloud or Kubernetes adapters.
  • Predict target/context, artifact and rollout state before mutation, then verify every prediction independently.
  • Create a healthy local Kubernetes revision from an exact artifact digest and record the target-side revision.
  • Inject one controlled bad rollout, preserve first-failure evidence, and roll back to a known-good revision without rebuilding.
  • Produce a compact evidence packet that links GitHub run identity to artifact, environment, target, rollout, health and recovery state.

1. Checkpoint scenario and deployment contract

The checkpoint uses a disposable public repository and a kind cluster so the entire mandatory path is free. The contract is provider-neutral: input exact artifact digest + environment + target + health policy; output deployed digest + target deployment/revision ID + health result + rollback handle. The local adapter implements that contract with Kubernetes; an AWS/Azure/GCP adapter would preserve the same fields while changing only authentication and target API syntax.

2. Preflight and explicit safety guards

  • Repository must be disposable and must not contain customer/production credentials.
  • Docker must be available on ubuntu-24.04; kind and kubectl are downloaded at pinned versions and checksummed.
  • Create/inspect the lab-k8s environment. If true reviewer protection is unavailable, document the simulation limitation.
  • The workflow has no cloud secret and no repository write permission.
  • Cluster name includes the current run ID and is the only cluster the cleanup step may delete.
  • The intentionally bad image points to localhost.invalid, ensuring the failure remains bounded to the disposable cluster.

Do not adapt the intentional failure to a real registry, production cluster or shared namespace. The exercise is designed so the failure has no external side effect beyond one disposable local cluster.

3. Predict state changes before running

Prediction Expected change Independent verification
P1 target identity context becomes kind-gha-checkpoint-RUN_ID kubectl config current-context exact comparison
P2 artifact identity revision 1 Pod template carries the build SHA-256 Deployment annotation + downloaded SHA256SUMS
P3 rollout state revision 1 becomes Available and HTTP returns source SHA rollout status + external curl
P4 failure state revision 2 cannot pull the intentionally invalid image rollout timeout + Pod status/events preserved
P5 recovery state rollback restores the known-good Pod template and health rollout history + annotation + curl after undo

4. Exact checkpoint workflow

The first job produces the subject and transfers it. The deployment job verifies the digest before creating the target. The bad rollout is allowed to fail as an individual step so evidence and rollback can execute, but the workflow records that failure explicitly rather than pretending it never happened.

name: Checkpoint - provider-neutral deployment contract
on:
  workflow_dispatch:
    inputs:
      approve_deploy:
        description: "I confirm this is the disposable lab target"
        required: true
        type: boolean
        default: false
      inject_bad_rollout:
        description: "Create a failing second revision before rollback"
        required: true
        type: boolean
        default: true

permissions: {}

jobs:
  subject:
    name: Create and fingerprint subject
    runs-on: ubuntu-24.04
    permissions:
      contents: read
    outputs:
      digest: ${{ steps.fp.outputs.digest }}
    steps:
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
      - id: fp
        shell: bash
        run: |
          set -euo pipefail
          mkdir -p dist
          printf '<h1>checkpoint</h1>
<p>source=%s</p>
<p>run=%s</p>
'             "$GITHUB_SHA" "$GITHUB_RUN_ID" > dist/index.html
          digest=$(sha256sum dist/index.html | awk '{print $1}')
          echo "$digest  index.html" > dist/SHA256SUMS
          echo "digest=$digest" >> "$GITHUB_OUTPUT"
      - uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
        with:
          name: cd-checkpoint-${{ github.run_id }}-${{ github.run_attempt }}
          path: dist/
          retention-days: 7

  deploy:
    name: Deploy exact subject and prove recovery
    needs: subject
    if: ${{ inputs.approve_deploy }}
    runs-on: ubuntu-24.04
    environment: lab-k8s
    concurrency:
      group: checkpoint-lab-k8s
      cancel-in-progress: false
    permissions: {}
    env:
      DIGEST: ${{ needs.subject.outputs.digest }}
      CLUSTER: gha-checkpoint-${{ github.run_id }}
      NS: checkpoint
      KIND_NODE_IMAGE: kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5
    steps:
      - uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
        with:
          name: cd-checkpoint-${{ github.run_id }}-${{ github.run_attempt }}
          path: dist

      - name: Verify transfer and record prediction boundary
        shell: bash
        run: |
          set -euo pipefail
          (cd dist && sha256sum -c SHA256SUMS)
          actual=$(sha256sum dist/index.html | awk '{print $1}')
          test "$actual" = "$DIGEST"
          printf 'PREDICTION 1: target context will become kind-%s
' "$CLUSTER"
          printf 'PREDICTION 2: first Deployment revision will carry artifact digest %s
' "$DIGEST"
          printf 'PREDICTION 3: bad second revision will fail health but first-failure evidence remains
'

      - name: Install pinned local target tools
        shell: bash
        run: |
          set -euo pipefail
          curl -fsSLo kind "https://kind.sigs.k8s.io/dl/v0.33.0/kind-linux-amd64"
          echo "aee6151561422756b764a4ae28e7f44cda5af5a9eead3cc9985112b1de8d8e0d  kind" | sha256sum -c -
          install -m 0755 kind "$RUNNER_TEMP/kind"
          echo "$RUNNER_TEMP" >> "$GITHUB_PATH"
          curl -fsSLo kubectl "https://dl.k8s.io/release/v1.37.0/bin/linux/amd64/kubectl"
          curl -fsSLo kubectl.sha256 "https://dl.k8s.io/release/v1.37.0/bin/linux/amd64/kubectl.sha256"
          echo "$(cat kubectl.sha256)  kubectl" | sha256sum -c -
          install -m 0755 kubectl "$RUNNER_TEMP/kubectl"

      - name: Create exact disposable target
        shell: bash
        run: |
          set -euo pipefail
          cat > kind.yaml <<'EOF'
          kind: Cluster
          apiVersion: kind.x-k8s.io/v1alpha4
          nodes:
          - role: control-plane
            extraPortMappings:
            - containerPort: 30080
              hostPort: 18080
              listenAddress: "127.0.0.1"
              protocol: TCP
          EOF
          kind create cluster --name "$CLUSTER" --image "$KIND_NODE_IMAGE" --config kind.yaml --wait 120s
          test "$(kubectl config current-context)" = "kind-$CLUSTER"
          kubectl create namespace "$NS"

      - name: Roll out revision 1 using verified content
        id: good
        shell: bash
        run: |
          set -euo pipefail
          docker pull nginx:alpine
          image=$(docker inspect --format='{{index .RepoDigests 0}}' nginx:alpine)
          short="${DIGEST:0:12}"
          cm="content-$short"
          kubectl -n "$NS" create configmap "$cm" --from-file=index.html=dist/index.html
          cat > deploy.yaml <<EOF
          apiVersion: apps/v1
          kind: Deployment
          metadata: {name: web, namespace: $NS}
          spec:
            replicas: 1
            selector: {matchLabels: {app: web}}
            template:
              metadata:
                labels: {app: web}
                annotations:
                  academy.example/artifact-sha256: "$DIGEST"
              spec:
                containers:
                - name: web
                  image: "$image"
                  readinessProbe:
                    httpGet: {path: /, port: 80}
                    periodSeconds: 2
                  volumeMounts:
                  - {name: content, mountPath: /usr/share/nginx/html}
                volumes:
                - name: content
                  configMap: {name: "$cm"}
          ---
          apiVersion: v1
          kind: Service
          metadata: {name: web, namespace: $NS}
          spec:
            type: NodePort
            selector: {app: web}
            ports:
            - {port: 80, targetPort: 80, nodePort: 30080}
          EOF
          kubectl apply -f deploy.yaml
          kubectl -n "$NS" rollout status deploy/web --timeout=120s
          curl -fsS --retry 10 --retry-delay 2 http://127.0.0.1:18080/ | tee health-good.txt
          grep -F "source=$GITHUB_SHA" health-good.txt
          rev=$(kubectl -n "$NS" get deploy web -o jsonpath='{{.metadata.annotations.deployment\.kubernetes\.io/revision}}')
          echo "good_revision=$rev" >> "$GITHUB_OUTPUT"
          echo "runtime_image=$image" >> "$GITHUB_OUTPUT"

      - name: Intentionally create bad revision 2
        if: ${{ inputs.inject_bad_rollout }}
        id: bad
        shell: bash
        continue-on-error: true
        run: |
          set -euo pipefail
          kubectl -n "$NS" set image deploy/web web=localhost.invalid/checkpoint@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
          set +e
          kubectl -n "$NS" rollout status deploy/web --timeout=25s
          rc=$?
          set -e
          kubectl -n "$NS" get pods -o wide
          kubectl -n "$NS" get events --sort-by=.lastTimestamp | tail -40
          test "$rc" -ne 0
          echo "expected_failure_observed=true" >> "$GITHUB_OUTPUT"
          exit 1

      - name: Preserve first-failure evidence before recovery
        if: ${{ always() && inputs.inject_bad_rollout }}
        shell: bash
        run: |
          set +e
          mkdir -p evidence
          kubectl config current-context > evidence/context.txt 2>&1
          kubectl -n "$NS" rollout history deploy/web > evidence/history-before-rollback.txt 2>&1
          kubectl -n "$NS" describe deploy/web > evidence/failed-deployment.txt 2>&1
          kubectl -n "$NS" get pods -o yaml > evidence/failed-pods.yaml 2>&1
          kubectl -n "$NS" get events --sort-by=.lastTimestamp > evidence/events-before-rollback.txt 2>&1

      - name: Roll back to known-good revision and verify recovery
        if: ${{ always() && inputs.inject_bad_rollout }}
        shell: bash
        run: |
          set -euo pipefail
          kubectl -n "$NS" rollout undo deploy/web --to-revision="${{ steps.good.outputs.good_revision }}"
          kubectl -n "$NS" rollout status deploy/web --timeout=120s
          curl -fsS --retry 10 --retry-delay 2 http://127.0.0.1:18080/ | tee evidence/health-after-rollback.txt
          grep -F "source=$GITHUB_SHA" evidence/health-after-rollback.txt
          test "$(kubectl -n "$NS" get deploy web -o jsonpath='{{.spec.template.metadata.annotations.academy\.example/artifact-sha256}}')" = "$DIGEST"

      - name: Build evidence packet
        if: ${{ always() }}
        shell: bash
        run: |
          set +e
          mkdir -p evidence
          cat > evidence/contract.txt <<EOF
          run_id=$GITHUB_RUN_ID
          run_attempt=$GITHUB_RUN_ATTEMPT
          source_sha=$GITHUB_SHA
          artifact_sha256=$DIGEST
          environment=lab-k8s
          expected_context=kind-$CLUSTER
          namespace=$NS
          good_revision=${{ steps.good.outputs.good_revision }}
          runtime_image=${{ steps.good.outputs.runtime_image }}
          kind=v0.33.0
          kubectl=v1.37.0
          node_image=kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5
          cloud_credentials=none
          EOF
          kubectl -n "$NS" rollout history deploy/web > evidence/history-final.txt 2>&1
          kubectl -n "$NS" get deploy web -o yaml > evidence/deployment-final.yaml 2>&1
          kubectl -n "$NS" get all -o wide > evidence/resources-final.txt 2>&1

      - uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
        if: ${{ always() }}
        with:
          name: cd-evidence-${{ github.run_id }}-${{ github.run_attempt }}
          path: evidence/
          retention-days: 7

      - name: Cleanup exact cluster
        if: ${{ always() }}
        shell: bash
        run: kind delete cluster --name "$CLUSTER" || true

5. Expected observations

  1. The subject job prints one SHA-256 and uploads an Actions artifact.
  2. The deploy job recomputes the same digest before cluster creation.
  3. The context assertion succeeds only for the run-specific kind cluster.
  4. Revision 1 becomes Available and the NodePort returns the page containing the exact source SHA.
  5. Revision 2 times out because the deliberately invalid image cannot be pulled; events identify the image-pull failure.
  6. Failure-state YAML/history/events are written before rollback.
  7. Rollback restores the known-good revision/template and the final HTTP health check succeeds.
  8. The evidence artifact persists after the cluster is deleted.

6. Evidence packet: link every boundary

Boundary Required evidence
GitHub event/run event name, run ID, attempt, source SHA, workflow file/revision
Artifact SHA256SUMS, recomputed digest, Actions artifact name
Environment/governance lab-k8s, protection availability/approval or simulation note, concurrency group
Runner/tooling ubuntu-24.04, kind v0.33.0 checksum, kubectl v1.37.0, pinned node image digest
Target exact context, cluster name, namespace
Runtime content digest annotation plus resolved Nginx repository digest
Rollout known-good revision, failed revision/history, Deployment/Pod state
Health successful external response before failure and after rollback
Recovery failed events captured before rollout undo, final known-good identity
Limitations no real cloud identity; environment reviewer rule may be simulated depending on plan/visibility

7. Map the checkpoint to cloud adapters without changing the contract

Contract field AWS adapter Azure adapter Google Cloud adapter Kubernetes lab
target_id account + region + service ARN/name subscription + resource ID project + region + resource context + namespace + Deployment
identity OIDC-assumed IAM role Entra federated service principal WIF principal/service account ephemeral kind kubeconfig
deployment_id service deployment/task revision deployment/resource revision service revision/operation Deployment revision
health service/task/LB health resource + endpoint health revision/traffic + endpoint health rollout + readiness + HTTP
rollback_handle previous task/version/revision previous slot/revision/version previous service revision known-good Deployment revision

8. Wrong artifact or wrong target must fail before rollout

Repeat the mental exercise without executing it: if the downloaded SHA-256 differs from the release evidence, the deployment job stops before cluster creation. If the context does not equal the run-specific kind target, it stops before namespace/resource mutation. Those two guards are more important than a sophisticated rollback because they prevent unauthorized or unverifiable changes from starting.

9. Why the failed revision remains part of the record

The checkpoint does not delete the failed Pod/revision immediately. It first captures the Deployment description, Pod YAML and events. Only then does it execute kubectl rollout undo. The original GitHub step remains a failure in the run evidence even though the later recovery succeeds. This is the same incident-safe principle used throughout the course: recovery should restore service without falsifying history.

10. Cleanup and rollback boundaries

The final cleanup deletes only $CLUSTER, whose name is derived from the run ID. The evidence artifact remains for seven days. In a real cloud checkpoint, cleanup would require an exact disposable resource ID guard and should be a separate consciously authorized operation; never use broad “delete all lab resources” selectors in shared accounts.

11. Verification checklist

  • Source SHA and artifact SHA-256 are present and linked.
  • Environment/concurrency state is recorded separately from target health.
  • No cloud admin/static credential exists in the mandatory path.
  • kind binary and node image identities are pinned; kubectl checksum is verified.
  • Kubernetes context is asserted before mutation.
  • Revision 1 serves the intended source.
  • Revision 2 failure evidence is preserved before rollback.
  • Rollback restores both digest annotation and endpoint health.
  • Evidence artifact survives target cleanup.

12. What Chapter 26 adds — and the bridge to Chapter 27

Chapter 26 extends the operating model from release identity to deployment identity: verified bytes enter a governed environment, obtain a short-lived or ephemeral target identity, create one observable rollout, and close with health/recovery evidence. Chapter 27 will apply the same state-separation discipline to infrastructure-as-code plans and guarded state changes, where stale plans and backend/workspace identity become the next critical boundaries.

13. Checkpoint summary

You now have a portable continuous-delivery contract that can survive a provider change because artifact identity, governance, target identity, rollout, health and rollback are explicit evidence fields rather than hidden inside a vendor-specific script.

Next lesson

Infrastructure as Code, Terraform Plans, Policy Checks, and Deployment Workflows: Core Concepts and Mental Model

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why does the checkpoint keep artifact SHA-256 and Kubernetes revision as separate fields?

The bad revision fails image pull. Why not delete the Pod immediately?

What would change when replacing kind with AWS ECS?

If the environment is approved but the context assertion fails, should the workflow deploy?

What is the Chapter 27 handoff?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.