Checkpoint Lab — Continuous Delivery to AWS, Azure, Google Cloud, and Kubernetes
Prove a provider-neutral deployment contract with an exact artifact digest, environment boundary, disposable Kubernetes rollout, health check, and evidence-preserving rollback.
Learning objectives
- Define one provider-neutral deployment contract whose inputs/outputs can be implemented by AWS, Azure, Google Cloud or Kubernetes adapters.
- Predict target/context, artifact and rollout state before mutation, then verify every prediction independently.
- Create a healthy local Kubernetes revision from an exact artifact digest and record the target-side revision.
- Inject one controlled bad rollout, preserve first-failure evidence, and roll back to a known-good revision without rebuilding.
- Produce a compact evidence packet that links GitHub run identity to artifact, environment, target, rollout, health and recovery state.
1. Checkpoint scenario and deployment contract
The checkpoint uses a disposable public repository and a kind cluster so the entire mandatory path is free. The contract is provider-neutral: input exact artifact digest + environment + target + health policy; output deployed digest + target deployment/revision ID + health result + rollback handle. The local adapter implements that contract with Kubernetes; an AWS/Azure/GCP adapter would preserve the same fields while changing only authentication and target API syntax.
2. Preflight and explicit safety guards
- Repository must be disposable and must not contain customer/production credentials.
-
Docker must be available on
ubuntu-24.04; kind and kubectl are downloaded at pinned versions and checksummed. -
Create/inspect the
lab-k8senvironment. If true reviewer protection is unavailable, document the simulation limitation. - The workflow has no cloud secret and no repository write permission.
- Cluster name includes the current run ID and is the only cluster the cleanup step may delete.
-
The intentionally bad image points to
localhost.invalid, ensuring the failure remains bounded to the disposable cluster.
Do not adapt the intentional failure to a real registry, production cluster or shared namespace. The exercise is designed so the failure has no external side effect beyond one disposable local cluster.
3. Predict state changes before running
| Prediction | Expected change | Independent verification |
|---|---|---|
| P1 target identity |
context becomes kind-gha-checkpoint-RUN_ID
|
kubectl config current-context exact comparison
|
| P2 artifact identity | revision 1 Pod template carries the build SHA-256 |
Deployment annotation + downloaded SHA256SUMS
|
| P3 rollout state | revision 1 becomes Available and HTTP returns source SHA | rollout status + external curl |
| P4 failure state | revision 2 cannot pull the intentionally invalid image | rollout timeout + Pod status/events preserved |
| P5 recovery state | rollback restores the known-good Pod template and health | rollout history + annotation + curl after undo |
4. Exact checkpoint workflow
The first job produces the subject and transfers it. The deployment job verifies the digest before creating the target. The bad rollout is allowed to fail as an individual step so evidence and rollback can execute, but the workflow records that failure explicitly rather than pretending it never happened.
name: Checkpoint - provider-neutral deployment contract
on:
workflow_dispatch:
inputs:
approve_deploy:
description: "I confirm this is the disposable lab target"
required: true
type: boolean
default: false
inject_bad_rollout:
description: "Create a failing second revision before rollback"
required: true
type: boolean
default: true
permissions: {}
jobs:
subject:
name: Create and fingerprint subject
runs-on: ubuntu-24.04
permissions:
contents: read
outputs:
digest: ${{ steps.fp.outputs.digest }}
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- id: fp
shell: bash
run: |
set -euo pipefail
mkdir -p dist
printf '<h1>checkpoint</h1>
<p>source=%s</p>
<p>run=%s</p>
' "$GITHUB_SHA" "$GITHUB_RUN_ID" > dist/index.html
digest=$(sha256sum dist/index.html | awk '{print $1}')
echo "$digest index.html" > dist/SHA256SUMS
echo "digest=$digest" >> "$GITHUB_OUTPUT"
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: cd-checkpoint-${{ github.run_id }}-${{ github.run_attempt }}
path: dist/
retention-days: 7
deploy:
name: Deploy exact subject and prove recovery
needs: subject
if: ${{ inputs.approve_deploy }}
runs-on: ubuntu-24.04
environment: lab-k8s
concurrency:
group: checkpoint-lab-k8s
cancel-in-progress: false
permissions: {}
env:
DIGEST: ${{ needs.subject.outputs.digest }}
CLUSTER: gha-checkpoint-${{ github.run_id }}
NS: checkpoint
KIND_NODE_IMAGE: kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5
steps:
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
with:
name: cd-checkpoint-${{ github.run_id }}-${{ github.run_attempt }}
path: dist
- name: Verify transfer and record prediction boundary
shell: bash
run: |
set -euo pipefail
(cd dist && sha256sum -c SHA256SUMS)
actual=$(sha256sum dist/index.html | awk '{print $1}')
test "$actual" = "$DIGEST"
printf 'PREDICTION 1: target context will become kind-%s
' "$CLUSTER"
printf 'PREDICTION 2: first Deployment revision will carry artifact digest %s
' "$DIGEST"
printf 'PREDICTION 3: bad second revision will fail health but first-failure evidence remains
'
- name: Install pinned local target tools
shell: bash
run: |
set -euo pipefail
curl -fsSLo kind "https://kind.sigs.k8s.io/dl/v0.33.0/kind-linux-amd64"
echo "aee6151561422756b764a4ae28e7f44cda5af5a9eead3cc9985112b1de8d8e0d kind" | sha256sum -c -
install -m 0755 kind "$RUNNER_TEMP/kind"
echo "$RUNNER_TEMP" >> "$GITHUB_PATH"
curl -fsSLo kubectl "https://dl.k8s.io/release/v1.37.0/bin/linux/amd64/kubectl"
curl -fsSLo kubectl.sha256 "https://dl.k8s.io/release/v1.37.0/bin/linux/amd64/kubectl.sha256"
echo "$(cat kubectl.sha256) kubectl" | sha256sum -c -
install -m 0755 kubectl "$RUNNER_TEMP/kubectl"
- name: Create exact disposable target
shell: bash
run: |
set -euo pipefail
cat > kind.yaml <<'EOF'
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
nodes:
- role: control-plane
extraPortMappings:
- containerPort: 30080
hostPort: 18080
listenAddress: "127.0.0.1"
protocol: TCP
EOF
kind create cluster --name "$CLUSTER" --image "$KIND_NODE_IMAGE" --config kind.yaml --wait 120s
test "$(kubectl config current-context)" = "kind-$CLUSTER"
kubectl create namespace "$NS"
- name: Roll out revision 1 using verified content
id: good
shell: bash
run: |
set -euo pipefail
docker pull nginx:alpine
image=$(docker inspect --format='{{index .RepoDigests 0}}' nginx:alpine)
short="${DIGEST:0:12}"
cm="content-$short"
kubectl -n "$NS" create configmap "$cm" --from-file=index.html=dist/index.html
cat > deploy.yaml <<EOF
apiVersion: apps/v1
kind: Deployment
metadata: {name: web, namespace: $NS}
spec:
replicas: 1
selector: {matchLabels: {app: web}}
template:
metadata:
labels: {app: web}
annotations:
academy.example/artifact-sha256: "$DIGEST"
spec:
containers:
- name: web
image: "$image"
readinessProbe:
httpGet: {path: /, port: 80}
periodSeconds: 2
volumeMounts:
- {name: content, mountPath: /usr/share/nginx/html}
volumes:
- name: content
configMap: {name: "$cm"}
---
apiVersion: v1
kind: Service
metadata: {name: web, namespace: $NS}
spec:
type: NodePort
selector: {app: web}
ports:
- {port: 80, targetPort: 80, nodePort: 30080}
EOF
kubectl apply -f deploy.yaml
kubectl -n "$NS" rollout status deploy/web --timeout=120s
curl -fsS --retry 10 --retry-delay 2 http://127.0.0.1:18080/ | tee health-good.txt
grep -F "source=$GITHUB_SHA" health-good.txt
rev=$(kubectl -n "$NS" get deploy web -o jsonpath='{{.metadata.annotations.deployment\.kubernetes\.io/revision}}')
echo "good_revision=$rev" >> "$GITHUB_OUTPUT"
echo "runtime_image=$image" >> "$GITHUB_OUTPUT"
- name: Intentionally create bad revision 2
if: ${{ inputs.inject_bad_rollout }}
id: bad
shell: bash
continue-on-error: true
run: |
set -euo pipefail
kubectl -n "$NS" set image deploy/web web=localhost.invalid/checkpoint@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
set +e
kubectl -n "$NS" rollout status deploy/web --timeout=25s
rc=$?
set -e
kubectl -n "$NS" get pods -o wide
kubectl -n "$NS" get events --sort-by=.lastTimestamp | tail -40
test "$rc" -ne 0
echo "expected_failure_observed=true" >> "$GITHUB_OUTPUT"
exit 1
- name: Preserve first-failure evidence before recovery
if: ${{ always() && inputs.inject_bad_rollout }}
shell: bash
run: |
set +e
mkdir -p evidence
kubectl config current-context > evidence/context.txt 2>&1
kubectl -n "$NS" rollout history deploy/web > evidence/history-before-rollback.txt 2>&1
kubectl -n "$NS" describe deploy/web > evidence/failed-deployment.txt 2>&1
kubectl -n "$NS" get pods -o yaml > evidence/failed-pods.yaml 2>&1
kubectl -n "$NS" get events --sort-by=.lastTimestamp > evidence/events-before-rollback.txt 2>&1
- name: Roll back to known-good revision and verify recovery
if: ${{ always() && inputs.inject_bad_rollout }}
shell: bash
run: |
set -euo pipefail
kubectl -n "$NS" rollout undo deploy/web --to-revision="${{ steps.good.outputs.good_revision }}"
kubectl -n "$NS" rollout status deploy/web --timeout=120s
curl -fsS --retry 10 --retry-delay 2 http://127.0.0.1:18080/ | tee evidence/health-after-rollback.txt
grep -F "source=$GITHUB_SHA" evidence/health-after-rollback.txt
test "$(kubectl -n "$NS" get deploy web -o jsonpath='{{.spec.template.metadata.annotations.academy\.example/artifact-sha256}}')" = "$DIGEST"
- name: Build evidence packet
if: ${{ always() }}
shell: bash
run: |
set +e
mkdir -p evidence
cat > evidence/contract.txt <<EOF
run_id=$GITHUB_RUN_ID
run_attempt=$GITHUB_RUN_ATTEMPT
source_sha=$GITHUB_SHA
artifact_sha256=$DIGEST
environment=lab-k8s
expected_context=kind-$CLUSTER
namespace=$NS
good_revision=${{ steps.good.outputs.good_revision }}
runtime_image=${{ steps.good.outputs.runtime_image }}
kind=v0.33.0
kubectl=v1.37.0
node_image=kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5
cloud_credentials=none
EOF
kubectl -n "$NS" rollout history deploy/web > evidence/history-final.txt 2>&1
kubectl -n "$NS" get deploy web -o yaml > evidence/deployment-final.yaml 2>&1
kubectl -n "$NS" get all -o wide > evidence/resources-final.txt 2>&1
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
if: ${{ always() }}
with:
name: cd-evidence-${{ github.run_id }}-${{ github.run_attempt }}
path: evidence/
retention-days: 7
- name: Cleanup exact cluster
if: ${{ always() }}
shell: bash
run: kind delete cluster --name "$CLUSTER" || true
5. Expected observations
- The subject job prints one SHA-256 and uploads an Actions artifact.
- The deploy job recomputes the same digest before cluster creation.
- The context assertion succeeds only for the run-specific kind cluster.
- Revision 1 becomes Available and the NodePort returns the page containing the exact source SHA.
- Revision 2 times out because the deliberately invalid image cannot be pulled; events identify the image-pull failure.
- Failure-state YAML/history/events are written before rollback.
- Rollback restores the known-good revision/template and the final HTTP health check succeeds.
- The evidence artifact persists after the cluster is deleted.
6. Evidence packet: link every boundary
| Boundary | Required evidence |
|---|---|
| GitHub event/run | event name, run ID, attempt, source SHA, workflow file/revision |
| Artifact |
SHA256SUMS, recomputed digest, Actions artifact
name
|
| Environment/governance |
lab-k8s, protection availability/approval or
simulation note, concurrency group
|
| Runner/tooling |
ubuntu-24.04, kind v0.33.0 checksum, kubectl
v1.37.0, pinned node image digest
|
| Target | exact context, cluster name, namespace |
| Runtime | content digest annotation plus resolved Nginx repository digest |
| Rollout | known-good revision, failed revision/history, Deployment/Pod state |
| Health | successful external response before failure and after rollback |
| Recovery |
failed events captured before rollout undo,
final known-good identity
|
| Limitations | no real cloud identity; environment reviewer rule may be simulated depending on plan/visibility |
7. Map the checkpoint to cloud adapters without changing the contract
| Contract field | AWS adapter | Azure adapter | Google Cloud adapter | Kubernetes lab |
|---|---|---|---|---|
| target_id | account + region + service ARN/name | subscription + resource ID | project + region + resource | context + namespace + Deployment |
| identity | OIDC-assumed IAM role | Entra federated service principal | WIF principal/service account | ephemeral kind kubeconfig |
| deployment_id | service deployment/task revision | deployment/resource revision | service revision/operation | Deployment revision |
| health | service/task/LB health | resource + endpoint health | revision/traffic + endpoint health | rollout + readiness + HTTP |
| rollback_handle | previous task/version/revision | previous slot/revision/version | previous service revision | known-good Deployment revision |
8. Wrong artifact or wrong target must fail before rollout
Repeat the mental exercise without executing it: if the downloaded SHA-256 differs from the release evidence, the deployment job stops before cluster creation. If the context does not equal the run-specific kind target, it stops before namespace/resource mutation. Those two guards are more important than a sophisticated rollback because they prevent unauthorized or unverifiable changes from starting.
9. Why the failed revision remains part of the record
The checkpoint does not delete the failed Pod/revision immediately.
It first captures the Deployment description, Pod YAML and events.
Only then does it execute kubectl rollout undo. The
original GitHub step remains a failure in the run evidence even
though the later recovery succeeds. This is the same incident-safe
principle used throughout the course: recovery should restore
service without falsifying history.
10. Cleanup and rollback boundaries
The final cleanup deletes only $CLUSTER, whose name is
derived from the run ID. The evidence artifact remains for seven
days. In a real cloud checkpoint, cleanup would require an exact
disposable resource ID guard and should be a separate consciously
authorized operation; never use broad “delete all lab resources”
selectors in shared accounts.
11. Verification checklist
- Source SHA and artifact SHA-256 are present and linked.
- Environment/concurrency state is recorded separately from target health.
- No cloud admin/static credential exists in the mandatory path.
- kind binary and node image identities are pinned; kubectl checksum is verified.
- Kubernetes context is asserted before mutation.
- Revision 1 serves the intended source.
- Revision 2 failure evidence is preserved before rollback.
- Rollback restores both digest annotation and endpoint health.
- Evidence artifact survives target cleanup.
12. What Chapter 26 adds — and the bridge to Chapter 27
Chapter 26 extends the operating model from release identity to deployment identity: verified bytes enter a governed environment, obtain a short-lived or ephemeral target identity, create one observable rollout, and close with health/recovery evidence. Chapter 27 will apply the same state-separation discipline to infrastructure-as-code plans and guarded state changes, where stale plans and backend/workspace identity become the next critical boundaries.
13. Checkpoint summary
You now have a portable continuous-delivery contract that can survive a provider change because artifact identity, governance, target identity, rollout, health and rollback are explicit evidence fields rather than hidden inside a vendor-specific script.
Knowledge check
Why does the checkpoint keep artifact SHA-256 and Kubernetes revision as separate fields?
The digest identifies bytes/content; the revision identifies a target-side rollout history entry. One does not prove the other.
The bad revision fails image pull. Why not delete the Pod immediately?
Because Pod status/events are first-failure evidence that explains the causal layer. Capture them before rollback/cleanup.
What would change when replacing kind with AWS ECS?
The authentication and target API/revision/health implementation changes; the provider-neutral inputs/outputs—artifact digest, environment, target identity, deployment ID, health, rollback handle—remain.
If the environment is approved but the context assertion fails, should the workflow deploy?
No. Environment approval authorizes proceeding toward a target; it does not override a target-identity mismatch.
What is the Chapter 27 handoff?
Use the same exact-revision, identity, evidence and guarded-mutation model for IaC plan/apply and external infrastructure state.
Official references and version notes
- GitHub OIDC overview — Why short-lived federation is preferred to long-lived cloud secrets.
- OIDC reference — Current claims, immutable subject format, audiences, reusable-workflow behavior and token-request permissions.
- OIDC in AWS — Current AWS trust conditions and official authentication action pattern.
- OIDC in Azure — Current Microsoft Entra workload identity federation pattern.
- OIDC in Google Cloud — Current Workload Identity Federation integration.
- Managing environments — Environment protection, deployment branches/tags, secrets and target governance.
- Deployments and environments — Conceptual distinction between deployment records, environment gates and external target health.
- kind quick start — Pinned local Kubernetes cluster tool used by the free lab.
- kind local registry — Current localhost registry/network model when digest-addressable local images are needed.
- Kubernetes Deployments — Rollout, revision and rollback model.
- kubectl rollout — Status, history, undo and restart operations.
- actions/checkout v7.0.1 — Pinned source checkout used only in the subject job.
- actions/upload-artifact v7.0.1 — Pinned transfer/evidence action.
- actions/download-artifact v8.0.1 — Pinned deployment-subject retrieval action.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.