Chapter 26Lesson 04~190 minutes

Continuous Delivery to AWS, Azure, Google Cloud, and Kubernetes: Diagnostics, Failure Modes, and Production Practices

Diagnose deployment failures by preserving source, artifact, identity, context, rollout, health, and external-target evidence before correction or rollback.

DiagnosticsWrong targetDigest driftFirst failureRecovery

Learning objectives

  • Use an evidence-first sequence that separates GitHub, credential, runner, provider, Kubernetes and application-health failures.
  • Diagnose wildcard trust, wrong context, digest drift and missing health verification without broadening permissions.
  • Preserve first-failure run/attempt and target evidence before rollback or cleanup.
  • Explain why automatic rollback can hide the original failure if evidence is not captured first.
  • Apply the least destructive correction and rerun the smallest equivalent scope.

1. Evidence-first diagnostic sequence

  1. Preserve the first failure: run ID, attempt, source/workflow SHA, logs, artifact digest and any target-side deployment ID.
  2. Confirm selection: event/ref, workflow revision, environment name and evaluated conditions.
  3. Confirm permissions/identity: job permissions, OIDC audience/subject expectation, provider principal/account/project/subscription.
  4. Confirm runner/tooling: runner image, CLI versions, network reachability and exact action SHAs.
  5. Confirm target: cloud resource ID or Kubernetes context/cluster/namespace before mutation.
  6. Confirm subject: deployed package/image/content digest equals the verified release digest.
  7. Confirm rollout: provider deployment state, Kubernetes events/revision/ready replicas.
  8. Confirm health: endpoint/application checks independent of controller success.
  9. Correct minimally: trust condition, target selector, digest or health configuration—not every layer at once.
  10. Rerun smallest equivalent scope only after the first failure is retained.

2. Intentionally broken deployment pattern

The following fragment is deliberately unsafe. Read it as evidence of multiple independent design failures rather than a recipe.

# BROKEN production pattern — do not use.
permissions: write-all
steps:
  - run: echo "$CLOUD_ADMIN_JSON" > admin.json
  - run: kubectl config use-context production-ish
  - run: kubectl set image deployment/web web=ghcr.io/example/web:latest
  - run: kubectl rollout undo deployment/web || true

The broad token is unrelated to the actual Kubernetes problem, the admin cloud secret is long-lived and written to disk, the context name is not verified against an exact cluster identity, the mutable latest tag can drift from the verified artifact, and || true can hide rollback failure. Fixing only one line would not make the chain trustworthy.

3. Failure: wildcard OIDC trust

Symptom: an unexpected repository or environment can exchange a GitHub OIDC token for the production role. Preserve provider audit logs and the accepted claims. The causal layer is provider trust, not the GitHub runner. Replace organization-wide or repository-wide wildcards with the actual repository/environment/ref conditions supported by the provider, and retain least-privilege role permissions as a second boundary.

Do not log a real JWT to debug this. Use provider denial messages, bounded claim metadata, synthetic claim simulation and the current GitHub OIDC reference.

4. Failure: cloud administrator credential stored as a GitHub secret

Symptom: the deploy “works,” but compromise of any privileged job exposes a reusable account-wide key. This is a design failure even if no run is red. Migrate to OIDC federation, narrow the provider role to one target/action set, place the deployment job behind the environment boundary, and revoke the static key after validation. Do not retain it indefinitely as an undocumented fallback.

5. Failure: verified digest differs from deployed image tag

Symptom: release evidence says sha256:AAA, but target inspection reports sha256:BBB. Preserve both identities before changing the target. The failure is artifact-selection/promotion, not “Kubernetes flakiness.” Resolve/deploy the exact verified digest, then verify the target reports that same digest. Never retag latest and call the mismatch repaired without recording the change.

6. Failure: kubectl targets the wrong cluster or namespace

A successful kubectl apply can be the most dangerous outcome when the context is wrong. Capture kubectl config current-context, cluster-info, namespace and resource UID. Compare them to the deployment contract. If they differ, stop; do not apply “just to test.” Select the exact intended context and re-run only the deployment scope after confirming no unauthorized mutation occurred.

7. Failure: rollout is green, application is unhealthy

Kubernetes may report a Deployment available while the application returns wrong content or a dependency path is broken. A cloud service may report a completed deployment while its public route fails. Keep controller rollout status and application health as separate evidence. Add readiness/startup probes where appropriate and an external or provider-level health check that verifies the promoted identity/behavior.

8. Failure: automatic rollback erases the scene

Automatic rollback can reduce user impact, but if it executes before evidence capture you may lose the failed Pod template, image ID, event stream or provider revision. A safer workflow first records the failed target revision, deployed digest, events/log references and health response, then performs a bounded rollback to the known-good revision and verifies recovery. The failed GitHub run should remain failed; rollback success is a separate recovery result.

Recovery is not a reason to rewrite history. Preserve the original run/attempt and target-side failure record. A second recovery job or controlled operator action may restore service while the first failure stays visible.

9. Provider denial, API and network failures

Symptom Likely layer Evidence Least-destructive response
OIDC exchange denied trust/audience/subject provider error + expected claim policy repair exact trust condition; do not broaden to wildcard
Authenticated, mutation denied provider IAM/RBAC principal + denied resource/action grant only required operation on exact target
API timeout/429 provider/network/rate limit request ID, status, Retry-After if present bounded retry/backoff; do not blind-loop
kubectl Unauthorized cluster auth/RBAC context, credential source, API response refresh short-lived auth / fix scoped RBAC
rollout timeout workload/runtime events, pod status, image pull/probes fix workload cause, preserve revision, then rerun
health mismatch application/external target endpoint/body/version header rollback or repair app config; controller success remains true but insufficient

10. Safer production shape

# Safer shape: identity + target + digest are explicit.
permissions:
  contents: read
  id-token: write

environment: production
concurrency:
  group: deploy-production
  cancel-in-progress: false

# after provider authentication and exact target selection:
# test "$ACTUAL_PROVIDER_ACCOUNT" = "$EXPECTED_PROVIDER_ACCOUNT"
# test "$VERIFIED_IMAGE" = "ghcr.io/example/web@sha256:<expected>"
# kubectl config current-context
# kubectl set image deployment/web web="$VERIFIED_IMAGE"
# kubectl rollout status deployment/web --timeout=5m
# health check, then record target revision

The comments are deliberate: each deployment system needs its own exact identity assertions. The YAML does not pretend one generic command can prove AWS account ID, Azure subscription, GCP project and Kubernetes context simultaneously.

11. Diagnostics summary

Deployment diagnosis is a chain-of-custody problem. Preserve the first failure, identify the exact layer where observed state diverges from the contract, make one bounded correction, then prove both target state and health.

Next lesson

Checkpoint Lab — Continuous Delivery to AWS, Azure, Google Cloud, and Kubernetes

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

A deployment job authenticates successfully but receives HTTP 403 on the target API. Should you widen OIDC trust?

Why can kubectl rollout status succeed while delivery is still bad?

What must be preserved before automatic rollback?

Verified image is @sha256:AAA, target runs @sha256:BBB. What layer failed?

Why is “use a cloud admin secret so deploy always works” an invalid recovery strategy?

Official references and version notes

  • GitHub OIDC overview — Why short-lived federation is preferred to long-lived cloud secrets.
  • OIDC reference — Current claims, immutable subject format, audiences, reusable-workflow behavior and token-request permissions.
  • OIDC in AWS — Current AWS trust conditions and official authentication action pattern.
  • OIDC in Azure — Current Microsoft Entra workload identity federation pattern.
  • OIDC in Google Cloud — Current Workload Identity Federation integration.
  • Managing environments — Environment protection, deployment branches/tags, secrets and target governance.
  • Deployments and environments — Conceptual distinction between deployment records, environment gates and external target health.
  • kind quick start — Pinned local Kubernetes cluster tool used by the free lab.
  • kind local registry — Current localhost registry/network model when digest-addressable local images are needed.
  • Kubernetes Deployments — Rollout, revision and rollback model.
  • kubectl rollout — Status, history, undo and restart operations.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.