Continuous Delivery to AWS, Azure, Google Cloud, and Kubernetes: Diagnostics, Failure Modes, and Production Practices
Diagnose deployment failures by preserving source, artifact, identity, context, rollout, health, and external-target evidence before correction or rollback.
Learning objectives
- Use an evidence-first sequence that separates GitHub, credential, runner, provider, Kubernetes and application-health failures.
- Diagnose wildcard trust, wrong context, digest drift and missing health verification without broadening permissions.
- Preserve first-failure run/attempt and target evidence before rollback or cleanup.
- Explain why automatic rollback can hide the original failure if evidence is not captured first.
- Apply the least destructive correction and rerun the smallest equivalent scope.
1. Evidence-first diagnostic sequence
- Preserve the first failure: run ID, attempt, source/workflow SHA, logs, artifact digest and any target-side deployment ID.
- Confirm selection: event/ref, workflow revision, environment name and evaluated conditions.
- Confirm permissions/identity: job permissions, OIDC audience/subject expectation, provider principal/account/project/subscription.
- Confirm runner/tooling: runner image, CLI versions, network reachability and exact action SHAs.
- Confirm target: cloud resource ID or Kubernetes context/cluster/namespace before mutation.
- Confirm subject: deployed package/image/content digest equals the verified release digest.
- Confirm rollout: provider deployment state, Kubernetes events/revision/ready replicas.
- Confirm health: endpoint/application checks independent of controller success.
- Correct minimally: trust condition, target selector, digest or health configuration—not every layer at once.
- Rerun smallest equivalent scope only after the first failure is retained.
2. Intentionally broken deployment pattern
The following fragment is deliberately unsafe. Read it as evidence of multiple independent design failures rather than a recipe.
# BROKEN production pattern — do not use.
permissions: write-all
steps:
- run: echo "$CLOUD_ADMIN_JSON" > admin.json
- run: kubectl config use-context production-ish
- run: kubectl set image deployment/web web=ghcr.io/example/web:latest
- run: kubectl rollout undo deployment/web || true
The broad token is unrelated to the actual Kubernetes problem, the
admin cloud secret is long-lived and written to disk, the context
name is not verified against an exact cluster identity, the mutable
latest tag can drift from the verified artifact, and
|| true can hide rollback failure. Fixing only one line
would not make the chain trustworthy.
3. Failure: wildcard OIDC trust
Symptom: an unexpected repository or environment can exchange a GitHub OIDC token for the production role. Preserve provider audit logs and the accepted claims. The causal layer is provider trust, not the GitHub runner. Replace organization-wide or repository-wide wildcards with the actual repository/environment/ref conditions supported by the provider, and retain least-privilege role permissions as a second boundary.
Do not log a real JWT to debug this. Use provider denial messages, bounded claim metadata, synthetic claim simulation and the current GitHub OIDC reference.
4. Failure: cloud administrator credential stored as a GitHub secret
Symptom: the deploy “works,” but compromise of any privileged job exposes a reusable account-wide key. This is a design failure even if no run is red. Migrate to OIDC federation, narrow the provider role to one target/action set, place the deployment job behind the environment boundary, and revoke the static key after validation. Do not retain it indefinitely as an undocumented fallback.
5. Failure: verified digest differs from deployed image tag
Symptom: release evidence says sha256:AAA, but target
inspection reports sha256:BBB. Preserve both identities
before changing the target. The failure is
artifact-selection/promotion, not “Kubernetes flakiness.”
Resolve/deploy the exact verified digest, then verify the target
reports that same digest. Never retag latest and call
the mismatch repaired without recording the change.
6. Failure: kubectl targets the wrong cluster or namespace
A successful kubectl apply can be the most dangerous
outcome when the context is wrong. Capture
kubectl config current-context, cluster-info, namespace
and resource UID. Compare them to the deployment contract. If they
differ, stop; do not apply “just to test.” Select the exact intended
context and re-run only the deployment scope after confirming no
unauthorized mutation occurred.
7. Failure: rollout is green, application is unhealthy
Kubernetes may report a Deployment available while the application returns wrong content or a dependency path is broken. A cloud service may report a completed deployment while its public route fails. Keep controller rollout status and application health as separate evidence. Add readiness/startup probes where appropriate and an external or provider-level health check that verifies the promoted identity/behavior.
8. Failure: automatic rollback erases the scene
Automatic rollback can reduce user impact, but if it executes before evidence capture you may lose the failed Pod template, image ID, event stream or provider revision. A safer workflow first records the failed target revision, deployed digest, events/log references and health response, then performs a bounded rollback to the known-good revision and verifies recovery. The failed GitHub run should remain failed; rollback success is a separate recovery result.
Recovery is not a reason to rewrite history. Preserve the original run/attempt and target-side failure record. A second recovery job or controlled operator action may restore service while the first failure stays visible.
9. Provider denial, API and network failures
| Symptom | Likely layer | Evidence | Least-destructive response |
|---|---|---|---|
| OIDC exchange denied | trust/audience/subject | provider error + expected claim policy | repair exact trust condition; do not broaden to wildcard |
| Authenticated, mutation denied | provider IAM/RBAC | principal + denied resource/action | grant only required operation on exact target |
| API timeout/429 | provider/network/rate limit | request ID, status, Retry-After if present | bounded retry/backoff; do not blind-loop |
| kubectl Unauthorized | cluster auth/RBAC | context, credential source, API response | refresh short-lived auth / fix scoped RBAC |
| rollout timeout | workload/runtime | events, pod status, image pull/probes | fix workload cause, preserve revision, then rerun |
| health mismatch | application/external target | endpoint/body/version header | rollback or repair app config; controller success remains true but insufficient |
10. Safer production shape
# Safer shape: identity + target + digest are explicit.
permissions:
contents: read
id-token: write
environment: production
concurrency:
group: deploy-production
cancel-in-progress: false
# after provider authentication and exact target selection:
# test "$ACTUAL_PROVIDER_ACCOUNT" = "$EXPECTED_PROVIDER_ACCOUNT"
# test "$VERIFIED_IMAGE" = "ghcr.io/example/web@sha256:<expected>"
# kubectl config current-context
# kubectl set image deployment/web web="$VERIFIED_IMAGE"
# kubectl rollout status deployment/web --timeout=5m
# health check, then record target revision
The comments are deliberate: each deployment system needs its own exact identity assertions. The YAML does not pretend one generic command can prove AWS account ID, Azure subscription, GCP project and Kubernetes context simultaneously.
11. Diagnostics summary
Deployment diagnosis is a chain-of-custody problem. Preserve the first failure, identify the exact layer where observed state diverges from the contract, make one bounded correction, then prove both target state and health.
Knowledge check
A deployment job authenticates successfully but receives HTTP 403 on the target API. Should you widen OIDC trust?
No. Authentication already succeeded. Inspect the provider role/resource authorization and grant only the missing exact action if justified.
Why can kubectl rollout status succeed while
delivery is still bad?
It proves controller rollout availability, not application correctness or external routing. A separate health check is required.
What must be preserved before automatic rollback?
The failed deployment/revision identity, digest, events/log references, health evidence and originating GitHub run/attempt.
Verified image is @sha256:AAA, target runs
@sha256:BBB. What layer failed?
Artifact selection/promotion. Do not troubleshoot runner capacity or OIDC until the deployed subject matches the verified digest.
Why is “use a cloud admin secret so deploy always works” an invalid recovery strategy?
It destroys least privilege, creates a reusable high-value credential and hides the actual trust/authorization problem.
Official references and version notes
- GitHub OIDC overview — Why short-lived federation is preferred to long-lived cloud secrets.
- OIDC reference — Current claims, immutable subject format, audiences, reusable-workflow behavior and token-request permissions.
- OIDC in AWS — Current AWS trust conditions and official authentication action pattern.
- OIDC in Azure — Current Microsoft Entra workload identity federation pattern.
- OIDC in Google Cloud — Current Workload Identity Federation integration.
- Managing environments — Environment protection, deployment branches/tags, secrets and target governance.
- Deployments and environments — Conceptual distinction between deployment records, environment gates and external target health.
- kind quick start — Pinned local Kubernetes cluster tool used by the free lab.
- kind local registry — Current localhost registry/network model when digest-addressable local images are needed.
- Kubernetes Deployments — Rollout, revision and rollback model.
- kubectl rollout — Status, history, undo and restart operations.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.