Capstone: Build a Secure Reusable Enterprise CI/CD Platform: Diagnostics, Failure Modes, and Production Practices
Diagnose end-to-end CI/CD failures by preserving first evidence and separating workflow, runner, identity, artifact, deployment, API and governance state before recovery.
Learning objectives
- Preserve first-failure evidence before mutation or rerun.
- Diagnose mutable dependencies, broad permissions, untrusted runner access and artifact identity loss.
- Choose rerun, new run, compensation or rollback from observed state.
- Use known artifact digests for rollback rather than rebuilding.
- Treat missing observability/audit evidence as an operational defect.
1. Diagnose the chain, not the final red box
A production delivery can fail before a runner starts, inside a build step, while uploading or downloading evidence, at an environment gate, during OIDC exchange, during target mutation, or after mutation when health verification fails. Treating all of these as “workflow failure” destroys the information needed for safe recovery. The first job is evidence preservation.
Record the original run ID, attempt, source SHA, event, workflow revision, job conclusions and artifact/deployment identifiers before changing configuration or rerunning. If the external target changed, inspect it read-only before cleanup. A rerun can repeat side effects, and a new commit is a new source identity even if the engineer calls it “the same retry.”
2. Evidence-first diagnostic sequence
- Preserve: run ID/attempt, URLs, first-failure logs, artifacts and target readback.
- Identity: confirm event, ref, source SHA and caller/reusable workflow revisions.
- Evaluation: confirm conditions, typed inputs, permissions and secrets/OIDC availability.
- Graph/queue: confirm which jobs existed, were skipped, queued, cancelled or blocked by environment/concurrency.
- Runner: record label, image/OS, architecture and exact tool versions.
- Step/runtime: isolate shell/action/runtime/network error without broadening privileges.
- Evidence services: inspect outputs, artifact/cache IDs, digests and attestation verification.
- Deployment: inspect environment decision, deployment record and external target state.
- Correct minimally: change only the causal layer, then rerun the smallest equivalent scope or create a new run when event/source/input changes.
3. Failure: a mutable dependency entered the trusted path
Suppose the platform caller references
vendor/action@v2. The app source SHA has not changed,
but a later run now executes different dependency bytes. The failure
could appear as a changed output, permission request or network
call. Do not “fix” the symptom by pinning whatever commit happens to
work. Preserve the run, resolve the tag to its exact commit, review
the upstream delta, map a trusted release to a full SHA and update
through the platform’s normal dependency process.
Production rule: full-length commit SHA pinning is currently the only immutable action reference GitHub recommends. A Marketplace badge or major tag is not release provenance.
4. Failure: broad token or secret inheritance
A reusable workflow that accepts implicit secret inheritance or a
caller that grants broad write permissions expands the blast radius
of every step in the called workflow. Diagnose the required
operation first: source checkout needs contents: read;
attestation needs contents: read,
id-token: write and attestations: write; a
simulated local deployment needs no repository write permission.
Separate privileged jobs and pass only typed non-secret outputs
where possible.
If a secret may have been printed or sent to an unintended host, treat that as credential exposure: preserve evidence, revoke/rotate at the provider, remove the value from the workflow and investigate downstream use. Log redaction alone does not make exposure safe.
5. Failure: untrusted PR code reaches a privileged runner or deploy job
The first question is not “why did checkout fail?” but “why was
untrusted code eligible for this trust zone?” Fork pull-request CI
should remain on isolated hosted runners with minimal permissions. A
privileged self-hosted runner, environment secret or cloud identity
belongs behind a trust transition that never executes
attacker-controlled code. Do not repair a blocked
pull_request_target checkout by opting into unsafe
checkout just to make the build green.
The least destructive fix is architectural: keep untrusted build/test in the low-trust lane, produce non-authoritative evidence, and allow a trusted workflow to consume only validated metadata or immutable artifacts under an explicit policy.
6. Failure: artifact identity is lost between build and deploy
An artifact named app is not enough. If deploy
downloads by a mutable name, rebuilds from source, or trusts a log
line, the production bytes may not be the CI-reviewed bytes.
Preserve the artifact ID, artifact-service digest, file SHA-256,
source SHA and attestation result. Download the exact artifact
record, verify the file digest, and only then promote it.
# Read-only evidence pattern after downloading the exact artifact record:
sha256sum dist/app.tgz
gh attestation verify dist/app.tgz -R OWNER/REPOSITORY
# Compare the verified subject digest to the CI output captured before deployment.
7. Failure: rollback rebuilds instead of promoting a known artifact
Rebuilding “the previous tag” during an incident reopens dependency, runner-image and toolchain uncertainty. A rollback target should be a previously verified artifact digest with known provenance and known target compatibility. If that artifact is no longer retained, say rollback evidence is incomplete; do not disguise a rebuild as equivalent.
The recovery action can be a forward fix, idempotent re-apply, compensating operation or rollback. Choose based on the external state already changed. Concurrency prevents overlapping GitHub jobs in one group, but it is not a universal external lock; target-native serialization or idempotency may still be required.
8. Intentionally broken incident and interpretation
Consider a deployment run where the artifact was copied to the target, the controlled incident step exited 42, and the job failed before health verification. The evidence says mutation occurred, health is unknown, and the deployment record is failed. A blind rerun with the same injected-failure input repeats the failure. Deleting the target would erase evidence and may worsen the incident.
run_id=8123456789
attempt=1
source_sha=7b7e...c91
artifact_id=412345678
subject_sha256=91e4...c21
environment=capstone-production
side_effect=target copy completed
health=not evaluated
failure_step=Controlled incident injection
exit_code=42
The repair is to inspect the target digest first, preserve the
failed deployment evidence, then start a new dispatch with the same
source revision and inject-failure=false. That new run
is correlated to the first by source SHA and artifact digest, not
mislabeled as attempt 2. If the target digest is wrong, stop and
choose rollback/compensation instead.
9. No first-failure evidence or audit trail is itself a production defect
If logs have expired, deployment IDs were never recorded, mutable dependencies were used and no target digest exists, recovery confidence is low. The corrective action is not only to fix today’s deployment; add retention/export, structured summaries, artifact/provenance records and policy/adoption telemetry so the next incident has a reconstructable chain. Enterprise audit logs can strengthen this, but every plan can still record run IDs, workflow SHAs, artifact identities and target verification in application-owned evidence.
10. From incident handling to production-readiness proof
Lesson 5 turns these controls into an acceptance dossier. The checkpoint is passed only when the platform can demonstrate what it trusts, what it changes, how evidence proves each transition, how it fails closed, how it recovers, what it costs and which residual risks remain.
Knowledge check
A deployment failed after copying bytes but before health checks. What should happen first?
Preserve the run/attempt and inspect the target read-only, including its exact artifact digest. Do not rerun or delete the target before knowing the side-effect state.
Why is changing a workflow file and then calling the next run “attempt 2” incorrect?
A source/workflow revision change creates a new run identity. Rerun attempts repeat the original run’s event/source identity.
What is the least safe response to a mutable action tag causing drift?
Blindly moving to another tag or branch. Resolve and review the exact upstream commit, then pin an approved full SHA.
Why can concurrency alone not guarantee deployment idempotency?
It serializes matching GitHub jobs, but external systems can have other writers and partial side effects. Target-native idempotency/locking may still be required.
What evidence makes a rollback target trustworthy?
A known retained artifact digest, provenance/source identity, prior verification/deployment record and target compatibility—not a recipe to rebuild old source.
Official references and version notes
- Workflow syntax for GitHub Actions — Current workflow/job/permissions/runner syntax and hosted-runner behavior.
- Reusing workflow configurations — Current reusable-workflow access, nesting and call-tree limits.
- Secure use reference — Least privilege, untrusted-input handling and full-SHA dependency guidance.
- Deployments and environments — Environment approvals, secrets, protection rules and deployment boundaries.
- OpenID Connect reference — OIDC claim semantics including immutable subject claims introduced in 2026.
- Artifact attestations — Provenance model and verification expectations.
- Using artifact attestations — Current permissions and actions/attest workflow pattern.
- GitHub-hosted runners reference — Current runner labels, images, hardware and billing boundaries.
- Runner groups — Runner-group trust boundary and access-control model.
- GitHub Actions billing and usage — Current public/private hosted-runner and usage accounting model.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.