Checkpoint Lab — Capstone: Build a Secure Reusable Enterprise CI/CD Platform
Deliver a working disposable platform and a production-readiness dossier that proves control effectiveness, residual risks, rollback readiness and a sustainable post-course operating plan.
Learning objectives
- Deliver a working disposable capstone and production-readiness dossier.
- Predict and independently verify important state transitions.
- Prove build-once artifact identity and provenance.
- Prove deployment authorization, target verification and incident-safe recovery.
- Document cost assumptions, governance, residual risks and a post-course operating plan.
1. Checkpoint mission: prove the platform, not just the YAML
The final checkpoint produces two deliverables: a working disposable platform demonstration and a production-readiness dossier. The demo proves the mechanics. The dossier explains why those mechanics are trustworthy, where they are intentionally simulated, what evidence each control creates, which risks remain and how the platform will be operated after the course.
Use only repositories/resources created for this lab. The primary path uses two public disposable GitHub repositories so standard hosted runners, environments and artifact attestations are available without paid infrastructure. The fallback local path must record that it cannot prove GitHub-hosted run identity, real environment approval, OIDC issuance or GitHub attestation verification.
2. Predict state changes before execution
| Prediction | Before | After successful run | Independent verification |
|---|---|---|---|
| CI identity | no run | one run bound to exact app SHA + platform SHA | run metadata + committed caller refs |
| Artifact | no release record | one artifact ID + service digest + file SHA | upload outputs + downloaded file hash |
| Provenance | none | attestation bound to release bytes on public path | gh attestation verify |
| Deployment authorization | environment idle | deployment job passes environment protection | environment/deployment record |
| Target state | no promoted file | simulated target contains exact artifact digest | readback SHA + health check |
| Recovery drill | no incident | failed run preserved, then separate recovery run succeeds | two run IDs + same intended source/artifact correlation |
| Governance | policy unevaluated | policy digest + pass/fail result | custom action output + workflow log |
Write your predictions to
evidence/predictions.md before triggering anything. At
least two predictions must be falsifiable—for example, “a PR cannot
create a deploy job” and “the deployment file hash equals the
reusable CI subject SHA.”
3. Production-readiness dossier structure
production-readiness/
00-scope-and-assumptions.md
01-event-source-and-workflow-identity.md
02-platform-dependencies-and-permissions.md
03-runner-trust-and-toolchain.md
04-ci-test-and-job-graph.md
05-artifact-and-provenance.md
06-deployment-authorization-and-target.md
07-observability-and-first-failure.md
08-recovery-and-rollback.md
09-usage-cost-and-capacity.md
10-governance-exceptions-and-audit.md
11-residual-risks.md
12-post-course-operating-plan.md
evidence/
predictions.md
runs.json
action-pins.txt
policy-result.txt
artifact-identity.json
attestation-verification.txt
deployment-evidence.txt
recovery-correlation.json
Do not place secrets, raw OIDC tokens or private infrastructure details in the dossier. Record credential names/scopes, trust-policy predicates and claim subsets, not bearer material.
4. Build and execute the acceptance sequence
- Platform release: commit the reusable CI workflow and policy action; capture the exact platform SHA and a human-readable v1 mapping.
-
Caller release: generate the app workflow with
literal platform/action SHAs and top-level
permissions: {}. - PR proof: open a disposable PR and prove policy + CI execute while deployment does not exist.
-
Main/manual proof: dispatch with
deploy=trueandinject-failure=false; record run/attempt, source SHA, job graph, runner/tool versions, artifact outputs, attestation result and deployment record. -
Incident proof: dispatch with
inject-failure=true; preserve first-run evidence after the side effect. - Recovery proof: inspect target/digest, then create a new safe dispatch. Correlate the failed and recovered runs without overwriting the first evidence.
- Policy negative test: on a temporary branch, introduce one mutable external action reference and prove the custom policy action blocks it; revert the branch after preserving denial evidence.
5. Record the dependency and permission manifest
Verified on 2026-09-10
actions/checkout v7.0.1 = 3d3c42e5aac5ba805825da76410c181273ba90b1
actions/setup-python v7.0.0 = 5fda3b95a4ea91299a34e894583c3862153e4b97
actions/upload-artifact v7.0.1 = 043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
actions/download-artifact v8.0.1 = 3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c
actions/attest v4.2.2 = 1e69f48acb82d1966a394da916b4c1698aa569d6
actions/cache v6.1.0 = 55cc8345863c7cc4c66a329aec7e433d2d1c52a9
hosted runner label = ubuntu-24.04
REST API version used for explicit API examples = 2026-03-10
reusable-workflow limit = 10 connected levels / 50 unique callees
OIDC default for repos created after 2026-07-15 = immutable owner/repository ID subject
Add your learner-generated PLATFORM_SHA, app source
SHA, workflow file SHA and any other executable dependency you
introduced. If a version changed after this course was generated,
update the dossier to the version actually executed and re-run the
compatibility/security review.
6. Prove build-once identity and provenance
The same release bytes must survive the transition from CI to deployment. Capture the reusable workflow outputs, download by exact artifact ID, verify the subject SHA-256 and then verify the attestation on the public GitHub path. If attestation is unavailable in your repository visibility/plan, the local fallback records a signed-attestation gap as residual risk rather than fabricating proof.
# After downloading the artifact from the exact successful run:
sha256sum dist/app.tgz
gh attestation verify dist/app.tgz -R OWNER/gha-capstone-app | tee production-readiness/evidence/attestation-verification.txt
# Compare the file digest to the CI output recorded before deploy.
7. Prove authorization, OIDC boundary and target health
For the GitHub path, preserve the environment/deployment record and the allowlisted OIDC claim output. The environment decision proves authorization sequencing; the OIDC claim subset proves which GitHub identity would be presented to a provider; the simulated target readback proves which bytes were promoted. These are three different controls and should appear as three different evidence items.
A production cloud deployment would replace
capstone.local with a provider-specific audience and an
external trust policy constrained to the exact
repository/environment/workflow identity. The lab intentionally does
not provision AWS, Azure, Google Cloud or Kubernetes. Provider
credentials and managed infrastructure are optional extensions, not
graduation requirements.
8. Incident drill and rollback readiness
The controlled failure must leave attempt-one evidence intact. Record the failure step, exit code, deployment record, artifact digest and target readback. Explain why a rerun repeats the original input and why the safe recovery is a new dispatch or a compensation/rollback chosen from observed state. Then identify a rollback artifact by exact digest. If no prior artifact exists in the tiny lab, state that the production rollback requirement is not satisfied until a previous known-good release has been retained.
{
"incident_run_id": 8123456789,
"incident_attempt": 1,
"source_sha": "<exact source SHA>",
"artifact_sha256": "<exact subject digest>",
"side_effect_observed": true,
"health_before_recovery": "not-evaluated",
"recovery_run_id": 8123456790,
"recovery_kind": "new workflow_dispatch",
"final_health": "verified",
"rollback_target": "<previous known-good digest or explicitly unavailable>"
}
9. Cost, governance and residual-risk review
| Control | Evidence | Residual question |
|---|---|---|
| Hosted runner isolation | runner label/image/tool versions | does any workload truly require self-hosted/private networking? |
| Least privilege | top-level deny + per-job permissions | can any reusable component obtain authority it does not need? |
| Dependency integrity | full SHA pins + release mapping | who reviews and upgrades the pins? |
| Artifact integrity | ID + service digest + file SHA + attestation | how long are rollback artifacts/provenance retained? |
| Deployment gate | environment + concurrency + target readback | who approves production and how is emergency access audited? |
| OIDC | actual claim subset + planned trust predicates | is provider policy constrained to repo/environment/workflow/audience? |
| Recovery | first-failure packet + exact rollback target | are side effects idempotent across partial failures? |
| Cost/capacity | run durations/queue/usage assumptions | what SLO justifies larger or self-hosted runners? |
| Governance | policy result + exception model | which controls require higher-plan enforcement versus repository convention? |
Residual risk is not failure. Undocumented residual risk is failure. Examples may include no independent security review of the custom policy action, no real provider federation test, short artifact retention, no enterprise audit export, no self-hosted runner isolation test or no previous release available for rollback. Assign each accepted risk an owner and next review date.
10. Post-course operating plan
- Weekly: review platform workflow/action dependency updates, failed/retried runs and stale exceptions.
- Monthly: sample caller pins/adoption, artifact/provenance verification, environment permissions and runner access.
- Quarterly: rehearse rollback/credential-revocation, review OIDC trust policies, runner images/groups, retention and cost/capacity.
- On incident: preserve first evidence, freeze unnecessary mutation, reconcile exact external state, recover with the least destructive action and record follow-up controls.
- On platform release: publish version→SHA mapping, compatibility-test representative callers, stage adoption and retain a known rollback version.
- On GitHub platform change: re-verify runner images, action runtimes, API versioning, permission keys, OIDC claims, artifact behavior and plan boundaries before changing production assumptions.
11. Graduation criteria and course close
The capstone is complete when another engineer can answer, from evidence rather than memory: what source and workflow code ran; which runner and executable dependencies were trusted; which permissions and identities were available; which tests passed; which exact bytes were produced; how provenance was verified; who or what authorized deployment; which target changed; whether it became healthy; what the first failure looked like; how recovery was chosen; what the run cost/capacity assumptions were; and which policy or exception allowed the delivery.
That operating model is the durable outcome of the course. GitHub Actions syntax will continue to evolve, but the discipline of explicit identity, least privilege, immutable evidence, guarded side effects, causal diagnosis and verifiable governance remains the foundation for production CI/CD.
Knowledge check
What is the minimum evidence that proves build-once promotion?
The build run/source identity, exact artifact record, subject file digest, and a deployment readback showing the same bytes were promoted; provenance verification strengthens the chain.
Why should the capstone record both the platform SHA and app SHA?
They identify different executable inputs: application source/workflow intent and centralized reusable automation. Reproducibility needs both.
A public PR passes CI but the deploy job never appears. Is that necessarily a failure?
No. In the capstone it is the intended trust boundary: pull requests test but cannot request the protected deployment path.
When is a residual risk acceptable in the dossier?
When it is explicit, scoped, justified, assigned an owner, paired with compensating controls where needed, and has a review/remediation plan.
What is the post-course response to a GitHub platform change?
Re-verify the affected assumptions and executable versions against current primary documentation, update pins/contracts deliberately, and preserve compatibility evidence.
Official references and version notes
- Workflow syntax for GitHub Actions — Current workflow/job/permissions/runner syntax and hosted-runner behavior.
- Reusing workflow configurations — Current reusable-workflow access, nesting and call-tree limits.
- Secure use reference — Least privilege, untrusted-input handling and full-SHA dependency guidance.
- Deployments and environments — Environment approvals, secrets, protection rules and deployment boundaries.
- OpenID Connect reference — OIDC claim semantics including immutable subject claims introduced in 2026.
- Artifact attestations — Provenance model and verification expectations.
- Using artifact attestations — Current permissions and actions/attest workflow pattern.
- GitHub-hosted runners reference — Current runner labels, images, hardware and billing boundaries.
- Runner groups — Runner-group trust boundary and access-control model.
- GitHub Actions billing and usage — Current public/private hosted-runner and usage accounting model.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.