Chapter 38Lesson 05~200 minutes

Production Capstone: Build, Secure, Scale, Observe, and Govern a Complete GitLab Delivery Platform: Final Operational Review and Handoff

Complete the GitLab CI/CD course with an operational handoff: architecture inventory, source-to-target proof, SLOs, runbooks, upgrade/deprecation watch, exception register, residual risks and next actions.

Final handoffOperational reviewEvidence packetCourse complete

Learning objectives

  • Produce a machine-readable architecture/configuration inventory without secret material.
  • Verify the final source/artifact/target evidence chain and distinguish producer SHA from current repository HEAD.
  • Define SLOs, dashboards, alerts and recovery runbooks with owners.
  • Create an upgrade/deprecation watch list, exception/residual-risk register and production next actions.
  • Package and hash the final handoff evidence and state what post-course production practice requires.

Final handoff boundary: this lesson produces a production-readiness packet and operating model, not a claim that the disposable lab itself is production infrastructure. Paid/admin/cloud capabilities remain optional mappings. The final review must list every simulation, unresolved dependency, experimental feature and residual risk instead of hiding them.

1. Final operational review: prove the platform as a system

The final review is a handoff to the engineers who will operate the delivery platform after the course. A useful handoff must be executable: architecture and trust boundaries, source/configuration inventory, runner fleet assumptions, identity matrix, artifact/evidence flows, environments and external dependencies, policy and exception state, SLOs/dashboards, runbooks, upgrade/deprecation ownership, residual risks and concrete next actions.

Do not hand off only a diagram or a YAML file. Those describe intent. The packet also needs observations from the last known-good run and the recovery drill.

2. Preflight and version assumptions

Record the environment before finalizing the packet.

LAB="${LAB:-${TMPDIR:-/tmp}/gitlab-ch38-capstone}"
cd "$LAB"
mkdir -p handoff
{
  printf 'checked_at_utc=%s\n' "$(date -u +%FT%TZ)"
  printf 'git=%s\n' "$(git --version)"
  printf 'python=%s\n' "$(python --version 2>&1)"
  printf 'source_sha=%s\n' "$(git rev-parse HEAD)"
  glab --version 2>/dev/null | head -1 | sed 's/^/glab=/' || true
  gitlab-runner --version 2>/dev/null | head -1 | sed 's/^/runner=/' || true
} | tee handoff/versions.env

Reference assumption used by this lesson: GitLab 19.3.2 is the patched 19.3 baseline as of 2026-09-13, and Runner 19.3 is the corresponding monthly Runner release. A real handoff must replace those reference values with observed production versions and exact patch levels.

3. Make final predictions before sealing evidence

Prediction Expected final state Independent verification
P1 — source/config Handoff identifies exact source SHA and configuration inventory Git SHA + file hashes + optional CI Lint/merged config
P2 — artifact Release bundle digest matches provenance and deployed target sha256sum -c + provenance JSON + target JSON
P3 — recovery Failure evidence remains distinct from recovered state Failure/recovery directories + hash manifest
P4 — governance Every exception is owned, scoped and expires; policy simulations are labeled Exception register + limitations note
P5 — operations SLOs/runbooks/upgrade watch have owners and next review date Handoff documents + issue/ticket IDs in production

4. Write the architecture and configuration inventory

Create a machine-readable inventory that points to the evidence instead of duplicating secrets or mutable state.

{
  "service": "Atlas Relay (training)",
  "trust_zones": ["developer-input", "pipeline-control", "runner-execution", "retained-and-external-state"],
  "configuration": [".gitlab-ci.yml", "ci/base.yml"],
  "pipeline_sources": ["merge_request_event", "default-branch push"],
  "runner_classes": ["untrusted-validation", "trusted-release-deploy"],
  "identities": ["CI_JOB_TOKEN where supported", "OIDC id_token for optional external federation"],
  "evidence": ["tests", "synthetic security", "artifact digest", "provenance-shaped JSON", "target verification"],
  "environments": ["staging simulation"],
  "governance": ["policy simulation", "exception register", "failure/recovery manifest"]
}

In production, replace descriptive labels with project IDs/URLs, policy project/ref, component refs, runner IDs, environment names, registry/package locations, external provider account/cluster identifiers and dashboards. Do not put token values or secret material into the inventory.

5. Verify source → artifact → deployment chain one last time

This is the core reproducibility proof. The same artifact digest must connect producer evidence to target state.

sha256sum -c dist/SHA256SUMS
python - <<'PY' | tee handoff/final-chain-verification.json
import json, hashlib, subprocess
prov=json.load(open('evidence/provenance.json'))
target=json.load(open('target/staging.json'))
source=subprocess.check_output(['git','rev-parse','HEAD'], text=True).strip()
actual=hashlib.sha256(open(prov['subject'],'rb').read()).hexdigest()
result={
 'source_sha':source,
 'provenance_source_sha':prov['source_sha'],
 'artifact_sha256':actual,
 'provenance_sha256':prov['sha256'],
 'target_digest':target['artifact_digest'],
 'target_healthy':target['healthy'],
}
result['chain_ok']=(actual==prov['sha256'] and target['artifact_digest']=='sha256:'+actual and target['healthy'])
print(json.dumps(result,indent=2))
assert result['chain_ok']
PY

The lab may have additional commits after the artifact build, so the current repository HEAD can differ from provenance_source_sha. That is not automatically a failure. The release identity is the producer source recorded in provenance; the handoff must distinguish “current repo HEAD” from “source SHA that built the deployed artifact.”

6. Define SLOs, dashboards and alert ownership

The local lab cannot provide GitLab queue metrics, so the handoff records the production measurements that must exist.

Signal / SLO Target example Data source Owner / response
Pipeline creation success 99.5% valid intended events create a pipeline GitLab pipeline/API metrics Platform team; investigate rules/config regressions
Queue time p95 < 60 s for validation GitLab CI analytics / Runner metrics Runner team; capacity/routing review
Validation critical path p95 < 8 min Pipeline/job timestamps App + platform; optimize measured bottleneck only
Artifact traceability 100% release artifacts have source SHA + digest Artifacts/registry + provenance ledger Release engineering; block promotion if missing
Deployment verification 100% deploys verify target digest and health Environment/deployment + target telemetry Service owner; rollback/roll-forward playbook
Exception hygiene 100% exceptions owned and unexpired Governance register/policy system Security/platform owner; expire automatically where possible

7. Handoff the runbooks, not tribal knowledge

At minimum keep these runbooks beside the platform configuration:

  • Pipeline creation/configuration failure: source/ref/SHA → CI Lint/merged config → workflow/job rules → new pipeline after config fix.
  • Pending/runner failure: job tags/protection → runner availability/version/executor → host/VM/container/network diagnostics → least-destructive routing/capacity repair.
  • Dependency/network failure: exact hostname → DNS → route/proxy → TLS trust → application endpoint; never disable TLS as the fix.
  • Artifact/report failure: producer job → artifact/report metadata → digest → transfer/retention/access; never rebuild the release artifact to hide a missing transfer.
  • Identity failure: identity type → token scope/allowlist or OIDC issuer/audience/claims → target trust decision; no broad PAT shortcut.
  • Deployment failure: authorization/approval → deploy job → environment/deployment record → external target version/digest/health → verified rollback or roll-forward.
  • Policy/governance failure: policy project/version/scope → effective configuration → exception identity/expiry → audit evidence → reviewed recovery.

8. Create an upgrade and deprecation watch list

GitLab and Runner evolve monthly, so “works today” is not a handoff strategy. Review release notes, patch releases, deprecations and the docs for every feature you depend on. The current baseline already demonstrates why: GitLab 19.3.2 is a critical patch release newer than the 19.3 monthly release.

Watch item Why it matters Cadence / trigger
GitLab patch releases Security/bug fixes can change safe supported baseline Each scheduled/critical patch release
Monthly GitLab/Runner release notes New CI syntax, executor behavior, APIs, policy/security features Before each minor upgrade
Deprecations/removals Old keywords/APIs/runner flows can break at major/minor boundaries Quarterly + before upgrade
Runner executor/autoscaling guidance Isolation/capacity/security recommendations evolve Before runner fleet changes
CI components/includes/inputs Resolution/context/validation behavior evolves Before shared platform component release
CI_JOB_TOKEN/OIDC Endpoint support, allowlists, claims and provider trust are security-sensitive Before identity-policy changes
Security reports/attestations Report schemas, tiers and experimental status change Before enabling scanner/attestation integration
Policies/audit/environments Tier/offering/precedence/approval behavior evolves Before governance rollout or renewal

9. Exception register and residual risks

Close the handoff with what remains imperfect. For this mandatory lab, the residual-risk list is explicit: no real GitLab pipeline IDs/jobs/runners were created; no actual merge-request pipeline was exercised; security evidence is synthetic, not a vulnerability scan; no real registry/package promotion exists; OIDC is a claim simulation; staging is a local file; protected environments/approvals and pipeline execution policies are simulated; audit APIs are not queried; provenance attestations are not used because the native feature is optional/experimental.

A production handoff should convert each applicable simulation into a tracked action with owner, due date, acceptance criterion and rollback plan.

10. Produce the final evidence packet

Package the evidence, configuration inventory, runbooks and assumptions without secrets. Hash the archive so reviewers can identify the exact handoff packet.

cat > handoff/limitations.md <<'MD'
# Capstone limitations
- Disposable local simulation; no production systems changed.
- Synthetic security evidence is not a vulnerability scan.
- OIDC, registry, protected environment, policy and audit integrations are optional mappings only.
- Replace reference GitLab/Runner versions with observed production versions before rollout.
MD

cat > handoff/next-actions.md <<'MD'
# Next actions for a real platform
1. Create an authorized disposable GitLab project and capture CI Lint/merged config plus pipeline/job IDs.
2. Define isolated runner trust classes and prove executor/network boundaries.
3. Publish reusable CI configuration with a reviewed versioning/pinning policy.
4. Choose a registry/package promotion mechanism and retain immutable artifact digest identity.
5. Configure workload identity with narrow OIDC trust where external systems support it.
6. Add protected-environment/policy/audit controls only where tier and operating model justify them.
7. Connect SLOs to dashboards/alerts and schedule monthly release/deprecation review.
MD

python - <<'PY'
from pathlib import Path
import hashlib,json
paths=[]
for root in ['evidence','handoff','governance']:
    for p in sorted(Path(root).rglob('*')):
        if p.is_file(): paths.append(p)
manifest=[]
for p in paths:
    manifest.append({'path':str(p),'sha256':hashlib.sha256(p.read_bytes()).hexdigest(),'bytes':p.stat().st_size})
Path('handoff/manifest.json').write_text(json.dumps(manifest,indent=2))
print('manifest files',len(manifest))
PY

tar -czf atlas-relay-handoff.tgz .gitlab-ci.yml ci scripts evidence governance handoff dist/SHA256SUMS target/staging.json
sha256sum atlas-relay-handoff.tgz | tee handoff/handoff-archive.sha256

11. Cleanup and rollback boundary

No external resource must be deleted because the mandatory lab created none. If you used an optional disposable GitLab project, first export/copy the evidence you are authorized to retain, then remove or archive only the disposable project/resources you created. Never delete shared runners, registries, environments, policy projects, audit destinations or cloud resources as “cleanup” unless their ownership and rollback are explicit.

# Local cleanup only, after copying atlas-relay-handoff.tgz somewhere safe.
cd /
case "$LAB" in
  */gitlab-ch38-capstone) rm -rf "$LAB" ;;
  *) echo 'Refusing cleanup: unexpected LAB path' >&2; exit 2 ;;
esac

12. Knowledge check

A release deployment job succeeded, but the target reports a different artifact digest. Is the pipeline “success” sufficient evidence?

Why is a privileged Docker runner not made safe simply by adding a “trusted” tag?

A shared CI component is included using a moving selector and the pipeline behavior changes without a project commit. Which evidence is missing?

An OIDC-authenticated deploy is denied because the audience is wrong. Should you replace it with a broad personal access token?

Why can a deployment approval and a healthy target be different states?

What is the most important post-course habit from this capstone?

13. What Chapter 38 adds to the production operating model

The complete course now forms one operating loop. Source changes become explicitly compiled pipelines. Jobs are routed to known runner trust classes. Variables/secrets and workload identities have bounded purpose. Reuse has version identity. Tests/security/report data are retained as evidence. Artifacts have immutable identity and promotion semantics. Deployments distinguish authorization, GitLab deployment state and external target health. Policies and exceptions expose their origin. Metrics make performance and reliability visible. Troubleshooting preserves first-failure evidence and recovery changes the smallest responsible layer.

Post-course production practice is therefore not “write more YAML.” It is operating this delivery system as a product: review platform changes, patch GitLab and Runner, test deprecations, maintain reusable components, rotate identities, validate runner isolation, exercise recovery, expire exceptions, tune capacity from measurements, and continuously prove that the source-to-target evidence chain still holds.

14. Course completion

You have moved from the first source-change mental model through configuration compilation, rules, DAGs, runners, artifacts, reuse, downstream pipelines, merge requests, environments, releases, security, SBOM/supply-chain evidence, OIDC, Kubernetes/IaC integration, runner fleets, automation, observability, optimization, governance and incident recovery—then integrated them into one evidence-driven production model.

The capstone packet is your final deliverable. Keep it as a template for future platform reviews, replacing every training assumption with observed production identifiers, versions, owners, SLOs, policy revisions and target-state proof.

Next lesson

Next: Post-course production practice

Carry the capstone operating model into real production work: keep dependencies current, review deprecations, rehearse recovery, measure SLOs, and revisit governance assumptions as GitLab and your delivery platform evolve.

Knowledge check

What makes the capstone platform independently verifiable rather than merely “green”?

If the pipeline succeeds but the external target is unhealthy, what conclusion is valid?

What belongs in the final operational handoff after the capstone?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Further reading — current official GitLab sources

Version-sensitive assumptions in this production capstone were checked against current official GitLab documentation on 2026-09-13. Re-check your exact GitLab, GitLab Runner, glab, executor, component, image/tool and external-provider versions before applying the operating model to production.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.