Chapter 38Lesson 03~160 minutes

Production Capstone: Build, Secure, Scale, Observe, and Govern a Complete GitLab Delivery Platform: Security, Governance, and Reliability Validation

Validate the capstone against runner isolation, least-privilege identity, pinned reuse, artifact provenance, deployment authorization, policy enforcement, auditability, observability and recovery criteria.

Security validationGovernanceReliabilityTier-aware controls

Learning objectives

  • Validate runner isolation as an execution boundary rather than a tag convention.
  • Choose identity types by purpose and prove narrow authorization/denial behavior.
  • Validate reusable configuration and artifact identity against mutation and provenance gaps.
  • Separate deployment approval, execution and target health; map paid controls to a free simulation path.
  • Produce a tier-aware production-readiness decision table and recovery checklist.

Validation baseline: current GitLab 19.3.x documentation was re-checked on 2026-09-13. CI/CD components and inputs are cross-tier; protected environments and deployment approvals are Premium/Ultimate; pipeline execution policies are Ultimate; project/group audit-event pages/APIs are Premium/Ultimate; GitLab provenance attestations are currently Ultimate and Experimental. The mandatory capstone therefore validates equivalent control objectives without requiring paid or experimental features.

1. Validation asks “what prevents or detects the wrong thing?”

Lesson 2 proved that the happy path works. Production readiness asks harder questions. Can untrusted code reach a privileged runner? Can a moving include change the pipeline without a project commit? Can a broad token cross a project boundary? Can a release artifact be swapped between build and deploy? Can the person triggering a deployment also satisfy its approval policy? Can a policy be bypassed or an exception live forever? Can operators detect drift and recover without destroying evidence?

For each control, name the asset, threat or failure mode, preventive control, detective evidence, owner and rollback. A checkbox without observable proof is not validation.

2. Validate runner trust boundaries

GitLab Runner executes user-controlled scripts. Current Runner security guidance warns that privileged Docker effectively disables container security mechanisms and can expose the host to privilege escalation. Therefore, “runner online” is not a security property. You need explicit trust classes.

Workload Recommended trust class Evidence
Fork/untrusted MR validation Non-privileged, isolated/ephemeral runner; no production secrets Runner ID/tags/protection/executor; no privileged mode; network boundary
Internal branch build Dedicated or appropriately isolated runner with minimal credentials Runner scope/tags + token/variable protection + workspace cleanup
Release/sign/publish Trusted protected runner; immutable tool/image references; narrow identity Protected ref/job evidence + runner identity + artifact digest/provenance
Production deploy Trusted deploy runner or external controller; environment authorization Allowed-to-deploy/approval evidence + OIDC/token scope + target verification

Do not “fix” isolation with tags alone. Tags route jobs; they do not sandbox a compromised host. Runner manager/executor/VM/container/network boundaries are separate controls.

3. Validate reusable configuration and dependency identity

Copied YAML hides origin and makes patching inconsistent. Reusable configuration improves maintainability only if its revision is reviewable. Current GitLab pipeline-security guidance recommends specific refs for included configuration when possible. Components add a catalog/reuse contract and typed inputs, but a moving semantic selector or tag is still different from a commit SHA.

Commit-SHA pin

Strong reproducibility for project includes. Update flow is explicit but maintenance automation is needed.

Protected release tag

Human-friendly and reviewable if tag mutation is controlled. Record resolved commit/digest in evidence.

Major/minor selector

Easier consumer upgrades but resolution can change over time. Appropriate only when that mutability is intentional and tested.

Copy/paste

No upstream identity after copy; fixes drift across repositories. Use only when ownership intentionally transfers to the consuming repo.

4. Validate identity by purpose, not convenience

CI_JOB_TOKEN, personal access tokens, deploy tokens, trigger tokens and OIDC ID tokens solve different problems. A production design should explain why each identity exists and which resource can accept it. Current job-token guidance uses an allowlist for cross-project access and still evaluates the triggering user’s permissions. OIDC id_tokens let a job present a short-lived signed identity to an external provider, but the provider’s trust policy remains an external control.

Need Preferred identity pattern Evidence / failure test
Same-project API/job operation CI_JOB_TOKEN where endpoint supports it Allowed endpoint succeeds; unrelated endpoint denied
Cross-project job access Job token + explicit target allowlist when supported Allowlist entry + denied access before/after removal
External cloud/Vault OIDC ID token with narrow aud and provider claims Issuer/audience/non-secret decoded claims + provider trust policy + short TTL
Human administration User identity/PAT only when automation identity cannot fit Owner, expiry, narrow scopes, rotation/revocation and audit trail
Deploy/package read-only automation Deploy token where appropriate Scope, project/group boundary, expiry/revocation test

5. Validate artifact identity and supply-chain evidence

The build artifact and the deployment authorization are different states. A deployment may be approved for “release 1.4” while the actual bytes differ unless you retain and verify a digest. The mandatory lab already produced dist/SHA256SUMS and evidence/provenance.json. Verify both before deployment and again at the target boundary.

cd "${LAB:-${TMPDIR:-/tmp}/gitlab-ch38-capstone}"
sha256sum -c dist/SHA256SUMS
python - <<'PY'
import json, hashlib
p=json.load(open('evidence/provenance.json'))
actual=hashlib.sha256(open(p['subject'],'rb').read()).hexdigest()
print({'declared':p['sha256'],'actual':actual,'source_sha':p['source_sha']})
assert actual == p['sha256']
PY

GitLab’s native attestations API and glab attestation can verify provenance-shaped attestations, but current documentation marks the feature Experimental and Ultimate. Treat it as an optional integration, not the mandatory production foundation of this course.

6. Validate deployment authorization separately from target health

Protected environments and deployment approvals can restrict who may deploy and who must approve, but they are Premium/Ultimate. Even with approvals, GitLab explicitly separates approval from running the deployment job, and neither approval nor job success proves the external target is healthy.

State Question Proof
Authorization Was this identity allowed to deploy this environment? Protected environment rule / simulated approval register
Approval Were required approvers satisfied and was self-approval allowed by policy? Deployment approval details / simulated signed approval
Execution Which job/runner/image/tool performed the change? Job ID, runner/executor/image/tool versions
Deployment record What environment/deployment does GitLab record? Environment/deployment API/UI record
External health What source/digest/version does the target actually serve? Independent target API/file/cluster/cloud read

7. Validate policy enforcement and exception lifecycle

Chapter 36 established the key rule: policy origin must remain visible. On Ultimate, pipeline execution policies can inject or override CI behavior from a security policy project. In the mandatory path, simulate the same control objective with a versioned policy file and an exception register.

{
  "policy_id": "capstone-release-v1",
  "policy_version": "git:deadbeef-replace-with-real-sha",
  "scope": "default-branch release",
  "required_controls": ["tests", "artifact-digest", "target-verification"],
  "exceptions": [
    {
      "id": "EX-2026-001",
      "owner": "platform-owner@example.invalid",
      "reason": "training example only",
      "expires_at": "2026-09-20T00:00:00Z",
      "scope": "synthetic security evidence only"
    }
  ]
}

An exception is not “turn control off.” It is explicit residual risk with owner, bounded scope, expiry and follow-up. If the product feature does not provide a generic waiver-expiry object, keep that lifecycle in a reviewed governance system rather than inventing a hidden field.

8. Validate SLOs and recovery readiness

Capture pipeline duration, queue time, job failure classes, runner saturation and deployment verification latency in production. The local lab cannot reproduce GitLab queue metrics, so write the assumptions explicitly. A good dashboard separates creation/compilation failures, pending/scheduling, runner/executor failures, script/tool/network failures, report/artifact ingestion, deployment authorization, deployment status and target health.

Recovery criteria should be binary enough to audit: configuration lints/compiles; the intended jobs exist; trusted work lands on trusted runners; artifact digest matches; deployment target reports the intended digest; the original failure remains preserved; policy/exception evidence points to a reviewed version; and SLOs return inside the accepted range.

9. Capstone design decision table

Decision Free/disposable default Production alternative Trust/tier prerequisite Observable evidence
Reuse include:local Catalog component pinned to protected version/SHA Component project access; all tiers Merged config + component ref/resolved revision
Untrusted execution Synthetic non-privileged runner inventory Ephemeral isolated runner fleet Runner infrastructure Runner ID/executor/privilege/network evidence
Secrets No real secret External provider via OIDC ID token Provider trust configuration Audience/claims + provider decision; no token value in logs
Promotion Local tar + SHA-256 Registry/package object by digest Registry/package permissions Producer job + object digest + consumer/deploy digest
Deployment gate Reviewed local approval JSON Protected environment + approvals Premium/Ultimate Approvers, blocked/approved state, deploy identity
Central policy Versioned local policy JSON Pipeline execution policy Ultimate Policy project/ref/scope + effective pipeline
Audit Hashed evidence ledger Project/group audit API / streaming Premium/Ultimate; streaming varies Event IDs/timestamps + external retention

10. Production-readiness validation checklist

  • Source, reusable config and policy versions are reviewable and recoverable.
  • Pipeline inputs/variables have explicit ownership; secrets are not printed or embedded in artifacts.
  • Untrusted and trusted workloads have justified runner isolation boundaries.
  • Artifact identity survives promotion; deploy does not silently rebuild release bytes.
  • Deployment authorization, deployment status and external target health are verified independently.
  • Policy scope and exception lifecycle are visible; audit retention meets the organization’s requirement.
  • Version/deprecation assumptions are timestamped and owned.
  • Recovery playbooks preserve first-failure evidence and specify smallest safe rerun scope.

11. Lesson summary

The capstone is now more than a working pipeline. Its trust boundaries, identities, artifacts, deployment controls, governance and operational evidence have explicit failure tests. Lesson 4 deliberately breaks representative layers and uses the Chapter 37 evidence-first recovery method to prove that the operating model survives failure.

Next lesson

Next: Production Capstone: Build, Secure, Scale, Observe, and Govern a Complete GitLab Delivery Platform: Failure Injection, Troubleshooting, and Recovery Drill

Continue with the next lesson in the course sequence and carry forward the evidence-first GitLab CI/CD operating model.

Knowledge check

What makes the capstone platform independently verifiable rather than merely “green”?

If the pipeline succeeds but the external target is unhealthy, what conclusion is valid?

What belongs in the final operational handoff after the capstone?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Further reading — current official GitLab sources

Version-sensitive assumptions in this production capstone were checked against current official GitLab documentation on 2026-09-13. Re-check your exact GitLab, GitLab Runner, glab, executor, component, image/tool and external-provider versions before applying the operating model to production.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.