Production Capstone: Build, Secure, Scale, Observe, and Govern a Complete GitLab Delivery Platform: Failure Injection, Troubleshooting, and Recovery Drill
Inject failures across configuration, runner routing, DNS/TLS, artifact integrity, identity and deployment verification, then preserve first-failure evidence and recover with the smallest safe changes.
Learning objectives
- Apply one evidence-first diagnostic sequence across six failure layers.
- Preserve known-good and first-failure evidence before repairing configuration or target state.
- Reject privileged-runner, TLS-disablement, broad-token, blind-retry and rebuild shortcuts.
- Detect artifact/identity/deployment mismatches without hiding the original failure.
- Produce a recovery manifest and decision table that can become an operations playbook.
Failure-injection safety: every fault in this
lesson is confined to the disposable
gitlab-ch38-capstone workspace or an optional
disposable GitLab project. Do not alter shared runner registration,
protected production environments, real OIDC trust, registries,
cloud/Kubernetes/IaC targets, enterprise policy projects,
package/releases, or live credentials for this drill.
1. Use one recovery sequence for every failure layer
The capstone deliberately reuses the Chapter 37 incident sequence: preserve pipeline/job IDs and first-failure evidence; confirm source/ref/SHA and compiled configuration; confirm rules/non-secret inputs; inspect graph/queue; confirm runner/executor/image/toolchain; inspect script/tool/network; inspect reports/artifacts/caches; inspect deployment/environment/external target; apply the least destructive correction; rerun only the smallest safe scope; independently verify recovery.
Do not skip earlier layers because a symptom “looks like network” or “looks like Runner.” A causal sequence prevents destructive guesswork.
2. Freeze the known-good baseline
Before injecting faults, hash the good state. This gives you a recovery reference and proves later that the broken evidence was not silently overwritten.
LAB="${LAB:-${TMPDIR:-/tmp}/gitlab-ch38-capstone}"
cd "$LAB"
mkdir -p evidence/failures evidence/recovery
git rev-parse HEAD | tee evidence/recovery/baseline-source.txt
sha256sum .gitlab-ci.yml ci/base.yml scripts/*.sh dist/atlas-relay.tgz target/staging.json \
| tee evidence/recovery/baseline-sha256.txt
cp target/staging.json evidence/recovery/baseline-target.json
3. Fault A — configuration compilation failure
Break indentation in the disposable YAML, preserve the invalid file and parser output, then repair only indentation. A real GitLab project should additionally use CI Lint because a generic YAML parser cannot validate GitLab-specific semantics.
cp .gitlab-ci.yml evidence/failures/A-before.gitlab-ci.yml
python - <<'PY'
p='.gitlab-ci.yml'; s=open(p).read();
s=s.replace(' stage: validate',' stage: validate',1)
open(p,'w').write(s)
PY
cp .gitlab-ci.yml evidence/failures/A-broken.gitlab-ci.yml
set +e
python - <<'PY' > evidence/failures/A-parser.txt 2>&1
import yaml
print(yaml.safe_load(open('.gitlab-ci.yml')))
PY
printf 'exit=%s\n' "$?" >> evidence/failures/A-parser.txt
set -e
cp evidence/failures/A-before.gitlab-ci.yml .gitlab-ci.yml
Real GitLab mapping: preserve the lint response and source SHA. If configuration changed, create a new pipeline tied to the repaired SHA rather than expecting an old job retry to recompile new YAML.
4. Fault B — trusted release job routed to the wrong runner class
A common anti-pattern is to solve a pending job by adding broader tags or enabling privileged execution everywhere. Instead, compare job requirements with runner capabilities and trust class. This local inventory simulates one untrusted runner and one trusted release runner.
cat > evidence/failures/B-runner-inventory.json <<'JSON'
[
{"id":41,"tags":["linux","untrusted"],"protected":false,"executor":"docker","privileged":false},
{"id":77,"tags":["linux","release"],"protected":true,"executor":"docker","privileged":false}
]
JSON
python - <<'PY' | tee evidence/failures/B-routing.txt
import json
r=json.load(open('evidence/failures/B-runner-inventory.json'))
job={'name':'build-release','required_tags':{'linux','release'},'trusted':True}
for x in r:
ok=job['required_tags'].issubset(set(x['tags'])) and (not job['trusted'] or x['protected'])
print(x['id'], 'eligible='+str(ok), 'privileged='+str(x['privileged']))
PY
The correct repair is a justified route to runner 77 or equivalent trusted isolation—not making runner 41 privileged. In a real incident preserve the pending job ID, required tags, runner status/version and protection state before editing routing.
5. Fault C — DNS/TLS evidence before application blame
Network failures can occur during image pulls, package downloads, external API calls or deployment verification. Never “fix” TLS by disabling certificate verification. The safe diagnostic order is DNS resolution, route/connectivity, TLS identity/trust, then application protocol.
python - <<'PY' | tee evidence/failures/C-network.txt
import socket, ssl
for host in ['localhost','does-not-exist.invalid']:
try:
print(host, socket.getaddrinfo(host, 443))
except Exception as e:
print(host, type(e).__name__, str(e))
print('default_ca_paths', ssl.get_default_verify_paths())
PY
The reserved .invalid domain should fail
deterministically. In production capture the exact hostname,
resolver result, proxy/no-proxy settings, certificate
chain/hostname/expiry and Runner/container CA state. Do not print
bearer tokens while debugging HTTP.
6. Fault D — artifact tampering detected before deployment
Preserve the original release bundle, tamper with a copy, and prove the digest gate rejects it. Do not rebuild to make the digest “match”—that would erase the original release identity.
cp dist/atlas-relay.tgz evidence/failures/D-original-artifact.tgz
cp dist/atlas-relay.tgz evidence/failures/D-tampered-artifact.tgz
printf 'tamper\n' >> evidence/failures/D-tampered-artifact.tgz
ORIGINAL="$(awk '{print $1}' dist/SHA256SUMS)"
TAMPERED="$(sha256sum evidence/failures/D-tampered-artifact.tgz | awk '{print $1}')"
printf 'expected=%s\nactual=%s\n' "$ORIGINAL" "$TAMPERED" | tee evidence/failures/D-digest.txt
test "$ORIGINAL" != "$TAMPERED"
7. Fault E — identity/audience mismatch is not a reason for a broad PAT
OIDC failures often come from issuer, audience, subject/claims or provider trust mismatch. Simulate a token claim set and reject the wrong audience. The safe repair is to align the narrowly scoped audience/trust rule, not to introduce a long-lived broad credential.
cat > evidence/failures/E-claims.json <<'JSON'
{"iss":"https://gitlab.example.invalid","aud":"https://wrong.example.invalid","project_path":"training/atlas-relay","ref":"main"}
JSON
set +e
python - <<'PY' | tee evidence/failures/E-authz.txt
import json
c=json.load(open('evidence/failures/E-claims.json'))
expected='https://staging.example.invalid'
print({'expected_aud':expected,'actual_aud':c['aud'],'authorized':c['aud']==expected})
raise SystemExit(0 if c['aud']==expected else 23)
PY
AUTH_RC=${PIPESTATUS[0]:-23}
set -e
printf 'authorization_exit=%s\n' "$AUTH_RC" >> evidence/failures/E-authz.txt
8. Fault F — rollback command succeeds but target verification fails
A rollback is an action, not proof. Preserve the target before mutation, write a deliberately wrong rollback result, detect it, then restore the known-good baseline target and verify source/digest/health.
cp target/staging.json evidence/failures/F-target-before.json
cat > target/staging.json <<'JSON'
{"environment":"staging","source_sha":"unexpected","artifact_digest":"sha256:unexpected","healthy":true}
JSON
set +e
python - <<'PY' | tee evidence/failures/F-verify-failed.txt
import json
x=json.load(open('target/staging.json')); b=json.load(open('evidence/recovery/baseline-target.json'))
ok=(x==b)
print({'rollback_verified':ok,'actual':x,'expected':b})
raise SystemExit(0 if ok else 31)
PY
set -e
cp evidence/recovery/baseline-target.json target/staging.json
python - <<'PY' | tee evidence/recovery/F-verify-recovered.txt
import json
x=json.load(open('target/staging.json')); b=json.load(open('evidence/recovery/baseline-target.json'))
print({'recovered':x==b,'target':x}); assert x==b
PY
9. Debugging must not create a second incident
Do not enable broad debug tracing on a job with real credentials merely to “see everything.” Current GitLab variable troubleshooting warns that debug trace can expose values. If extra tracing is unavoidable, use a disposable job with fake variables, restrict who can see logs, disable tracing immediately, and rotate any credential that may have appeared.
Never use these shortcuts: print all variables/tokens, disable TLS verification, grant a broad PAT, run untrusted code on a privileged runner, delete artifacts/logs to “clean up,” rerun blindly until green, rebuild a release artifact during incident recovery, or bypass policy/approvals without an explicit governed exception.
10. Build a first-failure and recovery manifest
Hash the preserved evidence and recovered state. This does not make the evidence cryptographically authoritative by itself, but it makes accidental mutation visible inside the lab packet.
python - <<'PY'
from pathlib import Path
import hashlib, json, datetime
files=[]
for p in sorted(Path('evidence').rglob('*')):
if p.is_file():
files.append({'path':str(p),'sha256':hashlib.sha256(p.read_bytes()).hexdigest(),'bytes':p.stat().st_size})
manifest={'generated_at':datetime.datetime.now(datetime.timezone.utc).isoformat(),'files':files}
Path('evidence/recovery-manifest.json').write_text(json.dumps(manifest,indent=2))
print(json.dumps({'file_count':len(files)},indent=2))
PY
11. Recovery playbook decision points
| Evidence says… | Smallest safe action | Do not do |
|---|---|---|
| Config invalid / effective config wrong | Fix reviewed config and create a new pipeline for the repaired SHA | Retry old job expecting new config |
| Job pending: no trusted runner matches | Fix routing/capacity/trust boundary | Make every runner privileged or remove protection |
| DNS/TLS failure | Repair resolver/proxy/CA/certificate cause and retry bounded idempotent job | Disable TLS verification |
| Artifact digest mismatch | Stop promotion; recover expected immutable artifact or rebuild as a new release identity | Overwrite digest or pretend tampered bytes are same release |
| OIDC audience/claim denied | Fix narrow trust contract or job audience | Replace with broad long-lived PAT |
| Deployment action succeeded but target wrong | Verify/repair external target and health; preserve deployment record | Declare recovery from job exit code alone |
12. Lesson summary
You injected failures across configuration, runner routing, network, artifact, identity and deployment layers; preserved first-failure evidence; rejected unsafe shortcuts; and restored the known-good target with explicit verification. Lesson 5 turns the entire capstone into an operational handoff that a platform team could own after the course ends.
Knowledge check
What makes the capstone platform independently verifiable rather than merely “green”?
It preserves a chain from source/MR/SHA through compiled configuration, runner and identity context, reports and artifact digests, deployment records, external health, and governance/recovery evidence.
If the pipeline succeeds but the external target is unhealthy, what conclusion is valid?
Only that the GitLab jobs completed successfully. Deployment authorization, deployment record, provider acceptance, rollout health, and target state are separate facts that need independent verification.
What belongs in the final operational handoff after the capstone?
Architecture and configuration inventory, SLOs/metrics, runbooks, upgrade/deprecation watch list, exception register, residual risks, recovery evidence, owners, and concrete next actions.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Further reading — current official GitLab sources
Version-sensitive assumptions in this production capstone were checked against current official GitLab documentation on 2026-09-13. Re-check your exact GitLab, GitLab Runner, glab, executor, component, image/tool and external-provider versions before applying the operating model to production.
- GitLab Docs — GitLab 19.3 release notes
- GitLab Docs — Critical patch release 19.3.2
- GitLab Docs — Release and maintenance policy
- GitLab Docs — Deprecations and removals
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — Use CI/CD configuration from other files
- GitLab Docs — CI/CD inputs
- GitLab Docs — CI/CD components
- GitLab Docs — Pipeline security
- GitLab Docs — Validate CI/CD configuration / CI Lint
- GitLab Docs — Runner security
- GitLab Docs — CI/CD job token
- GitLab Docs — OIDC authentication using ID tokens
- GitLab Docs — Job artifacts
- GitLab Docs — CI/CD artifact report types
- GitLab Docs — Environments and deployments
- GitLab Docs — Protected environments
- GitLab Docs — Deployment approvals
- GitLab Docs — Pipeline execution policies
- GitLab Docs — Audit events
- GitLab Docs — Audit events API
- GitLab Docs — Attestations API (Experimental)
- GitLab Docs — glab attestation (Experimental)
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.