Production Capstone: Build and Govern a Secure GitLab Software Delivery Platform: Failure Injection, Troubleshooting, and Recovery Drill
Inject four controlled failures across access, CI/runner/security, artifact/release, and integration/recovery domains; diagnose from preserved evidence and recover with the least destructive correction.
Learning objectives
- Inject at least four reversible failures across access/governance, CI/runner/security, artifact/release/deployment, and platform/integration domains.
- Begin every drill with a prediction and preserved evidence instead of changing settings immediately.
- Use the course-wide diagnostic sequence: scope resource → inspect immutable identity and permissions → inspect rules/config → inspect logs/audit/API/runner evidence → least destructive fix → independent verification.
- Practice a credential-incident response that starts with revoke/rotate/disable and containment, never with cosmetic history cleanup.
- Demonstrate that GitLab backup data, configuration/secrets, object storage, registry, external systems, and runner state can have different recovery lifecycles.
1. The failure-drill contract
Each drill follows the same six-step method. This repetition is intentional: incident response improves when engineers use a stable diagnostic sequence under pressure.
FAILURE DRILL CONTRACT
1. Predict expected state before injection.
2. Preserve baseline evidence and immutable identities.
3. Inject one reversible defect only.
4. Scope the failing resource and compare actual evidence to prediction.
5. Apply the least destructive correction; do not hide the original cause.
6. Independently verify recovery and record root cause, detection, TTR, prevention.
2. Create a four-domain incident fixture
from pathlib import Path
import json, shutil, hashlib
root=Path('gitlab-ch34-drills'); base=root/'baseline'; work=root/'working'; evidence=root/'evidence'
for p in [base,work,evidence]: p.mkdir(parents=True,exist_ok=True)
state={
'members':{'app-dev':'developer','release-bot':'developer','platform-owner':'owner'},
'inherited_members':{'external-contractor':'maintainer'},
'branch_rules':{'main':{'push':'no_one','merge':'maintainer'}},
'pipeline':{'source':'merge_request_event','commit':'c0ffee01','release_job':False},
'runner':{'class':'isolated-general','privileged':False,'production_network':False},
'variables':{'PROD_TOKEN':{'protected':True,'available':False}},
'job_token':{'allowlist':['asterline-lab/delivery-platform/app']},
'artifact':{'name':'asterline-app.tar.gz','sha256':'aaa111'},
'release':{'tag':'v0.1.0','commit':'c0ffee01','artifact_sha256':'aaa111'},
'webhook':{'signature_valid':True,'project':'asterline-lab/delivery-platform/app','processed_ids':['evt-001']},
'backup':{'gitlab_archive':'backup-19.3.tgz','config_bundle':'config-19.3.tgz','object_storage_snapshot':'obj-001','runner_config':'external-not-in-backup'}
}
(base/'state.json').write_text(json.dumps(state,indent=2))
shutil.copy2(base/'state.json',work/'state.json')
(evidence/'baseline.sha256').write_text(hashlib.sha256((base/'state.json').read_bytes()).hexdigest()+'\n')
print(root.resolve())
3. Drill 1 — inherited access defeats the intended least-privilege model
Prediction: only platform-owner has
Owner-like administrative scope; an external contractor should not
inherit Maintainer into the application project. Preserve
baseline/state.json, then inject the unexpected
inherited role.
from pathlib import Path
import json
p=Path('gitlab-ch34-drills/working/state.json')
s=json.loads(p.read_text())
# Injected defect already represented by this inherited role.
print('DIRECT MEMBERS :',s['members'])
print('INHERITED MEMBERS:',s['inherited_members'])
unexpected=[u for u,r in s['inherited_members'].items() if r in {'maintainer','owner'}]
print('UNEXPECTED ELEVATED INHERITANCE:',unexpected)
if not unexpected: raise SystemExit('drill not injected')
# Least-destructive correction in fixture: lower/remove the inherited grant at its source.
s['inherited_members']['external-contractor']='reporter'
p.write_text(json.dumps(s,indent=2))
print('RECOVERED:',s['inherited_members'])
In GitLab, fix inherited access at the group/share source rather than patching project symptoms blindly. Record which parent group or shared group caused access, who owned that boundary, and whether project-level changes can actually override the inheritance model.
4. Drill 2 — runner/credential trust is broader than branch intent
Prediction: merge-request code must not reach a production-network runner or a production credential. Inject a state where the runner is privileged and production-network connected, and the protected variable is incorrectly available.
from pathlib import Path
import json
p=Path('gitlab-ch34-drills/working/state.json'); s=json.loads(p.read_text())
s['runner']={'class':'shared-prod-host','privileged':True,'production_network':True}
s['variables']['PROD_TOKEN']['available']=True
print('INJECTED runner=',s['runner'])
print('INJECTED variable=',s['variables']['PROD_TOKEN'])
unsafe=s['pipeline']['source']=='merge_request_event' and (s['runner']['privileged'] or s['runner']['production_network'] or s['variables']['PROD_TOKEN']['available'])
print('UNSAFE TRUST INTERSECTION:',unsafe)
# Recovery: isolate runner and make credential unavailable for MR source.
s['runner']={'class':'isolated-general','privileged':False,'production_network':False}
s['variables']['PROD_TOKEN']['available']=False
p.write_text(json.dumps(s,indent=2))
assert not (s['runner']['privileged'] or s['runner']['production_network'] or s['variables']['PROD_TOKEN']['available'])
print('RECOVERED')
5. Credential incident simulation — contain first, investigate second
Now simulate a leaked credential using a nonfunctional placeholder. The response order matters. Once a credential might be exposed, removing it from a Git commit or job log does not revoke the authority already copied by an attacker.
SYNTHETIC INCIDENT
Credential identifier: TOKEN_FAKE_NOT_VALID
Observed in: synthetic job-log fixture
Scope: fictional package read/write
Status at detection: ACTIVE (fixture)
RESPONSE ORDER
1. REVOKE / ROTATE / DISABLE the credential.
2. CONTAIN dependent automation and suspicious sessions.
3. RECORD identity, scope, exposure window, and affected resources.
4. SEARCH audit/log/API evidence for misuse.
5. REPAIR source/history/log exposure if policy requires it.
6. VERIFY old credential is unusable and replacement is minimally scoped.
If the credential is a real provider token, use that provider’s revocation mechanism immediately. Do not spend the first incident minutes rewriting Git history.
6. Drill 3 — release label and artifact identity disagree
Prediction: release artifact digest equals the digest recorded for the exact built bytes. Inject a mismatch while leaving the human-friendly version labels unchanged.
from pathlib import Path
import json
p=Path('gitlab-ch34-drills/working/state.json'); s=json.loads(p.read_text())
s['release']['artifact_sha256']='bbb222' # injected mismatch
print('artifact digest:',s['artifact']['sha256'])
print('release digest :',s['release']['artifact_sha256'])
if s['artifact']['sha256']==s['release']['artifact_sha256']:
raise SystemExit('drill not injected')
print('STOP PROMOTION: immutable identity disagreement')
# Least-destructive fixture recovery after source-of-truth investigation.
s['release']['artifact_sha256']=s['artifact']['sha256']
p.write_text(json.dumps(s,indent=2))
assert s['release']['artifact_sha256']==s['artifact']['sha256']
print('RECOVERED: release references verified artifact digest')
In a real incident, do not overwrite evidence to match whichever value is convenient. Determine whether the tag, pipeline, artifact, package/image digest, or release record is wrong. If the bytes are wrong, rebuild/promote through the governed process. If metadata is wrong, correct it only after preserving the original state and authorization trail.
7. Drill 4 — webhook replay and recovery-scope mismatch
Prediction: an already processed webhook ID must not trigger the action again, and a GitLab application backup must not be assumed to include all external configuration, object storage, or runner state. The combined drill models an integration retry during an operational recovery.
from pathlib import Path
import json
p=Path('gitlab-ch34-drills/working/state.json'); s=json.loads(p.read_text())
event={'id':'evt-001','signature_valid':True,'project':'asterline-lab/delivery-platform/app'}
print('event=',event)
if event['id'] in s['webhook']['processed_ids']:
print('DUPLICATE: acknowledge without repeating side effect')
else:
s['webhook']['processed_ids'].append(event['id'])
required={'gitlab_archive','config_bundle','object_storage_snapshot','runner_config'}
present=set(s['backup'])
print('recovery inventory complete:',required <= present)
print('runner state note:',s['backup']['runner_config'])
assert s['backup']['runner_config']=='external-not-in-backup'
print('RECOVERY DESIGN: restore GitLab-compatible data, then separately rebuild/restore external runner configuration')
p.write_text(json.dumps(s,indent=2))
GitLab’s Self-Managed backup documentation explicitly separates data included in the backup archive from items such as configuration files, TLS/SSH material, Redis/Sidekiq state, and some object-storage responsibilities. Restore also requires a compatible target, including the same GitLab version/type for the documented restore path. Runner configuration is an external lifecycle and should be reproducibly rebuilt rather than assumed to appear after a GitLab database restore.
8. Run the structured diagnostic sequence across all four drills
| Drill | Root cause | Detection signal | Least destructive correction | Independent verification | Preventive control |
|---|---|---|---|---|---|
| Access | unexpected inherited Maintainer grant | membership snapshot shows contractor inheritance | change grant at parent/share source | effective membership no longer elevated | periodic inherited-access review |
| CI/runner | MR code intersected privileged/networked runner + prod credential | pipeline source + runner + variable evidence | remove credential from path; isolate/select safe runner | MR path cannot obtain credential/network trust | runner trust classes + protected variables/ref rules |
| Artifact/release | release metadata references different digest | digest comparison fails | stop promotion; rebuild or correct metadata from source of truth | tag/commit/pipeline/digest tuple agrees | immutable manifest required for release |
| Integration/recovery | duplicate event + incomplete recovery-lifecycle assumption | replayed webhook ID; backup inventory gap | idempotent event handling; separate recovery plans | duplicate produces no side effect; restore checklist covers external state | event ledger + tested recovery inventory |
9. Record time-to-recovery without turning it into a vanity metric
Time-to-recovery (TTR) is useful only with scope. Record when detection began, when containment occurred, when the corrective action was applied, and when independent verification completed. A fast but unverified “fix” is not recovery.
from pathlib import Path
import json
root=Path('gitlab-ch34-drills'); ev=root/'evidence'
records=[
{'drill':'access','detected_min':0,'contained_min':3,'corrected_min':7,'verified_min':10},
{'drill':'runner','detected_min':0,'contained_min':2,'corrected_min':8,'verified_min':12},
{'drill':'artifact','detected_min':0,'contained_min':1,'corrected_min':6,'verified_min':9},
{'drill':'integration_recovery','detected_min':0,'contained_min':4,'corrected_min':11,'verified_min':18},
]
for r in records: r['ttr_minutes']=r['verified_min']-r['detected_min']
(ev/'recovery_times.json').write_text(json.dumps(records,indent=2))
print(json.dumps(records,indent=2))
10. Verify recovered state and preserve the original baseline
from pathlib import Path
import json, hashlib
root=Path('gitlab-ch34-drills'); base=root/'baseline/state.json'; work=root/'working/state.json'; ev=root/'evidence'
expected=(ev/'baseline.sha256').read_text().strip()
assert hashlib.sha256(base.read_bytes()).hexdigest()==expected, 'baseline evidence was modified'
s=json.loads(work.read_text())
checks={
'contractor_not_elevated': s['inherited_members']['external-contractor'] not in {'maintainer','owner'},
'runner_not_privileged': not s['runner']['privileged'],
'runner_no_prod_network': not s['runner']['production_network'],
'prod_token_unavailable_to_mr': not s['variables']['PROD_TOKEN']['available'],
'release_digest_matches': s['release']['artifact_sha256']==s['artifact']['sha256'],
'duplicate_event_recorded': 'evt-001' in s['webhook']['processed_ids'],
'runner_recovery_is_external': s['backup']['runner_config']=='external-not-in-backup'
}
(ev/'recovery_verification.json').write_text(json.dumps(checks,indent=2))
print(json.dumps(checks,indent=2))
assert all(checks.values())
11. Cleanup / rollback boundaries
Do not delete the baseline before you have the final handoff record. The working fixture can be discarded after Lesson 5. A live disposable GitLab project should be archived/deleted only after confirming no valuable data, token, runner, package, release, or integration points to it.
from pathlib import Path
import shutil
p=Path('gitlab-ch34-drills')
print('REVIEW BEFORE DELETE:',p.resolve())
# shutil.rmtree(p) # Uncomment only after Lesson 5 evidence has been exported.
Knowledge check
Why should an access failure be corrected at the inheritance source?
Changing only the project symptom can leave the parent/shared-group grant intact or create confusing exceptions. The real authorization source must be identified and corrected.
What is the first response to a potentially leaked real credential?
Revoke, rotate, or disable it and contain dependent access. History or log cleanup can follow, but cannot make the already exposed credential safe.
Why does a release digest mismatch require stopping promotion?
It means the system cannot prove which bytes correspond to the reviewed/released identity. Continuing would knowingly cross a provenance boundary without evidence.
Why should webhook handlers be idempotent?
GitLab or the network can retry deliveries. The same valid event can arrive more than once, so duplicate processing must not repeat consequential side effects.
Why can a Self-Managed GitLab backup restore still leave the delivery platform incomplete?
GitLab application backups and external configuration, object storage, secrets, runners, external services, and cloud/Kubernetes state can have distinct backup and recovery lifecycles.
12. Lesson summary and bridge
- Four failures were injected across access, runner/credential, artifact/release, and integration/recovery domains.
- Every drill started with a prediction and preserved baseline evidence, then used the same diagnostic sequence.
- The credential incident began with revocation/containment semantics rather than cosmetic log/history cleanup.
- Recovery was not accepted until independent checks passed and the original evidence remained unchanged.
- The final lesson turns this architecture, validation, and drill evidence into an operational handoff package and production-readiness decision.
Next, you will conduct the final operational review, prove critical invariants, package runbooks/evidence/risks, and decide what is truly production-ready versus simulated or tier-dependent.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the reference release. Availability can vary by GitLab.com, Self-Managed, or Dedicated offering; Free, Premium, or Ultimate tier; namespace settings; administrator policy; runner type; and feature status. Re-check current documentation before applying a production design. The backup/restore drill intentionally stays synthetic unless the learner has a disposable Self-Managed instance. GitLab’s current restore documentation requires a compatible existing installation and exact version/type compatibility for the documented restore path.
- GitLab 19.3 release
- GitLab roles and permissions
- Protected branches
- Protection rules and permissions
- Merge requests
- Merge request approvals
- Code Owners
- CI/CD pipelines
- CI/CD variables
- CI_JOB_TOKEN
- Runner security
- Job artifacts
- Caching in GitLab CI/CD
- Protected environments
- Deployment approvals
- Container Registry
- Protected container repositories
- Releases
- Release evidence
- SAST
- Pipeline secret detection
- Security policies
- REST API
- Webhooks
- GitLab agent for Kubernetes
- Audit events
- Health check endpoints
- Back up GitLab
- Restore GitLab
- GitLab Duo Agent Platform
- Agent tool governance
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.