Checkpoint Lab — Runner Security, Isolation Boundaries, Privileged Containers, Fork Pipelines, Untrusted Code, and Threat Modeling
Threat-model and harden a disposable runner design, demonstrate a blocked attacker-like workload, prove trust reset through destruction/reimage simulation, and produce incident-evidence and rollback requirements.
Learning objectives
Checkpoint objectives
- Threat-model a runner design across trust, eligibility, executor privilege, network, credentials, persistence and incident evidence.
- Prove one trusted workload succeeds and one attacker-like workload is blocked before execution.
- Demonstrate that single-use worker destruction/reimage is an explicit trust-reset event.
- Produce an evidence packet that can be mapped to real GitLab pipeline/job/runner IDs and first-failure logs.
- Write migration/rollback requirements without weakening runner isolation.
1. Checkpoint scenario
Your platform team supports fork unit tests, protected release builds and a rare privileged image-build requirement. A recent review found that tag-based routing was treated as authorization. Design a hardened replacement, prove a fork cannot request the privileged pool, and document what must be preserved/destroyed after any suspicious privileged execution.
2. Current assumptions and preflight
| Item | Checkpoint assumption |
|---|---|
| Documentation date | 2026-09-12 |
| GitLab Runner baseline | 19.3.2 |
| GitLab CI/CD tier | Free-compatible core reasoning path |
| Python | 3.10+ standard library only |
| Credentials | None; only credential-class labels |
| Infrastructure | Local files/directories only |
| Production systems | Explicitly out of scope |
mkdir -p glci-ch31-checkpoint/{inputs,evidence,worker-sim}
cd glci-ch31-checkpoint
python3 --version
printf 'production_access=no\nreal_credentials=no\nprivileged_execution=no\n' > evidence/preflight.txt
3. Record predictions before execution
Write at least these two predictions to
evidence/predictions.txt:
-
A fork MR requesting
privileged-buildwill be denied even though the pool exists. -
A same-project protected push requesting
trusted-buildwill be authorized and its pool will be single-use with only registry-oriented authority.
Add a third prediction: destroying the simulated worker removes a residue marker; a new worker directory starts without that marker.
4. Implement the hardened policy
from dataclasses import dataclass
import json, sys
@dataclass(frozen=True)
class Pool:
name: str
trust: str
protected: bool
privileged: bool
single_use: bool
network: tuple
credential_class: str
POOLS = {
'untrusted-test': Pool('untrusted-test','untrusted',False,False,True,
('gitlab','dependency-mirror'), 'none'),
'trusted-build': Pool('trusted-build','protected',True,False,True,
('gitlab','registry'), 'registry-push-short-lived'),
'privileged-build': Pool('privileged-build','protected',True,True,True,
('gitlab','registry'), 'registry-push-short-lived'),
}
def classify(job):
fork = job['source_project_id'] != job['target_project_id']
if job['pipeline_source'] == 'merge_request_event' and fork:
return 'untrusted'
if job.get('ref_protected') and not fork:
return 'protected'
return 'untrusted'
def authorize(job):
pool=POOLS[job['requested_pool']]
trust=classify(job)
reasons=[]
if pool.protected and trust != 'protected':
reasons.append('protected pool requires same-project protected trust')
if trust == 'untrusted' and pool.privileged:
reasons.append('untrusted code cannot use privileged pool')
if trust == 'untrusted' and pool.credential_class != 'none':
reasons.append('untrusted code cannot receive privileged credentials')
if not pool.single_use and trust == 'untrusted':
reasons.append('untrusted pool must reset trust with a single-use worker')
return {
'job_id':job['job_id'], 'trust':trust, 'pool':pool.name,
'authorized':not reasons, 'reasons':reasons or ['policy satisfied'],
'protected':pool.protected, 'privileged':pool.privileged,
'single_use':pool.single_use, 'network':pool.network,
'credential_class':pool.credential_class,
}
if __name__ == '__main__':
job=json.load(open(sys.argv[1], encoding='utf-8'))
print(json.dumps(authorize(job), indent=2, sort_keys=True))
5. Create exact synthetic workload identities
{
"job_id": "attacker-fork-501",
"pipeline_source": "merge_request_event",
"source_project_id": 9002,
"target_project_id": 9001,
"ref_protected": false,
"requested_pool": "privileged-build",
"source_sha": "1111111111111111111111111111111111111111"
}
{
"job_id": "trusted-release-502",
"pipeline_source": "push",
"source_project_id": 9001,
"target_project_id": 9001,
"ref_protected": true,
"requested_pool": "trusted-build",
"source_sha": "2222222222222222222222222222222222222222"
}
Save these as inputs/attacker.json and
inputs/trusted.json. The exact SHA strings are
synthetic immutable identifiers, not real repository commits.
6. Evaluate both jobs and preserve decisions
python3 runner_policy.py inputs/attacker.json > evidence/attacker-decision.json
python3 runner_policy.py inputs/trusted.json > evidence/trusted-decision.json
python3 -m json.tool evidence/attacker-decision.json
python3 -m json.tool evidence/trusted-decision.json
Expected: the attacker case has authorized: false with
both protected-pool and privileged-workload reasons; the trusted
case has authorized: true. These are independent
authorization outcomes, not job-success statuses.
7. Turn the expected policy into executable verification
import json
from pathlib import Path
a=json.loads(Path('evidence/attacker-decision.json').read_text())
t=json.loads(Path('evidence/trusted-decision.json').read_text())
assert a['authorized'] is False, a
assert a['privileged'] is True, a
assert t['authorized'] is True, t
assert t['single_use'] is True, t
assert t['credential_class'] == 'registry-push-short-lived', t
print('policy_assertions=pass')
Save the output in evidence/assertions.txt. A real
GitLab equivalent would additionally preserve
CI_PIPELINE_SOURCE, CI_COMMIT_SHA,
pipeline/job IDs, runner ID/protection/tags/version/executor and the
compiled configuration.
8. Simulate worker residue and trust reset
This local directory stands in for a single-use worker. The point is the lifecycle invariant: after an untrusted/sensitive job, the execution boundary is destroyed rather than “declared clean.”
mkdir -p worker-sim/worker-001
printf 'synthetic-residue\n' > worker-sim/worker-001/residue.txt
sha256sum worker-sim/worker-001/residue.txt > evidence/residue-before.txt
# Preserve metadata first; then destroy the simulated worker.
printf 'worker=worker-001 action=destroy reason=single-use-trust-reset\n' > evidence/destruction-record.txt
rm -rf worker-sim/worker-001
test ! -e worker-sim/worker-001 && printf 'worker_001_destroyed=yes\n' >> evidence/destruction-record.txt
mkdir -p worker-sim/worker-002
test ! -e worker-sim/worker-002/residue.txt && printf 'worker_002_clean_start=yes\n' >> evidence/destruction-record.txt
In production, worker destruction does not revoke credentials that were already stolen. Credential expiry/revocation and incident response remain separate evidence.
9. Complete the threat model
| Threat | Preventive control | Detection/evidence | Recovery |
|---|---|---|---|
| Fork asks for privileged runner | Protected dedicated pool + trust-aware rules + no fallback | Source/target project IDs, compiled rules, tags, runner protection | Pause pool; fix routing; canary before resume |
| Shell host exposes another project | No untrusted Shell workloads; dedicated host if needed | Executor/host/build-dir inventory | Rotate exposed credentials; rebuild host |
| Docker socket exposes host daemon | No socket on untrusted pool | Trusted manager config/host mount inventory | Remove mount; rotate host/registry creds; reimage |
| Reusable Git state leaks prior content | Ephemeral worker or trusted-only fetch; cleanup defense in depth | Git strategy, worker reuse count, residue checks | Destroy/reimage worker; invalidate contaminated cache |
| Protected secret reaches malicious code | Protected resource eligibility + no secret on untrusted pool + short-lived identity | Ref/MR project identity, variable protection metadata, identity audit | Revoke/rotate credential; correct authorization boundary |
10. Incident-evidence requirements
- Exact source SHA, pipeline source, project/ref/MR source-target identity.
-
Merged CI configuration/rule decision and requested
CI_JOB_TAGS. - Pipeline/job IDs and first-failure trace.
-
Runner ID, scope, protected status,
CI_RUNNER_TAGS, Runner version19.3.2, executor and worker identity. - Privilege/socket/host-mount/device/network configuration from trusted manager state.
- Credential classes/scopes reachable by the job; never token values.
- Cache/worktree/image-reuse configuration and relevant metadata.
- Manager/provider/network logs correlated by timestamp.
- Credential revocation/rotation decisions.
- Worker isolation and destruction/reimage proof before returning capacity to service.
11. Production GitLab mapping
untrusted_tests:
image: alpine:3.22
tags: [untrusted, docker]
script:
- printf 'source=%s sha=%s\n' "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA"
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
trusted_release:
image: alpine:3.22
tags: [protected, deploy]
script:
- printf 'release-sha=%s\n' "$CI_COMMIT_SHA"
rules:
- if: '$CI_COMMIT_REF_PROTECTED == "true" && $CI_PIPELINE_SOURCE == "push"'
when: manual
- when: never
Before using this on a real project, independently verify that the release runner is protected, the untrusted runner has no protected/production credentials, tags cannot fall back to another unsafe pool, executor privilege/host mounts/network are correct, and fork parent-pipeline settings match your threat model.
12. Hardening rollout and rollback plan
| Stage | Go criterion | Rollback action |
|---|---|---|
| Canary untrusted pool | Fork-like synthetic job runs with no protected data/network and fresh worker | Pause new pool; return to prior restricted test path, not to privileged pool |
| Canary trusted pool | Protected synthetic job routes to expected protected runner and exact narrow identity | Pause trusted pool; restore prior protected runner configuration |
| Residue test | Single-use worker destruction and new-worker clean start proven | Stop expansion until lifecycle is fixed |
| Incident logging | Manager/worker/job evidence survives teardown | Do not enable sensitive jobs until evidence pipeline works |
| Traffic shift | Queue/error/security SLOs remain acceptable | Shift tags/routing back to prior safe pool |
13. Final evidence packet
sha256sum runner_policy.py inputs/*.json evidence/*.json > evidence/sha256.txt
printf 'pipeline_source_field=CI_PIPELINE_SOURCE\n' >> evidence/assumptions.txt
printf 'source_sha_field=CI_COMMIT_SHA\n' >> evidence/assumptions.txt
printf 'runner_baseline=19.3.2\n' >> evidence/assumptions.txt
printf 'real_tokens=none privileged_execution=none cloud_resources=none\n' >> evidence/assumptions.txt
find evidence -maxdepth 1 -type f -printf '%f\n' | sort
Review the packet before cleanup. It proves the local policy inputs/results and trust-reset simulation. It does not prove a real GitLab runner is configured securely; that requires platform and host evidence.
14. Cleanup
cd ..
rm -rf glci-ch31-checkpoint
printf 'cleanup=verified-local-only\n'
For a real incident, never delete the worker before evidence preservation and credential-response decisions. Cleanup must target exact runner/worker/cache resources created by the exercise or incident.
15. Verification checklist
- Predictions were written before evaluation.
- Fork attacker input is denied privileged/protected capacity.
- Trusted protected input is authorized only to the intended pool.
- Untrusted pool has no privileged credential class.
- Sensitive pools are single-use in the model.
- Trust reset preserves metadata before simulated destruction.
- Threat model covers fork, Shell, Docker socket, worktree reuse and protected-data leakage.
- Evidence packet contains exact synthetic inputs/results/digests and limitations.
- No real token, Docker socket, privileged container, cloud account or production resource was touched.
Knowledge check
What is the key proof in the attacker case?
The policy denies the fork before execution on privileged/protected capacity. A later job failure would be weaker because dangerous code would already have reached the boundary.
Why does the checkpoint preserve a destruction record before deleting worker-001?
Ephemeral infrastructure can erase the very evidence needed to prove what happened and when trust was reset.
Does worker destruction invalidate a token already exfiltrated by a malicious job?
No. Credential revocation/expiry is a separate identity-control action and must be recorded independently.
Why keep the untrusted and privileged-build pools separate even if both are single-use?
Single-use lifecycle limits residue, but privileged execution has a much larger per-job host blast radius. Trust and capability boundaries still need separate pools.
What does this lab not prove?
It does not prove a real GitLab runner, branch protection, host, network or credential setup is secure. It proves the modeled policy/evidence workflow and supplies a checklist for real verification.
16. What Chapter 31 adds to the production operating model
You can now treat runner security as a chain of independently verifiable controls: classify source trust, compile and inspect job rules, require explicit runner tags without confusing them with authorization, protect sensitive pools, minimize executor privilege, isolate network and credentials, distrust reusable state across trust classes, preserve first-failure evidence, and reset worker trust through verified destruction/reimage.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. The
current stable GitLab Runner patch used as this chapter's
reproducibility baseline is 19.3.2 (tagged 2026-09-10).
GitLab Runner security guidance treats self-managed runners as
remote-code-execution infrastructure, rates Shell executor as high
risk for untrusted builds, warns that privileged containers and
Docker-socket binding can collapse host isolation, and states that
GIT_STRATEGY: fetch on a shared environment is
appropriate only when all users are trusted. Since GitLab 18.1,
same-project merge-request pipelines can be allowed to use protected
variables/runners only when both source and target branches are
protected, the triggering user has suitable target-branch access,
and both branches belong to the same project; fork merge-request
pipelines cannot access those protected resources. The mandatory
exercises are local simulations: they require no runner registration
token, real secret, privileged container, Docker socket, cloud
account, or production network access. The checkpoint intentionally
stops at runner/executor trust. Chapter 32 adds pipeline dependency
pinning, signatures/provenance and trusted-builder identity on top
of this isolation foundation.
- Security for self-managed runners — official reference.
- Configure runners — official reference.
- Runner executors — official reference.
- Shell executor — official reference.
- Docker executor — official reference.
- Use Docker to build Docker images — official reference.
- Merge request pipelines and forks — official reference.
- CI/CD pipelines and protected runner behavior — official reference.
- Pipeline types — official reference.
- CI/CD variables — official reference.
- Predefined CI/CD variables — official reference.
- Advanced Runner configuration — official reference.
- Docker Autoscaler executor — official reference.
- GitLab Runner tags — official reference.
- GitLab 19.3 release notes — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.