Runner Security, Isolation Boundaries, Privileged Containers, Fork Pipelines, Untrusted Code, and Threat Modeling: Diagnostics, Failure Modes, Security, and Performance
Diagnose untrusted fork routing, Shell-executor residue, Docker-socket exposure, reused Git worktrees, and protected-data leakage with first-failure evidence before applying the smallest safe correction.
Learning objectives
- Use an evidence-first sequence from source/configuration through runner/executor and external state.
- Diagnose a fork routed to privileged capacity without executing a privileged container.
- Explain how Shell executor, Docker socket access and reusable worktrees expose cross-job or host state.
- Diagnose protected-data exposure as an identity/trust failure rather than a masking failure.
- Preserve forensic evidence and revoke authority before worker destruction or retries.
1. Evidence-first diagnostic sequence
Do not start by retrying. Preserve CI_PIPELINE_SOURCE,
CI_COMMIT_SHA, pipeline/job IDs and the first-failure
trace. Confirm the merged configuration and
workflow:rules/job rules. Inspect the job
graph and requested tags. Confirm the assigned runner
ID/scope/protection/tags/version/executor and worker identity. Then
inspect shell/tool/network behavior, artifacts/caches/reports, and
any external target. Apply the least destructive correction and
rerun only the smallest safe scope.
2. Failure: an untrusted fork appears eligible for a privileged runner
Suppose the CI configuration requests image-build and a
privileged runner advertises that tag but was not configured as
protected. A naive tag-only scheduler would accept the fork job. The
original failure is the runner trust policy, not the build script.
{
"pipeline_source": "merge_request_event",
"source_project_id": 2202,
"target_project_id": 1101,
"required_tags": ["image-build"],
"runner": {
"tags": ["image-build"],
"protected": false,
"privileged": true,
"ephemeral": false
}
}
Preserve the pipeline/job IDs and runner assignment, then stop new work on that runner. The safer correction is to move privileged building to a dedicated protected and single-use pool (or eliminate privilege), narrow tags/network/credentials, and route fork tests to a separate untrusted pool. Do not “repair” by trusting the fork.
3. Failure: Shell executor shares code or host secrets
GitLab states that Shell executor runs jobs with the runner user's permissions and can steal code from other projects on the same server. Evidence should include executor type, manager host identity, build directories, user/group permissions, home-directory credential files and which projects used the host. The correction is architectural: keep Shell executor for trusted builds only or move mixed/untrusted work to a stronger isolated/ephemeral boundary.
Do not print the contents of SSH keys, Docker config files or environment variables while investigating. Record path existence, ownership, mode and hashes where policy allows.
id
umask
find "$HOME" -maxdepth 1 -type f -printf '%f mode=%m owner=%u:%g\n' 2>/dev/null | sort
# Do not cat credential files or dump the environment.
4. Failure: host Docker socket is mounted
A runner can be configured privileged = false yet still
expose /var/run/docker.sock. GitLab documents that
shared Docker-daemon control effectively disables container security
and can permit host privilege escalation. Detect the configuration
from trusted runner-manager config—not from attacker-controlled job
code—and treat the host as potentially compromised if an untrusted
job had socket access.
# Evidence pattern to find in trusted runner configuration:
[runners.docker]
privileged = false
volumes = ["/var/run/docker.sock:/var/run/docker.sock", "/cache"]
Repair by removing host-daemon exposure from the untrusted pool and choosing an isolated image-build strategy. Rotate registry/cloud credentials reachable from that host and reimage/destroy the worker if necessary.
5. Failure: GIT_STRATEGY: fetch reuses hostile state
You can demonstrate the principle locally without GitLab. The first “job” leaves a marker in a reused worktree; the second sees it. A clean-start simulation removes the directory first.
rm -rf /tmp/ch31-worktree
mkdir -p /tmp/ch31-worktree/.git
printf 'left-by-job-A\n' > /tmp/ch31-worktree/.git/attacker-marker
# Reuse model (analogous to relying on old worktree state):
test -f /tmp/ch31-worktree/.git/attacker-marker && printf 'reuse_detected=yes\n'
# Clean-start model:
rm -rf /tmp/ch31-worktree
mkdir -p /tmp/ch31-worktree
if test -e /tmp/ch31-worktree/.git/attacker-marker; then
printf 'unexpected residue\n' >&2; exit 1
else
printf 'clean_start=yes\n'
fi
This does not emulate Git internals; it demonstrates the ownership problem. GitLab's real warning is stronger: reusable Git metadata and submodule reflogs can expose prior content. Use fetch only in trusted shared environments.
6. Failure: protected data reaches attacker-controlled code
Start by asking how the job became eligible for the data. Protected-variable masking is irrelevant if malicious code receives the variable legitimately. Check protected ref status, MR source/target project identity, the “allow MR pipelines to access protected variables and runners” setting, the triggering user's permissions and whether a parent-project fork MR was intentionally run.
Current GitLab behavior blocks protected resources for fork MR pipelines under the protected-MR resource feature. But a parent-project fork pipeline can still use parent resources/settings generally, so reviewers must inspect the fork branch CI configuration and code before triggering that context. Keep production secrets out of general parent runner hosts as a second line of defense.
7. Intentionally broken routing policy
This deliberately bad policy authorizes solely by tags:
def broken_eligible(job, runner):
# WRONG: capability matching is not a trust decision.
return set(job["required_tags"]).issubset(set(runner["tags"]))
job = {"trust": "untrusted-fork", "required_tags": ["image-build"]}
runner = {"tags": ["image-build"], "protected": False, "privileged": True}
print(broken_eligible(job, runner)) # True: dangerous false authorization
The repair introduces independent trust predicates: fork/same-project identity, protected-ref evidence, protected runner state, privilege restrictions and credential scope. Preserve the bad decision output as evidence; do not overwrite it with the repaired result.
8. Preserve the first failure before any retry
| Evidence | Why preserve it first |
|---|---|
| Pipeline/job IDs and source SHA | Retries create new job IDs and may execute a different mutable ref unless SHA is recorded. |
| Merged CI configuration/rule result | Shows why the sensitive job existed and which tags it requested. |
| Runner ID/version/protection/tags/executor | Shows which infrastructure accepted the job. |
| Manager logs/worker ID | Job trace alone may not show socket mounts, provisioning or host failures. |
| Non-secret network/provider events | Shows exfiltration/reachability or infrastructure changes. |
| Artifact/cache metadata | Can reveal whether malicious output was retained or propagated. |
| Credential scope/rotation record | Shows blast radius without storing credential values. |
| Destruction/reimage proof | Shows when the compromised worker stopped being trusted. |
9. Least-destructive incident response
- Pause or isolate the affected runner/pool so no new jobs are assigned.
- Preserve manager/provider/job evidence outside the worker.
- Identify exactly which credentials, networks, caches and projects were reachable.
- Revoke/rotate exposed credentials and invalidate suspicious sessions/tokens as appropriate.
- Destroy/reimage the worker from a trusted baseline; do not attempt to “clean” an unknown-compromised host in place.
- Repair routing/protection/executor/network policy.
- Run a synthetic canary first; restore workloads gradually.
10. Performance symptoms can be security signals
Unexpected cache misses, unusual image pulls, high outbound traffic, long after-script time, manager errors or repeated worker replacement can indicate both performance regressions and malicious behavior. Correlate queue/job metrics with runner logs and network/provider telemetry. Do not automatically add capacity to a pool showing unexplained resource consumption.
11. Shortcuts explicitly rejected
- Do not make a protected runner unprotected just to clear pending fork jobs.
- Do not mount the Docker socket to avoid designing an isolated image builder.
- Do not dump environment variables or token values for debugging.
- Do not disable TLS to make registry/network failures disappear.
- Do not blindly retry a job that may have repeated external side effects.
- Do not trust cache/worktree cleanup as proof that a compromised host is safe.
Knowledge check
A privileged runner accepted a fork MR because the tag matched. What was the primary control failure?
Runner eligibility/trust policy: the privileged pool was not constrained by protected/trust controls. The build script is downstream of that failure.
Why should a Docker-socket incident trigger host-level response?
The job had authority over the host Docker daemon, so the container boundary no longer contains the likely blast radius.
What does the local worktree demo prove and not prove?
It proves reusable directories carry state across jobs. It does not reproduce GitLab Runner or Git reflog behavior; GitLab documentation supplies that product-specific risk.
A protected variable appears in a malicious job. Is masking the first fix?
No. First determine why the job was authorized to receive protected data; fix the trust/protection/routing boundary and rotate the exposed credential.
Why preserve the original job ID before retrying?
A retry creates new execution evidence and can repeat side effects. The original ID anchors first-failure logs, runner assignment and source/configuration context.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. The
current stable GitLab Runner patch used as this chapter's
reproducibility baseline is 19.3.2 (tagged 2026-09-10).
GitLab Runner security guidance treats self-managed runners as
remote-code-execution infrastructure, rates Shell executor as high
risk for untrusted builds, warns that privileged containers and
Docker-socket binding can collapse host isolation, and states that
GIT_STRATEGY: fetch on a shared environment is
appropriate only when all users are trusted. Since GitLab 18.1,
same-project merge-request pipelines can be allowed to use protected
variables/runners only when both source and target branches are
protected, the triggering user has suitable target-branch access,
and both branches belong to the same project; fork merge-request
pipelines cannot access those protected resources. The mandatory
exercises are local simulations: they require no runner registration
token, real secret, privileged container, Docker socket, cloud
account, or production network access. The deliberately broken
examples are non-destructive local/configuration examples. Never
reproduce privileged/socket exposure on a shared or production
runner for training.
- Security for self-managed runners — official reference.
- Configure runners — official reference.
- Runner executors — official reference.
- Shell executor — official reference.
- Docker executor — official reference.
- Use Docker to build Docker images — official reference.
- Merge request pipelines and forks — official reference.
- CI/CD pipelines and protected runner behavior — official reference.
- Pipeline types — official reference.
- CI/CD variables — official reference.
- Predefined CI/CD variables — official reference.
- Advanced Runner configuration — official reference.
- Docker Autoscaler executor — official reference.
- GitLab Runner tags — official reference.
- GitLab 19.3 release notes — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.