Chapter 31Lesson 04~230 minutes

Runner Security, Isolation Boundaries, Privileged Containers, Fork Pipelines, Untrusted Code, and Threat Modeling: Diagnostics, Failure Modes, Security, and Performance

Diagnose untrusted fork routing, Shell-executor residue, Docker-socket exposure, reused Git worktrees, and protected-data leakage with first-failure evidence before applying the smallest safe correction.

DiagnosticsForksDocker socketWorkspace reuseIncident response

Learning objectives

  • Use an evidence-first sequence from source/configuration through runner/executor and external state.
  • Diagnose a fork routed to privileged capacity without executing a privileged container.
  • Explain how Shell executor, Docker socket access and reusable worktrees expose cross-job or host state.
  • Diagnose protected-data exposure as an identity/trust failure rather than a masking failure.
  • Preserve forensic evidence and revoke authority before worker destruction or retries.

1. Evidence-first diagnostic sequence

Do not start by retrying. Preserve CI_PIPELINE_SOURCE, CI_COMMIT_SHA, pipeline/job IDs and the first-failure trace. Confirm the merged configuration and workflow:rules/job rules. Inspect the job graph and requested tags. Confirm the assigned runner ID/scope/protection/tags/version/executor and worker identity. Then inspect shell/tool/network behavior, artifacts/caches/reports, and any external target. Apply the least destructive correction and rerun only the smallest safe scope.

Incident rule: if the runner host or privileged worker may be compromised, do not run “cleanup scripts” supplied by the suspect repository. Isolate it, preserve evidence externally, rotate exposed authority, and destroy/reimage from a trusted image.

2. Failure: an untrusted fork appears eligible for a privileged runner

Suppose the CI configuration requests image-build and a privileged runner advertises that tag but was not configured as protected. A naive tag-only scheduler would accept the fork job. The original failure is the runner trust policy, not the build script.

{
  "pipeline_source": "merge_request_event",
  "source_project_id": 2202,
  "target_project_id": 1101,
  "required_tags": ["image-build"],
  "runner": {
    "tags": ["image-build"],
    "protected": false,
    "privileged": true,
    "ephemeral": false
  }
}

Preserve the pipeline/job IDs and runner assignment, then stop new work on that runner. The safer correction is to move privileged building to a dedicated protected and single-use pool (or eliminate privilege), narrow tags/network/credentials, and route fork tests to a separate untrusted pool. Do not “repair” by trusting the fork.

3. Failure: Shell executor shares code or host secrets

GitLab states that Shell executor runs jobs with the runner user's permissions and can steal code from other projects on the same server. Evidence should include executor type, manager host identity, build directories, user/group permissions, home-directory credential files and which projects used the host. The correction is architectural: keep Shell executor for trusted builds only or move mixed/untrusted work to a stronger isolated/ephemeral boundary.

Do not print the contents of SSH keys, Docker config files or environment variables while investigating. Record path existence, ownership, mode and hashes where policy allows.

id
umask
find "$HOME" -maxdepth 1 -type f -printf '%f mode=%m owner=%u:%g\n' 2>/dev/null | sort
# Do not cat credential files or dump the environment.

4. Failure: host Docker socket is mounted

A runner can be configured privileged = false yet still expose /var/run/docker.sock. GitLab documents that shared Docker-daemon control effectively disables container security and can permit host privilege escalation. Detect the configuration from trusted runner-manager config—not from attacker-controlled job code—and treat the host as potentially compromised if an untrusted job had socket access.

# Evidence pattern to find in trusted runner configuration:
[runners.docker]
  privileged = false
  volumes = ["/var/run/docker.sock:/var/run/docker.sock", "/cache"]

Repair by removing host-daemon exposure from the untrusted pool and choosing an isolated image-build strategy. Rotate registry/cloud credentials reachable from that host and reimage/destroy the worker if necessary.

5. Failure: GIT_STRATEGY: fetch reuses hostile state

You can demonstrate the principle locally without GitLab. The first “job” leaves a marker in a reused worktree; the second sees it. A clean-start simulation removes the directory first.

rm -rf /tmp/ch31-worktree
mkdir -p /tmp/ch31-worktree/.git
printf 'left-by-job-A\n' > /tmp/ch31-worktree/.git/attacker-marker

# Reuse model (analogous to relying on old worktree state):
test -f /tmp/ch31-worktree/.git/attacker-marker && printf 'reuse_detected=yes\n'

# Clean-start model:
rm -rf /tmp/ch31-worktree
mkdir -p /tmp/ch31-worktree
if test -e /tmp/ch31-worktree/.git/attacker-marker; then
  printf 'unexpected residue\n' >&2; exit 1
else
  printf 'clean_start=yes\n'
fi

This does not emulate Git internals; it demonstrates the ownership problem. GitLab's real warning is stronger: reusable Git metadata and submodule reflogs can expose prior content. Use fetch only in trusted shared environments.

6. Failure: protected data reaches attacker-controlled code

Start by asking how the job became eligible for the data. Protected-variable masking is irrelevant if malicious code receives the variable legitimately. Check protected ref status, MR source/target project identity, the “allow MR pipelines to access protected variables and runners” setting, the triggering user's permissions and whether a parent-project fork MR was intentionally run.

Current GitLab behavior blocks protected resources for fork MR pipelines under the protected-MR resource feature. But a parent-project fork pipeline can still use parent resources/settings generally, so reviewers must inspect the fork branch CI configuration and code before triggering that context. Keep production secrets out of general parent runner hosts as a second line of defense.

7. Intentionally broken routing policy

This deliberately bad policy authorizes solely by tags:

def broken_eligible(job, runner):
    # WRONG: capability matching is not a trust decision.
    return set(job["required_tags"]).issubset(set(runner["tags"]))

job = {"trust": "untrusted-fork", "required_tags": ["image-build"]}
runner = {"tags": ["image-build"], "protected": False, "privileged": True}
print(broken_eligible(job, runner))  # True: dangerous false authorization

The repair introduces independent trust predicates: fork/same-project identity, protected-ref evidence, protected runner state, privilege restrictions and credential scope. Preserve the bad decision output as evidence; do not overwrite it with the repaired result.

8. Preserve the first failure before any retry

Evidence Why preserve it first
Pipeline/job IDs and source SHA Retries create new job IDs and may execute a different mutable ref unless SHA is recorded.
Merged CI configuration/rule result Shows why the sensitive job existed and which tags it requested.
Runner ID/version/protection/tags/executor Shows which infrastructure accepted the job.
Manager logs/worker ID Job trace alone may not show socket mounts, provisioning or host failures.
Non-secret network/provider events Shows exfiltration/reachability or infrastructure changes.
Artifact/cache metadata Can reveal whether malicious output was retained or propagated.
Credential scope/rotation record Shows blast radius without storing credential values.
Destruction/reimage proof Shows when the compromised worker stopped being trusted.

9. Least-destructive incident response

  1. Pause or isolate the affected runner/pool so no new jobs are assigned.
  2. Preserve manager/provider/job evidence outside the worker.
  3. Identify exactly which credentials, networks, caches and projects were reachable.
  4. Revoke/rotate exposed credentials and invalidate suspicious sessions/tokens as appropriate.
  5. Destroy/reimage the worker from a trusted baseline; do not attempt to “clean” an unknown-compromised host in place.
  6. Repair routing/protection/executor/network policy.
  7. Run a synthetic canary first; restore workloads gradually.

10. Performance symptoms can be security signals

Unexpected cache misses, unusual image pulls, high outbound traffic, long after-script time, manager errors or repeated worker replacement can indicate both performance regressions and malicious behavior. Correlate queue/job metrics with runner logs and network/provider telemetry. Do not automatically add capacity to a pool showing unexplained resource consumption.

11. Shortcuts explicitly rejected

  • Do not make a protected runner unprotected just to clear pending fork jobs.
  • Do not mount the Docker socket to avoid designing an isolated image builder.
  • Do not dump environment variables or token values for debugging.
  • Do not disable TLS to make registry/network failures disappear.
  • Do not blindly retry a job that may have repeated external side effects.
  • Do not trust cache/worktree cleanup as proof that a compromised host is safe.

Knowledge check

A privileged runner accepted a fork MR because the tag matched. What was the primary control failure?

Why should a Docker-socket incident trigger host-level response?

What does the local worktree demo prove and not prove?

A protected variable appears in a malicious job. Is masking the first fix?

Why preserve the original job ID before retrying?

Next lesson

Checkpoint Lab — Harden and prove a runner design

You will build a complete local threat model, test allowed and denied workload cases, simulate trust reset, and produce a production-ready evidence checklist.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. The current stable GitLab Runner patch used as this chapter's reproducibility baseline is 19.3.2 (tagged 2026-09-10). GitLab Runner security guidance treats self-managed runners as remote-code-execution infrastructure, rates Shell executor as high risk for untrusted builds, warns that privileged containers and Docker-socket binding can collapse host isolation, and states that GIT_STRATEGY: fetch on a shared environment is appropriate only when all users are trusted. Since GitLab 18.1, same-project merge-request pipelines can be allowed to use protected variables/runners only when both source and target branches are protected, the triggering user has suitable target-branch access, and both branches belong to the same project; fork merge-request pipelines cannot access those protected resources. The mandatory exercises are local simulations: they require no runner registration token, real secret, privileged container, Docker socket, cloud account, or production network access. The deliberately broken examples are non-destructive local/configuration examples. Never reproduce privileged/socket exposure on a shared or production runner for training.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.