Chapter 31Lesson 03~220 minutes

Runner Security, Isolation Boundaries, Privileged Containers, Fork Pipelines, Untrusted Code, and Threat Modeling: Configuration, Design Choices, and Tradeoffs

Choose shared versus dedicated runners, privileged versus rootless/non-privileged execution, persistent versus ephemeral workers, fork-routing strategy, and network reachability using explicit security, latency, cost, and recovery tradeoffs.

DesignPrivilegeEphemeralNetworkTradeoffs

Learning objectives

  • Choose runner tenancy and scope based on trust, not convenience alone.
  • Compare privileged, restricted-rootless and non-privileged execution without assuming containers are perfect sandboxes.
  • Decide when ephemeral single-use workers justify their cold-start/cost overhead.
  • Design fork contribution flows that keep unreviewed code away from protected data and sensitive networks.
  • Balance network reachability, caching and image reuse against cross-job contamination risk.

1. The design problem: every convenience feature changes a trust boundary

Runner architecture is full of performance/security trades: sharing a host improves utilization; persistent worktrees speed checkout; Docker socket binding speeds image workflows; broad network access avoids dependency failures; shared caches reduce build time. Each optimization also increases the amount of state or authority that untrusted code can reach. The correct design starts with workload trust classes and then chooses the smallest boundary that satisfies them.

Keep source identity explicit in every design review: record CI_PIPELINE_SOURCE and the exact CI_COMMIT_SHA alongside the compiled rules, requested tags and assigned runner. Without those fields, you cannot prove which code/trust context the runner architecture actually served.

2. Shared versus dedicated: separate runner scope from actual tenancy

GitLab runner scope can be instance, group or project-oriented, but scope alone does not prove physical isolation. A project runner may still share a manager or worker host with another project; an instance runner can still be backed by single-use workers. Document both the GitLab assignment scope and the execution tenancy.

Pattern Strength Weakness Good fit
Broad shared pool + non-privileged ephemeral workers High utilization and fresh workers Shared manager/network/cache boundaries still need hardening Untrusted tests with no sensitive credentials
Project/dedicated pool Narrow eligibility and clearer ownership Higher idle cost/operational overhead Sensitive builds or special hardware
Single-use VM per job Strong trust reset between jobs Cold start and infrastructure cost Privileged image build, external contributions, high-risk compilation
Persistent static host Low startup latency, simple tooling Highest residue and host-compromise risk Only tightly trusted/internal workloads with strong host hardening

3. Privileged versus rootless/non-privileged

GitLab's Docker executor security guidance favors non-privileged execution for ordinary jobs. Privileged mode grants capabilities that can expose the host. Docker socket binding is not a safe “non-privileged workaround”; the job controls the host Docker daemon. If container-image building needs extra capability, isolate that workload on a dedicated ephemeral pool and prefer a rootless/non-privileged builder where feasible.

GitLab also supports restricting which service images may use privileged mode through services_privileged and allowed_privileged_services, including rootless Docker-in-Docker patterns. This narrows exposure but does not make privileged execution harmless.

# Architecture example: narrower than globally privileged jobs.
# Still use only on an isolated, dedicated worker after validating the exact image source.
[runners.docker]
  privileged = false
  services_privileged = true
  allowed_privileged_services = [
    "docker.io/library/docker:*-dind-rootless",
    "docker.io/library/docker:dind-rootless"
  ]

4. Persistent versus ephemeral is a cost/forensics/isolation decision

Ephemeral workers reduce cross-job residue because you can destroy the entire worker after use. They also make local logs transient, so manager/provider logs must leave the worker. Persistent workers reduce cold start and can improve cache locality, but every reusable worktree, home directory, image layer and volume becomes a possible cross-job channel.

For high-risk or privileged workloads, a strong baseline is single-use worker + no production credentials on the host + restricted network + externalized logs + proven destruction. For ordinary trusted internal tests, reuse may be acceptable with cleanup and monitored boundaries.

5. Source reuse: fetch is a performance choice with a trust prerequisite

Git strategy State behavior Security interpretation
fetch Attempts to reuse the prior worktree Fast, but GitLab says to use it only when all users of the shared environment are trusted.
clone Removes an existing worktree and clones again Cleaner repository start, but other host/cache state can still persist.
empty Deletes/recreates build directory and skips Git checkout Useful artifact-only jobs that need a controlled empty directory.
Ephemeral worker Destroys the worker boundary after job Strongest reset when destruction is actually verified.

FF_ENABLE_JOB_CLEANUP adds Git repository cleanup on static runners, but it is defense in depth—not equivalent to reimaging a compromised host.

6. Fork execution versus an external-contribution path

There are two common models. In a permissive model, fork code gets CI feedback on a deliberately untrusted pool: no protected variables, no privileged runner, minimal internal network, read-only/fake fixtures. After review/merge, a separate trusted pipeline performs release/deployment. In a stricter model, external contributions run only static/lint/test jobs that need no secrets; privileged or integration jobs are replayed from reviewed code after it enters a trusted ref.

A parent-project fork MR pipeline should be an explicit exception, not the default shortcut for “use our fast runners,” because it combines untrusted fork configuration with parent project resources and the triggering member's permissions.

7. Network reachability versus usability

Most CI jobs need outbound dependency access. Very few need unrestricted routes to production databases, metadata services, internal admin panels or runner-management endpoints. Separate runner pools into network zones. Use egress allowlists/proxies where practical, filter metadata endpoints, and keep privileged/deployment pools on narrow service-specific routes.

Pool Suggested reachability Credentials
Untrusted test GitLab + approved dependency mirrors/Internet egress; no production/internal admin networks None beyond narrow job-scoped GitLab access
Trusted build GitLab + registry/package services Registry push or signing identity only as needed
Deployment GitLab + exact target API/cluster/account Short-lived deployment identity, environment-scoped
Privileged image build Source + registry; isolated from production control plane Registry push only; never production deploy credentials

8. Image reuse can become an authorization bypass

GitLab warns against Docker if-not-present pull policy on public/instance runners that mix users and private images. One user can authenticate and populate a private image locally; another job on the same runner could reuse that local image without the original authorization check. For mixed-trust shared runners, use a policy that forces the appropriate registry authorization behavior and avoid persistent host image state where possible.

9. Cache boundaries are not trust boundaries

Do not share writeable caches across trust classes just because the key matches a language lockfile. An attacker can poison executable package caches, compiler outputs or tool directories. Separate cache namespaces/backends for untrusted and protected pools, prefer read-only dependency mirrors where possible, and never place credentials in cache content.

10. Identity design: the runner host should not be a credential warehouse

Use job-scoped identities: CI_JOB_TOKEN for supported GitLab-to-GitLab access and explicit OIDC ID tokens for external federation. Avoid long-lived personal tokens and cloud keys on worker disks. A protected runner can be compromised by protected code; credentials should still be least-privilege and short-lived.

For a fork/untrusted pool, assume job code can read every environment variable supplied to it and can make network requests. “Masked” reduces accidental log exposure; it does not stop malicious code from exfiltrating a secret.

11. Worked design decision table

Workload Runner design Prerequisites Observable evidence Rollback
Public/fork unit tests Non-privileged, ephemeral untrusted pool; no protected data Free CI/CD; fork/MR rules; narrow network Fork/source project IDs, tags, runner protection=false, executor/worker ID, no secret scope Pause pool; route to safer restricted pool
Protected release packaging Protected, dedicated non-privileged pool Protected ref governance; runner protected; narrow tags CI_COMMIT_REF_PROTECTED, runner ID/protection/tags, artifact digest Pause release pool; use prior trusted package pipeline
Privileged image build Protected, dedicated single-use VM; privileged only if unavoidable Restricted maintainers; isolated network; external logs Worker one-use proof, privileged config, registry-only identity, destruction record Pause pool; switch to non-privileged builder or prior builder image
Production deployment Protected deployment pool, no arbitrary fork/MR execution Protected ref/environment authorization; short-lived external identity Pipeline/source SHA, runner ID, OIDC/deploy identity metadata, target health Stop new deploy jobs; rollback exact artifact/deployment

12. Security controls have performance and cost consequences

Single-use VMs add cold-start latency. Network proxies and image pulls add transfer time. Dedicated pools can sit idle. These are measurable costs, not reasons to collapse boundaries. Track queue time, provisioning time, image-pull/cache time and cost per successful job by trust class. Tune warm capacity or artifact distribution without silently converting a single-use protected worker into a shared persistent host.

13. Runner-hardening changes need rollback too

A safer runner change can still break builds. Canary a hardened pool with synthetic representative jobs, keep the previous configuration paused but available during a rollback window, verify queue/error/security evidence, then decommission. Never restore service by broadly enabling privileged mode or exposing the host Docker socket to all jobs.

Knowledge check

A project-scoped runner executes on a host shared with other projects. Is the job physically dedicated?

Why is Docker socket binding not a safe substitute for privileged mode on an untrusted runner?

When can GIT_STRATEGY: fetch be reasonable?

What is the cleanest trust boundary between fork validation and production deployment?

Why should a protected runner still use short-lived least-privilege credentials?

Next lesson

Diagnostics and realistic failure modes

Lesson 4 intentionally creates safe local evidence for the five required failure classes and shows how to diagnose the control layer without weakening runner protection.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. The current stable GitLab Runner patch used as this chapter's reproducibility baseline is 19.3.2 (tagged 2026-09-10). GitLab Runner security guidance treats self-managed runners as remote-code-execution infrastructure, rates Shell executor as high risk for untrusted builds, warns that privileged containers and Docker-socket binding can collapse host isolation, and states that GIT_STRATEGY: fetch on a shared environment is appropriate only when all users are trusted. Since GitLab 18.1, same-project merge-request pipelines can be allowed to use protected variables/runners only when both source and target branches are protected, the triggering user has suitable target-branch access, and both branches belong to the same project; fork merge-request pipelines cannot access those protected resources. The mandatory exercises are local simulations: they require no runner registration token, real secret, privileged container, Docker socket, cloud account, or production network access. Tier note: the core runner/pipeline/protected-ref concepts used here are part of normal GitLab CI/CD offerings. Organization-specific network controls, cloud workers and enterprise governance can vary; the mandatory reasoning path remains local/free-compatible.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.