Chapter 14Lesson 03~225 minutes

GitLab Runners, Executors, Tags, Registration, Autoscaling, Isolation, and Security: Configuration, Design Choices, and Tradeoffs

Choose runner ownership, scope, executor, protection, persistence, network, cache, and autoscaling strategies using explicit blast-radius, reliability, maintainability, and cost tradeoffs.

Runner designShell riskAutoscalingEphemeral workersBlast radiusCost

Learning objectives

  • Choose between GitLab-hosted and self-managed runners based on control, isolation, locality, maintenance, and compute constraints.
  • Compare shell, Docker, Kubernetes, Docker Autoscaler, and Instance executor designs without treating executor names as security guarantees.
  • Select instance/group/project runner scope and tag/protection policy according to blast radius and ownership.
  • Decide when persistent workers are acceptable and when ephemeral/autoscaled workers materially reduce risk.
  • Document a runner architecture decision that includes network, cache, upgrade, observability, and cost assumptions.
Availability baseline (verified 2026-08-21 against current GitLab documentation). GitLab Runner, project/group/instance runner scopes, tags, protected runners, the supported executor framework, and self-managed runners are available across Free/Premium/Ultimate. GitLab-hosted runners are a GitLab.com service and are also available for GitLab Dedicated under a separately provisioned Limited Availability offering; Self-Managed installations provide their own runner infrastructure. Hosted-runner compute can be quota/billing constrained, so every mandatory exercise has a no-runner/static or evidence-fixture path. Legacy runner registration tokens are deprecated and scheduled for removal in GitLab 20.0; this chapter teaches the modern runner creation workflow with runner authentication tokens.

1. Design from workload trust and required capability

Start with the workload, not the executor. Ask: who can modify the code that reaches this runner? Which secrets can jobs receive? Does the workload need internal network access, privileged container operations, GPUs, special hardware, licensed tooling, or deployment credentials? How much cross-job persistence is acceptable?

The runner architecture should make the expected workload easy and dangerous workloads structurally difficult.

2. GitLab-hosted versus self-managed

Dimension GitLab-hosted runner Self-managed runner
Operations GitLab manages fleet provisioning and base runner infrastructure. You install, patch, upgrade, monitor, scale, and decommission Runner/hosts.
Isolation model Hosted jobs normally receive fresh provider-managed infrastructure. Depends entirely on executor, host reuse, mounts, network, and your cleanup.
Network locality Internet/service-defined connectivity; private access requires supported networking patterns. Can be placed near internal build/deploy targets—but that expands blast radius.
Special hardware/software Limited to offered runner classes/images. Full control over hardware, images, drivers, licensed tools, and architecture.
Cost model Compute quota/credits/cost factors can apply; verify current service terms. Infrastructure + operations + idle capacity + cloud/network/storage cost.
Best fit General CI where managed ephemeral infrastructure is sufficient. Private-network, hardware-specific, compliance/locality, or specialized workloads.

3. Shell simplicity versus isolation risk

The Shell executor is operationally simple because commands run directly on the runner host. That same property creates the core security problem: job code shares the host account, filesystem, installed tools, network, and potentially leftovers from other jobs. Current GitLab documentation places Shell in maintenance mode and explicitly recommends using it only for trusted builds.

A non-privileged Docker executor creates a cleaner process/filesystem boundary and makes dependencies reproducible through images. But Docker is not a magic security boundary: privileged mode, host Docker socket binding, host PID mode, broad capabilities, and sensitive host-volume mounts can effectively hand the job the host.

4. Executor selection matrix

Need Prefer Why / caution
Ordinary reproducible Linux build Docker executor or hosted runner. Image-defined dependencies, clean job environment; keep non-privileged where possible.
Kubernetes-native build isolation/integration Kubernetes executor. Per-job pods; still govern cluster RBAC, service accounts, node trust, namespace and network policy.
Autoscaled container workload Docker Autoscaler. Current fleeting/taskscaler path; can make each instance ephemeral.
Autoscaled non-containerized host workload Instance executor. Job runs directly on autoscaled host; useful for VM-native tooling.
Legacy trusted host tooling Shell executor only with explicit trust decision. Persistent host and direct execution create high cross-job/host risk.
Unsupported custom environment Custom only after accepting maintenance-mode/support tradeoff. You own driver correctness, provisioning, cleanup, and security.

5. Scope is blast-radius architecture

Instance scope maximizes reuse and therefore requires the strongest trust policy. Group scope can align runner ownership with a platform team and a related project family. Project scope is easier to reason about for sensitive deployment or specialized workloads. More sharing can improve utilization but also increases the number of repositories/users whose code can reach the same execution fleet.

Design rule. Never broaden scope solely to improve utilization if the runner has privileged host capability, production network reachability, or powerful credentials. First reduce privilege/isolate the workload, then decide whether broader sharing is acceptable.

6. Tags route capability; protection gates trust

Use tags for capabilities such as linux, arm64, gpu, or a dedicated workload class. Do not rely on a tag named prod as a security boundary: a repository author who can set job tags can request that runner if scope and policy permit it.

For sensitive runners, combine narrow scope with protected-runner policy, protected branches/tags, reviewed pipeline configuration, limited secrets, and network segmentation. GitLab documents that protected runners can be restricted to protected branches/tags, with specific support for eligible merge-request pipelines when their protected-ref conditions are met.

7. Persistent versus ephemeral workers

Pattern Advantages Risks/costs Good use
Persistent runner host Fast warm state; simple; predictable local tools. Cross-job persistence, patch drift, cache/workspace residue, larger compromise window. Trusted internal workloads with strong cleanup and hardening.
Reusable autoscaled instance Amortizes startup cost over several jobs. A compromised job can affect later jobs on same worker. Moderate-trust jobs when bounded reuse is justified.
One job per ephemeral instance Strong reduction in cross-job persistence. Higher provisioning latency/cost; image bootstrap must be reliable. Untrusted or high-sensitivity workloads, especially privileged builds.
Provider-managed hosted ephemeral worker Low operational burden and clean lifecycle. Less infrastructure control; quota/service constraints. General GitLab.com CI.

The Docker Autoscaler documentation explicitly supports configurations such as capacity_per_instance=1 and max_use_count=1 so each job gets a disposable instance. That is a security design choice, not merely a scaling knob.

8. Autoscaling policy: latency, utilization, and failure containment

Autoscaling introduces a control loop: Runner asks for work, taskscaler manages desired capacity, a fleeting provider plugin creates/deletes cloud instances, and the executor runs jobs. The key knobs determine maximum instances, capacity per instance, reuse count, idle capacity, and idle time.

  • More idle instances reduce queue latency but increase idle cost.
  • More jobs per instance increase utilization but increase cross-job persistence/blast radius.
  • Fast scale-out depends on cloud quotas, image availability, network, DNS, and provider API health.
  • Dedicated autoscaling resources should not be controlled by multiple independent runner-manager configurations; competing controllers can produce unpredictable behavior and cost.

9. Network and cache are part of the runner threat model

Runner placement often exists because jobs need private package servers, test databases, registries, or deployment targets. That same connectivity becomes reachable by arbitrary job scripts unless filtered. Put runner hosts in a dedicated segment, restrict egress/metadata endpoints, and expose only required services.

Cache is optimization state, not trusted release evidence. Scope cache keys carefully, avoid putting secrets in caches, and assume an untrusted job may try to poison any cache it can write. Release artifacts/packages should have their own integrity and permission controls.

10. Privileged workloads deserve dedicated ephemeral capacity

Security-sensitive design. GitLab's runner security guidance warns that Docker privileged mode effectively disables important container security mechanisms and can permit host-level compromise. If privileged execution is unavoidable, isolate it onto dedicated runners, restrict them to trusted/protected workloads, and prefer ephemeral single-use virtual machines.

Do not mix ordinary fork/MR jobs and privileged Docker-in-Docker or deployment jobs on the same persistent host simply because both need Docker.

11. Upgrades, ownership, and decommissioning are design requirements

Every self-managed runner needs a named owner, supported Runner version policy, patch window, health check, token-rotation/revocation procedure, host inventory, and decommission checklist. GitLab distinguishes online, offline, stale, and never_contacted states; these are operational signals, not merely UI decoration.

Some centralized runner fleet statistics are tier-gated, but every team can still inventory runner IDs, versions, status, jobs, host logs, and local gitlab-runner --version without buying a tier.

12. Worked scenario: payment service CI

Workload Chosen runner design Reasoning
Unit tests from developer branches/MRs GitLab-hosted or non-privileged ephemeral Docker runner. Untrusted code; no internal secrets; reproducible image; minimize persistence.
Integration tests against a synthetic isolated database Dedicated group/project Docker runner with narrow network segment. Needs private test service but should not reach production.
Container build requiring privileged mode Dedicated ephemeral autoscaled runner restricted to reviewed/protected refs. Privileged workload gets its own blast radius; one job per instance preferred.
Production deployment Project-scoped protected deployment runner or identity-based deployment path. Narrow source trust, narrow network, short-lived credentials; do not share with ordinary CI.

The architecture deliberately uses more than one runner class because the workloads have different trust and capability requirements. “One powerful runner for everything” is cheaper to draw but harder to secure.

13. Runner architecture decision record

Record at least: workload owners; allowed projects/refs; required tags; protected/run-untagged policy; executor and privilege mode; host/worker lifetime; network destinations; cache backend/scope; credential sources; autoscaling limits; expected concurrency; upgrade owner; logging/metrics; token rotation; incident isolation; and decommission steps.

Knowledge check

Why is a runner tag not sufficient to protect a production deployment runner?

When is a Shell executor defensible?

What does max_use_count=1 accomplish in an autoscaled design?

Why can broader runner scope increase blast radius?

What is the principal security concern with Docker privileged mode?

Why should cache be treated differently from release evidence?

Summary

Runner architecture is a workload-trust decision. Choose scope, tags, protection, executor, persistence, network, cache, and autoscaling as one system. Prefer managed or non-privileged isolated capacity for ordinary work, narrow and protect sensitive runners, and make privileged execution ephemeral and dedicated when it cannot be avoided.

Official references

Next lesson

Diagnose the runner without weakening it

Lesson 4 builds a failure-first runbook for pending jobs, protected-runner mismatches, obsolete registration-token errors, stale/decommissioned runners, and privileged-host compromise risk.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.