GitLab Runners, Executors, Tags, Registration, Autoscaling, Isolation, and Security: Configuration, Design Choices, and Tradeoffs
Choose runner ownership, scope, executor, protection, persistence, network, cache, and autoscaling strategies using explicit blast-radius, reliability, maintainability, and cost tradeoffs.
Learning objectives
- Choose between GitLab-hosted and self-managed runners based on control, isolation, locality, maintenance, and compute constraints.
- Compare shell, Docker, Kubernetes, Docker Autoscaler, and Instance executor designs without treating executor names as security guarantees.
- Select instance/group/project runner scope and tag/protection policy according to blast radius and ownership.
- Decide when persistent workers are acceptable and when ephemeral/autoscaled workers materially reduce risk.
- Document a runner architecture decision that includes network, cache, upgrade, observability, and cost assumptions.
1. Design from workload trust and required capability
Start with the workload, not the executor. Ask: who can modify the code that reaches this runner? Which secrets can jobs receive? Does the workload need internal network access, privileged container operations, GPUs, special hardware, licensed tooling, or deployment credentials? How much cross-job persistence is acceptable?
The runner architecture should make the expected workload easy and dangerous workloads structurally difficult.
2. GitLab-hosted versus self-managed
| Dimension | GitLab-hosted runner | Self-managed runner |
|---|---|---|
| Operations | GitLab manages fleet provisioning and base runner infrastructure. | You install, patch, upgrade, monitor, scale, and decommission Runner/hosts. |
| Isolation model | Hosted jobs normally receive fresh provider-managed infrastructure. | Depends entirely on executor, host reuse, mounts, network, and your cleanup. |
| Network locality | Internet/service-defined connectivity; private access requires supported networking patterns. | Can be placed near internal build/deploy targets—but that expands blast radius. |
| Special hardware/software | Limited to offered runner classes/images. | Full control over hardware, images, drivers, licensed tools, and architecture. |
| Cost model | Compute quota/credits/cost factors can apply; verify current service terms. | Infrastructure + operations + idle capacity + cloud/network/storage cost. |
| Best fit | General CI where managed ephemeral infrastructure is sufficient. | Private-network, hardware-specific, compliance/locality, or specialized workloads. |
3. Shell simplicity versus isolation risk
The Shell executor is operationally simple because commands run directly on the runner host. That same property creates the core security problem: job code shares the host account, filesystem, installed tools, network, and potentially leftovers from other jobs. Current GitLab documentation places Shell in maintenance mode and explicitly recommends using it only for trusted builds.
A non-privileged Docker executor creates a cleaner process/filesystem boundary and makes dependencies reproducible through images. But Docker is not a magic security boundary: privileged mode, host Docker socket binding, host PID mode, broad capabilities, and sensitive host-volume mounts can effectively hand the job the host.
4. Executor selection matrix
| Need | Prefer | Why / caution |
|---|---|---|
| Ordinary reproducible Linux build | Docker executor or hosted runner. | Image-defined dependencies, clean job environment; keep non-privileged where possible. |
| Kubernetes-native build isolation/integration | Kubernetes executor. | Per-job pods; still govern cluster RBAC, service accounts, node trust, namespace and network policy. |
| Autoscaled container workload | Docker Autoscaler. | Current fleeting/taskscaler path; can make each instance ephemeral. |
| Autoscaled non-containerized host workload | Instance executor. | Job runs directly on autoscaled host; useful for VM-native tooling. |
| Legacy trusted host tooling | Shell executor only with explicit trust decision. | Persistent host and direct execution create high cross-job/host risk. |
| Unsupported custom environment | Custom only after accepting maintenance-mode/support tradeoff. | You own driver correctness, provisioning, cleanup, and security. |
5. Scope is blast-radius architecture
Instance scope maximizes reuse and therefore requires the strongest trust policy. Group scope can align runner ownership with a platform team and a related project family. Project scope is easier to reason about for sensitive deployment or specialized workloads. More sharing can improve utilization but also increases the number of repositories/users whose code can reach the same execution fleet.
7. Persistent versus ephemeral workers
| Pattern | Advantages | Risks/costs | Good use |
|---|---|---|---|
| Persistent runner host | Fast warm state; simple; predictable local tools. | Cross-job persistence, patch drift, cache/workspace residue, larger compromise window. | Trusted internal workloads with strong cleanup and hardening. |
| Reusable autoscaled instance | Amortizes startup cost over several jobs. | A compromised job can affect later jobs on same worker. | Moderate-trust jobs when bounded reuse is justified. |
| One job per ephemeral instance | Strong reduction in cross-job persistence. | Higher provisioning latency/cost; image bootstrap must be reliable. | Untrusted or high-sensitivity workloads, especially privileged builds. |
| Provider-managed hosted ephemeral worker | Low operational burden and clean lifecycle. | Less infrastructure control; quota/service constraints. | General GitLab.com CI. |
The Docker Autoscaler documentation explicitly supports
configurations such as capacity_per_instance=1 and
max_use_count=1 so each job gets a disposable instance.
That is a security design choice, not merely a scaling knob.
8. Autoscaling policy: latency, utilization, and failure containment
Autoscaling introduces a control loop: Runner asks for work, taskscaler manages desired capacity, a fleeting provider plugin creates/deletes cloud instances, and the executor runs jobs. The key knobs determine maximum instances, capacity per instance, reuse count, idle capacity, and idle time.
- More idle instances reduce queue latency but increase idle cost.
- More jobs per instance increase utilization but increase cross-job persistence/blast radius.
- Fast scale-out depends on cloud quotas, image availability, network, DNS, and provider API health.
- Dedicated autoscaling resources should not be controlled by multiple independent runner-manager configurations; competing controllers can produce unpredictable behavior and cost.
9. Network and cache are part of the runner threat model
Runner placement often exists because jobs need private package servers, test databases, registries, or deployment targets. That same connectivity becomes reachable by arbitrary job scripts unless filtered. Put runner hosts in a dedicated segment, restrict egress/metadata endpoints, and expose only required services.
Cache is optimization state, not trusted release evidence. Scope cache keys carefully, avoid putting secrets in caches, and assume an untrusted job may try to poison any cache it can write. Release artifacts/packages should have their own integrity and permission controls.
10. Privileged workloads deserve dedicated ephemeral capacity
Do not mix ordinary fork/MR jobs and privileged Docker-in-Docker or deployment jobs on the same persistent host simply because both need Docker.
11. Upgrades, ownership, and decommissioning are design requirements
Every self-managed runner needs a named owner, supported Runner
version policy, patch window, health check,
token-rotation/revocation procedure, host inventory, and
decommission checklist. GitLab distinguishes online,
offline, stale, and
never_contacted states; these are operational signals,
not merely UI decoration.
Some centralized runner fleet statistics are tier-gated, but every
team can still inventory runner IDs, versions, status, jobs, host
logs, and local gitlab-runner --version without buying
a tier.
12. Worked scenario: payment service CI
| Workload | Chosen runner design | Reasoning |
|---|---|---|
| Unit tests from developer branches/MRs | GitLab-hosted or non-privileged ephemeral Docker runner. | Untrusted code; no internal secrets; reproducible image; minimize persistence. |
| Integration tests against a synthetic isolated database | Dedicated group/project Docker runner with narrow network segment. | Needs private test service but should not reach production. |
| Container build requiring privileged mode | Dedicated ephemeral autoscaled runner restricted to reviewed/protected refs. | Privileged workload gets its own blast radius; one job per instance preferred. |
| Production deployment | Project-scoped protected deployment runner or identity-based deployment path. | Narrow source trust, narrow network, short-lived credentials; do not share with ordinary CI. |
The architecture deliberately uses more than one runner class because the workloads have different trust and capability requirements. “One powerful runner for everything” is cheaper to draw but harder to secure.
13. Runner architecture decision record
Record at least: workload owners; allowed projects/refs; required tags; protected/run-untagged policy; executor and privilege mode; host/worker lifetime; network destinations; cache backend/scope; credential sources; autoscaling limits; expected concurrency; upgrade owner; logging/metrics; token rotation; incident isolation; and decommission steps.
Knowledge check
Why is a runner tag not sufficient to protect a production deployment runner?
Tags route jobs by capability. Security requires scope, protected-ref policy, reviewed configuration, credentials, executor isolation, and network controls.
When is a Shell executor defensible?
For explicitly trusted builds on a hardened host where direct host execution is an accepted design tradeoff—not as the default for untrusted shared workloads.
What does max_use_count=1 accomplish in an autoscaled design?
It prevents worker reuse after a job, reducing cross-job persistence and compromise propagation.
Why can broader runner scope increase blast radius?
More projects and repository writers can potentially send code to the same execution infrastructure.
What is the principal security concern with Docker privileged mode?
It grants extremely broad container capabilities and can enable host compromise/container breakout, effectively weakening the isolation boundary.
Why should cache be treated differently from release evidence?
Cache is mutable performance state that can be overwritten/poisoned; release evidence should use explicit immutable identity/integrity and governed permissions.
Summary
Runner architecture is a workload-trust decision. Choose scope, tags, protection, executor, persistence, network, cache, and autoscaling as one system. Prefer managed or non-privileged isolated capacity for ordinary work, narrow and protect sensitive runners, and make privileged execution ephemeral and dedicated when it cannot be avoided.
Official references
- GitLab Docs — Get started with GitLab Runner
- GitLab Docs — Manage runners and runner scope
- GitLab Docs — Configure runners, tags, and protected runners
- GitLab Docs — New runner creation/registration workflow
- GitLab Docs — Register runners
- GitLab Docs — Runner commands and unregister behavior
- GitLab Docs — Runner executors
- GitLab Docs — Docker executor
- GitLab Docs — Shell executor
- GitLab Docs — Kubernetes executor
- GitLab Docs — Docker Autoscaler executor
- GitLab Docs — Instance executor
- GitLab Docs — Self-managed runner security
- GitLab Docs — Runner fleet scaling
- GitLab Docs — Instance-group autoscaler
- GitLab Docs — GitLab-hosted runners
- GitLab Docs — Hosted runners on Linux for GitLab.com
- GitLab Docs — Hosted runners for GitLab Dedicated
- GitLab Docs — Runners API
- GitLab Docs — Token overview / runner authentication tokens
- GitLab Docs — CI/CD YAML tags
- GitLab Docs — Protected branches and CI/CD
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.