Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning: Configuration, Design Choices, and Tradeoffs
Choose warm idle versus cold start, ephemeral versus reused workers, VM versus Kubernetes autoscaling, and centralized versus specialized pools using explicit trust, latency, cost, observability, and rollback criteria.
Learning objectives
- Choose warm idle versus cold-start capacity using queue SLO and cost evidence.
- Choose single-use ephemeral workers versus reuse based on trust and startup economics.
- Choose Kubernetes versus VM autoscaling based on execution and isolation requirements.
- Design centralized and specialized pools without collapsing trust boundaries.
- Define canary, upgrade and rollback controls for runner-manager changes.
1. Fleet design is a multi-objective control problem
There is no universally “best” executor or autoscaling policy. A pool that minimizes cost can violate feedback SLOs. A pool that reuses workers aggressively can weaken isolation. A fleet split into many specialized pools can improve least privilege but reduce utilization and increase operational overhead. The design must state which objective is primary for each trust/workload class.
Evidence still begins with CI_PIPELINE_SOURCE,
CI_COMMIT_SHA, pipeline/job ID and requested tags, then
follows the job through manager version/config, executor, worker
identity, queue/provisioning timing and teardown.
2. Warm idle capacity versus cold-start cost
Warm workers spend money while idle but reduce queue latency during bursts. Cold-only fleets save idle cost but turn VM or node provisioning into user-visible delay. Measure the distribution of queue duration and machine/pod creation time; then define an idle policy by time window and workload criticality rather than copying a fixed number from another organization.
| Policy | Latency | Cost | Best fit |
|---|---|---|---|
| Zero warm workers | Highest cold-start exposure | Lowest idle compute | Infrequent, non-interactive batch work. |
| Small fixed warm floor | Absorbs common bursts | Predictable idle cost | Developer feedback pools with moderate variance. |
| Time-window warm capacity | Low latency during working hours | Cost follows expected demand | Regional teams with clear demand windows. |
| Large warm pool | Lowest queue latency | Highest idle cost | Only when strict SLO/business value justifies it. |
3. Single-use ephemeral versus worker reuse
For Fleeting-backed workers, max_use_count = 1 and
capacity_per_instance = 1 make each instance a
single-job lifecycle. That reduces cross-job residue and simplifies
“trust reset” reasoning, at the cost of more provisioning and
image/bootstrap work. Reuse raises utilization and lowers startup
overhead, but increases the amount of cleanup you must trust.
4. Kubernetes versus VM autoscaling
| Question | Kubernetes executor | Docker Autoscaler / Instance |
|---|---|---|
| Execution unit | Fresh pod per job | VM instance with configurable jobs-per-instance/use count. |
| Underlying capacity | Kubernetes nodes/cluster autoscaler | Provider autoscaling group through Fleeting. |
| Isolation emphasis | Pod/security context plus node/cluster isolation | VM boundary; can be one job per VM. |
| Workload fit | Container-native jobs with pod resource controls | Docker jobs needing VM isolation, or host/device jobs with Instance executor. |
| Cold path | Pod scheduling + possible node provisioning + image pull | Instance provisioning + connector/bootstrap + Docker/image pull if applicable. |
| Operational owner | Cluster/RBAC/node lifecycle skills | Cloud VM image/network/IAM/Fleeting skills. |
Do not compare them only by raw startup time. Include skill ownership, failure domains, network policy, image provenance, cache topology and incident containment.
5. Centralized fleet versus specialized pools
A centralized pool improves utilization and operational consistency, but only when jobs have compatible trust and platform requirements. Specialized pools can isolate protected release jobs, Windows/macOS builds, GPU/device workloads, untrusted forks, high-memory builds or regulated workloads. Tags and runner protection are routing controls—not proof that the underlying worker boundary is secure.
7. Hard bounds versus adaptive targets
Hard controls such as max_instances, provider quota,
concurrent and limit cap damage/cost. Soft
controls such as idle policies target responsiveness. Treat “0 =
unlimited” defaults carefully: current
max_instances default is 0/unlimited, which is
unsuitable as an unexamined production cost policy.
8. Job polling is capacity too
A fleet can have idle workers yet underperform because managers do
not request jobs fast enough.
request_concurrency defaults to 1. Current Runner
guidance recommends increasing it for high-volume runners when
long-poll bottlenecks are observed; adaptive request concurrency can
tune within the configured ceiling. Record request metrics before
changing the value.
9. High availability without shared-controller conflict
GitLab recommends at least two runner managers for fault tolerance in an autoscaled service model. That does not mean two Docker Autoscaler managers may control the same autoscaling group. Give managers the same logical job tags if needed, but separate provider autoscaling resources and bookkeeping ownership.
10. Upgrade strategy: canary, drain, observe, expand
- Inventory manager and helper versions; record current config hashes.
- Create or select a canary pool with a small percentage of representative jobs.
-
Validate configuration; Runner 19.3 adds a
config lintcommand. - Pause/drain the canary manager where required and upgrade to a pinned version.
- Run synthetic plus representative workload; compare queue, duration, errors, artifacts/cache and worker teardown.
- Expand incrementally. Keep the prior package/image/chart and config as rollback material.
11. Docker Machine migration mapping
| Legacy concept | Migration question | Likely target |
|---|---|---|
| Machine driver/options | Which provider group/plugin and image reproduce required compute/network? | Fleeting plugin + dedicated provider autoscaling resource. |
IdleCount/IdleTime |
What warm-capacity SLO/cost policy do measurements justify? |
Autoscaler policy
idle_count/idle_time.
|
MaxBuilds |
How many jobs may one worker run before destruction? | max_use_count. |
| Machine count/concurrency | What is the hard cost ceiling and jobs-per-worker? |
max_instances,
capacity_per_instance,
concurrent/limit.
|
| Docker Machine host behavior | Do jobs need Docker semantics or full host access? | Docker Autoscaler or Instance executor. |
12. Decision table: choose the fleet deliberately
| Scenario | Choice | Prerequisites/trust | Observable proof |
|---|---|---|---|
| Public MR tests with arbitrary code | Single-use ephemeral Linux pool | No production secrets; restricted network; dedicated untrusted tags | worker use count 1, destruction evidence, no protected credentials. |
| Fast internal unit tests | Small warm Docker Autoscaler pool | Trusted project class; bounded VM group; cache policy | queue SLO, warm idle count, max instance/cost evidence. |
| Cloud-native integration tests | Kubernetes executor | Dedicated namespace/RBAC, pod resources, cluster autoscaling | pod scheduling + node scale events + job trace. |
| Host/device compiler job | Instance executor specialized pool | Dedicated images/devices; narrow job routing | instance identity, device exposure, teardown/reuse policy. |
| Legacy Docker Machine pool | Canary migration to Docker Autoscaler | Dedicated new provider group; Fleeting plugin; rollback pool | old/new queue/cost/error comparison and drain evidence. |
13. A simple cost envelope
For a VM pool, define a conservative hourly upper bound before rollout:
\[ C_{max/hour} = N_{max} \times C_{instance/hour} + C_{storage/network/cache} \]
This is intentionally simple; reserved/spot pricing, egress,
cache/object storage and node sharing complicate the real number.
The important control is that
N_ is explicit
and observable.
14. Tier and offering boundary
Runner executors, autoscaling, monitoring and the Kubernetes executor documented here are available across Free, Premium and Ultimate offerings. Cloud VM groups, Kubernetes clusters and provider credentials themselves are external infrastructure with their own cost and authorization. The core chapter never requires them.
Knowledge check
Which is safer for untrusted fork jobs:
max_use_count=1 or a reused worker with strong
cleanup scripts?
Single-use workers provide a simpler lifecycle boundary because the worker is destroyed instead of trusting cleanup to remove every residue. Other shared layers still need controls.
Why can specialized pools increase cost even if each pool autos-scales?
Partitioning demand reduces sharing of idle capacity, can create separate warm floors and increases management overhead. The isolation benefit may still justify it.
What is wrong with max_instances = 0 as an
unreviewed cost policy?
In current autoscaler configuration, zero means unlimited. A production fleet should normally have an explicit bounded ceiling aligned to quotas and budget.
Can the same runner tags be used by two fault-tolerant managers?
Yes, if they are intended to serve the same job class. For Docker Autoscaler, each manager/configuration still needs its own dedicated provider autoscaling resource.
What evidence should justify raising warm idle count?
Measured queue/provisioning latency and a clear feedback SLO, balanced against observed idle cost. Do not raise it from intuition alone.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12.
Current stable GitLab Runner patch verified for this chapter is
19.3.2 (tagged 2026-09-10). The Docker Machine executor
was deprecated in GitLab 17.5 and is scheduled for removal as a
supported feature in GitLab 20.0 (May 2027); GitLab directs users
toward the Instance or Docker Autoscaler executors. Docker
Autoscaler is GA and, together with the Instance executor, uses
Taskscaler/Fleeting. A Docker Autoscaler configuration must have its
own dedicated provider autoscaling resource and must not share that
resource with another manager or another
[[runners]] entry. The mandatory exercises in this
chapter are local simulations and require no cloud account, runner
registration token, privileged executor, or managed Kubernetes
cluster. GitLab Runner 19.3.2 is the chapter baseline. Pin
production upgrades to an explicitly tested patch and verify release
notes rather than treating “latest” as a fleet policy.
- Executors — official reference.
- GitLab Runner autoscaling — official reference.
- Docker Autoscaler executor — official reference.
- Instance executor — official reference.
- Fleeting — official reference.
- Plan and operate a runner fleet — official reference.
- Advanced Runner configuration — official reference.
- Monitor GitLab Runner usage — official reference.
- Kubernetes executor — official reference.
- GitLab Runner Helm chart — official reference.
- Docker Machine executor — official reference.
- GitLab deprecations and removals — official reference.
- Runner troubleshooting — official reference.
- GitLab Runner changelog — official reference.
- GitLab Runner tags — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.