Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning: Concepts, Architecture, and Mental Model
Treat runner fleets as capacity systems: connect queue demand, runner eligibility, manager concurrency, executor lifecycle, worker reuse, metrics, isolation, cost bounds, and upgrade/migration evidence.
Learning objectives
- Explain why a runner fleet is a queue-and-capacity system rather than “a group of servers.”
- Trace job arrival → eligibility/tags → runner manager → worker provisioning → executor → job → teardown/reuse → metrics feedback.
- Separate GitLab pipeline state, runner-manager state, worker state, provider/Kubernetes state, evidence and cost governance.
- Explain the current Instance/Docker Autoscaler/Fleeting model and the Docker Machine migration boundary.
- Use read-only evidence before changing concurrency, autoscaling, executor, or fleet topology.
1. The practical problem: green jobs can hide an unhealthy fleet
Chapter 29 treated infrastructure mutation as a reviewed state transition. Chapter 30 turns to the infrastructure that executes the pipeline itself. A fleet can produce green jobs while developers wait ten minutes in queue, while idle VMs burn money, while an old Runner version silently misses fixes, or while reused workers retain state from previous untrusted jobs. Conversely, a low-utilization fleet is not automatically wasteful if warm capacity is what keeps critical feedback inside an agreed latency objective.
The operating question is therefore not “how many runners do we have?” It is: which jobs arrive, which runners are eligible, how quickly can managers acquire work, how many workers can be provisioned, how long do they live, what isolation boundary do they provide, and what observable evidence proves cost and queue behavior?
2. Mental model: demand becomes a bounded worker lifecycle
A GitLab pipeline creates jobs with requirements such as tags, protection/trust context, platform and image. Eligible runner managers poll for work. The manager is not necessarily the worker: with Docker Autoscaler or Instance executors, it asks a Fleeting plugin to manage instances; with Kubernetes executor, it asks the Kubernetes API for a new job pod. The worker executes the job, sends logs/artifacts back, and is then reused or destroyed according to policy. Queue duration, provisioning latency, running jobs, errors, worker state and cost feed the next sizing decision.
The loop is the important part. Autoscaling is control feedback, not a one-time configuration choice.
3. State domains to keep separate
| State layer | Evidence to capture | Why it matters |
|---|---|---|
| Pipeline demand |
CI_PIPELINE_SOURCE, ref,
CI_COMMIT_SHA, pipeline/job IDs, job tags
|
Demand must be attributable to an exact pipeline and eligibility requirement. |
| Compiled configuration |
Merged YAML, rules result, job graph, tags,
image/services
|
Proves which jobs were created and what runner characteristics they requested. |
| Runner registration/eligibility | Runner ID, scope, protection, tags, paused/online status | Determines which manager can request and accept which jobs. |
| Runner manager |
Runner version, concurrent, per-runner
limit, request_concurrency, logs
|
The manager is a capacity controller and poller, distinct from the worker that executes a job. |
| Executor/worker | Executor type, worker/pod/instance identity, image, lifetime, reuse count | Defines the isolation and lifecycle boundary for untrusted build code. |
| Autoscaler |
Fleeting plugin/version, capacity_per_instance,
max_use_count, max_instances, idle
policy
|
These are the hard and soft controls that turn demand into infrastructure. |
| Queue/performance | Queue duration, jobs running, machine creation latency, request-concurrency pressure | Capacity must be tuned from observed demand and cold-start behavior, not guesswork. |
| Cost/governance | Maximum workers, instance size/rate assumptions, budgets, ownership, change record | A technically successful scale-out can still be an operational failure if cost is unbounded. |
| Logs/forensics | Manager logs, job trace, provider/autoscaler events, worker destruction evidence | Ephemeral compute disappears; evidence must leave the worker before teardown. |
| Upgrade/migration | Old/new Runner versions, canary pool, drain state, rollback image/config | Fleet changes need a reversible rollout independent of application pipelines. |
4. Read-only inspection before any fleet change
Start with evidence that does not register, delete, scale or restart anything. On a disposable/self-managed runner manager, typical read-only checks are:
gitlab-runner --version
gitlab-runner list
# If metrics listening is already enabled, read it without changing configuration:
curl -fsS http://127.0.0.1:9252/metrics | grep -E 'gitlab_runner_(version_info|concurrent|limit|request_concurrency|jobs|job_queue_duration_seconds|errors_total)' | head -80
In GitLab, record the runner ID/scope/tags/protection and the
waiting job’s CI_PIPELINE_SOURCE,
CI_COMMIT_SHA, pipeline/job IDs and requested tags. The
metrics endpoint is optional and must already be enabled; do not
expose it publicly merely for a lab.
5. Current executor map
| Executor | Worker model | Autoscaling boundary | Typical reason to choose it |
|---|---|---|---|
| Docker | Containers on a pre-existing Docker host | Scale the runner hosts by an external mechanism | Simple containerized jobs when host capacity is already managed. |
| Docker Autoscaler | Docker jobs on dynamically managed VM instances | Taskscaler + Fleeting + provider plugin | Docker semantics plus elastic VM isolation/capacity. |
| Instance | Jobs run directly on dynamically managed instances | Taskscaler + Fleeting + provider plugin | Full host/device/OS access or non-container workloads. |
| Kubernetes | A new pod per CI job | Kubernetes schedules pods; node/autoscaling capacity is a cluster concern | Cloud-native scheduling, pod-level isolation and resource controls. |
| Docker Machine (legacy) | Docker jobs on VMs created by Docker Machine | Legacy Docker Machine driver model | Migration source only; deprecated since 17.5 and scheduled for removal in 20.0. |
6. Docker Autoscaler and Instance: Fleeting is the provider boundary
Current GitLab Runner uses Taskscaler for autoscaling logic and
Fleeting as a plugin abstraction over AWS Auto Scaling Groups,
Google instance groups, Azure VM Scale Sets, and supported community
providers. Docker Autoscaler wraps Docker executor behavior;
Instance executor gives jobs host-level access. Both use
[runners.autoscaler].
concurrent = 12
[[runners]]
name = "example-ephemeral-linux"
executor = "docker-autoscaler"
limit = 12
request_concurrency = 4
[runners.docker]
image = "alpine:3.22"
[runners.autoscaler]
plugin = "aws"
capacity_per_instance = 1
max_use_count = 1
max_instances = 12
[[runners.autoscaler.policy]]
idle_count = 2
idle_time = "10m0s"
This is architecture syntax, not a runnable lab: it deliberately
omits provider credentials and provider-specific resource
identifiers. The hard cost/isolation controls are visible: at most
12 instances, one job per instance, one use per instance.
concurrent is global to the runner process;
limit is per runner configuration.
7. One autoscaling resource has one controller
GitLab explicitly requires every Docker Autoscaler configuration to
own a dedicated cloud autoscaling resource. Do not attach two runner
managers—or two [[runners]] entries—to the same
provider group. Each manager maintains instance bookkeeping;
competing controllers can issue conflicting scale commands, causing
job failures and cost surprises.
8. Kubernetes executor: one pod per job is not one node per job
The Kubernetes executor creates a pod for each GitLab CI job. That pod is a worker execution unit, while node capacity belongs to the cluster. If pods queue in Pending, the bottleneck may be CPU/memory requests, quotas, taints, topology, image pulls, or node-autoscaler latency—not Runner job polling. Current Runner documentation also includes pause-pod prewarming for Kubernetes executor; treat it as a latency/cost tradeoff and verify the exact Runner version before using it.
9. Four different capacity ceilings
| Control | What it limits | Failure symptom when too low |
|---|---|---|
concurrent |
Total simultaneous jobs across all
[[runners]] entries in one runner process
|
Managers have eligible work but process-wide slots are exhausted. |
limit |
Jobs for one runner configuration | A specific pool stays capped while other pools have spare process capacity. |
request_concurrency |
Concurrent requests for new jobs | Long-poll request bottleneck; worker capacity exists but jobs are acquired slowly. |
| Provider/cluster bounds |
Instances, pods, nodes, quota, max_instances
|
Runner wants capacity but infrastructure cannot provide it. |
Current Runner guidance notes that high-volume runners may need
request_concurrency above its default of 1. The
adaptive request-concurrency feature flag can tune effective
concurrency up to the configured ceiling; it cannot exceed a ceiling
of 1 unless you raise it.
10. Capacity planning starts with queue duration, not CPU alone
Runner exposes Prometheus metrics including current jobs, job duration, queue duration, manager errors, configured concurrency/limit/request concurrency, autoscaling machine state and creation duration. A healthy fleet tracks at least queue latency, provisioning latency, job duration, utilization, errors/retries and infrastructure cost per successful job.
11. Ephemeral is a lifecycle property, not a security slogan
For Fleeting-backed VM workers,
capacity_per_instance = 1 and
max_use_count = 1 provide a strong single-job instance
lifecycle when combined with a secure base image and teardown
verification. Reuse can reduce cold-start cost but increases
cross-job residue risk. Kubernetes creates a fresh job pod, yet
shared nodes, caches, volumes, service accounts and privileged
settings can still cross trust boundaries. Chapter 31 threat-models
those boundaries in depth.
12. Runner upgrades are fleet changes
Pin and record Runner versions. Current stable patch on this
chapter’s verification date is 19.3.2. Use a canary
pool, pause/drain it before manager upgrade where practical, verify
representative jobs and metrics, then expand. For the Kubernetes
Helm chart, GitLab recommends pausing the runner and waiting for
jobs to complete before upgrading. Keep the prior
package/image/chart and configuration as an explicit rollback
target.
13. Docker Machine migration is now an operational deadline
Docker Machine executor is deprecated and scheduled for removal in
GitLab 20.0 (May 2027). A migration should inventory
MachineOptions, instance image assumptions, caches,
network/security rules, idle policy, job concurrency, special
devices and provider quotas; then map those requirements to Docker
Autoscaler or Instance plus a Fleeting plugin. Do not perform an
in-place “syntax translation” without a capacity and rollback test.
14. Misconceptions to remove now
| Claim | Why it fails | Better operating model |
|---|---|---|
| “Autoscaling means unlimited capacity.” |
Provider quotas, max_instances, manager
concurrency and budgets still bound capacity.
|
Define hard worker and cost ceilings before testing burst behavior. |
| “Two managers on one ASG improve HA.” | For Docker Autoscaler, shared control of one autoscaling resource is unsupported and conflicts with bookkeeping. | Use separate autoscaling resources behind the same logical job tags. |
| “One pod per job means perfect isolation.” | Pods can share nodes, caches, credentials, network and cluster policy. | Treat pod lifecycle and node/trust isolation as different layers. |
| “If jobs are queued, add VMs.” |
Polling, tags, protection, concurrent,
limit or request_concurrency may
be the bottleneck.
|
Diagnose the causal layer from metrics and eligibility first. |
| “Ephemeral workers need no logs.” | Teardown destroys local evidence first. | Ship manager/provider/job evidence before worker destruction. |
Knowledge check
A fleet has idle VMs but GitLab jobs remain pending. What
should you inspect before increasing
max_instances?
Runner eligibility/tags/protection, process
concurrent, per-runner limit,
request_concurrency/long-poll evidence, manager
health and job requirements. Idle infrastructure does not prove
the manager can accept the job.
Why is max_use_count = 1 meaningful?
It schedules an autoscaled instance for removal after one use, strengthening the worker lifecycle boundary and reducing cross-job residue risk.
Can two Docker Autoscaler managers safely control the same AWS Auto Scaling Group?
No. GitLab requires a dedicated autoscaling resource per autoscaler configuration because competing controllers can issue conflicting state changes.
What is the difference between Kubernetes job-pod scaling and node scaling?
Runner creates a pod per job; the Kubernetes scheduler and cluster/node autoscaler determine whether underlying node capacity exists. They are separate control loops.
Why is queue duration a better first fleet SLO than average CPU utilization?
Queue duration directly represents developer wait for capacity. CPU can look efficient while jobs wait, or look low by design when warm capacity is preserving latency.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12.
Current stable GitLab Runner patch verified for this chapter is
19.3.2 (tagged 2026-09-10). The Docker Machine executor
was deprecated in GitLab 17.5 and is scheduled for removal as a
supported feature in GitLab 20.0 (May 2027); GitLab directs users
toward the Instance or Docker Autoscaler executors. Docker
Autoscaler is GA and, together with the Instance executor, uses
Taskscaler/Fleeting. A Docker Autoscaler configuration must have its
own dedicated provider autoscaling resource and must not share that
resource with another manager or another
[[runners]] entry. The mandatory exercises in this
chapter are local simulations and require no cloud account, runner
registration token, privileged executor, or managed Kubernetes
cluster. Runner 19.3 introduced
gitlab-runner config lint; use it during real fleet
changes to validate manager configuration syntax before rollout,
while still testing behavior on a canary.
- Executors — official reference.
- GitLab Runner autoscaling — official reference.
- Docker Autoscaler executor — official reference.
- Instance executor — official reference.
- Fleeting — official reference.
- Plan and operate a runner fleet — official reference.
- Advanced Runner configuration — official reference.
- Monitor GitLab Runner usage — official reference.
- Kubernetes executor — official reference.
- GitLab Runner Helm chart — official reference.
- Docker Machine executor — official reference.
- GitLab deprecations and removals — official reference.
- Runner troubleshooting — official reference.
- GitLab Runner changelog — official reference.
- GitLab Runner tags — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.