Chapter 30Lesson 01~190 minutes

Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning: Concepts, Architecture, and Mental Model

Treat runner fleets as capacity systems: connect queue demand, runner eligibility, manager concurrency, executor lifecycle, worker reuse, metrics, isolation, cost bounds, and upgrade/migration evidence.

Runner fleetCapacityAutoscalingFleetingMetrics

Learning objectives

  • Explain why a runner fleet is a queue-and-capacity system rather than “a group of servers.”
  • Trace job arrival → eligibility/tags → runner manager → worker provisioning → executor → job → teardown/reuse → metrics feedback.
  • Separate GitLab pipeline state, runner-manager state, worker state, provider/Kubernetes state, evidence and cost governance.
  • Explain the current Instance/Docker Autoscaler/Fleeting model and the Docker Machine migration boundary.
  • Use read-only evidence before changing concurrency, autoscaling, executor, or fleet topology.

1. The practical problem: green jobs can hide an unhealthy fleet

Chapter 29 treated infrastructure mutation as a reviewed state transition. Chapter 30 turns to the infrastructure that executes the pipeline itself. A fleet can produce green jobs while developers wait ten minutes in queue, while idle VMs burn money, while an old Runner version silently misses fixes, or while reused workers retain state from previous untrusted jobs. Conversely, a low-utilization fleet is not automatically wasteful if warm capacity is what keeps critical feedback inside an agreed latency objective.

The operating question is therefore not “how many runners do we have?” It is: which jobs arrive, which runners are eligible, how quickly can managers acquire work, how many workers can be provisioned, how long do they live, what isolation boundary do they provide, and what observable evidence proves cost and queue behavior?

Fleet invariant: runner eligibility, manager polling capacity, executor capacity, worker lifecycle and provider/Kubernetes capacity are separate controls. Increasing only one can leave the true bottleneck unchanged.

2. Mental model: demand becomes a bounded worker lifecycle

A GitLab pipeline creates jobs with requirements such as tags, protection/trust context, platform and image. Eligible runner managers poll for work. The manager is not necessarily the worker: with Docker Autoscaler or Instance executors, it asks a Fleeting plugin to manage instances; with Kubernetes executor, it asks the Kubernetes API for a new job pod. The worker executes the job, sends logs/artifacts back, and is then reused or destroyed according to policy. Queue duration, provisioning latency, running jobs, errors, worker state and cost feed the next sizing decision.

Causal runner-fleet lifecycle
flowchart TD A[Pipeline job + tags/trust] --> B[GitLab queue] B --> C[Eligible runner manager] C --> D{Capacity available?} D -->|VM fleet| E[Fleeting / provider instance] D -->|Kubernetes| F[Job pod / node capacity] E --> G[Executor prepares job] F --> G G --> H[Checkout + script + artifacts] H --> I{Reuse policy} I -->|max_use_count reached| J[Destroy worker] I -->|reuse allowed| K[Return worker idle] J --> L[Metrics + logs + cost] K --> L L --> C

The loop is the important part. Autoscaling is control feedback, not a one-time configuration choice.

3. State domains to keep separate

State layer Evidence to capture Why it matters
Pipeline demand CI_PIPELINE_SOURCE, ref, CI_COMMIT_SHA, pipeline/job IDs, job tags Demand must be attributable to an exact pipeline and eligibility requirement.
Compiled configuration Merged YAML, rules result, job graph, tags, image/services Proves which jobs were created and what runner characteristics they requested.
Runner registration/eligibility Runner ID, scope, protection, tags, paused/online status Determines which manager can request and accept which jobs.
Runner manager Runner version, concurrent, per-runner limit, request_concurrency, logs The manager is a capacity controller and poller, distinct from the worker that executes a job.
Executor/worker Executor type, worker/pod/instance identity, image, lifetime, reuse count Defines the isolation and lifecycle boundary for untrusted build code.
Autoscaler Fleeting plugin/version, capacity_per_instance, max_use_count, max_instances, idle policy These are the hard and soft controls that turn demand into infrastructure.
Queue/performance Queue duration, jobs running, machine creation latency, request-concurrency pressure Capacity must be tuned from observed demand and cold-start behavior, not guesswork.
Cost/governance Maximum workers, instance size/rate assumptions, budgets, ownership, change record A technically successful scale-out can still be an operational failure if cost is unbounded.
Logs/forensics Manager logs, job trace, provider/autoscaler events, worker destruction evidence Ephemeral compute disappears; evidence must leave the worker before teardown.
Upgrade/migration Old/new Runner versions, canary pool, drain state, rollback image/config Fleet changes need a reversible rollout independent of application pipelines.

4. Read-only inspection before any fleet change

Start with evidence that does not register, delete, scale or restart anything. On a disposable/self-managed runner manager, typical read-only checks are:

gitlab-runner --version
gitlab-runner list
# If metrics listening is already enabled, read it without changing configuration:
curl -fsS http://127.0.0.1:9252/metrics   | grep -E 'gitlab_runner_(version_info|concurrent|limit|request_concurrency|jobs|job_queue_duration_seconds|errors_total)'   | head -80

In GitLab, record the runner ID/scope/tags/protection and the waiting job’s CI_PIPELINE_SOURCE, CI_COMMIT_SHA, pipeline/job IDs and requested tags. The metrics endpoint is optional and must already be enabled; do not expose it publicly merely for a lab.

5. Current executor map

Executor Worker model Autoscaling boundary Typical reason to choose it
Docker Containers on a pre-existing Docker host Scale the runner hosts by an external mechanism Simple containerized jobs when host capacity is already managed.
Docker Autoscaler Docker jobs on dynamically managed VM instances Taskscaler + Fleeting + provider plugin Docker semantics plus elastic VM isolation/capacity.
Instance Jobs run directly on dynamically managed instances Taskscaler + Fleeting + provider plugin Full host/device/OS access or non-container workloads.
Kubernetes A new pod per CI job Kubernetes schedules pods; node/autoscaling capacity is a cluster concern Cloud-native scheduling, pod-level isolation and resource controls.
Docker Machine (legacy) Docker jobs on VMs created by Docker Machine Legacy Docker Machine driver model Migration source only; deprecated since 17.5 and scheduled for removal in 20.0.

6. Docker Autoscaler and Instance: Fleeting is the provider boundary

Current GitLab Runner uses Taskscaler for autoscaling logic and Fleeting as a plugin abstraction over AWS Auto Scaling Groups, Google instance groups, Azure VM Scale Sets, and supported community providers. Docker Autoscaler wraps Docker executor behavior; Instance executor gives jobs host-level access. Both use [runners.autoscaler].

concurrent = 12

[[runners]]
  name = "example-ephemeral-linux"
  executor = "docker-autoscaler"
  limit = 12
  request_concurrency = 4

  [runners.docker]
    image = "alpine:3.22"

  [runners.autoscaler]
    plugin = "aws"
    capacity_per_instance = 1
    max_use_count = 1
    max_instances = 12

    [[runners.autoscaler.policy]]
      idle_count = 2
      idle_time = "10m0s"

This is architecture syntax, not a runnable lab: it deliberately omits provider credentials and provider-specific resource identifiers. The hard cost/isolation controls are visible: at most 12 instances, one job per instance, one use per instance. concurrent is global to the runner process; limit is per runner configuration.

7. One autoscaling resource has one controller

GitLab explicitly requires every Docker Autoscaler configuration to own a dedicated cloud autoscaling resource. Do not attach two runner managers—or two [[runners]] entries—to the same provider group. Each manager maintains instance bookkeeping; competing controllers can issue conflicting scale commands, causing job failures and cost surprises.

8. Kubernetes executor: one pod per job is not one node per job

The Kubernetes executor creates a pod for each GitLab CI job. That pod is a worker execution unit, while node capacity belongs to the cluster. If pods queue in Pending, the bottleneck may be CPU/memory requests, quotas, taints, topology, image pulls, or node-autoscaler latency—not Runner job polling. Current Runner documentation also includes pause-pod prewarming for Kubernetes executor; treat it as a latency/cost tradeoff and verify the exact Runner version before using it.

9. Four different capacity ceilings

Control What it limits Failure symptom when too low
concurrent Total simultaneous jobs across all [[runners]] entries in one runner process Managers have eligible work but process-wide slots are exhausted.
limit Jobs for one runner configuration A specific pool stays capped while other pools have spare process capacity.
request_concurrency Concurrent requests for new jobs Long-poll request bottleneck; worker capacity exists but jobs are acquired slowly.
Provider/cluster bounds Instances, pods, nodes, quota, max_instances Runner wants capacity but infrastructure cannot provide it.

Current Runner guidance notes that high-volume runners may need request_concurrency above its default of 1. The adaptive request-concurrency feature flag can tune effective concurrency up to the configured ceiling; it cannot exceed a ceiling of 1 unless you raise it.

10. Capacity planning starts with queue duration, not CPU alone

Runner exposes Prometheus metrics including current jobs, job duration, queue duration, manager errors, configured concurrency/limit/request concurrency, autoscaling machine state and creation duration. A healthy fleet tracks at least queue latency, provisioning latency, job duration, utilization, errors/retries and infrastructure cost per successful job.

Do not optimize only for utilization. A fleet at 95% worker utilization can be developer-hostile if queue SLOs are violated; a fleet at 30% can be intentionally prewarmed for low-latency feedback.

11. Ephemeral is a lifecycle property, not a security slogan

For Fleeting-backed VM workers, capacity_per_instance = 1 and max_use_count = 1 provide a strong single-job instance lifecycle when combined with a secure base image and teardown verification. Reuse can reduce cold-start cost but increases cross-job residue risk. Kubernetes creates a fresh job pod, yet shared nodes, caches, volumes, service accounts and privileged settings can still cross trust boundaries. Chapter 31 threat-models those boundaries in depth.

12. Runner upgrades are fleet changes

Pin and record Runner versions. Current stable patch on this chapter’s verification date is 19.3.2. Use a canary pool, pause/drain it before manager upgrade where practical, verify representative jobs and metrics, then expand. For the Kubernetes Helm chart, GitLab recommends pausing the runner and waiting for jobs to complete before upgrading. Keep the prior package/image/chart and configuration as an explicit rollback target.

13. Docker Machine migration is now an operational deadline

Docker Machine executor is deprecated and scheduled for removal in GitLab 20.0 (May 2027). A migration should inventory MachineOptions, instance image assumptions, caches, network/security rules, idle policy, job concurrency, special devices and provider quotas; then map those requirements to Docker Autoscaler or Instance plus a Fleeting plugin. Do not perform an in-place “syntax translation” without a capacity and rollback test.

14. Misconceptions to remove now

Claim Why it fails Better operating model
“Autoscaling means unlimited capacity.” Provider quotas, max_instances, manager concurrency and budgets still bound capacity. Define hard worker and cost ceilings before testing burst behavior.
“Two managers on one ASG improve HA.” For Docker Autoscaler, shared control of one autoscaling resource is unsupported and conflicts with bookkeeping. Use separate autoscaling resources behind the same logical job tags.
“One pod per job means perfect isolation.” Pods can share nodes, caches, credentials, network and cluster policy. Treat pod lifecycle and node/trust isolation as different layers.
“If jobs are queued, add VMs.” Polling, tags, protection, concurrent, limit or request_concurrency may be the bottleneck. Diagnose the causal layer from metrics and eligibility first.
“Ephemeral workers need no logs.” Teardown destroys local evidence first. Ship manager/provider/job evidence before worker destruction.

Knowledge check

A fleet has idle VMs but GitLab jobs remain pending. What should you inspect before increasing max_instances?

Why is max_use_count = 1 meaningful?

Can two Docker Autoscaler managers safely control the same AWS Auto Scaling Group?

What is the difference between Kubernetes job-pod scaling and node scaling?

Why is queue duration a better first fleet SLO than average CPU utilization?

Next lesson

Continue to the guided workflow

Lesson 2 builds a completely local queue/capacity simulator, measures a burst, compares static and elastic designs, and turns the observations into a bounded Docker Machine migration plan.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Current stable GitLab Runner patch verified for this chapter is 19.3.2 (tagged 2026-09-10). The Docker Machine executor was deprecated in GitLab 17.5 and is scheduled for removal as a supported feature in GitLab 20.0 (May 2027); GitLab directs users toward the Instance or Docker Autoscaler executors. Docker Autoscaler is GA and, together with the Instance executor, uses Taskscaler/Fleeting. A Docker Autoscaler configuration must have its own dedicated provider autoscaling resource and must not share that resource with another manager or another [[runners]] entry. The mandatory exercises in this chapter are local simulations and require no cloud account, runner registration token, privileged executor, or managed Kubernetes cluster. Runner 19.3 introduced gitlab-runner config lint; use it during real fleet changes to validate manager configuration syntax before rollout, while still testing behavior on a canary.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.