Chapter 30Lesson 03~215 minutes

Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning: Configuration, Design Choices, and Tradeoffs

Choose warm idle versus cold start, ephemeral versus reused workers, VM versus Kubernetes autoscaling, and centralized versus specialized pools using explicit trust, latency, cost, observability, and rollback criteria.

DesignEphemeral workersKubernetesVM autoscalingCost

Learning objectives

  • Choose warm idle versus cold-start capacity using queue SLO and cost evidence.
  • Choose single-use ephemeral workers versus reuse based on trust and startup economics.
  • Choose Kubernetes versus VM autoscaling based on execution and isolation requirements.
  • Design centralized and specialized pools without collapsing trust boundaries.
  • Define canary, upgrade and rollback controls for runner-manager changes.

1. Fleet design is a multi-objective control problem

There is no universally “best” executor or autoscaling policy. A pool that minimizes cost can violate feedback SLOs. A pool that reuses workers aggressively can weaken isolation. A fleet split into many specialized pools can improve least privilege but reduce utilization and increase operational overhead. The design must state which objective is primary for each trust/workload class.

Evidence still begins with CI_PIPELINE_SOURCE, CI_COMMIT_SHA, pipeline/job ID and requested tags, then follows the job through manager version/config, executor, worker identity, queue/provisioning timing and teardown.

2. Warm idle capacity versus cold-start cost

Warm workers spend money while idle but reduce queue latency during bursts. Cold-only fleets save idle cost but turn VM or node provisioning into user-visible delay. Measure the distribution of queue duration and machine/pod creation time; then define an idle policy by time window and workload criticality rather than copying a fixed number from another organization.

Policy Latency Cost Best fit
Zero warm workers Highest cold-start exposure Lowest idle compute Infrequent, non-interactive batch work.
Small fixed warm floor Absorbs common bursts Predictable idle cost Developer feedback pools with moderate variance.
Time-window warm capacity Low latency during working hours Cost follows expected demand Regional teams with clear demand windows.
Large warm pool Lowest queue latency Highest idle cost Only when strict SLO/business value justifies it.

3. Single-use ephemeral versus worker reuse

For Fleeting-backed workers, max_use_count = 1 and capacity_per_instance = 1 make each instance a single-job lifecycle. That reduces cross-job residue and simplifies “trust reset” reasoning, at the cost of more provisioning and image/bootstrap work. Reuse raises utilization and lowers startup overhead, but increases the amount of cleanup you must trust.

Trust-first rule: do not justify reuse only with cost. State which caches, workspaces, credentials, Docker layers, devices and host services survive between jobs, and whether jobs from different trust classes can ever share them.

4. Kubernetes versus VM autoscaling

Question Kubernetes executor Docker Autoscaler / Instance
Execution unit Fresh pod per job VM instance with configurable jobs-per-instance/use count.
Underlying capacity Kubernetes nodes/cluster autoscaler Provider autoscaling group through Fleeting.
Isolation emphasis Pod/security context plus node/cluster isolation VM boundary; can be one job per VM.
Workload fit Container-native jobs with pod resource controls Docker jobs needing VM isolation, or host/device jobs with Instance executor.
Cold path Pod scheduling + possible node provisioning + image pull Instance provisioning + connector/bootstrap + Docker/image pull if applicable.
Operational owner Cluster/RBAC/node lifecycle skills Cloud VM image/network/IAM/Fleeting skills.

Do not compare them only by raw startup time. Include skill ownership, failure domains, network policy, image provenance, cache topology and incident containment.

5. Centralized fleet versus specialized pools

A centralized pool improves utilization and operational consistency, but only when jobs have compatible trust and platform requirements. Specialized pools can isolate protected release jobs, Windows/macOS builds, GPU/device workloads, untrusted forks, high-memory builds or regulated workloads. Tags and runner protection are routing controls—not proof that the underlying worker boundary is secure.

6. Tags are capacity partition keys

unit_test:
  tags: [linux, untrusted]
  script: ./ci/test.sh

release_sign:
  tags: [linux, release-protected]
  rules:
    - if: '$CI_COMMIT_TAG'
  script: ./ci/sign.sh

These jobs should not merely have different labels; their eligible runners, credentials, network access and worker lifecycles should match the trust classification. Chapter 31 expands this into a full threat model.

7. Hard bounds versus adaptive targets

Hard controls such as max_instances, provider quota, concurrent and limit cap damage/cost. Soft controls such as idle policies target responsiveness. Treat “0 = unlimited” defaults carefully: current max_instances default is 0/unlimited, which is unsuitable as an unexamined production cost policy.

8. Job polling is capacity too

A fleet can have idle workers yet underperform because managers do not request jobs fast enough. request_concurrency defaults to 1. Current Runner guidance recommends increasing it for high-volume runners when long-poll bottlenecks are observed; adaptive request concurrency can tune within the configured ceiling. Record request metrics before changing the value.

9. High availability without shared-controller conflict

GitLab recommends at least two runner managers for fault tolerance in an autoscaled service model. That does not mean two Docker Autoscaler managers may control the same autoscaling group. Give managers the same logical job tags if needed, but separate provider autoscaling resources and bookkeeping ownership.

10. Upgrade strategy: canary, drain, observe, expand

  1. Inventory manager and helper versions; record current config hashes.
  2. Create or select a canary pool with a small percentage of representative jobs.
  3. Validate configuration; Runner 19.3 adds a config lint command.
  4. Pause/drain the canary manager where required and upgrade to a pinned version.
  5. Run synthetic plus representative workload; compare queue, duration, errors, artifacts/cache and worker teardown.
  6. Expand incrementally. Keep the prior package/image/chart and config as rollback material.

11. Docker Machine migration mapping

Legacy concept Migration question Likely target
Machine driver/options Which provider group/plugin and image reproduce required compute/network? Fleeting plugin + dedicated provider autoscaling resource.
IdleCount/IdleTime What warm-capacity SLO/cost policy do measurements justify? Autoscaler policy idle_count/idle_time.
MaxBuilds How many jobs may one worker run before destruction? max_use_count.
Machine count/concurrency What is the hard cost ceiling and jobs-per-worker? max_instances, capacity_per_instance, concurrent/limit.
Docker Machine host behavior Do jobs need Docker semantics or full host access? Docker Autoscaler or Instance executor.

12. Decision table: choose the fleet deliberately

Scenario Choice Prerequisites/trust Observable proof
Public MR tests with arbitrary code Single-use ephemeral Linux pool No production secrets; restricted network; dedicated untrusted tags worker use count 1, destruction evidence, no protected credentials.
Fast internal unit tests Small warm Docker Autoscaler pool Trusted project class; bounded VM group; cache policy queue SLO, warm idle count, max instance/cost evidence.
Cloud-native integration tests Kubernetes executor Dedicated namespace/RBAC, pod resources, cluster autoscaling pod scheduling + node scale events + job trace.
Host/device compiler job Instance executor specialized pool Dedicated images/devices; narrow job routing instance identity, device exposure, teardown/reuse policy.
Legacy Docker Machine pool Canary migration to Docker Autoscaler Dedicated new provider group; Fleeting plugin; rollback pool old/new queue/cost/error comparison and drain evidence.

13. A simple cost envelope

For a VM pool, define a conservative hourly upper bound before rollout:

\[ C_{max/hour} = N_{max} \times C_{instance/hour} + C_{storage/network/cache} \]

This is intentionally simple; reserved/spot pricing, egress, cache/object storage and node sharing complicate the real number. The important control is that N_ is explicit and observable.

14. Tier and offering boundary

Runner executors, autoscaling, monitoring and the Kubernetes executor documented here are available across Free, Premium and Ultimate offerings. Cloud VM groups, Kubernetes clusters and provider credentials themselves are external infrastructure with their own cost and authorization. The core chapter never requires them.

Knowledge check

Which is safer for untrusted fork jobs: max_use_count=1 or a reused worker with strong cleanup scripts?

Why can specialized pools increase cost even if each pool autos-scales?

What is wrong with max_instances = 0 as an unreviewed cost policy?

Can the same runner tags be used by two fault-tolerant managers?

What evidence should justify raising warm idle count?

Next lesson

Continue to diagnostics

Lesson 4 treats cost runaway, controller conflicts, stale versions, lost ephemeral evidence and cross-job residue as causal failures—not as reasons for blind capacity increases.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Current stable GitLab Runner patch verified for this chapter is 19.3.2 (tagged 2026-09-10). The Docker Machine executor was deprecated in GitLab 17.5 and is scheduled for removal as a supported feature in GitLab 20.0 (May 2027); GitLab directs users toward the Instance or Docker Autoscaler executors. Docker Autoscaler is GA and, together with the Instance executor, uses Taskscaler/Fleeting. A Docker Autoscaler configuration must have its own dedicated provider autoscaling resource and must not share that resource with another manager or another [[runners]] entry. The mandatory exercises in this chapter are local simulations and require no cloud account, runner registration token, privileged executor, or managed Kubernetes cluster. GitLab Runner 19.3.2 is the chapter baseline. Pin production upgrades to an explicitly tested patch and verify release notes rather than treating “latest” as a fleet policy.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.