Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning: Guided Hands-On Workflow and Core Operations
Model a bursty workload locally, compare static, Docker Autoscaler/Instance, and Kubernetes runner designs, calculate bounded capacity, simulate failure, and produce a Docker Machine migration plan without provisioning cloud resources.
Learning objectives
- Create a synthetic workload trace and calculate queue latency under hard capacity bounds.
- Compare static workers, VM-based autoscaling and Kubernetes pod/node scaling as distinct designs.
-
Translate observed workload requirements into
concurrent,limit, request and worker bounds. - Build a Docker Machine migration worksheet without touching a cloud provider.
- Preserve source, model assumptions, outputs and cleanup evidence.
1. Scenario: size a disposable Linux test pool without buying infrastructure
You operate a fictional project glci-ch30-lab. A 12-job
burst arrives over 90 seconds. The mandatory path uses Python only;
it does not register a runner, create VMs, create pods, or call
cloud APIs. The simulator is intentionally simple: it is a decision
aid, not a replacement for production telemetry.
2. Assumptions to record before simulation
| Item | Lab assumption | Production evidence |
|---|---|---|
| GitLab Runner semantics | Runner 19.3.2 documentation |
gitlab-runner --version and fleet inventory.
|
| Workload | 12 synthetic Linux jobs | Measured queue/job duration histograms by trust/platform pool. |
| Static design | 2 immediately available slots |
Existing runner count,
concurrent/limit, host capacity.
|
| Elastic design | Hard maximum 8 workers, 20 s cold start |
max_instances, creation-duration histogram,
provider quota.
|
| Trust | Synthetic, non-secret jobs | Pipeline source/ref/protection and runner protection/tag rules. |
| Cost | Relative worker-seconds only | Provider rates, storage/network/cache cost and idle policy. |
3. Local preflight
mkdir -p glci-ch30-lab/evidence
cd glci-ch30-lab
python3 --version
printf 'source=local-simulation
' > evidence/context.txt
git rev-parse HEAD >> evidence/context.txt 2>/dev/null || true
If you later run the simulator in a throwaway GitLab pipeline,
extend the context with CI_PIPELINE_SOURCE,
CI_COMMIT_SHA, CI_PIPELINE_ID and
CI_JOB_ID.
4. Create the synthetic workload trace
job_id,arrival_s,duration_s,pool
j01,0,45,linux
j02,3,30,linux
j03,5,90,linux
j04,8,35,linux
j05,12,40,linux
j06,15,120,linux
j07,18,25,linux
j08,20,55,linux
j09,24,30,linux
j10,28,75,linux
j11,70,30,linux
j12,90,45,linux
Save it as workload.csv. Arrival time and duration are
deliberately deterministic so results can be reproduced.
5. Create the bounded queue model
from dataclasses import dataclass
from heapq import heappush, heappop
import csv, sys
@dataclass
class Job:
id: str
arrival: int
duration: int
pool: str
def read_jobs(path):
with open(path, newline='', encoding='utf-8') as f:
rows = csv.DictReader(f)
return [Job(r['job_id'], int(r['arrival_s']), int(r['duration_s']), r['pool']) for r in rows]
def simulate(jobs, capacity, boot_s=0):
# Greedy bounded-capacity queue model. Each slot becomes available at a timestamp.
slots = []
out = []
for job in sorted(jobs, key=lambda j: (j.arrival, j.id)):
# Reuse an already-created slot when one is free; otherwise create until hard capacity.
if len(slots) < capacity:
available = job.arrival + boot_s
else:
available = heappop(slots)
available = max(available, job.arrival)
start = max(job.arrival, available)
finish = start + job.duration
heappush(slots, finish)
out.append((job.id, job.arrival, start, finish, start-job.arrival))
return out
def summarize(rows):
waits = [r[4] for r in rows]
return {
'jobs': len(rows),
'max_wait_s': max(waits, default=0),
'avg_wait_s': round(sum(waits)/len(waits), 1) if waits else 0,
'finish_s': max((r[3] for r in rows), default=0),
}
if __name__ == '__main__':
jobs = read_jobs(sys.argv[1])
cap = int(sys.argv[2])
boot = int(sys.argv[3]) if len(sys.argv) > 3 else 0
rows = simulate(jobs, cap, boot)
for row in rows:
print(','.join(map(str, row)))
print(summarize(rows), file=sys.stderr)
Save it as capacity_model.py. The model has one hard
control: slot capacity. A cold-start delay can be added to represent
worker provisioning. It does not model image pulls, provider
throttling, retries or heterogeneous machines, so those limitations
belong in the evidence packet.
6. Baseline: two static workers
python3 capacity_model.py workload.csv 2 0 > evidence/static.csv 2> evidence/static-summary.txt
cat evidence/static-summary.txt
Interpret max_wait_s and avg_wait_s as
developer-facing queue pressure. The baseline gives you a comparison
before introducing autoscaling complexity.
7. Simulate bounded elastic capacity
python3 capacity_model.py workload.csv 8 20 > evidence/elastic.csv 2> evidence/elastic-summary.txt
cat evidence/elastic-summary.txt
The model pessimistically applies cold-start time when it creates
capacity. In a real Runner autoscaler, idle workers, creation
overlap, provider API latency and manager behavior change the shape.
Keep 8 as a hard capacity hypothesis, not as a
recommendation.
8. Map simulation controls to Runner controls
| Simulation concept | Runner/infra control | Evidence |
|---|---|---|
| Fleet-wide job ceiling | Runner process concurrent |
gitlab_runner_concurrent metric + config.
|
| Pool ceiling | Per-runner limit |
gitlab_runner_limit + pool config. |
| Job request parallelism | request_concurrency |
request-concurrency metric/exceeded counter + long-poll symptoms. |
| VM count ceiling | max_instances |
Autoscaler config + machine-state metrics + provider group limits. |
| Jobs per VM | capacity_per_instance |
Autoscaler config + worker/job mapping. |
| Lifetime/reuse | max_use_count |
Worker lifecycle logs and destruction events. |
| Warm capacity |
autoscaler idle_count/idle_time
policy
|
Idle machine metric + queue latency + cost. |
9. Compare three production architectures
| Design | Scale unit | Strong point | Primary caveat |
|---|---|---|---|
| Static runner hosts | Pre-provisioned host/container slots | Predictable and simple | Idle cost and manual capacity management. |
| Docker Autoscaler / Instance | Fleeting-managed VM instances | Explicit hard VM/reuse bounds and strong ephemeral lifecycle options | Cloud/plugin/provider complexity and cold start. |
| Kubernetes executor | Job pod, with node capacity managed by Kubernetes | Native pod scheduling/resources; one pod per job | Node-autoscaler latency, cluster RBAC/trust and shared-node effects. |
10. Build a Docker Machine migration inventory
Create evidence/migration.csv with one row per legacy
runner pool:
legacy_pool,current_runner_version,job_tags,platform,machine_driver,idle_count,idle_time,max_builds,cache,network,devices,target_executor,rollback
linux-general,18.x,"linux,docker",linux,amazonec2,2,600,1,s3,private-subnet,none,docker-autoscaler,keep-old-pool-paused
host-tools,18.x,"linux,host",linux,google,0,0,1,none,private-subnet,"usb?",instance,keep-old-pool-paused
Do not merely copy MachineOptions into a new file. For
each pool, decide whether Docker semantics are required (Docker
Autoscaler) or host access is required (Instance), identify the
provider group and Fleeting plugin, define hard
max_instances, and preserve a paused rollback pool
until canary evidence is accepted.
11. Progressive migration plan
- Inventory: exact Runner version, tags, protection, executor, machine/provider settings, caches, network, volumes/devices and queue SLO.
- Design: dedicated new autoscaling resource, pinned Runner/plugin versions, instance image, hard bounds, logging and metrics.
-
Canary: new runner tags such as
fleet-v2-canary; representative synthetic jobs only. - Shadow measure: compare queue/provisioning/job/error/cost metrics, not just success rate.
- Traffic shift: change runner eligibility/tags in controlled increments.
- Drain legacy: pause old runner for new work; let in-flight jobs finish.
- Rollback window: retain old config/resource intact but paused until new fleet proves stable.
- Decommission: remove only the identified Docker Machine resources after evidence and ownership checks.
12. Optional GitLab pipeline for the simulator
capacity_model:
image: python:3.12-alpine
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
script:
- printf 'source=%s sha=%s pipeline=%s job=%s\n' "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID" | tee evidence/gitlab-context.txt
- python3 capacity_model.py workload.csv 2 0 > evidence/static.csv 2> evidence/static-summary.txt
- python3 capacity_model.py workload.csv 8 20 > evidence/elastic.csv 2> evidence/elastic-summary.txt
artifacts:
when: always
expire_in: 1 week
paths: [evidence/]
The image tag is convenient for a disposable simulation. A production training repository can pin an image digest after verifying the upstream image for its architecture.
13. Challenge: the queue stays high even after doubling
max_instances
You observe max_instances=16, only four machines
running, manager CPU low, and repeated
request_concurrency saturation. Which layer should you
change first?
Reasoning target: do not raise provider capacity
again. Inspect request_concurrency, long-poll behavior,
concurrent/limit, eligibility and manager
errors. Capacity that cannot be requested/assigned is not a provider
shortage.
14. Cleanup
cd ..
rm -rf glci-ch30-lab
printf 'cleanup=local-files-removed
'
No runner registration, cloud resource, Kubernetes object or external state was created by the mandatory path.
Knowledge check
Why run a static baseline before an autoscaling simulation?
It gives a reproducible queue-latency baseline so you can quantify what complexity/cold-start tradeoff the elastic design actually improves.
Why does the migration inventory include caches, network and devices?
Executor migration changes more than VM creation. Hidden host/network/cache/device assumptions are common reasons a syntactically correct migration fails real jobs.
A Kubernetes job pod is Pending. Which layer owns the next diagnosis?
Start with Kubernetes scheduling/node capacity, requests/limits, quota, taints and cluster autoscaling—not GitLab pipeline compilation—unless evidence shows the pod was never created.
What makes the simulated hard maximum useful even though it is not real autoscaling?
It forces a capacity/cost ceiling into the design and lets you evaluate queue behavior under bounded resources rather than assuming infinite scale.
Why keep the old Docker Machine pool paused rather than immediately delete it?
A paused, intact pool is a bounded rollback option while the new executor proves representative workload, queue, isolation and cost behavior.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12.
Current stable GitLab Runner patch verified for this chapter is
19.3.2 (tagged 2026-09-10). The Docker Machine executor
was deprecated in GitLab 17.5 and is scheduled for removal as a
supported feature in GitLab 20.0 (May 2027); GitLab directs users
toward the Instance or Docker Autoscaler executors. Docker
Autoscaler is GA and, together with the Instance executor, uses
Taskscaler/Fleeting. A Docker Autoscaler configuration must have its
own dedicated provider autoscaling resource and must not share that
resource with another manager or another
[[runners]] entry. The mandatory exercises in this
chapter are local simulations and require no cloud account, runner
registration token, privileged executor, or managed Kubernetes
cluster. The Python simulator intentionally omits provider APIs. Its
outputs are capacity hypotheses; validate any real fleet with Runner
and provider/Kubernetes metrics before rollout.
- Executors — official reference.
- GitLab Runner autoscaling — official reference.
- Docker Autoscaler executor — official reference.
- Instance executor — official reference.
- Fleeting — official reference.
- Plan and operate a runner fleet — official reference.
- Advanced Runner configuration — official reference.
- Monitor GitLab Runner usage — official reference.
- Kubernetes executor — official reference.
- GitLab Runner Helm chart — official reference.
- Docker Machine executor — official reference.
- GitLab deprecations and removals — official reference.
- Runner troubleshooting — official reference.
- GitLab Runner changelog — official reference.
- GitLab Runner tags — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.