Checkpoint Lab — Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning
Capacity-test a hypothetical runner fleet from a synthetic workload trace, set hard scaling/cost bounds, inject burst and manager failures, and produce a measurable migration and rollback dossier.
Learning objectives
Checkpoint objectives
- Capacity-test a hypothetical runner fleet from a reproducible workload trace.
- Set explicit hard worker and cost bounds before burst testing.
- Inject burst and manager-capacity failures and diagnose the correct layer.
- Produce a Docker Machine migration and rollback plan with measurable gates.
- Deliver an evidence packet without provisioning any cloud/Kubernetes resources.
1. Checkpoint scenario
You own the fictional linux-general runner pool. It
currently uses Docker Machine, must be migrated before GitLab 20.0,
and has a target queue SLO of
95% of jobs starting within 45 seconds. You must
evaluate a bounded replacement, inject one burst and one manager
failure, and produce a rollback plan. No real runner, autoscaling
group, token, cloud account or Kubernetes cluster is required.
2. Assumptions and preflight
| Item | Checkpoint value |
|---|---|
| Documentation/Runner baseline | 2026-09-12 / GitLab Runner 19.3.2 |
| Legacy source | Docker Machine executor; deprecated since 17.5; removal scheduled for 20.0 |
| Target candidate | Docker Autoscaler, one job/use per instance |
| Hard worker bound | 8 instances maximum |
| Warm target | 1 idle instance during active window |
| Simulated cold start | 20 seconds |
| Queue SLO | 95% start within 45 seconds |
| Cost model | Relative worker-seconds plus explicit maximum instance count |
mkdir -p glci-ch30-checkpoint/evidence
cd glci-ch30-checkpoint
python3 --version
printf 'runner_doc_baseline=19.3.2
max_instances=8
queue_slo_s=45
' > evidence/assumptions.txt
3. Predict before you measure
Write two or more predictions to
evidence/predictions.txt. Example:
P1: Static capacity=2 will breach the 45 s queue target during the opening burst.
P2: Bounded elastic capacity=8 with 20 s cold start will reduce queue delay but cannot exceed 8 simultaneous slots.
P3: Reducing manager-effective capacity to 1 will create queue delay even if provider capacity is available.
P4: max_instances=8 prevents the model from “solving” burst latency with unlimited cost.
4. Workload trace
job_id,arrival_s,duration_s,pool
j01,0,45,linux
j02,3,30,linux
j03,5,90,linux
j04,8,35,linux
j05,12,40,linux
j06,15,120,linux
j07,18,25,linux
j08,20,55,linux
j09,24,30,linux
j10,28,75,linux
j11,70,30,linux
j12,90,45,linux
Save as workload.csv. Hash it so later results are tied
to the exact input:
sha256sum workload.csv | tee evidence/workload.sha256
5. Capacity model
from dataclasses import dataclass
from heapq import heappush, heappop
import csv, sys
@dataclass
class Job:
id: str
arrival: int
duration: int
pool: str
def read_jobs(path):
with open(path, newline='', encoding='utf-8') as f:
rows = csv.DictReader(f)
return [Job(r['job_id'], int(r['arrival_s']), int(r['duration_s']), r['pool']) for r in rows]
def simulate(jobs, capacity, boot_s=0):
# Greedy bounded-capacity queue model. Each slot becomes available at a timestamp.
slots = []
out = []
for job in sorted(jobs, key=lambda j: (j.arrival, j.id)):
# Reuse an already-created slot when one is free; otherwise create until hard capacity.
if len(slots) < capacity:
available = job.arrival + boot_s
else:
available = heappop(slots)
available = max(available, job.arrival)
start = max(job.arrival, available)
finish = start + job.duration
heappush(slots, finish)
out.append((job.id, job.arrival, start, finish, start-job.arrival))
return out
def summarize(rows):
waits = [r[4] for r in rows]
return {
'jobs': len(rows),
'max_wait_s': max(waits, default=0),
'avg_wait_s': round(sum(waits)/len(waits), 1) if waits else 0,
'finish_s': max((r[3] for r in rows), default=0),
}
if __name__ == '__main__':
jobs = read_jobs(sys.argv[1])
cap = int(sys.argv[2])
boot = int(sys.argv[3]) if len(sys.argv) > 3 else 0
rows = simulate(jobs, cap, boot)
for row in rows:
print(','.join(map(str, row)))
print(summarize(rows), file=sys.stderr)
Save it as capacity_model.py, then hash the model:
sha256sum capacity_model.py | tee evidence/model.sha256
6. Run the static baseline
python3 capacity_model.py workload.csv 2 0 > evidence/baseline.csv 2> evidence/baseline-summary.txt
cat evidence/baseline-summary.txt
Record whether the baseline breaches the SLO. Do not change the trace after seeing the output; a changed workload is a new experiment.
7. Run the bounded autoscaler candidate
python3 capacity_model.py workload.csv 8 20 > evidence/candidate.csv 2> evidence/candidate-summary.txt
cat evidence/candidate-summary.txt
The candidate maximum is the same hard bound you would later encode
as max_instances=8. The simulation does not claim that
eight is correct for production; it proves the experiment cannot
expand beyond its declared capacity hypothesis.
8. Inject a burst
cat >> workload.csv <<'EOF'
j13,30,60,linux
j14,31,60,linux
j15,32,60,linux
j16,33,60,linux
EOF
sha256sum workload.csv | tee evidence/workload-burst.sha256
python3 capacity_model.py workload.csv 8 20 > evidence/burst.csv 2> evidence/burst-summary.txt
cat evidence/burst-summary.txt
Compare queue delay with the hard cap unchanged. A production response might tune warm capacity, worker count or job architecture, but the cost ceiling remains explicit.
9. Inject a manager-capacity failure
Simulate a manager/polling/concurrency bottleneck by allowing only one effective slot despite an eight-instance provider ceiling:
python3 capacity_model.py workload.csv 1 0 > evidence/manager-bottleneck.csv 2> evidence/manager-bottleneck-summary.txt
cat evidence/manager-bottleneck-summary.txt
Interpretation: available provider scale is irrelevant when the
manager layer cannot acquire/dispatch enough jobs. In a real
incident, inspect concurrent, limit,
request_concurrency, manager health and eligibility
rather than raising provider quota blindly.
10. Verify the SLO from raw rows
import csv
from pathlib import Path
for name in ['baseline','candidate','burst','manager-bottleneck']:
p = Path('evidence') / f'{name}.csv'
waits=[]
with p.open() as f:
for row in csv.reader(f):
waits.append(int(row[4]))
ok=sum(w <= 45 for w in waits)
ratio=ok/len(waits) if waits else 1
print(f'{name}: within_45s={ok}/{len(waits)} ratio={ratio:.3f} pass={ratio >= .95}')
Save the output as evidence/slo.txt. This turns a vague
“seems faster” claim into a testable objective.
11. Candidate production configuration sketch
concurrent = 16
[[runners]]
name = "linux-general-v2-canary"
executor = "docker-autoscaler"
limit = 8
request_concurrency = 4
[runners.docker]
image = "alpine:3.22"
[runners.autoscaler]
plugin = "aws"
capacity_per_instance = 1
max_use_count = 1
max_instances = 8
[[runners.autoscaler.policy]]
idle_count = 1
idle_time = "10m0s"
This is a design artifact only. A real deployment additionally requires a dedicated provider autoscaling resource, pinned Fleeting plugin, secure instance image, connector/provider credentials, cache/network configuration, observability and a registration/authentication flow. Never paste a real runner token into documentation or artifacts.
12. Migration gates and rollback plan
| Gate | Pass criterion | Rollback trigger/action |
|---|---|---|
| Config validation | New manager config validates; exact Runner/plugin/image versions recorded | Do not register/route jobs until corrected. |
| Canary functionality | Representative synthetic jobs pass with expected executor/worker identity | Pause v2 canary; resume unchanged legacy pool. |
| Queue SLO | 95% start within 45 s for agreed canary trace/window | Shift traffic back; tune only after causal analysis. |
| Isolation | single-use worker destruction proven; no cross-job residue | Stop untrusted routing immediately. |
| Cost bound | instance count never exceeds 8 and provider quotas/budget alarms align | Pause scale-out/new work and investigate. |
| Evidence | manager/job/provider logs and metrics correlate by ID/timestamp | Do not expand traffic until observability works. |
| Drain legacy | old pool paused; zero in-flight jobs before decommission | Unpause legacy if v2 regression appears during rollback window. |
13. Required evidence packet
- Workload and model SHA-256 digests.
-
Assumptions: Runner
19.3.2, executor candidate, cold-start and hard capacity bounds. - Baseline, candidate, burst and manager-bottleneck raw rows/summaries.
- SLO calculation and prediction-versus-observation notes.
-
Candidate
concurrent/limit/request_concurrency/max_instances/max_use_countdesign. - Docker Machine migration inventory and dedicated-resource ownership rule.
- Canary, drain, rollback and decommission criteria.
- Limitations: no real provider API, queue long polling, image-pull/cache/network effects or cloud price model.
-
If later executed in GitLab:
CI_PIPELINE_SOURCE,CI_COMMIT_SHA, pipeline/job IDs, runner ID/version/executor and worker identity.
14. Cleanup and rollback
cd ..
rm -rf glci-ch30-checkpoint
printf 'cleanup=verified-local-only
'
The mandatory checkpoint created only local text/code/evidence files. A real migration cleanup must target the exact legacy runner registrations, provider groups, VM images, caches and credentials after the rollback window—never wildcard-delete “old runners.”
15. Verification checklist
- At least two predictions were recorded before execution.
- Trace and model were hashed before comparison.
- Static baseline and bounded candidate both produced raw queue rows.
- Burst kept the same hard maximum instead of using unlimited scale.
- Manager bottleneck was diagnosed as manager capacity, not provider shortage.
- SLO was calculated from raw wait times.
- Candidate architecture uses single-use workers and explicit hard bounds.
- Docker Autoscaler ownership is one dedicated provider resource per autoscaler configuration.
- Migration includes canary, drain, rollback and decommission gates.
- No real token, cloud credential, production runner log or provider resource was used.
Knowledge check
Why hash both workload.csv and capacity_model.py?
The result depends on both demand input and model logic. Hashes let reviewers prove they are comparing outputs from the same experiment definition.
Why keep max_instances=8 fixed during the burst
test?
It tests behavior under the declared cost/capacity ceiling. Raising the ceiling during the test hides whether the design meets requirements within governance bounds.
The provider could create eight workers, but the one-slot manager simulation queues badly. What is the lesson?
Capacity is layered. Provider capacity cannot compensate for manager acquisition/dispatch ceilings such as concurrent/limit/request behavior or eligibility.
What rollback state must exist before decommissioning Docker Machine?
A proven new fleet plus an explicit rollback window. During migration, the legacy pool can remain intact but paused so it can be resumed if canary/traffic-shift criteria fail.
What would make this checkpoint insufficient for a real cost approval?
It uses relative local timing and no provider rates, quotas, cache/network cost or real provisioning metrics. Production approval requires those environment-specific data.
16. What Chapter 30 adds to the production operating model
You can now operate runners as a measurable service: jobs are classified and routed to explicit pools; managers and workers have distinct capacity controls; autoscaling is bounded by cost and isolation policy; queue and provisioning metrics drive tuning; ephemeral evidence survives teardown; upgrades are canaried; and Docker Machine migration has a tested rollback path instead of a deadline-only scramble.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12.
Current stable GitLab Runner patch verified for this chapter is
19.3.2 (tagged 2026-09-10). The Docker Machine executor
was deprecated in GitLab 17.5 and is scheduled for removal as a
supported feature in GitLab 20.0 (May 2027); GitLab directs users
toward the Instance or Docker Autoscaler executors. Docker
Autoscaler is GA and, together with the Instance executor, uses
Taskscaler/Fleeting. A Docker Autoscaler configuration must have its
own dedicated provider autoscaling resource and must not share that
resource with another manager or another
[[runners]] entry. The mandatory exercises in this
chapter are local simulations and require no cloud account, runner
registration token, privileged executor, or managed Kubernetes
cluster. The checkpoint deliberately proves operating logic without
cloud side effects. Before a real Docker Machine migration, verify
provider-specific Fleeting plugin documentation, your exact
GitLab/Runner versions, image/bootstrap compatibility, quotas and
cost controls.
- Executors — official reference.
- GitLab Runner autoscaling — official reference.
- Docker Autoscaler executor — official reference.
- Instance executor — official reference.
- Fleeting — official reference.
- Plan and operate a runner fleet — official reference.
- Advanced Runner configuration — official reference.
- Monitor GitLab Runner usage — official reference.
- Kubernetes executor — official reference.
- GitLab Runner Helm chart — official reference.
- Docker Machine executor — official reference.
- GitLab deprecations and removals — official reference.
- Runner troubleshooting — official reference.
- GitLab Runner changelog — official reference.
- GitLab Runner tags — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.