Chapter 30Lesson 05~275 minutes

Checkpoint Lab — Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning

Capacity-test a hypothetical runner fleet from a synthetic workload trace, set hard scaling/cost bounds, inject burst and manager failures, and produce a measurable migration and rollback dossier.

CheckpointCapacity testHard boundsMigrationEvidence

Learning objectives

Checkpoint objectives

  • Capacity-test a hypothetical runner fleet from a reproducible workload trace.
  • Set explicit hard worker and cost bounds before burst testing.
  • Inject burst and manager-capacity failures and diagnose the correct layer.
  • Produce a Docker Machine migration and rollback plan with measurable gates.
  • Deliver an evidence packet without provisioning any cloud/Kubernetes resources.

1. Checkpoint scenario

You own the fictional linux-general runner pool. It currently uses Docker Machine, must be migrated before GitLab 20.0, and has a target queue SLO of 95% of jobs starting within 45 seconds. You must evaluate a bounded replacement, inject one burst and one manager failure, and produce a rollback plan. No real runner, autoscaling group, token, cloud account or Kubernetes cluster is required.

Checkpoint boundary: synthetic workload only. Do not register runners, mutate provider groups, create Kubernetes resources or expose production runner configuration. Every “provider” action is a documented simulation.

2. Assumptions and preflight

Item Checkpoint value
Documentation/Runner baseline 2026-09-12 / GitLab Runner 19.3.2
Legacy source Docker Machine executor; deprecated since 17.5; removal scheduled for 20.0
Target candidate Docker Autoscaler, one job/use per instance
Hard worker bound 8 instances maximum
Warm target 1 idle instance during active window
Simulated cold start 20 seconds
Queue SLO 95% start within 45 seconds
Cost model Relative worker-seconds plus explicit maximum instance count
mkdir -p glci-ch30-checkpoint/evidence
cd glci-ch30-checkpoint
python3 --version
printf 'runner_doc_baseline=19.3.2
max_instances=8
queue_slo_s=45
'   > evidence/assumptions.txt

3. Predict before you measure

Write two or more predictions to evidence/predictions.txt. Example:

P1: Static capacity=2 will breach the 45 s queue target during the opening burst.
P2: Bounded elastic capacity=8 with 20 s cold start will reduce queue delay but cannot exceed 8 simultaneous slots.
P3: Reducing manager-effective capacity to 1 will create queue delay even if provider capacity is available.
P4: max_instances=8 prevents the model from “solving” burst latency with unlimited cost.

4. Workload trace

job_id,arrival_s,duration_s,pool
j01,0,45,linux
j02,3,30,linux
j03,5,90,linux
j04,8,35,linux
j05,12,40,linux
j06,15,120,linux
j07,18,25,linux
j08,20,55,linux
j09,24,30,linux
j10,28,75,linux
j11,70,30,linux
j12,90,45,linux

Save as workload.csv. Hash it so later results are tied to the exact input:

sha256sum workload.csv | tee evidence/workload.sha256

5. Capacity model

from dataclasses import dataclass
from heapq import heappush, heappop
import csv, sys

@dataclass
class Job:
    id: str
    arrival: int
    duration: int
    pool: str


def read_jobs(path):
    with open(path, newline='', encoding='utf-8') as f:
        rows = csv.DictReader(f)
        return [Job(r['job_id'], int(r['arrival_s']), int(r['duration_s']), r['pool']) for r in rows]


def simulate(jobs, capacity, boot_s=0):
    # Greedy bounded-capacity queue model. Each slot becomes available at a timestamp.
    slots = []
    out = []
    for job in sorted(jobs, key=lambda j: (j.arrival, j.id)):
        # Reuse an already-created slot when one is free; otherwise create until hard capacity.
        if len(slots) < capacity:
            available = job.arrival + boot_s
        else:
            available = heappop(slots)
            available = max(available, job.arrival)
        start = max(job.arrival, available)
        finish = start + job.duration
        heappush(slots, finish)
        out.append((job.id, job.arrival, start, finish, start-job.arrival))
    return out


def summarize(rows):
    waits = [r[4] for r in rows]
    return {
        'jobs': len(rows),
        'max_wait_s': max(waits, default=0),
        'avg_wait_s': round(sum(waits)/len(waits), 1) if waits else 0,
        'finish_s': max((r[3] for r in rows), default=0),
    }

if __name__ == '__main__':
    jobs = read_jobs(sys.argv[1])
    cap = int(sys.argv[2])
    boot = int(sys.argv[3]) if len(sys.argv) > 3 else 0
    rows = simulate(jobs, cap, boot)
    for row in rows:
        print(','.join(map(str, row)))
    print(summarize(rows), file=sys.stderr)

Save it as capacity_model.py, then hash the model:

sha256sum capacity_model.py | tee evidence/model.sha256

6. Run the static baseline

python3 capacity_model.py workload.csv 2 0   > evidence/baseline.csv 2> evidence/baseline-summary.txt
cat evidence/baseline-summary.txt

Record whether the baseline breaches the SLO. Do not change the trace after seeing the output; a changed workload is a new experiment.

7. Run the bounded autoscaler candidate

python3 capacity_model.py workload.csv 8 20   > evidence/candidate.csv 2> evidence/candidate-summary.txt
cat evidence/candidate-summary.txt

The candidate maximum is the same hard bound you would later encode as max_instances=8. The simulation does not claim that eight is correct for production; it proves the experiment cannot expand beyond its declared capacity hypothesis.

8. Inject a burst

cat >> workload.csv <<'EOF'
j13,30,60,linux
j14,31,60,linux
j15,32,60,linux
j16,33,60,linux
EOF
sha256sum workload.csv | tee evidence/workload-burst.sha256
python3 capacity_model.py workload.csv 8 20   > evidence/burst.csv 2> evidence/burst-summary.txt
cat evidence/burst-summary.txt

Compare queue delay with the hard cap unchanged. A production response might tune warm capacity, worker count or job architecture, but the cost ceiling remains explicit.

9. Inject a manager-capacity failure

Simulate a manager/polling/concurrency bottleneck by allowing only one effective slot despite an eight-instance provider ceiling:

python3 capacity_model.py workload.csv 1 0   > evidence/manager-bottleneck.csv 2> evidence/manager-bottleneck-summary.txt
cat evidence/manager-bottleneck-summary.txt

Interpretation: available provider scale is irrelevant when the manager layer cannot acquire/dispatch enough jobs. In a real incident, inspect concurrent, limit, request_concurrency, manager health and eligibility rather than raising provider quota blindly.

10. Verify the SLO from raw rows

import csv
from pathlib import Path

for name in ['baseline','candidate','burst','manager-bottleneck']:
    p = Path('evidence') / f'{name}.csv'
    waits=[]
    with p.open() as f:
        for row in csv.reader(f):
            waits.append(int(row[4]))
    ok=sum(w <= 45 for w in waits)
    ratio=ok/len(waits) if waits else 1
    print(f'{name}: within_45s={ok}/{len(waits)} ratio={ratio:.3f} pass={ratio >= .95}')

Save the output as evidence/slo.txt. This turns a vague “seems faster” claim into a testable objective.

11. Candidate production configuration sketch

concurrent = 16

[[runners]]
  name = "linux-general-v2-canary"
  executor = "docker-autoscaler"
  limit = 8
  request_concurrency = 4

  [runners.docker]
    image = "alpine:3.22"

  [runners.autoscaler]
    plugin = "aws"
    capacity_per_instance = 1
    max_use_count = 1
    max_instances = 8

    [[runners.autoscaler.policy]]
      idle_count = 1
      idle_time = "10m0s"

This is a design artifact only. A real deployment additionally requires a dedicated provider autoscaling resource, pinned Fleeting plugin, secure instance image, connector/provider credentials, cache/network configuration, observability and a registration/authentication flow. Never paste a real runner token into documentation or artifacts.

12. Migration gates and rollback plan

Gate Pass criterion Rollback trigger/action
Config validation New manager config validates; exact Runner/plugin/image versions recorded Do not register/route jobs until corrected.
Canary functionality Representative synthetic jobs pass with expected executor/worker identity Pause v2 canary; resume unchanged legacy pool.
Queue SLO 95% start within 45 s for agreed canary trace/window Shift traffic back; tune only after causal analysis.
Isolation single-use worker destruction proven; no cross-job residue Stop untrusted routing immediately.
Cost bound instance count never exceeds 8 and provider quotas/budget alarms align Pause scale-out/new work and investigate.
Evidence manager/job/provider logs and metrics correlate by ID/timestamp Do not expand traffic until observability works.
Drain legacy old pool paused; zero in-flight jobs before decommission Unpause legacy if v2 regression appears during rollback window.

13. Required evidence packet

  • Workload and model SHA-256 digests.
  • Assumptions: Runner 19.3.2, executor candidate, cold-start and hard capacity bounds.
  • Baseline, candidate, burst and manager-bottleneck raw rows/summaries.
  • SLO calculation and prediction-versus-observation notes.
  • Candidate concurrent/limit/request_concurrency/max_instances/max_use_count design.
  • Docker Machine migration inventory and dedicated-resource ownership rule.
  • Canary, drain, rollback and decommission criteria.
  • Limitations: no real provider API, queue long polling, image-pull/cache/network effects or cloud price model.
  • If later executed in GitLab: CI_PIPELINE_SOURCE, CI_COMMIT_SHA, pipeline/job IDs, runner ID/version/executor and worker identity.

14. Cleanup and rollback

cd ..
rm -rf glci-ch30-checkpoint
printf 'cleanup=verified-local-only
'

The mandatory checkpoint created only local text/code/evidence files. A real migration cleanup must target the exact legacy runner registrations, provider groups, VM images, caches and credentials after the rollback window—never wildcard-delete “old runners.”

15. Verification checklist

  • At least two predictions were recorded before execution.
  • Trace and model were hashed before comparison.
  • Static baseline and bounded candidate both produced raw queue rows.
  • Burst kept the same hard maximum instead of using unlimited scale.
  • Manager bottleneck was diagnosed as manager capacity, not provider shortage.
  • SLO was calculated from raw wait times.
  • Candidate architecture uses single-use workers and explicit hard bounds.
  • Docker Autoscaler ownership is one dedicated provider resource per autoscaler configuration.
  • Migration includes canary, drain, rollback and decommission gates.
  • No real token, cloud credential, production runner log or provider resource was used.

Knowledge check

Why hash both workload.csv and capacity_model.py?

Why keep max_instances=8 fixed during the burst test?

The provider could create eight workers, but the one-slot manager simulation queues badly. What is the lesson?

What rollback state must exist before decommissioning Docker Machine?

What would make this checkpoint insufficient for a real cost approval?

16. What Chapter 30 adds to the production operating model

You can now operate runners as a measurable service: jobs are classified and routed to explicit pools; managers and workers have distinct capacity controls; autoscaling is bounded by cost and isolation policy; queue and provisioning metrics drive tuning; ephemeral evidence survives teardown; upgrades are canaried; and Docker Machine migration has a tested rollback path instead of a deadline-only scramble.

Next chapter

Runner Security, Isolation Boundaries, Privileged Containers, Fork Pipelines, Untrusted Code, and Threat Modeling

Chapter 31 keeps the same fleet model but changes the primary question from capacity to adversarial trust: what happens when the job itself is malicious, comes from a fork, reaches privileged execution, or can observe host/network/device credentials?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Current stable GitLab Runner patch verified for this chapter is 19.3.2 (tagged 2026-09-10). The Docker Machine executor was deprecated in GitLab 17.5 and is scheduled for removal as a supported feature in GitLab 20.0 (May 2027); GitLab directs users toward the Instance or Docker Autoscaler executors. Docker Autoscaler is GA and, together with the Instance executor, uses Taskscaler/Fleeting. A Docker Autoscaler configuration must have its own dedicated provider autoscaling resource and must not share that resource with another manager or another [[runners]] entry. The mandatory exercises in this chapter are local simulations and require no cloud account, runner registration token, privileged executor, or managed Kubernetes cluster. The checkpoint deliberately proves operating logic without cloud side effects. Before a real Docker Machine migration, verify provider-specific Fleeting plugin documentation, your exact GitLab/Runner versions, image/bootstrap compatibility, quotas and cost controls.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.