Chapter 30Lesson 02~250 minutes

Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning: Guided Hands-On Workflow and Core Operations

Model a bursty workload locally, compare static, Docker Autoscaler/Instance, and Kubernetes runner designs, calculate bounded capacity, simulate failure, and produce a Docker Machine migration plan without provisioning cloud resources.

Hands-onQueue modelBurst testMigrationFree/local

Learning objectives

  • Create a synthetic workload trace and calculate queue latency under hard capacity bounds.
  • Compare static workers, VM-based autoscaling and Kubernetes pod/node scaling as distinct designs.
  • Translate observed workload requirements into concurrent, limit, request and worker bounds.
  • Build a Docker Machine migration worksheet without touching a cloud provider.
  • Preserve source, model assumptions, outputs and cleanup evidence.

1. Scenario: size a disposable Linux test pool without buying infrastructure

You operate a fictional project glci-ch30-lab. A 12-job burst arrives over 90 seconds. The mandatory path uses Python only; it does not register a runner, create VMs, create pods, or call cloud APIs. The simulator is intentionally simple: it is a decision aid, not a replacement for production telemetry.

Safety boundary: use synthetic workload data and local files only. Do not paste real registration tokens, cloud credentials, provider group names, customer workload traces or production runner logs into this lab.

2. Assumptions to record before simulation

Item Lab assumption Production evidence
GitLab Runner semantics Runner 19.3.2 documentation gitlab-runner --version and fleet inventory.
Workload 12 synthetic Linux jobs Measured queue/job duration histograms by trust/platform pool.
Static design 2 immediately available slots Existing runner count, concurrent/limit, host capacity.
Elastic design Hard maximum 8 workers, 20 s cold start max_instances, creation-duration histogram, provider quota.
Trust Synthetic, non-secret jobs Pipeline source/ref/protection and runner protection/tag rules.
Cost Relative worker-seconds only Provider rates, storage/network/cache cost and idle policy.

3. Local preflight

mkdir -p glci-ch30-lab/evidence
cd glci-ch30-lab
python3 --version
printf 'source=local-simulation
' > evidence/context.txt
git rev-parse HEAD >> evidence/context.txt 2>/dev/null || true

If you later run the simulator in a throwaway GitLab pipeline, extend the context with CI_PIPELINE_SOURCE, CI_COMMIT_SHA, CI_PIPELINE_ID and CI_JOB_ID.

4. Create the synthetic workload trace

job_id,arrival_s,duration_s,pool
j01,0,45,linux
j02,3,30,linux
j03,5,90,linux
j04,8,35,linux
j05,12,40,linux
j06,15,120,linux
j07,18,25,linux
j08,20,55,linux
j09,24,30,linux
j10,28,75,linux
j11,70,30,linux
j12,90,45,linux

Save it as workload.csv. Arrival time and duration are deliberately deterministic so results can be reproduced.

5. Create the bounded queue model

from dataclasses import dataclass
from heapq import heappush, heappop
import csv, sys

@dataclass
class Job:
    id: str
    arrival: int
    duration: int
    pool: str


def read_jobs(path):
    with open(path, newline='', encoding='utf-8') as f:
        rows = csv.DictReader(f)
        return [Job(r['job_id'], int(r['arrival_s']), int(r['duration_s']), r['pool']) for r in rows]


def simulate(jobs, capacity, boot_s=0):
    # Greedy bounded-capacity queue model. Each slot becomes available at a timestamp.
    slots = []
    out = []
    for job in sorted(jobs, key=lambda j: (j.arrival, j.id)):
        # Reuse an already-created slot when one is free; otherwise create until hard capacity.
        if len(slots) < capacity:
            available = job.arrival + boot_s
        else:
            available = heappop(slots)
            available = max(available, job.arrival)
        start = max(job.arrival, available)
        finish = start + job.duration
        heappush(slots, finish)
        out.append((job.id, job.arrival, start, finish, start-job.arrival))
    return out


def summarize(rows):
    waits = [r[4] for r in rows]
    return {
        'jobs': len(rows),
        'max_wait_s': max(waits, default=0),
        'avg_wait_s': round(sum(waits)/len(waits), 1) if waits else 0,
        'finish_s': max((r[3] for r in rows), default=0),
    }

if __name__ == '__main__':
    jobs = read_jobs(sys.argv[1])
    cap = int(sys.argv[2])
    boot = int(sys.argv[3]) if len(sys.argv) > 3 else 0
    rows = simulate(jobs, cap, boot)
    for row in rows:
        print(','.join(map(str, row)))
    print(summarize(rows), file=sys.stderr)

Save it as capacity_model.py. The model has one hard control: slot capacity. A cold-start delay can be added to represent worker provisioning. It does not model image pulls, provider throttling, retries or heterogeneous machines, so those limitations belong in the evidence packet.

6. Baseline: two static workers

python3 capacity_model.py workload.csv 2 0   > evidence/static.csv 2> evidence/static-summary.txt
cat evidence/static-summary.txt

Interpret max_wait_s and avg_wait_s as developer-facing queue pressure. The baseline gives you a comparison before introducing autoscaling complexity.

7. Simulate bounded elastic capacity

python3 capacity_model.py workload.csv 8 20   > evidence/elastic.csv 2> evidence/elastic-summary.txt
cat evidence/elastic-summary.txt

The model pessimistically applies cold-start time when it creates capacity. In a real Runner autoscaler, idle workers, creation overlap, provider API latency and manager behavior change the shape. Keep 8 as a hard capacity hypothesis, not as a recommendation.

8. Map simulation controls to Runner controls

Simulation concept Runner/infra control Evidence
Fleet-wide job ceiling Runner process concurrent gitlab_runner_concurrent metric + config.
Pool ceiling Per-runner limit gitlab_runner_limit + pool config.
Job request parallelism request_concurrency request-concurrency metric/exceeded counter + long-poll symptoms.
VM count ceiling max_instances Autoscaler config + machine-state metrics + provider group limits.
Jobs per VM capacity_per_instance Autoscaler config + worker/job mapping.
Lifetime/reuse max_use_count Worker lifecycle logs and destruction events.
Warm capacity autoscaler idle_count/idle_time policy Idle machine metric + queue latency + cost.

9. Compare three production architectures

Design Scale unit Strong point Primary caveat
Static runner hosts Pre-provisioned host/container slots Predictable and simple Idle cost and manual capacity management.
Docker Autoscaler / Instance Fleeting-managed VM instances Explicit hard VM/reuse bounds and strong ephemeral lifecycle options Cloud/plugin/provider complexity and cold start.
Kubernetes executor Job pod, with node capacity managed by Kubernetes Native pod scheduling/resources; one pod per job Node-autoscaler latency, cluster RBAC/trust and shared-node effects.

10. Build a Docker Machine migration inventory

Create evidence/migration.csv with one row per legacy runner pool:

legacy_pool,current_runner_version,job_tags,platform,machine_driver,idle_count,idle_time,max_builds,cache,network,devices,target_executor,rollback
linux-general,18.x,"linux,docker",linux,amazonec2,2,600,1,s3,private-subnet,none,docker-autoscaler,keep-old-pool-paused
host-tools,18.x,"linux,host",linux,google,0,0,1,none,private-subnet,"usb?",instance,keep-old-pool-paused

Do not merely copy MachineOptions into a new file. For each pool, decide whether Docker semantics are required (Docker Autoscaler) or host access is required (Instance), identify the provider group and Fleeting plugin, define hard max_instances, and preserve a paused rollback pool until canary evidence is accepted.

11. Progressive migration plan

  1. Inventory: exact Runner version, tags, protection, executor, machine/provider settings, caches, network, volumes/devices and queue SLO.
  2. Design: dedicated new autoscaling resource, pinned Runner/plugin versions, instance image, hard bounds, logging and metrics.
  3. Canary: new runner tags such as fleet-v2-canary; representative synthetic jobs only.
  4. Shadow measure: compare queue/provisioning/job/error/cost metrics, not just success rate.
  5. Traffic shift: change runner eligibility/tags in controlled increments.
  6. Drain legacy: pause old runner for new work; let in-flight jobs finish.
  7. Rollback window: retain old config/resource intact but paused until new fleet proves stable.
  8. Decommission: remove only the identified Docker Machine resources after evidence and ownership checks.

12. Optional GitLab pipeline for the simulator

capacity_model:
  image: python:3.12-alpine
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
  script:
    - printf 'source=%s sha=%s pipeline=%s job=%s\n' "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID" | tee evidence/gitlab-context.txt
    - python3 capacity_model.py workload.csv 2 0 > evidence/static.csv 2> evidence/static-summary.txt
    - python3 capacity_model.py workload.csv 8 20 > evidence/elastic.csv 2> evidence/elastic-summary.txt
  artifacts:
    when: always
    expire_in: 1 week
    paths: [evidence/]

The image tag is convenient for a disposable simulation. A production training repository can pin an image digest after verifying the upstream image for its architecture.

13. Challenge: the queue stays high even after doubling max_instances

You observe max_instances=16, only four machines running, manager CPU low, and repeated request_concurrency saturation. Which layer should you change first?

Reasoning target: do not raise provider capacity again. Inspect request_concurrency, long-poll behavior, concurrent/limit, eligibility and manager errors. Capacity that cannot be requested/assigned is not a provider shortage.

14. Cleanup

cd ..
rm -rf glci-ch30-lab
printf 'cleanup=local-files-removed
'

No runner registration, cloud resource, Kubernetes object or external state was created by the mandatory path.

Knowledge check

Why run a static baseline before an autoscaling simulation?

Why does the migration inventory include caches, network and devices?

A Kubernetes job pod is Pending. Which layer owns the next diagnosis?

What makes the simulated hard maximum useful even though it is not real autoscaling?

Why keep the old Docker Machine pool paused rather than immediately delete it?

Next lesson

Continue to design tradeoffs

Lesson 3 turns measurements into architecture choices: warm versus cold capacity, reuse versus single-use workers, Kubernetes versus VM fleets, and centralized versus specialized runner pools.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Current stable GitLab Runner patch verified for this chapter is 19.3.2 (tagged 2026-09-10). The Docker Machine executor was deprecated in GitLab 17.5 and is scheduled for removal as a supported feature in GitLab 20.0 (May 2027); GitLab directs users toward the Instance or Docker Autoscaler executors. Docker Autoscaler is GA and, together with the Instance executor, uses Taskscaler/Fleeting. A Docker Autoscaler configuration must have its own dedicated provider autoscaling resource and must not share that resource with another manager or another [[runners]] entry. The mandatory exercises in this chapter are local simulations and require no cloud account, runner registration token, privileged executor, or managed Kubernetes cluster. The Python simulator intentionally omits provider APIs. Its outputs are capacity hypotheses; validate any real fleet with Runner and provider/Kubernetes metrics before rollout.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.