Chapter 30Lesson 04~220 minutes

Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning: Diagnostics, Failure Modes, Security, and Performance

Diagnose cost runaway, shared autoscaling ownership, stale Runner versions, lost ephemeral logs, long-poll bottlenecks, and retained untrusted state with an evidence-first sequence and bounded corrections.

DiagnosticsLong pollingLogsVersion driftIsolation

Learning objectives

  • Use an evidence-first diagnostic sequence from pipeline identity through manager, worker, provider and governance state.
  • Diagnose unbounded cost, conflicting autoscaler managers, stale Runner versions, lost logs and retained worker residue.
  • Differentiate queue, polling, executor, provider/Kubernetes and script failures.
  • Apply the least destructive correction while preserving original evidence.
  • Recover capacity without broad credentials, privileged shortcuts or unbounded scale.

1. Evidence-first sequence

Before changing anything, preserve the first-failure evidence. Record CI_PIPELINE_SOURCE, ref, CI_COMMIT_SHA, pipeline/job IDs and job tags. Confirm merged configuration and rule result. Check whether the job is created, pending, assigned or already executing. Then record runner manager ID/version, concurrent/limit/request_concurrency, executor, worker identity/image and queue/provisioning metrics. Only after that inspect provider/Kubernetes state, job script/tool/network, artifacts/caches and teardown evidence.

Diagnosis order prevents layer confusion
flowchart TD A[Preserve pipeline/job IDs + first failure] --> B[Source/ref/SHA + merged config] B --> C[Rules + job eligibility/tags] C --> D[Queue + manager polling] D --> E[Runner version/concurrency/executor] E --> F[Worker/pod/instance provisioning] F --> G[Script/tool/network] G --> H[Artifacts/cache/logs] H --> I[Teardown + provider state + cost] I --> J[Smallest safe correction]

2. Failure: unbounded autoscaling cost

Symptom: a burst or retry storm drives machine count upward faster than expected. Preserve: job arrival IDs, retry causes, machine-state metrics, creation rate, provider group events and current autoscaler config. Cause candidates: max_instances=0, excessive capacity_per_instance assumptions, retry loops, duplicate pipelines, broken teardown, or a provider group whose own scaling policy competes with Runner.

Do not “fix” cost by deleting arbitrary instances while jobs are running. First bound new growth through the approved fleet control plane, preserve job/provider evidence, drain/cancel only identified disposable work when authorized, and fix the causal retry/demand problem.

3. Failure: two managers control the same autoscaling group

Symptom: unexpected instance creation/removal, jobs losing workers, inconsistent manager views, cost spikes. Docker Autoscaler documentation explicitly forbids sharing one autoscaling resource across multiple managers or multiple runner configurations. The correction is architectural: assign a dedicated provider group to each autoscaler configuration. Do not attempt to “synchronize” two controllers with scripts.

4. Failure: stale Runner version

Symptom: old bugs, missing metrics/features, manager/helper mismatch or behavior that no longer matches current GitLab guidance. Capture gitlab-runner --version, GitLab version, helper image version, executor/plugin versions and the first failing job before upgrade. GitLab troubleshooting recommends checking Runner and GitLab versions early. Upgrade by canary/drain, not by replacing every manager simultaneously.

5. Failure: ephemeral worker disappears with the evidence

Symptom: a failed VM/pod is destroyed and only the GitLab job trace remains; provider/bootstrap/kernel or manager-side evidence is gone. Design logging before incidents: centralize runner-manager logs, provider autoscaler events and Kubernetes events where authorized; retain correlation fields such as runner ID, job ID, instance/pod ID and timestamps. Avoid placing secrets in logs.

# Read-only examples on an authorized manager; choose the command matching your deployment.
journalctl --unit=gitlab-runner.service -n 100 --no-pager
# containerized manager:
docker logs --tail 100 gitlab-runner
# Kubernetes manager pod:
kubectl logs -n gitlab-runner deploy/gitlab-runner --tail=100

These commands inspect manager logs. They do not recover logs that were never shipped from a destroyed worker.

6. Failure: reused worker retains untrusted state

Symptom: a later job can observe workspace files, Docker state, caches, sockets, services, credentials or devices left by an earlier job. Preserve the worker/job mapping and prove the residue in a synthetic lab rather than exposing real secrets. The strongest correction for mixed/untrusted workloads is often a trust-boundary change—single-use ephemeral worker or dedicated pool—not a longer cleanup script.

7. Intentionally broken configuration: infinite cost ceiling

concurrent = 200

[[runners]]
  name = "broken-autoscaler"
  executor = "docker-autoscaler"
  limit = 200
  request_concurrency = 1

  [runners.autoscaler]
    plugin = "aws"
    capacity_per_instance = 1
    max_use_count = 1
    max_instances = 0

Interpretation: max_instances=0 is unlimited, concurrent/limit allow very high demand, and request concurrency 1 can simultaneously throttle job acquisition. The configuration can be both too permissive for cost and too restrictive for polling. Fix each control for its actual purpose.

concurrent = 24

[[runners]]
  name = "bounded-autoscaler"
  executor = "docker-autoscaler"
  limit = 12
  request_concurrency = 4

  [runners.autoscaler]
    plugin = "aws"
    capacity_per_instance = 1
    max_use_count = 1
    max_instances = 12

The numbers are examples only. Production values must come from queue/provisioning/cost evidence and provider quotas.

8. Pending jobs: a causal matrix

Observation Likely layer Evidence before change
No eligible runner shown / tag mismatch Eligibility/config Job tags, runner tags/scope/protection, rule result.
Eligible manager, request concurrency saturated Manager polling request metrics, long-poll warnings, manager logs.
Manager accepted job, waits for instance Autoscaler/provider machine states, creation duration/errors, provider quota/events.
Kubernetes pod exists but Pending Kubernetes scheduler/node capacity pod events, requests/limits, quota, taints, node-autoscaler state.
Worker starts but script stalls Job/tool/network job trace, process/network/tool logs; not an autoscaling change.

9. Security-sensitive actions

Runner registration/authentication tokens, cloud IAM credentials, Kubernetes RBAC, privileged Docker/Kubernetes settings, device exposure, provider-group mutation and broad network access are control-plane changes. Use fake placeholders in documentation and narrowly scoped identities in real systems. Never print tokens, disable TLS, run untrusted jobs on a privileged shared runner, or grant broad cloud administrator roles as a troubleshooting shortcut.

10. Performance diagnosis without over-scaling

Track queue duration and job duration separately. A large queue with stable job duration indicates capacity/acquisition pressure. A stable queue with growing job duration points toward job/tool/cache/network performance. Machine creation duration isolates the cold-start path. Error counters and provider throttling can explain why requested capacity never becomes usable.

11. Docker Machine migration failure: old and new pools overlap incorrectly

If both pools use the same broad tags during canary, you lose control over which executor ran the job and cannot compare results. Use explicit canary tags and record runner ID/executor for each representative job. When shifting traffic, change one eligibility dimension at a time and keep the legacy pool paused—not deleted—until rollback criteria expire.

12. Upgrade rollback

If a canary Runner upgrade causes failures, preserve manager logs and representative job IDs, pause the canary from new work, let safe in-flight jobs finish, and revert to the recorded prior package/image/chart/config. Do not rebuild an old manager from memory. The rollback artifact is part of fleet governance.

13. Minimal incident packet

  • Pipeline source/ref/SHA, pipeline/job IDs and job tags.
  • Runner ID, manager hostname/pod, Runner version, executor and plugin/image versions.
  • concurrent, limit, request_concurrency and autoscaler hard bounds.
  • Queue, job duration, request concurrency, machine/pod creation and error metrics.
  • Worker/pod/instance identity and lifetime/use count.
  • Manager/provider/Kubernetes events and exact timestamps.
  • Cost-impact estimate and any authorized containment action.
  • Correction, canary evidence and rollback status.

Knowledge check

A job is pending and there are zero autoscaled instances. Does that prove max_instances is too low?

What is the first architectural correction when two Docker Autoscaler managers control the same provider group?

Why is deleting a failed ephemeral VM immediately poor diagnostics?

A job takes 20 minutes but queue duration is near zero. Is autoscaling the first suspect?

Why can a canary migration fail if old and new pools share identical broad tags from the start?

Next lesson

Continue to the checkpoint

Lesson 5 turns the chapter into an operator dossier: bounded sizing from a trace, burst and manager-failure injection, migration gates, rollback criteria and cleanup proof.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Current stable GitLab Runner patch verified for this chapter is 19.3.2 (tagged 2026-09-10). The Docker Machine executor was deprecated in GitLab 17.5 and is scheduled for removal as a supported feature in GitLab 20.0 (May 2027); GitLab directs users toward the Instance or Docker Autoscaler executors. Docker Autoscaler is GA and, together with the Instance executor, uses Taskscaler/Fleeting. A Docker Autoscaler configuration must have its own dedicated provider autoscaling resource and must not share that resource with another manager or another [[runners]] entry. The mandatory exercises in this chapter are local simulations and require no cloud account, runner registration token, privileged executor, or managed Kubernetes cluster. Runner 19.3.2 is used as the current-version evidence point; always check the exact GitLab/Runner versions in the affected environment before applying configuration advice.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.