Runner Fleets, Autoscaling, Docker Machine Migration, Kubernetes Runners, Ephemeral Workers, and Capacity Planning: Diagnostics, Failure Modes, Security, and Performance
Diagnose cost runaway, shared autoscaling ownership, stale Runner versions, lost ephemeral logs, long-poll bottlenecks, and retained untrusted state with an evidence-first sequence and bounded corrections.
Learning objectives
- Use an evidence-first diagnostic sequence from pipeline identity through manager, worker, provider and governance state.
- Diagnose unbounded cost, conflicting autoscaler managers, stale Runner versions, lost logs and retained worker residue.
- Differentiate queue, polling, executor, provider/Kubernetes and script failures.
- Apply the least destructive correction while preserving original evidence.
- Recover capacity without broad credentials, privileged shortcuts or unbounded scale.
1. Evidence-first sequence
Before changing anything, preserve the first-failure evidence.
Record CI_PIPELINE_SOURCE, ref,
CI_COMMIT_SHA, pipeline/job IDs and job tags. Confirm
merged configuration and rule result. Check whether the job is
created, pending, assigned or already executing. Then record runner
manager ID/version,
concurrent/limit/request_concurrency, executor,
worker identity/image and queue/provisioning metrics. Only after
that inspect provider/Kubernetes state, job script/tool/network,
artifacts/caches and teardown evidence.
2. Failure: unbounded autoscaling cost
Symptom: a burst or retry storm drives machine
count upward faster than expected. Preserve: job
arrival IDs, retry causes, machine-state metrics, creation rate,
provider group events and current autoscaler config.
Cause candidates: max_instances=0,
excessive capacity_per_instance assumptions, retry
loops, duplicate pipelines, broken teardown, or a provider group
whose own scaling policy competes with Runner.
3. Failure: two managers control the same autoscaling group
Symptom: unexpected instance creation/removal, jobs losing workers, inconsistent manager views, cost spikes. Docker Autoscaler documentation explicitly forbids sharing one autoscaling resource across multiple managers or multiple runner configurations. The correction is architectural: assign a dedicated provider group to each autoscaler configuration. Do not attempt to “synchronize” two controllers with scripts.
4. Failure: stale Runner version
Symptom: old bugs, missing metrics/features,
manager/helper mismatch or behavior that no longer matches current
GitLab guidance. Capture gitlab-runner --version,
GitLab version, helper image version, executor/plugin versions and
the first failing job before upgrade. GitLab troubleshooting
recommends checking Runner and GitLab versions early. Upgrade by
canary/drain, not by replacing every manager simultaneously.
5. Failure: ephemeral worker disappears with the evidence
Symptom: a failed VM/pod is destroyed and only the GitLab job trace remains; provider/bootstrap/kernel or manager-side evidence is gone. Design logging before incidents: centralize runner-manager logs, provider autoscaler events and Kubernetes events where authorized; retain correlation fields such as runner ID, job ID, instance/pod ID and timestamps. Avoid placing secrets in logs.
# Read-only examples on an authorized manager; choose the command matching your deployment.
journalctl --unit=gitlab-runner.service -n 100 --no-pager
# containerized manager:
docker logs --tail 100 gitlab-runner
# Kubernetes manager pod:
kubectl logs -n gitlab-runner deploy/gitlab-runner --tail=100
These commands inspect manager logs. They do not recover logs that were never shipped from a destroyed worker.
6. Failure: reused worker retains untrusted state
Symptom: a later job can observe workspace files, Docker state, caches, sockets, services, credentials or devices left by an earlier job. Preserve the worker/job mapping and prove the residue in a synthetic lab rather than exposing real secrets. The strongest correction for mixed/untrusted workloads is often a trust-boundary change—single-use ephemeral worker or dedicated pool—not a longer cleanup script.
7. Intentionally broken configuration: infinite cost ceiling
concurrent = 200
[[runners]]
name = "broken-autoscaler"
executor = "docker-autoscaler"
limit = 200
request_concurrency = 1
[runners.autoscaler]
plugin = "aws"
capacity_per_instance = 1
max_use_count = 1
max_instances = 0
Interpretation: max_instances=0 is unlimited,
concurrent/limit allow very high demand,
and request concurrency 1 can simultaneously throttle job
acquisition. The configuration can be both
too permissive for cost and
too restrictive for polling. Fix each control for its
actual purpose.
concurrent = 24
[[runners]]
name = "bounded-autoscaler"
executor = "docker-autoscaler"
limit = 12
request_concurrency = 4
[runners.autoscaler]
plugin = "aws"
capacity_per_instance = 1
max_use_count = 1
max_instances = 12
The numbers are examples only. Production values must come from queue/provisioning/cost evidence and provider quotas.
8. Pending jobs: a causal matrix
| Observation | Likely layer | Evidence before change |
|---|---|---|
| No eligible runner shown / tag mismatch | Eligibility/config | Job tags, runner tags/scope/protection, rule result. |
| Eligible manager, request concurrency saturated | Manager polling | request metrics, long-poll warnings, manager logs. |
| Manager accepted job, waits for instance | Autoscaler/provider | machine states, creation duration/errors, provider quota/events. |
| Kubernetes pod exists but Pending | Kubernetes scheduler/node capacity | pod events, requests/limits, quota, taints, node-autoscaler state. |
| Worker starts but script stalls | Job/tool/network | job trace, process/network/tool logs; not an autoscaling change. |
9. Security-sensitive actions
Runner registration/authentication tokens, cloud IAM credentials, Kubernetes RBAC, privileged Docker/Kubernetes settings, device exposure, provider-group mutation and broad network access are control-plane changes. Use fake placeholders in documentation and narrowly scoped identities in real systems. Never print tokens, disable TLS, run untrusted jobs on a privileged shared runner, or grant broad cloud administrator roles as a troubleshooting shortcut.
10. Performance diagnosis without over-scaling
Track queue duration and job duration separately. A large queue with stable job duration indicates capacity/acquisition pressure. A stable queue with growing job duration points toward job/tool/cache/network performance. Machine creation duration isolates the cold-start path. Error counters and provider throttling can explain why requested capacity never becomes usable.
11. Docker Machine migration failure: old and new pools overlap incorrectly
If both pools use the same broad tags during canary, you lose control over which executor ran the job and cannot compare results. Use explicit canary tags and record runner ID/executor for each representative job. When shifting traffic, change one eligibility dimension at a time and keep the legacy pool paused—not deleted—until rollback criteria expire.
12. Upgrade rollback
If a canary Runner upgrade causes failures, preserve manager logs and representative job IDs, pause the canary from new work, let safe in-flight jobs finish, and revert to the recorded prior package/image/chart/config. Do not rebuild an old manager from memory. The rollback artifact is part of fleet governance.
13. Minimal incident packet
- Pipeline source/ref/SHA, pipeline/job IDs and job tags.
- Runner ID, manager hostname/pod, Runner version, executor and plugin/image versions.
-
concurrent,limit,request_concurrencyand autoscaler hard bounds. - Queue, job duration, request concurrency, machine/pod creation and error metrics.
- Worker/pod/instance identity and lifetime/use count.
- Manager/provider/Kubernetes events and exact timestamps.
- Cost-impact estimate and any authorized containment action.
- Correction, canary evidence and rollback status.
Knowledge check
A job is pending and there are zero autoscaled instances. Does
that prove max_instances is too low?
No. First prove eligibility, manager polling, concurrent/limit/request settings and provider errors. The runner may never have acquired or been able to provision the job.
What is the first architectural correction when two Docker Autoscaler managers control the same provider group?
Separate ownership: give each autoscaler configuration its own dedicated provider autoscaling resource.
Why is deleting a failed ephemeral VM immediately poor diagnostics?
It can destroy provider/worker evidence before you preserve the first failure and correlate it to a job/manager. Teardown should occur after evidence capture or through established centralized logging.
A job takes 20 minutes but queue duration is near zero. Is autoscaling the first suspect?
No. The job was scheduled promptly. Inspect script/tool/network/cache performance and runner resource constraints before adding fleet capacity.
Why can a canary migration fail if old and new pools share identical broad tags from the start?
You cannot reliably attribute jobs/results to one executor, so comparison and rollback evidence become ambiguous.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12.
Current stable GitLab Runner patch verified for this chapter is
19.3.2 (tagged 2026-09-10). The Docker Machine executor
was deprecated in GitLab 17.5 and is scheduled for removal as a
supported feature in GitLab 20.0 (May 2027); GitLab directs users
toward the Instance or Docker Autoscaler executors. Docker
Autoscaler is GA and, together with the Instance executor, uses
Taskscaler/Fleeting. A Docker Autoscaler configuration must have its
own dedicated provider autoscaling resource and must not share that
resource with another manager or another
[[runners]] entry. The mandatory exercises in this
chapter are local simulations and require no cloud account, runner
registration token, privileged executor, or managed Kubernetes
cluster. Runner 19.3.2 is used as the current-version evidence
point; always check the exact GitLab/Runner versions in the affected
environment before applying configuration advice.
- Executors — official reference.
- GitLab Runner autoscaling — official reference.
- Docker Autoscaler executor — official reference.
- Instance executor — official reference.
- Fleeting — official reference.
- Plan and operate a runner fleet — official reference.
- Advanced Runner configuration — official reference.
- Monitor GitLab Runner usage — official reference.
- Kubernetes executor — official reference.
- GitLab Runner Helm chart — official reference.
- Docker Machine executor — official reference.
- GitLab deprecations and removals — official reference.
- Runner troubleshooting — official reference.
- GitLab Runner changelog — official reference.
- GitLab Runner tags — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.