Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling: Diagnostics, Failure Modes, and Production Practices
Ephemeral runners disappear quickly, which makes “delete the pod and retry” especially dangerous as a troubleshooting habit. This lesson keeps the failed run, queue state and cluster evidence intact long enough to locate the causal layer.
Learning objectives
- Apply a fixed evidence-first diagnostic sequence from event/ref/SHA through external target state.
- Recognize public/untrusted workload placement as a trust failure rather than a scheduling issue.
- Diagnose credential exposure, routing mismatch, missing logs and capacity bounds without blind reruns.
- Explain why pod deletion is not credential revocation and why unbounded autoscaling is an operational defect.
- Repair the smallest causal layer and rerun only the smallest equivalent scope.
1. Freeze evidence before touching the cluster
Record the workflow run ID, attempt, event/ref/SHA, workflow
revision, job ID, requested runs-on, queue timestamps
and current conclusion. Then snapshot read-only Kubernetes state:
Helm versions, AutoScalingRunnerSet/EphemeralRunnerSet resources,
pods, events, controller/listener logs and relevant quota/admission
status.
helm list -A
kubectl -n arc-runners get pods -o wide
kubectl -n arc-runners get events --sort-by=.lastTimestamp
kubectl -n arc-runners get autoscalingrunnersets,ephemeralrunnersets 2>/dev/null || true
kubectl -n arc-systems get pods -o wide
Do not begin with kubectl delete pod, Helm reinstall or
credential rotation unless the incident itself requires containment.
Those actions change the evidence.
2. Evidence-first diagnostic sequence
- Preserve run/attempt and first-failure evidence.
- Confirm event, ref, SHA and exact workflow revision.
- Confirm evaluated conditions, permissions and inputs.
- Inspect job graph, requested scale-set name/labels and queue state.
- Confirm listener demand, desired runner count and max bound.
- Inspect ARC controller reconciliation and Kubernetes events.
- Confirm runner pod image, service account, node, network and JIT registration.
- Inspect the failing step/action/runtime only after runner assignment is proven.
- Inspect artifacts/caches/deployments/external target separately.
- Apply the least destructive correction and rerun the smallest equivalent scope.
3. Failure: untrusted workflow beside production workloads
Symptom: scheduling succeeds; the job is green. The defect is architectural: unreviewed fork code executed in a runner pod that shares nodes/network/RBAC with production workloads.
Evidence: event is a fork PR or otherwise untrusted source; runner cluster has routes/service accounts that reach production. Repair: route untrusted jobs to GitHub-hosted or a separately isolated fleet, and restrict the trusted scale set by group/repository policy. This is not fixed by changing a label.
4. Failure: plaintext credential on a CLI
Symptom: a Helm command embeds a PAT or private key
in --set. Even if the install succeeds, shell
history/process inspection may have captured it.
Containment: preserve the command context without copying the secret value, revoke/rotate the credential, remove it from history if policy allows, and switch to a Kubernetes Secret referenced by name. Do not rely on GitHub log masking to protect local shell history.
5. Failure: assuming scale-set labels behave like classic host labels
Symptom: a workflow requests
[self-hosted, linux, gpu] copied from a classic runner
tutorial, but the ARC scale set was configured only by name or with
a different runnerScaleSetLabels set. The job remains
queued.
Evidence: compare workflow
runs-on with the scale-set name/labels shown by current
configuration. Repair: use the
installation/scale-set name directly or configure the documented
scale-set labels consistently. Do not add arbitrary labels to
ephemeral pods and assume GitHub knows about them.
6. Failure: the pod vanished with the useful logs
Symptom: a runner crashed, then Kubernetes deleted/recreated it. GitHub shows a job failure but not the controller/listener cause.
Repair: externalize controller, listener and
runner/container logs with run/job/pod correlation fields. For
immediate investigation, capture
kubectl logs --previous when available and Kubernetes
events before retention expires. Production logging is an
architecture requirement, not an afterthought.
7. Failure: unbounded scale under a burst
Symptom: many jobs arrive; ARC legitimately creates
more runners than budget/cluster policy expected because
maxRunners was omitted and no external quota existed.
Repair: set an explicit ceiling and cluster resource quota, then decide whether queue latency is acceptable. During maintenance, setting both min and max to zero is a documented drain mechanism. Do not “fix” cost by abruptly deleting running pods.
8. Failure: pod deletion mistaken for revocation
Symptom: an incident response deletes the runner pod but leaves a compromised external cloud credential, GitHub App installation or Kubernetes Secret valid.
Repair: enumerate credential owners and revoke/rotate each affected credential at its authoritative system. Kubernetes pod lifecycle is only one containment boundary.
9. Intentionally broken configuration
This example is inert teaching material; do not deploy it. It combines an unbounded scale policy, shared production namespace and privileged DinD.
# BROKEN — DO NOT DEPLOY
githubConfigUrl: "https://github.com/example/shared-org"
runnerScaleSetName: shared-ci
# maxRunners omitted
containerMode:
type: dind
template:
metadata:
namespace: production
There are three independent defects: trust/isolation, privilege and capacity. A successful job would not make this configuration acceptable.
10. Minimal repaired design
# Conceptual safer baseline
githubConfigUrl: "https://github.com/example/disposable-or-trusted-scope"
githubConfigSecret: arc-github-auth
runnerGroup: trusted-builds
runnerScaleSetName: arc-trusted-builds
minRunners: 0
maxRunners: 4
# No dind unless a reviewed workload actually requires it.
Place the scale set in a dedicated runner namespace/cluster with bounded RBAC/network policy and centralized logs. This does not eliminate all risk; it makes trust and capacity explicit.
11. Production runbook checkpoints
- Queue grows: verify scale-set match/listener before Kubernetes.
- Desired runners grow but pods do not: inspect controller/Kubernetes events/quota.
- Pods run but jobs do not: inspect JIT registration/network/GitHub connectivity.
- Jobs fail after assignment: inspect runner image/tool/action/runtime.
- Pods disappear too early: inspect external log pipeline and controller lifecycle.
- Incident involves credential exposure: revoke at credential owner; do not stop at pod deletion.
-
Capacity exceeds budget: bound
maxRunnersand quota; preserve queued/running work evidence before changing policy.
Knowledge check
Why is “delete the pod and rerun” a bad first diagnostic step?
It destroys or changes the first-failure evidence and may simply recreate the same causal defect.
A job stays queued but the workflow file is valid. What ARC-specific check comes before Kubernetes scheduling?
Confirm that runs-on matches the scale-set name/current labels and that the listener sees the job as available demand.
Does a green workflow prove a shared production runner cluster is safe?
No. Trust/isolation is an architecture property; successful scheduling/execution does not prove least privilege.
What should you do after discovering a PAT in shell history?
Treat it as exposed: preserve non-secret evidence, revoke/rotate it, remove local traces according to policy, and replace CLI literals with a referenced secret mechanism.
Why is maxRunners an incident-control setting?
It prevents a burst or malicious queue from scaling compute beyond an explicit ceiling, limiting cost and resource blast radius.
Official references and version notes
- GitHub Docs — Actions Runner Controller — control-plane architecture, listener/JIT registration and ephemeral runner lifecycle.
- GitHub Docs — Deploy runner scale sets — Helm deployment, runner groups, min/max runners, pod templates, container modes and security guidance.
-
GitHub Docs — Use ARC in a workflow
— scale-set names and scale-set labels in
runs-on. - GitHub Docs — Troubleshoot ARC — controller/listener/runner diagnostic workflow.
- ARC release 0.14.2 — pinned release used for concrete examples in this chapter.
- GitHub Actions runner container image — minimal runner image published with runner releases.
Version-sensitive behavior was rechecked against current
GitHub-maintained documentation and repositories on
2026-09-09. The latest public ARC
runner-scale-set release found during authoring is
0.14.2, published 2026-05-22. Its release notes
updated the bundled runner to v2.334.0; the runner binary has an
independent release cadence, so do not infer that the ARC chart
version and newest standalone runner version are identical.
Current docs prefer Helm for ARC deployment, recommend production
workload isolation and retained
controller/listener/ephemeral-runner logs, support bounded
minRunners/maxRunners, and document
scale-set names plus runnerScaleSetLabels for
routing. Current container modes include dind,
kubernetes, and kubernetes-novolume;
DinD is privileged, and Kubernetes mode changes the Kubernetes
API/service-account trust boundary.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.