Chapter 10Lesson 04~190 minutes

Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling: Diagnostics, Failure Modes, and Production Practices

Ephemeral runners disappear quickly, which makes “delete the pod and retry” especially dangerous as a troubleshooting habit. This lesson keeps the failed run, queue state and cluster evidence intact long enough to locate the causal layer.

DiagnosticsFirst failureSecurityScale boundsIncident response

Learning objectives

  • Apply a fixed evidence-first diagnostic sequence from event/ref/SHA through external target state.
  • Recognize public/untrusted workload placement as a trust failure rather than a scheduling issue.
  • Diagnose credential exposure, routing mismatch, missing logs and capacity bounds without blind reruns.
  • Explain why pod deletion is not credential revocation and why unbounded autoscaling is an operational defect.
  • Repair the smallest causal layer and rerun only the smallest equivalent scope.

1. Freeze evidence before touching the cluster

Record the workflow run ID, attempt, event/ref/SHA, workflow revision, job ID, requested runs-on, queue timestamps and current conclusion. Then snapshot read-only Kubernetes state: Helm versions, AutoScalingRunnerSet/EphemeralRunnerSet resources, pods, events, controller/listener logs and relevant quota/admission status.

helm list -A
kubectl -n arc-runners get pods -o wide
kubectl -n arc-runners get events --sort-by=.lastTimestamp
kubectl -n arc-runners get autoscalingrunnersets,ephemeralrunnersets 2>/dev/null || true
kubectl -n arc-systems get pods -o wide

Do not begin with kubectl delete pod, Helm reinstall or credential rotation unless the incident itself requires containment. Those actions change the evidence.

2. Evidence-first diagnostic sequence

  1. Preserve run/attempt and first-failure evidence.
  2. Confirm event, ref, SHA and exact workflow revision.
  3. Confirm evaluated conditions, permissions and inputs.
  4. Inspect job graph, requested scale-set name/labels and queue state.
  5. Confirm listener demand, desired runner count and max bound.
  6. Inspect ARC controller reconciliation and Kubernetes events.
  7. Confirm runner pod image, service account, node, network and JIT registration.
  8. Inspect the failing step/action/runtime only after runner assignment is proven.
  9. Inspect artifacts/caches/deployments/external target separately.
  10. Apply the least destructive correction and rerun the smallest equivalent scope.

3. Failure: untrusted workflow beside production workloads

Symptom: scheduling succeeds; the job is green. The defect is architectural: unreviewed fork code executed in a runner pod that shares nodes/network/RBAC with production workloads.

Evidence: event is a fork PR or otherwise untrusted source; runner cluster has routes/service accounts that reach production. Repair: route untrusted jobs to GitHub-hosted or a separately isolated fleet, and restrict the trusted scale set by group/repository policy. This is not fixed by changing a label.

4. Failure: plaintext credential on a CLI

Symptom: a Helm command embeds a PAT or private key in --set. Even if the install succeeds, shell history/process inspection may have captured it.

Containment: preserve the command context without copying the secret value, revoke/rotate the credential, remove it from history if policy allows, and switch to a Kubernetes Secret referenced by name. Do not rely on GitHub log masking to protect local shell history.

5. Failure: assuming scale-set labels behave like classic host labels

Symptom: a workflow requests [self-hosted, linux, gpu] copied from a classic runner tutorial, but the ARC scale set was configured only by name or with a different runnerScaleSetLabels set. The job remains queued.

Evidence: compare workflow runs-on with the scale-set name/labels shown by current configuration. Repair: use the installation/scale-set name directly or configure the documented scale-set labels consistently. Do not add arbitrary labels to ephemeral pods and assume GitHub knows about them.

6. Failure: the pod vanished with the useful logs

Symptom: a runner crashed, then Kubernetes deleted/recreated it. GitHub shows a job failure but not the controller/listener cause.

Repair: externalize controller, listener and runner/container logs with run/job/pod correlation fields. For immediate investigation, capture kubectl logs --previous when available and Kubernetes events before retention expires. Production logging is an architecture requirement, not an afterthought.

7. Failure: unbounded scale under a burst

Symptom: many jobs arrive; ARC legitimately creates more runners than budget/cluster policy expected because maxRunners was omitted and no external quota existed.

Repair: set an explicit ceiling and cluster resource quota, then decide whether queue latency is acceptable. During maintenance, setting both min and max to zero is a documented drain mechanism. Do not “fix” cost by abruptly deleting running pods.

8. Failure: pod deletion mistaken for revocation

Symptom: an incident response deletes the runner pod but leaves a compromised external cloud credential, GitHub App installation or Kubernetes Secret valid.

Repair: enumerate credential owners and revoke/rotate each affected credential at its authoritative system. Kubernetes pod lifecycle is only one containment boundary.

9. Intentionally broken configuration

This example is inert teaching material; do not deploy it. It combines an unbounded scale policy, shared production namespace and privileged DinD.

# BROKEN — DO NOT DEPLOY
githubConfigUrl: "https://github.com/example/shared-org"
runnerScaleSetName: shared-ci
# maxRunners omitted
containerMode:
  type: dind
template:
  metadata:
    namespace: production

There are three independent defects: trust/isolation, privilege and capacity. A successful job would not make this configuration acceptable.

10. Minimal repaired design

# Conceptual safer baseline
githubConfigUrl: "https://github.com/example/disposable-or-trusted-scope"
githubConfigSecret: arc-github-auth
runnerGroup: trusted-builds
runnerScaleSetName: arc-trusted-builds
minRunners: 0
maxRunners: 4
# No dind unless a reviewed workload actually requires it.

Place the scale set in a dedicated runner namespace/cluster with bounded RBAC/network policy and centralized logs. This does not eliminate all risk; it makes trust and capacity explicit.

11. Production runbook checkpoints

  • Queue grows: verify scale-set match/listener before Kubernetes.
  • Desired runners grow but pods do not: inspect controller/Kubernetes events/quota.
  • Pods run but jobs do not: inspect JIT registration/network/GitHub connectivity.
  • Jobs fail after assignment: inspect runner image/tool/action/runtime.
  • Pods disappear too early: inspect external log pipeline and controller lifecycle.
  • Incident involves credential exposure: revoke at credential owner; do not stop at pod deletion.
  • Capacity exceeds budget: bound maxRunners and quota; preserve queued/running work evidence before changing policy.

Knowledge check

Why is “delete the pod and rerun” a bad first diagnostic step?

A job stays queued but the workflow file is valid. What ARC-specific check comes before Kubernetes scheduling?

Does a green workflow prove a shared production runner cluster is safe?

What should you do after discovering a PAT in shell history?

Why is maxRunners an incident-control setting?

Next chapter concept

Checkpoint: prove one ephemeral runner lifecycle

Lesson 5 packages the chapter into a security/operations dossier with predictions, a deterministic simulation or optional live scale set, lifecycle evidence, cleanup and rollback.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against current GitHub-maintained documentation and repositories on 2026-09-09. The latest public ARC runner-scale-set release found during authoring is 0.14.2, published 2026-05-22. Its release notes updated the bundled runner to v2.334.0; the runner binary has an independent release cadence, so do not infer that the ARC chart version and newest standalone runner version are identical. Current docs prefer Helm for ARC deployment, recommend production workload isolation and retained controller/listener/ephemeral-runner logs, support bounded minRunners/maxRunners, and document scale-set names plus runnerScaleSetLabels for routing. Current container modes include dind, kubernetes, and kubernetes-novolume; DinD is privileged, and Kubernetes mode changes the Kubernetes API/service-account trust boundary.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.