Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling: Core Concepts and Mental Model
Chapter 09 treated one self-hosted runner as privileged infrastructure. Chapter 10 asks what changes when that infrastructure becomes a fleet: GitHub job demand must be translated into short-lived Kubernetes runner pods without turning autoscaling into an unbounded or unauditable privilege system.
Learning objectives
- Explain the end-to-end path from a queued GitHub job to listener acknowledgement, JIT runner registration, pod execution and deletion.
- Distinguish GitHub-owned queue/scale-set state from Kubernetes controller, namespace, service-account, secret and pod state.
- Record the exact ARC/chart/controller/listener versions, scale-set scope/name/group and min/max capacity before execution.
- Explain why ephemeral runner deletion reduces persistence risk but does not replace credential revocation, log retention or cluster isolation.
- Recognize the production security boundary created by a shared Kubernetes cluster and by container modes such as DinD and Kubernetes mode.
1. The problem is not “how do I install a Helm chart?”
A single persistent self-hosted runner has an obvious host to patch, monitor and clean. Autoscaling changes the operational shape. A GitHub job can appear at any moment, a listener must observe demand, Kubernetes must create enough runner capacity, each runner must register just in time, execute exactly one bounded job, send evidence back, and disappear. Meanwhile the cluster, controller credentials, service accounts and network routes persist.
The useful mental shift is: ARC is a control plane that converts trusted GitHub job demand into ephemeral compute. Helm is only one packaging mechanism for that control plane.
2. Causal model: demand → listener → pod → evidence → deletion
Every arrow below changes a different state. A production incident is diagnosable only when those states are kept separate.
flowchart TD A[Workflow job becomes eligible] --> B[GitHub Actions scale-set service] B --> C[Scale-set listener long-polls for demand] C --> D[Listener patches desired EphemeralRunnerSet replicas] D --> E[ARC controller requests JIT runner configuration] E --> F[Kubernetes creates ephemeral runner pod] F --> G[Runner registers and accepts one job] G --> H[Job logs and status return to GitHub] H --> I[Runner deregisters and pod is deleted] C --> J[Controller / listener logs and metrics] F --> J
GitHub owns the queued job, scale-set registration and job assignment. The listener observes available work and asks Kubernetes for capacity. The ARC controllers reconcile desired resources. Kubernetes owns pod scheduling and service-account/network policy. The runner process executes workflow code and streams job state back. Deleting the runner pod ends that compute instance, but controller credentials, Kubernetes secrets, external logs and downstream credentials have separate lifetimes.
3. Name every control-plane object before changing it
| Object/state | Owner | Evidence to record | Why it matters |
|---|---|---|---|
| ARC chart/controller | Kubernetes / platform team | chart version, image tag/digest, namespace | Defines reconciliation behavior and compatibility. |
| Runner scale set | GitHub Actions service | scope URL, scale-set name, runner group, labels | Determines which jobs can target the fleet. |
| Listener | ARC + GitHub | pod UID, logs, connection health, desired runner count | Bridges queued demand to Kubernetes capacity. |
| Ephemeral runner pod | Kubernetes | pod UID/node, service account, image, start/finish | Actual execution boundary for arbitrary workflow code. |
| GitHub auth secret | Kubernetes + credential owner | secret name/owner/rotation date, never value | Allows ARC to manage scale-set registration. |
| Job | GitHub Actions | event/ref/SHA, run ID/attempt, job ID, conclusion | Demand identity must remain tied to source revision. |
| External telemetry | Operations system | sink name, retention, correlation fields | Ephemeral pods may disappear before investigation starts. |
4. Read-only preflight: prove what exists
Before installing anything, inspect the local toolchain and current cluster context. The commands below do not create or mutate resources.
kubectl config current-context
kubectl version --client
helm version
kubectl get namespaces
kubectl get crd | grep -E 'actions.github.com|autoscaling' || true
helm list -A
kubectl get pods -A
For a real ARC environment, add read-only inventory of the
controller namespace and scale-set namespace. Record only secret
names and metadata; never run
kubectl get secret ... -o yaml into a course log or
ticket.
5. GitHub state and Kubernetes state are different clocks
A job can be queued in GitHub while Kubernetes has zero runner pods. A listener can be healthy while pod scheduling is blocked by quota. A runner pod can be running while JIT registration fails. A Kubernetes pod can terminate while the GitHub job is still finalizing logs. The correct diagnosis starts by asking which clock stopped advancing.
| Observed symptom | First state to inspect | Do not assume |
|---|---|---|
| Job waiting for runner | scale-set match + listener demand | that Kubernetes is broken |
| Desired replicas increased, no pod | Kubernetes events/quota/scheduler | that GitHub did not dispatch |
| Pod running, job never starts | runner registration/JIT/network | that pod Running means runner Online |
| Job finished, pod remains | EphemeralRunner reconciliation | that cleanup is instantaneous |
| Pod vanished, no incident evidence | external log pipeline | that GitHub job logs contain controller diagnostics |
6. Autoscaling is bounded capacity management
Current ARC charts expose minRunners and
maxRunners. With both omitted, capacity can grow with
assigned demand and scale back to zero. That is useful for
elasticity but not a sensible default for a beginner production
design because cost, cluster quota and blast radius become implicit.
runnerScaleSetName: arc-course-lab
minRunners: 0
maxRunners: 2
This lab bound means “keep no idle runners, and never create more than two runners.” Setting both values to zero is also a documented drain technique: new runner pods are not created while maintenance is in progress.
7. Ephemeral does not mean anonymous
Each runner is registered with a JIT configuration, receives a specific job, streams status, and is then removed. That short lifetime is a strong isolation primitive because normal job state is not intentionally reused. But the job still inherits the pod service account, network routes, mounted volumes, container runtime privileges and any injected secrets. Those are the meaningful security inputs.
Use the scale-set name or current scale-set labels for routing. Do not confuse routing metadata with authorization: the runner group and GitHub configuration scope decide which repositories or organizations may use the set, while Kubernetes RBAC/network policy constrain what a runner pod can reach.
8. Container mode changes the privilege model
Current ARC supports Docker-in-Docker and Kubernetes-based container modes. DinD provides familiar Docker semantics but uses a privileged Docker daemon. Kubernetes mode uses runner container hooks and the Kubernetes API to create workload pods. The latter can avoid a Docker daemon, but it grants the runner-side mechanism Kubernetes permissions and introduces storage/service-account design.
Do not treat either mode as “just another build feature.” A privileged DinD sidecar and a Kubernetes service account able to create pods are both high-impact capabilities. Production runner clusters should be isolated from unrelated production workloads, with dedicated namespaces, bounded RBAC, network policy and external logs.
9. Version pinning is part of the execution input
At authoring time the latest public runner-scale-set release is 0.14.2. Pin the chart version in examples and record the controller image version/digest after installation. The ARC release notes and the standalone Actions runner can move independently; a newer runner binary does not imply a newer ARC chart, and vice versa.
helm list -A
kubectl -n arc-systems get deploy,pods -o wide
kubectl -n arc-runners get pods -o wide
For production, add image digests, CRD versions, Kubernetes server version and the runner container image digest to the evidence packet.
10. Common wrong models
- “ARC autoscaling is serverless.” The runner pods are ephemeral, but the Kubernetes cluster, controllers, listener credentials, RBAC and network remain your infrastructure.
- “Pod deletion revokes everything.” External credentials, GitHub App installations and leaked tokens have separate revocation paths.
- “A label is an isolation boundary.” Routing labels do not replace runner-group access or cluster isolation.
- “No idle runners means no cost.” The cluster/control plane, controller/listener and telemetry still consume resources.
- “The GitHub job log is enough.” Controller, listener, scheduler and pod events can disappear independently.
Knowledge check
Who converts available GitHub job demand into a desired number of runner replicas?
The ARC scale-set listener observes demand and updates the EphemeralRunnerSet desired replicas through the Kubernetes API; controllers then reconcile runner pods.
Does a Kubernetes pod in Running state prove that the GitHub job was assigned?
No. Pod scheduling, runner JIT registration and GitHub job assignment are separate states.
What does maxRunners protect?
It bounds how many runners the scale set may create, limiting capacity, cost and blast radius. It does not by itself authorize repositories or secure the cluster.
Why must ephemeral runner logs be externalized?
The pod may be deleted immediately after the job; controller/listener/runner diagnostics can otherwise disappear before an incident is investigated.
Is pod deletion sufficient credential revocation?
No. Any external credential or persistent GitHub/Kubernetes secret has its own lifetime and must be rotated or revoked separately when required.
Official references and version notes
- GitHub Docs — Actions Runner Controller — control-plane architecture, listener/JIT registration and ephemeral runner lifecycle.
- GitHub Docs — Deploy runner scale sets — Helm deployment, runner groups, min/max runners, pod templates, container modes and security guidance.
-
GitHub Docs — Use ARC in a workflow
— scale-set names and scale-set labels in
runs-on. - GitHub Docs — Troubleshoot ARC — controller/listener/runner diagnostic workflow.
- ARC release 0.14.2 — pinned release used for concrete examples in this chapter.
- GitHub Actions runner container image — minimal runner image published with runner releases.
Version-sensitive behavior was rechecked against current
GitHub-maintained documentation and repositories on
2026-09-09. The latest public ARC
runner-scale-set release found during authoring is
0.14.2, published 2026-05-22. Its release notes
updated the bundled runner to v2.334.0; the runner binary has an
independent release cadence, so do not infer that the ARC chart
version and newest standalone runner version are identical.
Current docs prefer Helm for ARC deployment, recommend production
workload isolation and retained
controller/listener/ephemeral-runner logs, support bounded
minRunners/maxRunners, and document
scale-set names plus runnerScaleSetLabels for
routing. Current container modes include dind,
kubernetes, and kubernetes-novolume;
DinD is privileged, and Kubernetes mode changes the Kubernetes
API/service-account trust boundary.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.