Chapter 10Lesson 01~165 minutes

Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling: Core Concepts and Mental Model

Chapter 09 treated one self-hosted runner as privileged infrastructure. Chapter 10 asks what changes when that infrastructure becomes a fleet: GitHub job demand must be translated into short-lived Kubernetes runner pods without turning autoscaling into an unbounded or unauditable privilege system.

ARCRunner scale setsEphemeral runnersKubernetesControl plane

Learning objectives

  • Explain the end-to-end path from a queued GitHub job to listener acknowledgement, JIT runner registration, pod execution and deletion.
  • Distinguish GitHub-owned queue/scale-set state from Kubernetes controller, namespace, service-account, secret and pod state.
  • Record the exact ARC/chart/controller/listener versions, scale-set scope/name/group and min/max capacity before execution.
  • Explain why ephemeral runner deletion reduces persistence risk but does not replace credential revocation, log retention or cluster isolation.
  • Recognize the production security boundary created by a shared Kubernetes cluster and by container modes such as DinD and Kubernetes mode.

1. The problem is not “how do I install a Helm chart?”

A single persistent self-hosted runner has an obvious host to patch, monitor and clean. Autoscaling changes the operational shape. A GitHub job can appear at any moment, a listener must observe demand, Kubernetes must create enough runner capacity, each runner must register just in time, execute exactly one bounded job, send evidence back, and disappear. Meanwhile the cluster, controller credentials, service accounts and network routes persist.

The useful mental shift is: ARC is a control plane that converts trusted GitHub job demand into ephemeral compute. Helm is only one packaging mechanism for that control plane.

2. Causal model: demand → listener → pod → evidence → deletion

Every arrow below changes a different state. A production incident is diagnosable only when those states are kept separate.

ARC runner-scale-set control and workload lifecycle
flowchart TD
  A[Workflow job becomes eligible] --> B[GitHub Actions scale-set service]
  B --> C[Scale-set listener long-polls for demand]
  C --> D[Listener patches desired EphemeralRunnerSet replicas]
  D --> E[ARC controller requests JIT runner configuration]
  E --> F[Kubernetes creates ephemeral runner pod]
  F --> G[Runner registers and accepts one job]
  G --> H[Job logs and status return to GitHub]
  H --> I[Runner deregisters and pod is deleted]
  C --> J[Controller / listener logs and metrics]
  F --> J

GitHub owns the queued job, scale-set registration and job assignment. The listener observes available work and asks Kubernetes for capacity. The ARC controllers reconcile desired resources. Kubernetes owns pod scheduling and service-account/network policy. The runner process executes workflow code and streams job state back. Deleting the runner pod ends that compute instance, but controller credentials, Kubernetes secrets, external logs and downstream credentials have separate lifetimes.

3. Name every control-plane object before changing it

Object/state Owner Evidence to record Why it matters
ARC chart/controller Kubernetes / platform team chart version, image tag/digest, namespace Defines reconciliation behavior and compatibility.
Runner scale set GitHub Actions service scope URL, scale-set name, runner group, labels Determines which jobs can target the fleet.
Listener ARC + GitHub pod UID, logs, connection health, desired runner count Bridges queued demand to Kubernetes capacity.
Ephemeral runner pod Kubernetes pod UID/node, service account, image, start/finish Actual execution boundary for arbitrary workflow code.
GitHub auth secret Kubernetes + credential owner secret name/owner/rotation date, never value Allows ARC to manage scale-set registration.
Job GitHub Actions event/ref/SHA, run ID/attempt, job ID, conclusion Demand identity must remain tied to source revision.
External telemetry Operations system sink name, retention, correlation fields Ephemeral pods may disappear before investigation starts.

4. Read-only preflight: prove what exists

Before installing anything, inspect the local toolchain and current cluster context. The commands below do not create or mutate resources.

kubectl config current-context
kubectl version --client
helm version
kubectl get namespaces
kubectl get crd | grep -E 'actions.github.com|autoscaling' || true
helm list -A
kubectl get pods -A

For a real ARC environment, add read-only inventory of the controller namespace and scale-set namespace. Record only secret names and metadata; never run kubectl get secret ... -o yaml into a course log or ticket.

5. GitHub state and Kubernetes state are different clocks

A job can be queued in GitHub while Kubernetes has zero runner pods. A listener can be healthy while pod scheduling is blocked by quota. A runner pod can be running while JIT registration fails. A Kubernetes pod can terminate while the GitHub job is still finalizing logs. The correct diagnosis starts by asking which clock stopped advancing.

Observed symptom First state to inspect Do not assume
Job waiting for runner scale-set match + listener demand that Kubernetes is broken
Desired replicas increased, no pod Kubernetes events/quota/scheduler that GitHub did not dispatch
Pod running, job never starts runner registration/JIT/network that pod Running means runner Online
Job finished, pod remains EphemeralRunner reconciliation that cleanup is instantaneous
Pod vanished, no incident evidence external log pipeline that GitHub job logs contain controller diagnostics

6. Autoscaling is bounded capacity management

Current ARC charts expose minRunners and maxRunners. With both omitted, capacity can grow with assigned demand and scale back to zero. That is useful for elasticity but not a sensible default for a beginner production design because cost, cluster quota and blast radius become implicit.

runnerScaleSetName: arc-course-lab
minRunners: 0
maxRunners: 2

This lab bound means “keep no idle runners, and never create more than two runners.” Setting both values to zero is also a documented drain technique: new runner pods are not created while maintenance is in progress.

7. Ephemeral does not mean anonymous

Each runner is registered with a JIT configuration, receives a specific job, streams status, and is then removed. That short lifetime is a strong isolation primitive because normal job state is not intentionally reused. But the job still inherits the pod service account, network routes, mounted volumes, container runtime privileges and any injected secrets. Those are the meaningful security inputs.

Use the scale-set name or current scale-set labels for routing. Do not confuse routing metadata with authorization: the runner group and GitHub configuration scope decide which repositories or organizations may use the set, while Kubernetes RBAC/network policy constrain what a runner pod can reach.

8. Container mode changes the privilege model

Current ARC supports Docker-in-Docker and Kubernetes-based container modes. DinD provides familiar Docker semantics but uses a privileged Docker daemon. Kubernetes mode uses runner container hooks and the Kubernetes API to create workload pods. The latter can avoid a Docker daemon, but it grants the runner-side mechanism Kubernetes permissions and introduces storage/service-account design.

Security boundary

Do not treat either mode as “just another build feature.” A privileged DinD sidecar and a Kubernetes service account able to create pods are both high-impact capabilities. Production runner clusters should be isolated from unrelated production workloads, with dedicated namespaces, bounded RBAC, network policy and external logs.

9. Version pinning is part of the execution input

At authoring time the latest public runner-scale-set release is 0.14.2. Pin the chart version in examples and record the controller image version/digest after installation. The ARC release notes and the standalone Actions runner can move independently; a newer runner binary does not imply a newer ARC chart, and vice versa.

helm list -A
kubectl -n arc-systems get deploy,pods -o wide
kubectl -n arc-runners get pods -o wide

For production, add image digests, CRD versions, Kubernetes server version and the runner container image digest to the evidence packet.

10. Common wrong models

  • “ARC autoscaling is serverless.” The runner pods are ephemeral, but the Kubernetes cluster, controllers, listener credentials, RBAC and network remain your infrastructure.
  • “Pod deletion revokes everything.” External credentials, GitHub App installations and leaked tokens have separate revocation paths.
  • “A label is an isolation boundary.” Routing labels do not replace runner-group access or cluster isolation.
  • “No idle runners means no cost.” The cluster/control plane, controller/listener and telemetry still consume resources.
  • “The GitHub job log is enough.” Controller, listener, scheduler and pod events can disappear independently.

Knowledge check

Who converts available GitHub job demand into a desired number of runner replicas?

Does a Kubernetes pod in Running state prove that the GitHub job was assigned?

What does maxRunners protect?

Why must ephemeral runner logs be externalized?

Is pod deletion sufficient credential revocation?

Next chapter concept

From the mental model to a disposable lab

Lesson 2 turns the lifecycle into observable evidence with a mandatory local simulation and an optional pinned ARC deployment in a disposable Kubernetes/GitHub environment.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against current GitHub-maintained documentation and repositories on 2026-09-09. The latest public ARC runner-scale-set release found during authoring is 0.14.2, published 2026-05-22. Its release notes updated the bundled runner to v2.334.0; the runner binary has an independent release cadence, so do not infer that the ARC chart version and newest standalone runner version are identical. Current docs prefer Helm for ARC deployment, recommend production workload isolation and retained controller/listener/ephemeral-runner logs, support bounded minRunners/maxRunners, and document scale-set names plus runnerScaleSetLabels for routing. Current container modes include dind, kubernetes, and kubernetes-novolume; DinD is privileged, and Kubernetes mode changes the Kubernetes API/service-account trust boundary.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.