Chapter 10Lesson 03~175 minutes

Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling: Configuration, Design Patterns, and Trade-Offs

Autoscaling runner platforms fail when teams pick a topology from a feature checklist instead of from workload trust and operational constraints. This lesson turns the control-plane model into explicit architecture decisions.

Design trade-offsIsolationDinDKubernetes modeCapacity

Learning objectives

  • Compare GitHub-hosted runners with ARC-managed self-hosted scale sets using ownership and trust boundaries.
  • Choose persistent versus ephemeral capacity and understand cold-start versus residual-state trade-offs.
  • Compare DinD, Kubernetes and no-volume Kubernetes container modes without hiding their privilege implications.
  • Separate runner-group authorization, scale-set routing labels and Kubernetes isolation controls.
  • Justify ARC versus a custom Runner Scale Set Client with observable operational evidence.

1. Start with workload trust, not Kubernetes enthusiasm

Ask five questions before choosing ARC: Who authors the workflow? Can pull requests contain untrusted code? What internal network/data must jobs reach? How quickly must capacity appear? Who will patch and monitor the Kubernetes control plane? If the answer to the last question is “nobody,” GitHub-hosted runners are usually the safer default.

2. GitHub-hosted versus ARC-managed runners

Dimension GitHub-hosted ARC scale set
Host lifecycle GitHub provisions/tears down You operate Kubernetes/controller/pod images
Network reach Public GitHub-hosted networking unless special features Can reach private networks you expose to cluster
Patch burden Mostly GitHub Cluster, ARC, images, policies, secrets are yours
Isolation Fresh GitHub-managed environment per job Ephemeral pods, but cluster/control-plane persists
Customization Image/tool constraints Custom images, pod specs, hardware/node pools
Evidence runner/image metadata + GitHub logs GitHub evidence plus Kubernetes/controller/listener telemetry

3. Persistent fleet versus ephemeral scale set

Persistent runners reduce cold start and can keep expensive tool caches locally, but they create cross-job residual-state risk and require explicit cleanup. Ephemeral ARC runners are designed for one-job lifetimes, making state reuse less likely and autoscaling simpler. The trade-off is startup latency and dependence on external artifact/cache services.

Production default for mixed workloads: prefer ephemeral runners and move durable outputs to explicit artifacts/caches rather than relying on runner disk.

4. DinD versus Kubernetes mode

Choice Benefit Security/operations cost Use when
No container mode simplest runner pod no container-job/container-action support via ARC mode jobs execute directly in runner image
dind familiar Docker semantics privileged Docker daemon; larger attack surface Docker-native workloads justify privilege in isolated cluster
kubernetes container jobs become Kubernetes pods runner hooks need Kubernetes API permissions and storage design cluster-native isolation/governance is available
kubernetes-novolume avoids shared RWX volume dependency lifecycle-hook semantics require validation cluster lacks suitable shared storage
Do not disable safeguards casually

Current Kubernetes mode can require jobs to declare a job container. Disabling that requirement can expose the runner pod’s Kubernetes API privileges, including pod creation and secret access. Treat that change as a security architecture decision, not a compatibility fix.

5. Shared cluster versus isolated runner cluster

GitHub explicitly recommends isolating production ARC workloads. Workflows execute arbitrary code. If runner pods share a cluster with production services, a container escape, over-broad service account, node compromise, insecure storage class or permissive network policy can turn CI code into production access.

A strong design uses a dedicated runner cluster or at minimum dedicated nodes/namespaces, default-deny network policy, bounded RBAC, no production service-account mounts, admission policy, resource quotas and centralized logs. Labels alone do none of this.

6. Runner group, scale-set name and labels solve different problems

Runner groups control which repositories/organizations may access runner resources. A scale-set name is a routable target and can be used directly in runs-on. Current ARC also supports runnerScaleSetLabels so multiple labels can target a scale set. Those labels are fleet routing metadata; they are not equivalent to arbitrary per-host labels on classic persistent runners, and they are not an isolation mechanism.

runnerGroup: arc-lab-group
runnerScaleSetName: arc-linux-build
runnerScaleSetLabels:
  - linux
  - build
  - isolated-ci

7. Capacity policy: warm floor, hard ceiling, drain

minRunners is an idle-capacity floor. maxRunners is a hard scale ceiling. Omitting both allows demand-driven growth with scale-to-zero, which can be appropriate only when cluster quota and cost bounds exist elsewhere. A safer production policy normally sets an explicit ceiling.

Policy Values Trade-off
Scale to zero min 0, bounded max lowest idle cost, cold starts
Warm pool min > 0, bounded max faster pickup, idle resource cost
Maintenance drain min 0, max 0 no new runner pods while existing work drains
Unbounded chart default both omitted simple elasticity, requires external quota/cost guardrails

8. ARC versus a custom Runner Scale Set Client

ARC is the supported Kubernetes operator path and already implements listener/controller/JIT runner lifecycle. GitHub also documents a Runner Scale Set Client for platform teams building custom autoscaling integrations. Choose the custom client only when Kubernetes is the wrong substrate or when an infrastructure provider needs a different control plane and is prepared to own scheduling, isolation, logs, upgrades and failure recovery.

The custom client is not “less infrastructure”; it transfers more of the control plane into your code and operations.

9. Upgrade compatibility is a state transition

Pin controller and runner-scale-set chart versions separately in GitOps. Before upgrading, record current CRDs, chart versions, controller image, runner image, Kubernetes version, values, workload queue depth and rollback package. Test chart/CRD changes in a disposable cluster. Do not let latest silently change the runner image underneath a regulated release process.

10. Worked decision: internal build fleet

Requirement Decision Prerequisite Observable evidence
Private package mirror reachable only on internal network ARC dedicated runner network segment network policy + successful bounded probe
Untrusted public fork PRs GitHub-hosted or separate untrusted fleet no production trust on that fleet routing policy + denied internal reach
Docker build needs privileged DinD dedicated node pool/cluster explicit risk acceptance pod security spec + node isolation
Peak 20 jobs, budget ceiling 6 concurrent runners maxRunners=6 queue latency accepted scale-set values + queue metrics
Incident investigation requirement external logs central sink and retention policy controller/listener/runner correlations by run/job

Knowledge check

When is GitHub-hosted usually safer than ARC?

Why is DinD a security-sensitive choice?

Do runnerScaleSetLabels replace runner groups?

What does maxRunners trade away?

Why might a custom Runner Scale Set Client be harder to operate than ARC?

Next chapter concept

Failure diagnosis under ephemeral infrastructure

Lesson 4 turns common ARC misconfigurations into an evidence-first diagnostic runbook without deleting the evidence you need to understand the original failure.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against current GitHub-maintained documentation and repositories on 2026-09-09. The latest public ARC runner-scale-set release found during authoring is 0.14.2, published 2026-05-22. Its release notes updated the bundled runner to v2.334.0; the runner binary has an independent release cadence, so do not infer that the ARC chart version and newest standalone runner version are identical. Current docs prefer Helm for ARC deployment, recommend production workload isolation and retained controller/listener/ephemeral-runner logs, support bounded minRunners/maxRunners, and document scale-set names plus runnerScaleSetLabels for routing. Current container modes include dind, kubernetes, and kubernetes-novolume; DinD is privileged, and Kubernetes mode changes the Kubernetes API/service-account trust boundary.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.