Checkpoint Lab — Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling
The checkpoint is intentionally small: one queued synthetic job, at most two runners, no production credentials and no external deployment. The goal is not to demonstrate throughput; it is to prove that the control plane is bounded, observable, isolated and reversible.
Learning objectives
- Predict the GitHub, listener, Kubernetes and runner state transitions before execution.
- Prove one queued job creates one ephemeral runner and returns to zero idle runner pods.
- Produce a security/operations evidence packet without storing secret values.
- Demonstrate cleanup and credential revocation as separate actions.
- Translate the lab into production controls and connect autoscaling to Chapter 11 job dependencies/concurrency/cancellation.
1. Checkpoint scenario and hard boundaries
You are the platform engineer for a fictional repository
octo-lab/arc-course-lab. The only workload is a manual
synthetic probe. Mandatory completion uses the local simulator.
Optional live completion uses a local disposable Kubernetes cluster
and a disposable GitHub repository you own.
Production Kubernetes clusters, employer/customer repositories, real cloud credentials, production registries, shared self-hosted runners or organization-wide policies for this checkpoint.
2. Current assumptions to record
| Input | Checkpoint value |
|---|---|
| Authoring date | 2026-09-09 |
| ARC chart baseline | 0.14.2 for optional live path |
| Scale-set name | arc-ch10-lab |
| Capacity |
minRunners: 0, maxRunners: 2
|
| Job |
one manual synthetic probe; permissions: {}
|
| External side effects | none |
| Credentials | none for simulation; disposable GitHub App/authorized lab credential only for optional live path |
| Logs | simulation JSONL or externalized controller/listener/pod evidence |
3. Preflight: prove the environment is disposable
Simulation path:
python --version
pwd
# Work only under a new temporary arc-ch10-checkpoint directory.
Optional live path:
kubectl config current-context
# Require exactly kind-arc-course-lab (or your documented disposable equivalent).
helm version
kubectl get nodes
helm list -A
If the context points to a shared or production cluster, the checkpoint stops here.
4. Write predictions before execution
| Prediction | Evidence that will verify it |
|---|---|
| Idle runner count starts at zero. |
simulator preflight row or kubectl get pods
|
| One queued job increases desired capacity to one, not two. | listener/simulator demand row |
| The runner gets one job identity and no deployment credential. | run/job log + permissions declaration |
| After success, runner pod/state is deleted and idle count returns to zero. | final JSONL row or Kubernetes watch/events |
| Controller/auth state outlives the runner pod until explicit cleanup. | Helm/secret metadata before cleanup |
5. Exact bounded scale-set configuration
For a live path, use a values file that references a pre-created Kubernetes Secret by name and pins the chart at install time.
githubConfigUrl: "https://github.com/OWNER/DISPOSABLE_REPO"
githubConfigSecret: arc-github-auth
runnerScaleSetName: arc-ch10-lab
minRunners: 0
maxRunners: 2
Do not place a PAT/private key inline in this file or in shell arguments committed to the repository.
6. Exact synthetic workflow
name: ch10-arc-checkpoint
on:
workflow_dispatch:
permissions: {}
jobs:
prove-ephemeral-runner:
runs-on: arc-ch10-lab
steps:
- name: Emit non-sensitive identity evidence
shell: bash
run: |
printf 'run=%s attempt=%s sha=%s\n' "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA"
printf 'runner=%s os=%s arch=%s\n' "$RUNNER_NAME" "$RUNNER_OS" "$RUNNER_ARCH"
printf 'checkpoint=ok\n'
The workflow intentionally has no checkout, cache, artifact, package, environment, cloud call or deployment. That isolates the runner lifecycle from unrelated failure layers.
7. Execute exactly one path
Mandatory simulator: reuse the deterministic
arc_sim.py from Lesson 2 and retain
arc-evidence.jsonl.
Optional live: install both 0.14.2 charts into the
disposable cluster, dispatch one workflow run, watch runner pods and
capture controller/listener evidence. Do not dispatch a second run
until the first packet is complete.
8. Reconcile predictions against actual state
Fill the packet with actual values rather than expected values:
source_sha: <actual>
run_id: <actual-or-simulated-id>
run_attempt: 1
scale_set: arc-ch10-lab
min_runners: 0
max_runners: 2
desired_peak: <actual>
runner_pod_uid: <actual-or-simulated-id>
job_conclusion: success
idle_runner_count_after: 0
credential_values_recorded: false
external_side_effects: none
If desired capacity exceeded one for one queued job, or the runner remains after the documented reconciliation window, do not “massage” the report. Preserve the discrepancy and diagnose it.
9. Security and operations dossier
The checkpoint packet should contain:
- Source identity: event, ref, SHA, workflow revision, run ID/attempt.
- ARC identity: chart/controller/listener version, scale-set name/group/labels.
- Kubernetes identity: context, namespaces, service accounts, pod UID/node/image digest.
- Capacity: min/max, queue count, desired/actual peak, scale-down proof.
- Trust: repository/event allowed to use the fleet, no public fork/untrusted route.
- Credentials: owner/name/rotation/revocation evidence only; never values.
- Isolation: network/RBAC/node/namespace assumptions and what was not tested.
- Logs: sink/retention or simulator JSONL; correlation fields.
- Rollback: exact releases/cluster/credential resources to remove.
10. Negative control: prove the ceiling matters
In simulation only, change queued_jobs to 5 while
max_runners remains 2. The desired runner count must
stop at 2. This demonstrates capacity protection without launching
five real jobs or consuming real infrastructure.
Restore the original script after capturing the negative-control evidence.
11. Cleanup and revocation
- Preserve the final evidence packet outside ephemeral runner storage.
- Verify the disposable Kubernetes context.
- Uninstall the scale-set and controller releases.
- Delete the local disposable cluster.
- Delete/revoke the lab GitHub App installation/PAT independently.
- Delete the disposable repository if created solely for the course.
- Remove local private-key/temp files using the credential owner’s policy.
-
Verify no
arc-ch10-labrunner remains registered and no local cluster remains.
12. Verification checklist
- Five states are evidenced: demand, desired capacity, runner creation/registration, job completion, runner deletion.
-
maxRunners=2is recorded and negative-control simulation respects it. - No secret value appears in workflow logs, values files, shell history captured in the packet or artifacts.
- Run ID/attempt/SHA are preserved for live path; equivalent synthetic IDs/limitations are explicit for simulation.
- Controller/listener/runner evidence is retained independently of pod lifetime.
- Isolation assumptions are explicit; no claim of production security is inferred from a local lab.
- Cleanup and credential revocation are separately verified.
13. What Chapter 10 adds to the operating model
Chapter 09 established that self-hosted runners are privileged compute. Chapter 10 adds a fleet control contract: demand is bounded, runner identity is ephemeral, capacity is explicit, GitHub and Kubernetes states are separable, credentials live outside runner pods, logs survive pod deletion, and rollback removes both infrastructure and authorization.
Chapter 11 moves upward from runner capacity to the
job graph: needs, concurrency,
cancellation and deployment serialization. Autoscaling can create
capacity, but it cannot decide which jobs should run together, which
should wait, or which older work should be canceled safely.
Knowledge check
One job was queued and desiredPeak became 2. What should the checkpoint do?
Preserve the discrepancy and diagnose why demand/minimum/current runner count produced two desired runners; do not report success merely because the job passed.
Why is the negative-control run performed only in simulation?
It proves the maxRunners ceiling without creating unnecessary real jobs, cost or infrastructure side effects.
What evidence must outlive an ephemeral runner pod?
At minimum the run/job identity plus controller/listener/runner diagnostics, Kubernetes events/identity and the security/operations dossier.
Does deleting the kind cluster prove GitHub credentials were revoked?
No. Credential revocation is a separate action at the GitHub App/PAT owner.
What problem does Chapter 11 solve that ARC does not?
ARC supplies execution capacity; job dependencies, concurrency, cancellation and serialization determine which work is eligible to consume that capacity and in what order.
Official references and version notes
- GitHub Docs — Actions Runner Controller — control-plane architecture, listener/JIT registration and ephemeral runner lifecycle.
- GitHub Docs — Deploy runner scale sets — Helm deployment, runner groups, min/max runners, pod templates, container modes and security guidance.
-
GitHub Docs — Use ARC in a workflow
— scale-set names and scale-set labels in
runs-on. - GitHub Docs — Troubleshoot ARC — controller/listener/runner diagnostic workflow.
- ARC release 0.14.2 — pinned release used for concrete examples in this chapter.
- GitHub Actions runner container image — minimal runner image published with runner releases.
Version-sensitive behavior was rechecked against current
GitHub-maintained documentation and repositories on
2026-09-09. The latest public ARC
runner-scale-set release found during authoring is
0.14.2, published 2026-05-22. Its release notes
updated the bundled runner to v2.334.0; the runner binary has an
independent release cadence, so do not infer that the ARC chart
version and newest standalone runner version are identical.
Current docs prefer Helm for ARC deployment, recommend production
workload isolation and retained
controller/listener/ephemeral-runner logs, support bounded
minRunners/maxRunners, and document
scale-set names plus runnerScaleSetLabels for
routing. Current container modes include dind,
kubernetes, and kubernetes-novolume;
DinD is privileged, and Kubernetes mode changes the Kubernetes
API/service-account trust boundary. The mandatory checkpoint uses
a no-credential faithful simulation; live ARC is optional and must
be disposable.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.