Chapter 10Lesson 05~260 minutes

Checkpoint Lab — Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling

The checkpoint is intentionally small: one queued synthetic job, at most two runners, no production credentials and no external deployment. The goal is not to demonstrate throughput; it is to prove that the control plane is bounded, observable, isolated and reversible.

CheckpointSecurity dossierEphemeral lifecycleRollbackChapter 11 bridge

Learning objectives

  • Predict the GitHub, listener, Kubernetes and runner state transitions before execution.
  • Prove one queued job creates one ephemeral runner and returns to zero idle runner pods.
  • Produce a security/operations evidence packet without storing secret values.
  • Demonstrate cleanup and credential revocation as separate actions.
  • Translate the lab into production controls and connect autoscaling to Chapter 11 job dependencies/concurrency/cancellation.

1. Checkpoint scenario and hard boundaries

You are the platform engineer for a fictional repository octo-lab/arc-course-lab. The only workload is a manual synthetic probe. Mandatory completion uses the local simulator. Optional live completion uses a local disposable Kubernetes cluster and a disposable GitHub repository you own.

Never use

Production Kubernetes clusters, employer/customer repositories, real cloud credentials, production registries, shared self-hosted runners or organization-wide policies for this checkpoint.

2. Current assumptions to record

Input Checkpoint value
Authoring date 2026-09-09
ARC chart baseline 0.14.2 for optional live path
Scale-set name arc-ch10-lab
Capacity minRunners: 0, maxRunners: 2
Job one manual synthetic probe; permissions: {}
External side effects none
Credentials none for simulation; disposable GitHub App/authorized lab credential only for optional live path
Logs simulation JSONL or externalized controller/listener/pod evidence

3. Preflight: prove the environment is disposable

Simulation path:

python --version
pwd
# Work only under a new temporary arc-ch10-checkpoint directory.

Optional live path:

kubectl config current-context
# Require exactly kind-arc-course-lab (or your documented disposable equivalent).
helm version
kubectl get nodes
helm list -A

If the context points to a shared or production cluster, the checkpoint stops here.

4. Write predictions before execution

Prediction Evidence that will verify it
Idle runner count starts at zero. simulator preflight row or kubectl get pods
One queued job increases desired capacity to one, not two. listener/simulator demand row
The runner gets one job identity and no deployment credential. run/job log + permissions declaration
After success, runner pod/state is deleted and idle count returns to zero. final JSONL row or Kubernetes watch/events
Controller/auth state outlives the runner pod until explicit cleanup. Helm/secret metadata before cleanup

5. Exact bounded scale-set configuration

For a live path, use a values file that references a pre-created Kubernetes Secret by name and pins the chart at install time.

githubConfigUrl: "https://github.com/OWNER/DISPOSABLE_REPO"
githubConfigSecret: arc-github-auth
runnerScaleSetName: arc-ch10-lab
minRunners: 0
maxRunners: 2

Do not place a PAT/private key inline in this file or in shell arguments committed to the repository.

6. Exact synthetic workflow

name: ch10-arc-checkpoint
on:
  workflow_dispatch:
permissions: {}
jobs:
  prove-ephemeral-runner:
    runs-on: arc-ch10-lab
    steps:
      - name: Emit non-sensitive identity evidence
        shell: bash
        run: |
          printf 'run=%s attempt=%s sha=%s\n'             "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA"
          printf 'runner=%s os=%s arch=%s\n'             "$RUNNER_NAME" "$RUNNER_OS" "$RUNNER_ARCH"
          printf 'checkpoint=ok\n'

The workflow intentionally has no checkout, cache, artifact, package, environment, cloud call or deployment. That isolates the runner lifecycle from unrelated failure layers.

7. Execute exactly one path

Mandatory simulator: reuse the deterministic arc_sim.py from Lesson 2 and retain arc-evidence.jsonl. Optional live: install both 0.14.2 charts into the disposable cluster, dispatch one workflow run, watch runner pods and capture controller/listener evidence. Do not dispatch a second run until the first packet is complete.

8. Reconcile predictions against actual state

Fill the packet with actual values rather than expected values:

source_sha: <actual>
run_id: <actual-or-simulated-id>
run_attempt: 1
scale_set: arc-ch10-lab
min_runners: 0
max_runners: 2
desired_peak: <actual>
runner_pod_uid: <actual-or-simulated-id>
job_conclusion: success
idle_runner_count_after: 0
credential_values_recorded: false
external_side_effects: none

If desired capacity exceeded one for one queued job, or the runner remains after the documented reconciliation window, do not “massage” the report. Preserve the discrepancy and diagnose it.

9. Security and operations dossier

The checkpoint packet should contain:

  • Source identity: event, ref, SHA, workflow revision, run ID/attempt.
  • ARC identity: chart/controller/listener version, scale-set name/group/labels.
  • Kubernetes identity: context, namespaces, service accounts, pod UID/node/image digest.
  • Capacity: min/max, queue count, desired/actual peak, scale-down proof.
  • Trust: repository/event allowed to use the fleet, no public fork/untrusted route.
  • Credentials: owner/name/rotation/revocation evidence only; never values.
  • Isolation: network/RBAC/node/namespace assumptions and what was not tested.
  • Logs: sink/retention or simulator JSONL; correlation fields.
  • Rollback: exact releases/cluster/credential resources to remove.

10. Negative control: prove the ceiling matters

In simulation only, change queued_jobs to 5 while max_runners remains 2. The desired runner count must stop at 2. This demonstrates capacity protection without launching five real jobs or consuming real infrastructure.

Restore the original script after capturing the negative-control evidence.

11. Cleanup and revocation

  1. Preserve the final evidence packet outside ephemeral runner storage.
  2. Verify the disposable Kubernetes context.
  3. Uninstall the scale-set and controller releases.
  4. Delete the local disposable cluster.
  5. Delete/revoke the lab GitHub App installation/PAT independently.
  6. Delete the disposable repository if created solely for the course.
  7. Remove local private-key/temp files using the credential owner’s policy.
  8. Verify no arc-ch10-lab runner remains registered and no local cluster remains.

12. Verification checklist

  • Five states are evidenced: demand, desired capacity, runner creation/registration, job completion, runner deletion.
  • maxRunners=2 is recorded and negative-control simulation respects it.
  • No secret value appears in workflow logs, values files, shell history captured in the packet or artifacts.
  • Run ID/attempt/SHA are preserved for live path; equivalent synthetic IDs/limitations are explicit for simulation.
  • Controller/listener/runner evidence is retained independently of pod lifetime.
  • Isolation assumptions are explicit; no claim of production security is inferred from a local lab.
  • Cleanup and credential revocation are separately verified.

13. What Chapter 10 adds to the operating model

Chapter 09 established that self-hosted runners are privileged compute. Chapter 10 adds a fleet control contract: demand is bounded, runner identity is ephemeral, capacity is explicit, GitHub and Kubernetes states are separable, credentials live outside runner pods, logs survive pod deletion, and rollback removes both infrastructure and authorization.

Chapter 11 moves upward from runner capacity to the job graph: needs, concurrency, cancellation and deployment serialization. Autoscaling can create capacity, but it cannot decide which jobs should run together, which should wait, or which older work should be canceled safely.

Knowledge check

One job was queued and desiredPeak became 2. What should the checkpoint do?

Why is the negative-control run performed only in simulation?

What evidence must outlive an ephemeral runner pod?

Does deleting the kind cluster prove GitHub credentials were revoked?

What problem does Chapter 11 solve that ARC does not?

Next chapter concept

Job Dependencies, needs, Concurrency, Cancellation, and Deployment Serialization

Next, move from fleet capacity to job-graph control: dependencies, concurrency groups, cancellation and serialized deployment behavior.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against current GitHub-maintained documentation and repositories on 2026-09-09. The latest public ARC runner-scale-set release found during authoring is 0.14.2, published 2026-05-22. Its release notes updated the bundled runner to v2.334.0; the runner binary has an independent release cadence, so do not infer that the ARC chart version and newest standalone runner version are identical. Current docs prefer Helm for ARC deployment, recommend production workload isolation and retained controller/listener/ephemeral-runner logs, support bounded minRunners/maxRunners, and document scale-set names plus runnerScaleSetLabels for routing. Current container modes include dind, kubernetes, and kubernetes-novolume; DinD is privileged, and Kubernetes mode changes the Kubernetes API/service-account trust boundary. The mandatory checkpoint uses a no-credential faithful simulation; live ARC is optional and must be disposable.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.