Chapter 16Lesson 05~215 minutes

Checkpoint Lab — GitHub-Hosted and Self-Hosted Runners, Labels, Groups, Scaling, and Runner Security

This checkpoint makes runner governance measurable. You will compare two explicit GitHub-hosted images, capture queue/runtime and environment evidence, verify that hosted execution did not create self-hosted registrations, then design a production self-hosted architecture whose lifecycle, routing, networking, patching, logging, and de-registration are testable controls rather than assumptions.

Checkpoint labQueue evidenceTrust-zone designRunner lifecycle

Learning objectives

  • Run the same observation workflow on two explicit standard GitHub-hosted Ubuntu labels and compare exact source, environment, queue, and runtime evidence.
  • Predict runner/resource changes before dispatch and verify that hosted execution creates no repository self-hosted runner registrations.
  • Produce a production self-hosted architecture with separate trust zones, ephemeral lifecycle, labels/groups, patching, external logs, capacity signals, and de-registration.
  • Optionally validate the architecture with a one-job ephemeral runner on a disposable private repository/VM without exposing enrollment credentials.
  • Define a runner operating checklist that separates CI correctness, security, operability, performance, and cost.

Checkpoint assumptions: Mandatory path: GitHub.com, GitHub Free, a fresh public personal repository, standard ubuntu-22.04/ubuntu-24.04 hosted runners, and GitHub CLI. Optional path: a separate private disposable repository and a truly disposable isolated VM. No organization, paid larger runner, internal network, cloud credential, self-hosted runner, or secret is required to pass the checkpoint.

1. Scenario and preflight

You are defining a runner policy for Atlas Delivery. The team wants reproducible CI across two Ubuntu generations and is considering self-hosted runners for future proprietary build requirements. Before they register infrastructure, you must produce evidence about hosted images and a design that proves how self-hosted work would be isolated.

gh --version
gh auth status --active --hostname github.com

OWNER="$(gh api -H "X-GitHub-Api-Version: 2026-03-10" user --jq .login)"
REPO="$OWNER/atlas-c16-checkpoint"

gh repo view "$REPO" --json nameWithOwner >/dev/null 2>&1 && {
  echo "Checkpoint repository already exists; choose a fresh name." >&2
  exit 1
} || true

gh repo create "$REPO" --public --clone --add-readme
cd atlas-c16-checkpoint
DEFAULT_BRANCH="$(gh repo view "$REPO" --json defaultBranchRef --jq .defaultBranchRef.name)"

gh api -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/actions/runners?per_page=100" --jq '{total_count,runners}'

Preflight must show zero self-hosted runner registrations. If it does not, stop and use a fresh repository; the checkpoint depends on a clean runner inventory.

2. Write predictions before any run exists

Record at least these predictions in RUNNER_CHECKPOINT.md before dispatch:

Prediction Before action Expected after action Independent proof
P1 — source identity One workflow commit SHA Both hosted runs use the same head_sha Run REST metadata + local Git SHA.
P2 — hosted image selection No run 22.04 run and 24.04 run report different requested/image OS identity Safe workflow logs + job labels.
P3 — runner inventory Repository self-hosted count = 0 Still 0 after both standard hosted runs Self-hosted runner REST inventory.
P4 — timing No queue/runtime evidence Each run has created/start and job start/complete timestamps Versioned run/job REST endpoints.
P5 — optional ephemeral VM No registration Runner appears online/ephemeral, executes one job, then disappears Runner inventory before/during/after + VM destruction evidence.

3. Build the controlled hosted-runner workflow

name: Chapter 16 checkpoint runner comparison

on:
  workflow_dispatch:
    inputs:
      runner:
        description: Controlled hosted runner image
        required: true
        type: choice
        options:
          - ubuntu-22.04
          - ubuntu-24.04

permissions: {}

jobs:
  inspect:
    runs-on: "${{ inputs.runner }}"
    timeout-minutes: 10
    steps:
      - name: Safe identity evidence
        env:
          REQUESTED_RUNNER: "${{ inputs.runner }}"
        run: |
          printf 'requested_runner=%s\n' "$REQUESTED_RUNNER"
          printf 'runner_os=%s\n' "$RUNNER_OS"
          printf 'runner_arch=%s\n' "$RUNNER_ARCH"
          printf 'runner_name=%s\n' "$RUNNER_NAME"
          printf 'runner_environment=%s\n' "$RUNNER_ENVIRONMENT"
          printf 'source_sha=%s\n' "$GITHUB_SHA"

      - name: Environment observations
        run: |
          cat /etc/os-release
          uname -a
          printf 'cpu_count='; getconf _NPROCESSORS_ONLN
          free -h
          git --version
          python3 --version

This intentionally uses no checkout, action marketplace dependency, secret, network probe, or write permission. Runner comparison is the only experiment.

4. Commit once and bind both runs to that revision

mkdir -p .github/workflows
# Save YAML as .github/workflows/ch16-checkpoint.yml

git add .github/workflows/ch16-checkpoint.yml RUNNER_CHECKPOINT.md
git commit -m "ci: add Chapter 16 runner checkpoint"
git push origin "$DEFAULT_BRANCH"
CHECKPOINT_SHA="$(git rev-parse HEAD)"
printf 'checkpoint_sha=%s\n' "$CHECKPOINT_SHA"

gh workflow view ch16-checkpoint.yml -R "$REPO" --yaml

No runner is active because of the commit itself; the workflow is manual. Both future dispatches must use this exact SHA.

5. Run Ubuntu 22.04, then Ubuntu 24.04

gh workflow run ch16-checkpoint.yml -R "$REPO" \
  --ref "$DEFAULT_BRANCH" -f runner=ubuntu-22.04
sleep 3
RUN22="$(gh run list -R "$REPO" --workflow ch16-checkpoint.yml \
  --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$RUN22" -R "$REPO" --exit-status

gh workflow run ch16-checkpoint.yml -R "$REPO" \
  --ref "$DEFAULT_BRANCH" -f runner=ubuntu-24.04
sleep 3
RUN24="$(gh run list -R "$REPO" --workflow ch16-checkpoint.yml \
  --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$RUN24" -R "$REPO" --exit-status

printf 'run22=%s run24=%s\n' "$RUN22" "$RUN24"

Do not dispatch both simultaneously for this checkpoint; sequential dispatch reduces ambiguity when measuring IDs/timestamps.

6. Verify P1, P2, and P4 from independent evidence

for RUN_ID in "$RUN22" "$RUN24"; do
  echo "===== RUN $RUN_ID LOG ====="
  gh run view "$RUN_ID" -R "$REPO" --log

  echo "===== RUN $RUN_ID API ====="
  gh api -H "X-GitHub-Api-Version: 2026-03-10" \
    "repos/$REPO/actions/runs/$RUN_ID" \
    --jq '{id,head_sha,status,conclusion,created_at,run_started_at,updated_at}'

  gh api -H "X-GitHub-Api-Version: 2026-03-10" \
    "repos/$REPO/actions/runs/$RUN_ID/jobs?filter=latest&per_page=100" \
    --jq '.jobs[] | {name,status,conclusion,runner_name,runner_group_name,labels,started_at,completed_at}'
done

printf 'expected_sha=%s\n' "$CHECKPOINT_SHA"

Both API objects must report the checkpoint SHA. The logs should show the requested runner label and corresponding Ubuntu release. Record tool versions as observed facts; do not turn a one-run difference into a permanent platform guarantee.

7. Verify P3: hosted execution did not create managed runner state in your repository

gh api -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/actions/runners?per_page=100" \
  --jq '{total_count,runners:[.runners[]|{id,name,status,busy,ephemeral,labels:[.labels[].name]}]}'

Expected: total_count: 0. If not, the repository was not a clean checkpoint and you must investigate the unexpected self-hosted registration rather than ignoring it.

8. Build a timing comparison without overclaiming performance

For each run, copy created_at, run_started_at, job started_at, and completed_at into the checkpoint document. Compute approximate queue delay and job duration locally. Two tiny observations cannot establish which image is “faster”; they only prove how to capture scheduling evidence.

Runner label | Run ID | Head SHA | Created | Run started | Job started | Completed | Queue note | Tool observations
ubuntu-22.04 | ... | ... | ... | ... | ... | ... | ... | ...
ubuntu-24.04 | ... | ... | ... | ... | ... | ... | ... | ... | ...

9. Design the production self-hosted architecture

Add this architecture section to RUNNER_CHECKPOINT.md. It is mandatory even if you do not perform the optional VM exercise.

Concept / workflow diagram
              flowchart LR
                PR["Public / fork PR"] --> H["Standard hosted CI No internal reach"]
                M["Trusted protected branch"] --> BG["Group: trusted-build"]
                BG --> E["Ephemeral build runner label: isolated-ci"]
                E --> PKG["Approved package mirror"]
                D["Protected deploy job"] --> DG["Group: deploy-prod"]
                DG --> DEP["Ephemeral deploy runner label: deploy-prod"]
                DEP --> API["Deployment API only"]
                E --> LOG["External runner logs"]
                DEP --> LOG
            

A production runner design separates source trust and privilege. Untrusted PRs never enter self-hosted internal zones. Trusted build and deployment runners are separate ephemeral fleets with distinct group/label/network permissions, and runner application logs survive instance destruction in external storage.

Your design must state:

  • Trust classification: public/fork PR, internal PR, protected branch, deployment.
  • Scope/access: repository runner for one repo or organization group with explicit repository access.
  • Labels: stable capabilities/trust classes—not machine names.
  • Lifecycle: one job per runner, automatic de-registration, infrastructure destruction, no reused workspace.
  • Network: deny-by-default; build mirror only for build; deployment API only for deploy; no broad corporate subnet route.
  • Machine identity/secrets: minimum per zone; prefer short-lived/OIDC patterns where later deployment lessons permit.
  • Patching: immutable runner image pipeline, OS/tool patch cadence, runner version monitoring; rebuild rather than hand-patch snowflakes.
  • Logging: workflow run/job IDs plus external runner application logs retained beyond VM deletion.
  • Scaling: queue/capacity signal, maximum concurrency, scale-to-zero/low baseline, 24-hour queue timeout treated as hard backstop—not SLO.
  • De-registration: verify runner disappears after job; admin force-remove stale identity if a host is lost; destroy VM/disk.

10. Optional P5: validate one disposable ephemeral runner

Only perform this if you can create a dedicated throwaway VM and a separate private disposable repository. Follow the current GitHub “New self-hosted runner” page for the exact platform download and registration commands. Keep the short-lived registration token local. Add --ephemeral and a custom label ch16-disposable during configuration.

name: Chapter 16 optional ephemeral validation

on:
  workflow_dispatch:

permissions: {}

jobs:
  one_job_only:
    runs-on: [self-hosted, ch16-disposable]
    timeout-minutes: 10
    steps:
      - run: |
          printf 'runner_name=%s\n' "$RUNNER_NAME"
          printf 'runner_os=%s\n' "$RUNNER_OS"
          printf 'runner_arch=%s\n' "$RUNNER_ARCH"
          uname -a

Observe runner inventory before dispatch, while online, and after completion. Expected: the runner processes one job and is automatically de-registered. Then destroy the VM. If anything about the runner registration, network, or token handling is unclear, skip this optional validation—the architecture exercise is sufficient.

11. Production runner handoff checklist

Control Question that must have an answer
Correctness Which exact OS/image/tool versions executed the required job and source SHA?
Security Can untrusted input reach this runner? What network and machine identity can it reach?
Routing Which groups/repositories and label combinations can select this fleet?
Lifecycle Is the runner one-job ephemeral? Who proves destruction and de-registration?
Patching How are OS/tool/runner updates tested and rolled out?
Observability Where do workflow and runner application logs survive instance deletion?
Reliability What queue SLO/max concurrency/autoscaling signal is used?
Cost Who sees utilization/per-minute/infra cost and reviews oversized capacity?
Incident response How is a runner/group quarantined immediately and stale identity removed?

12. Verification checklist and cleanup/rollback

  • Exactly two hosted runs used the same checkpoint SHA.
  • One requested ubuntu-22.04; the other requested ubuntu-24.04.
  • Safe logs contain selected runner/OS/tool evidence but no whole environment/context or secrets.
  • Run/job APIs captured queue/runtime timestamps and runner labels/names.
  • Self-hosted repository inventory remained zero on the mandatory public repository.
  • The architecture document includes trust zones, groups/labels, ephemeral lifecycle, network, patching, logs, scaling, cost, and de-registration.
  • If optional VM was used, its runner de-registered and the VM/disk was destroyed.
gh workflow disable ch16-checkpoint.yml -R "$REPO"
gh workflow list -R "$REPO" --all --json name,path,state

gh repo archive "$REPO" --yes
gh repo view "$REPO" --json nameWithOwner,isArchived,url

Rollback: For a repeatable run, create a fresh disposable repository rather than unarchiving and layering new runner state over old evidence. If an optional self-hosted runner is still visible after its host is destroyed, remove the stale registration before doing anything else.

13. What Chapter 16 adds to the production GitHub operating model

Chapters 13–15 defined workflow identity, activation/data flow, and job/container/service topology. Chapter 16 adds the compute trust plane: every job must route to an execution environment whose lifecycle, repository access, labels/groups, network, machine identity, image/patch level, logs, capacity, and cost are known and reviewable.

Chapter 17 moves back up one layer to the data produced by those jobs: artifacts, dependency caches, logs, test reports, retention, and the integrity/sensitivity of workflow data.

14. Checkpoint summary

You compared two controlled GitHub-hosted runner labels against one source revision, captured exact runner/timing evidence, proved no self-hosted registration was created by hosted execution, and designed an ephemeral self-hosted architecture with explicit trust zones and cleanup. That closes the runner control loop: select, observe, isolate, scale, patch, log, de-register, and verify.

Knowledge check

Why is P3—self-hosted inventory stays zero—important in a hosted-runner lab?

The Ubuntu 24.04 run is faster once. Can you conclude it is the better runner?

A production self-hosted build runner is ephemeral but can reach the entire corporate network. What remains wrong?

An organization puts build and production deployment runners in one broad group and distinguishes them only by labels. What governance improvement is stronger?

An ephemeral runner disappeared from GitHub after the job. What two evidence/cleanup items should still exist?

What does Chapter 17 add next?

Next lesson

Next: Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data: Concepts, Architecture, and Mental Model

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.