Chapter 09Lesson 05~240 minutes

Checkpoint Lab — Self-Hosted Runners, Runner Groups, Isolation, and Operational Security

The checkpoint produces a complete runner-lifecycle evidence packet rather than merely proving that a self-hosted job can turn green. You will use disposable compute, a dedicated unprivileged account and a private lab repository, predict the registration/routing/persistence changes, run one trusted synthetic workload, inspect and clean controlled residual state, preserve diagnostics, deregister the runner and document the controls required before any organizational use.

CheckpointLeast privilegeEvidence packetTeardownProduction controls

Learning objectives

  • Provision or faithfully simulate a disposable self-hosted runner with least-privilege host and repository boundaries.
  • Predict and verify runner registration, routing, job execution and residual-state changes.
  • Preserve run, runner, host-user, update and diagnostic evidence without recording real registration credentials.
  • Clean controlled residue, deregister the runner and prove both GitHub and host-side teardown.
  • Produce a production-readiness control list covering trust, groups, network, patching, logs and one-job lifecycle.

1. Checkpoint charter

Field Required checkpoint boundary
Repository Disposable private learner-owned repository.
Compute Disposable local VM/throwaway host; simulation allowed if unavailable.
Account Dedicated unprivileged runner user; no personal credentials.
Runner scope Repository only; custom label chapter09-lab.
Workflow Manual trusted synthetic job; permissions: {}; no checkout/action dependency.
Network GitHub control-plane connectivity only; no production/internal target required.
Evidence Run ID/attempt/SHA, runner name/status/version/labels, service user, selected diagnostics hashes, residue/cleanup/deregistration proof.
Teardown Runner removed from GitHub and lab compute wiped/reverted.

2. Predict state changes before registration

  1. After registration, GitHub will show one repository-scoped runner with default OS/architecture labels plus chapter09-lab.
  2. When run.sh connects, status will become Idle; during the job it will become Active.
  3. The trusted job will execute under the dedicated lab OS account, not a GitHub-hosted VM.
  4. A synthetic file written under the lab account’s home will remain after the job until explicitly cleaned or the VM is destroyed.
  5. After removal, GitHub routing identity will disappear; after VM wipe, local runner/workspace/credential state will disappear.

3. Setup and registration

Follow the exact account-generated package instructions in Settings → Actions → Runners → New self-hosted runner. Record the package hash and ./run.sh --version. Register only the disposable repository using a locally entered time-limited token:

read -rsp 'Runner registration token: ' RUNNER_REG_TOKEN; echo
./config.sh \
  --url https://github.com/OWNER/gha-runner-lab \
  --token "$RUNNER_REG_TOKEN" \
  --name ch09-checkpoint-01 \
  --labels chapter09-lab \
  --work _work \
  --unattended
unset RUNNER_REG_TOKEN
./run.sh

If the host contains any unexpected secret, personal SSH key, production mount or sensitive internal route, abort and rebuild the lab on cleaner compute.

4. Trusted checkpoint workflow

name: chapter09-checkpoint
on: workflow_dispatch
permissions: {}

jobs:
  trusted-lab:
    runs-on: [self-hosted, linux, x64, chapter09-lab]
    steps:
      - name: Record safe execution identity
        shell: bash
        run: |
          set -euo pipefail
          printf 'runner=%s os=%s arch=%s\n' \
            '${{ runner.name }}' '${{ runner.os }}' '${{ runner.arch }}'
          printf 'run_id=%s attempt=%s sha=%s\n' \
            "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA"
          id
          printf 'workspace=%s\n' "$GITHUB_WORKSPACE"

      - name: Create controlled residual state
        shell: bash
        run: |
          set -euo pipefail
          state="$HOME/.gha-ch09-checkpoint"
          mkdir -p "$state"
          printf 'run=%s\n' "$GITHUB_RUN_ID" > "$state/sentinel.txt"
          sha256sum "$state/sentinel.txt"

This is the only workload. It does not fetch repository code, invoke Docker, call external APIs or access secrets.

5. Inspect residual state and diagnostics

After the job completes, keep the runner stopped from accepting additional unrelated work while you inspect the lab:

test -f "$HOME/.gha-ch09-checkpoint/sentinel.txt"
wc -c "$HOME/.gha-ch09-checkpoint/sentinel.txt"
ls -1 _diag/Runner_*.log 2>/dev/null | tail -n 3
ls -1 _diag/Worker_*.log 2>/dev/null | tail -n 3
./run.sh --version

Record filenames, hashes and relevant non-sensitive diagnostic excerpts. Do not copy .credentials, registration tokens or whole environment dumps into the evidence packet.

6. Clean the controlled residue

rm -rf "$HOME/.gha-ch09-checkpoint"
test ! -e "$HOME/.gha-ch09-checkpoint" && echo 'lab_residue_removed=true'

This cleanup proves the synthetic file was removed. It does not prove a compromised persistent host is trustworthy. Production integrity should come from immutable/rebuilt compute, especially for one-job ephemeral/JIT designs.

7. Deregister and destroy the lab

Use the current runner removal flow in GitHub Settings, enter the generated token locally, and run the provided removal command. Verify the runner disappears from GitHub Settings. Retain reviewed evidence, then delete the runner installation directory and destroy or revert the disposable VM.

# Shape of the local removal command; use the token generated by GitHub.
read -rsp 'Runner removal token: ' RUNNER_REMOVE_TOKEN; echo
./config.sh remove --token "$RUNNER_REMOVE_TOKEN"
unset RUNNER_REMOVE_TOKEN

8. Evidence packet

Evidence Required content
Workflow identity Workflow path/hash, event, source SHA, run ID, attempt.
Runner identity Name, repository scope, group if any, labels, status transitions, runner version.
Host boundary OS/architecture, dedicated user/UID/groups, note that no admin/root or Docker socket was used.
Network boundary Documented GitHub outbound requirement and confirmation that no production target was needed.
Residual state Sentinel path/hash before cleanup and proof of removal.
Diagnostics Reviewed Runner/Worker log filenames/hashes; no credential files.
Teardown Runner removed from GitHub; installation/VM wiped or reverted.
Limitations One disposable trusted-job lab does not validate a production runner fleet.

9. Controls required before real organizational use

  • Private/trusted source boundary; no public fork code on trusted self-hosted infrastructure.
  • Runner groups with selected repository/workflow access where organization/enterprise sharing is used.
  • Dedicated least-privilege service account; no ambient personal/admin credentials.
  • Network segmentation and explicit destination policy.
  • Ephemeral/JIT one-job compute for scalable production use where practical.
  • Clean immutable image/bootstrap process and defined host/runner/tool update SLA.
  • No unnecessary Docker socket or privileged host mounts.
  • Externally retained Runner/Worker diagnostics for ephemeral fleets.
  • Credential inventory, rotation/revocation and incident isolation/rebuild procedure.
  • Capacity/queue monitoring with routing labels documented as capability, not trust.

10. Simulation-only completion path

If you used no real runner, your evidence packet must clearly say so. Simulate repository scope, runner status, labels, a one-job state file and teardown in a local directory, then explain which claims remain unverified: GitHub registration, live job routing, status transitions and actual runner diagnostics.

11. Chapter 09 production operating contract

Chapter 09 adds a self-hosted runner governance contract: runner scope and group authorization, job-source trust, least-privilege service identity, network reachability, persistent-state handling, runner/host patch ownership, external diagnostics, controlled cleanup and incident-safe rebuild are all part of CI/CD correctness—not infrastructure details outside the pipeline.

Chapter 10 moves to Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling, where the same trust and lifecycle rules must survive automated provisioning and fleet-scale scheduling.

Next lesson

Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling: Core Concepts and Mental Model

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

What two independent teardown facts must the checkpoint prove?

Why does deleting the sentinel not prove the host is safe after a real compromise?

Which credentials belong in the checkpoint evidence packet?

What should a production organization add before sharing runners broadly?

What does Chapter 10 add to the Chapter 09 model?

Official references and version notes

Version and compatibility note

Version-sensitive self-hosted-runner behavior was rechecked against current primary GitHub documentation on 2026-09-09. The latest public actions/runner release visible at verification time is v2.337.0, released through a progressive rollout; the repository/organization “New self-hosted runner” page remains the authoritative version/download instruction for the specific account. The current reference lists x64 support on Linux/macOS/Windows, Arm64 on Linux/macOS/Windows in public preview, and Arm32 on Linux. GitHub currently lists Ubuntu 20.04+, Debian 10+, several other Linux families, Windows 10/11 and Windows Server 2016/2019/2022, and macOS 11+ as supported runner hosts. Self-hosted runners connect outbound to GitHub over HTTPS port 443 and must remain able to reach the documented Actions domains; no inbound job-listener port is required. By default the runner application self-updates, but the operating system and all other host software remain the operator’s responsibility. If automatic runner updates are disabled, the runner must be updated within 30 days of a new available release; critical security updates can block new jobs until applied. GitHub recommends ephemeral self-hosted runners for autoscaling and warns that ephemeral/JIT runner logs should be forwarded to external storage before production use. Runner groups can restrict repository/workflow access; labels select capabilities but are not an authorization boundary. The checkpoint requires no cloud account, enterprise feature, package publication, production credential or managed Kubernetes. Runner groups/ARC are discussed as optional/future controls.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.