Checkpoint Lab — Self-Hosted Runners, Runner Groups, Isolation, and Operational Security
The checkpoint produces a complete runner-lifecycle evidence packet rather than merely proving that a self-hosted job can turn green. You will use disposable compute, a dedicated unprivileged account and a private lab repository, predict the registration/routing/persistence changes, run one trusted synthetic workload, inspect and clean controlled residual state, preserve diagnostics, deregister the runner and document the controls required before any organizational use.
Learning objectives
- Provision or faithfully simulate a disposable self-hosted runner with least-privilege host and repository boundaries.
- Predict and verify runner registration, routing, job execution and residual-state changes.
- Preserve run, runner, host-user, update and diagnostic evidence without recording real registration credentials.
- Clean controlled residue, deregister the runner and prove both GitHub and host-side teardown.
- Produce a production-readiness control list covering trust, groups, network, patching, logs and one-job lifecycle.
1. Checkpoint charter
| Field | Required checkpoint boundary |
|---|---|
| Repository | Disposable private learner-owned repository. |
| Compute | Disposable local VM/throwaway host; simulation allowed if unavailable. |
| Account | Dedicated unprivileged runner user; no personal credentials. |
| Runner scope |
Repository only; custom label chapter09-lab.
|
| Workflow |
Manual trusted synthetic job; permissions: {};
no checkout/action dependency.
|
| Network | GitHub control-plane connectivity only; no production/internal target required. |
| Evidence | Run ID/attempt/SHA, runner name/status/version/labels, service user, selected diagnostics hashes, residue/cleanup/deregistration proof. |
| Teardown | Runner removed from GitHub and lab compute wiped/reverted. |
2. Predict state changes before registration
-
After registration, GitHub will show one repository-scoped runner
with default OS/architecture labels plus
chapter09-lab. -
When
run.shconnects, status will become Idle; during the job it will become Active. - The trusted job will execute under the dedicated lab OS account, not a GitHub-hosted VM.
- A synthetic file written under the lab account’s home will remain after the job until explicitly cleaned or the VM is destroyed.
- After removal, GitHub routing identity will disappear; after VM wipe, local runner/workspace/credential state will disappear.
3. Setup and registration
Follow the exact account-generated package instructions in
Settings → Actions → Runners → New self-hosted runner. Record the package hash and ./run.sh --version.
Register only the disposable repository using a locally entered
time-limited token:
read -rsp 'Runner registration token: ' RUNNER_REG_TOKEN; echo
./config.sh \
--url https://github.com/OWNER/gha-runner-lab \
--token "$RUNNER_REG_TOKEN" \
--name ch09-checkpoint-01 \
--labels chapter09-lab \
--work _work \
--unattended
unset RUNNER_REG_TOKEN
./run.sh
If the host contains any unexpected secret, personal SSH key, production mount or sensitive internal route, abort and rebuild the lab on cleaner compute.
4. Trusted checkpoint workflow
name: chapter09-checkpoint
on: workflow_dispatch
permissions: {}
jobs:
trusted-lab:
runs-on: [self-hosted, linux, x64, chapter09-lab]
steps:
- name: Record safe execution identity
shell: bash
run: |
set -euo pipefail
printf 'runner=%s os=%s arch=%s\n' \
'${{ runner.name }}' '${{ runner.os }}' '${{ runner.arch }}'
printf 'run_id=%s attempt=%s sha=%s\n' \
"$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA"
id
printf 'workspace=%s\n' "$GITHUB_WORKSPACE"
- name: Create controlled residual state
shell: bash
run: |
set -euo pipefail
state="$HOME/.gha-ch09-checkpoint"
mkdir -p "$state"
printf 'run=%s\n' "$GITHUB_RUN_ID" > "$state/sentinel.txt"
sha256sum "$state/sentinel.txt"
This is the only workload. It does not fetch repository code, invoke Docker, call external APIs or access secrets.
5. Inspect residual state and diagnostics
After the job completes, keep the runner stopped from accepting additional unrelated work while you inspect the lab:
test -f "$HOME/.gha-ch09-checkpoint/sentinel.txt"
wc -c "$HOME/.gha-ch09-checkpoint/sentinel.txt"
ls -1 _diag/Runner_*.log 2>/dev/null | tail -n 3
ls -1 _diag/Worker_*.log 2>/dev/null | tail -n 3
./run.sh --version
Record filenames, hashes and relevant non-sensitive diagnostic
excerpts. Do not copy .credentials, registration tokens
or whole environment dumps into the evidence packet.
6. Clean the controlled residue
rm -rf "$HOME/.gha-ch09-checkpoint"
test ! -e "$HOME/.gha-ch09-checkpoint" && echo 'lab_residue_removed=true'
This cleanup proves the synthetic file was removed. It does not prove a compromised persistent host is trustworthy. Production integrity should come from immutable/rebuilt compute, especially for one-job ephemeral/JIT designs.
7. Deregister and destroy the lab
Use the current runner removal flow in GitHub Settings, enter the generated token locally, and run the provided removal command. Verify the runner disappears from GitHub Settings. Retain reviewed evidence, then delete the runner installation directory and destroy or revert the disposable VM.
# Shape of the local removal command; use the token generated by GitHub.
read -rsp 'Runner removal token: ' RUNNER_REMOVE_TOKEN; echo
./config.sh remove --token "$RUNNER_REMOVE_TOKEN"
unset RUNNER_REMOVE_TOKEN
8. Evidence packet
| Evidence | Required content |
|---|---|
| Workflow identity | Workflow path/hash, event, source SHA, run ID, attempt. |
| Runner identity | Name, repository scope, group if any, labels, status transitions, runner version. |
| Host boundary | OS/architecture, dedicated user/UID/groups, note that no admin/root or Docker socket was used. |
| Network boundary | Documented GitHub outbound requirement and confirmation that no production target was needed. |
| Residual state | Sentinel path/hash before cleanup and proof of removal. |
| Diagnostics | Reviewed Runner/Worker log filenames/hashes; no credential files. |
| Teardown | Runner removed from GitHub; installation/VM wiped or reverted. |
| Limitations | One disposable trusted-job lab does not validate a production runner fleet. |
9. Controls required before real organizational use
- Private/trusted source boundary; no public fork code on trusted self-hosted infrastructure.
- Runner groups with selected repository/workflow access where organization/enterprise sharing is used.
- Dedicated least-privilege service account; no ambient personal/admin credentials.
- Network segmentation and explicit destination policy.
- Ephemeral/JIT one-job compute for scalable production use where practical.
- Clean immutable image/bootstrap process and defined host/runner/tool update SLA.
- No unnecessary Docker socket or privileged host mounts.
- Externally retained Runner/Worker diagnostics for ephemeral fleets.
- Credential inventory, rotation/revocation and incident isolation/rebuild procedure.
- Capacity/queue monitoring with routing labels documented as capability, not trust.
10. Simulation-only completion path
If you used no real runner, your evidence packet must clearly say so. Simulate repository scope, runner status, labels, a one-job state file and teardown in a local directory, then explain which claims remain unverified: GitHub registration, live job routing, status transitions and actual runner diagnostics.
11. Chapter 09 production operating contract
Chapter 09 adds a self-hosted runner governance contract: runner scope and group authorization, job-source trust, least-privilege service identity, network reachability, persistent-state handling, runner/host patch ownership, external diagnostics, controlled cleanup and incident-safe rebuild are all part of CI/CD correctness—not infrastructure details outside the pipeline.
Chapter 10 moves to Runner Scale Sets, Actions Runner Controller, and Kubernetes Autoscaling, where the same trust and lifecycle rules must survive automated provisioning and fleet-scale scheduling.
Knowledge check
What two independent teardown facts must the checkpoint prove?
The runner identity is removed from GitHub routing, and the local runner/host state is wiped or reverted.
Why does deleting the sentinel not prove the host is safe after a real compromise?
A compromise can persist through processes, tools, credentials or other state not covered by that cleanup; rebuild from a known clean image is stronger.
Which credentials belong in the checkpoint evidence packet?
None of the actual registration/removal token values. Record only that the time-limited flow was used and preserve non-sensitive identity/status evidence.
What should a production organization add before sharing runners broadly?
Runner-group repository/workflow access controls, clean lifecycle, network segmentation, external logs, least-privilege identities and patch/incident procedures.
What does Chapter 10 add to the Chapter 09 model?
Automated runner fleet provisioning/autoscaling through scale sets/ARC while preserving the same trust, isolation, update, logging and cleanup controls.
Official references and version notes
- Self-hosted runners concept — current responsibility boundary, hierarchy scopes, persistence model and maintenance ownership.
- Self-hosted runners reference — current supported operating systems/architectures, routing, communication, ephemeral/JIT guidance and update requirements.
- Adding self-hosted runners — current repository/organization/enterprise registration workflow and time-limited registration-token process.
-
Using self-hosted runners in a workflow
— current
runs-onlabel/group selection semantics. - Runner groups — runner groups as access-control boundaries and plan-dependent organization features.
- Managing access to self-hosted runners using groups — selected repository/workflow access policies and group governance.
- Secure use reference — current self-hosted runner hardening, public-repository warning, JIT guidance and trust-boundary risks.
- Compromised runners — current impact model for malicious workflow code, secrets, tokens and cross-repository credentials.
-
Monitoring and troubleshooting self-hosted runners
— current runner status,
_diagRunner/Worker logs and update diagnostics. - Configuring the runner as a service — current Linux/macOS/Windows service-mode procedures and status checks.
- Removing self-hosted runners — current deregistration/offline behavior, automatic stale-runner removal and local cleanup guidance.
- Actions Runner releases — public runner release history and progressive-release note.
Version-sensitive self-hosted-runner behavior was rechecked
against current primary GitHub documentation on
2026-09-09. The latest public
actions/runner release visible at verification time
is v2.337.0, released through a progressive
rollout; the repository/organization “New self-hosted runner” page
remains the authoritative version/download instruction for the
specific account. The current reference lists x64 support on
Linux/macOS/Windows, Arm64 on Linux/macOS/Windows in public
preview, and Arm32 on Linux. GitHub currently lists Ubuntu 20.04+,
Debian 10+, several other Linux families, Windows 10/11 and
Windows Server 2016/2019/2022, and macOS 11+ as supported runner
hosts. Self-hosted runners connect outbound to GitHub over HTTPS
port 443 and must remain able to reach the documented Actions
domains; no inbound job-listener port is required. By default the
runner application self-updates, but the operating system and all
other host software remain the operator’s responsibility. If
automatic runner updates are disabled, the runner must be updated
within 30 days of a new available release; critical security
updates can block new jobs until applied. GitHub recommends
ephemeral self-hosted runners for autoscaling and warns that
ephemeral/JIT runner logs should be forwarded to external storage
before production use. Runner groups can restrict
repository/workflow access; labels select capabilities but are not
an authorization boundary. The checkpoint requires no cloud
account, enterprise feature, package publication, production
credential or managed Kubernetes. Runner groups/ARC are discussed
as optional/future controls.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.