Chapter 16Lesson 01~175 minutes

GitHub-Hosted and Self-Hosted Runners, Labels, Groups, Scaling, and Runner Security: Concepts, Architecture, and Mental Model

Chapter 15 showed that job topology determines where state and dependencies flow. Chapter 16 moves one layer lower: the machine that actually receives a job is part of the security boundary. A runner can be disposable GitHub-managed compute or a machine your team owns with network reach, cached credentials, mutable tools, and a lifecycle that must be governed.

Runner trustHosted vs self-hostedEphemeral lifecycleRouting & scaling

Learning objectives

  • Explain the lifecycle and ownership boundary of standard GitHub-hosted, larger GitHub-hosted, and self-hosted runners.
  • Predict how runs-on, labels, groups, repository/organization scope, online state, and idleness determine routing.
  • Explain why persistent self-hosted runners create workspace, credential, network, and lateral-movement risk that ephemeral runners reduce but do not eliminate.
  • Separate runner registration identity, workflow GITHUB_TOKEN, repository secrets, and machine/network credentials.
  • Reason about queueing, capacity, concurrency, autoscaling, patch ownership, and cost as observable runner controls.

Availability: The mandatory chapter path uses GitHub.com, GitHub Free, a disposable public personal repository, and standard GitHub-hosted runners. Standard hosted runners are currently free and unlimited for public repositories. Self-hosted registration is optional and uses a fresh private disposable repository plus a dedicated disposable VM. Larger runners are optional because GitHub currently limits them to organizations/enterprises on GitHub Team or Enterprise Cloud and bills their execution per minute.

1. The problem: “runs-on” is a trust decision, not a hardware preference

Chapter 15 treated a job as an isolated scheduling unit. That abstraction is useful only if you also understand the machine beneath it. A job that lands on a GitHub-managed fresh VM has a different residue, network, patching, and ownership model from a job routed to a workstation, long-lived VM, Kubernetes pod, or datacenter host that your team registered as self-hosted.

A runner can see everything the job process can see: checked-out code, temporary files, job-scoped credentials made available by the workflow, and network endpoints reachable from the machine. Therefore runner selection participates directly in the blast radius of a malicious pull request, a compromised dependency, a leaked secret, or an administrator mistake.

2. Mental model: GitHub routes a hosted job to an eligible trust zone

Concept / workflow diagram
              flowchart TD
                W["Workflow job runs-on requirement"] --> R["GitHub Actions routing"]
                R --> H["Standard GitHub-hosted pool GitHub-managed lifecycle"]
                R --> L["Larger hosted runner/group Org policy + billed capacity"]
                R --> S["Self-hosted scope/group Labels + online/idle state"]
                S --> E["Ephemeral runner One job then de-register"]
                S --> P["Persistent runner Reusable host"]
                E --> N["Reachable network/resources"]
                P --> N
            

The workflow declares eligibility requirements; GitHub matches those requirements to runner inventory and access policy. The final machine’s lifecycle and network are independent security dimensions. An eligible runner is not automatically a safe runner for every workload.

The arrows do not imply that GitHub “trusts” one runner more than another. They show routing. Your governance must define which repositories, events, branches, and job classes are allowed to reach each execution zone.

3. Standard GitHub-hosted runners: managed, disposable execution

For the standard public-repository labels used in this chapter, GitHub provisions managed execution environments. Except for the documented single-CPU container-backed option, standard hosted runners are fresh virtual machines for jobs. GitHub owns the base image and runner infrastructure lifecycle.

Do not interpret ubuntu-latest as “the newest Ubuntu release that exists.” GitHub documents -latest as its latest stable runner image. If release reproducibility matters, select an explicit image label such as ubuntu-22.04 or ubuntu-24.04 and still record runtime/tool versions because hosted images evolve over time.

Property Standard GitHub-hosted Operational consequence
Machine lifecycle GitHub-managed fresh execution environment per job for normal VM labels Low residue between unrelated jobs; no host cleanup runbook for the learner.
OS/tool image GitHub-maintained image selected by workflow label Tool versions can change as images are updated; inspect rather than assume.
Network GitHub-hosted network unless an eligible private-networking feature is configured Internal reach is not implicit.
Capacity GitHub-hosted service supplies runner capacity subject to Actions limits/policy No VM fleet autoscaler to operate for standard hosted runners.
Cost Standard runners free/unlimited for public repositories Private-repo billing/quotas are plan dependent; verify current plan.

4. Self-hosted runners: you own the machine lifecycle

A self-hosted runner is a machine you deploy and register with GitHub Actions. GitHub provides the runner application and control-plane integration, but you own the operating system, packages, network routes, machine identity, disk cleanup, monitoring, and most patching. GitHub notes that the runner application can auto-update by default; that does not patch your OS, Docker daemon, language runtimes, browsers, compilers, or local agents.

Self-hosted runners are free to use from the GitHub Actions usage perspective, but the compute, storage, network, administration, observability, and incident-response costs are yours.

5. Labels, groups, scope, and state are routing predicates

Self-hosted runners receive default labels such as self-hosted, operating system, and architecture unless configured otherwise. Custom labels can express capabilities such as gpu, arm64-build, or—more usefully for governance—a deliberately named trust/capability class such as isolated-ci. When runs-on specifies multiple labels, an eligible runner must match all of them.

Runner groups add an organization/enterprise access boundary around larger or self-hosted runners. A group controls which repositories/workflows may use a set of runners; labels then narrow capability inside the allowed set. Custom groups are organization/plan/policy dependent, so the mandatory personal-repository lab uses labels only and teaches groups with a fixture.

jobs:
  build:
    # Eligible only if one runner satisfies every listed label.
    runs-on: [self-hosted, linux, x64, isolated-ci]

  deploy_fixture_only:
    # Organization runner-group pattern; not executed in the mandatory lab.
    runs-on:
      group: trusted-deploy
      labels: linux-x64

Routing also depends on runner state. GitHub currently assigns a matching online/idle self-hosted runner; if it does not accept the job within 60 seconds, the job is re-queued. If no eligible runner becomes available, a queued job can wait up to 24 hours before failing.

6. Persistent versus ephemeral: registration lifetime changes the risk

Lifecycle Persistent self-hosted runner Ephemeral self-hosted runner
Job count Can receive many jobs over time Registered for one job; GitHub de-registers it after the job.
Residue risk Workspace/tool/process/credential residue must be controlled between jobs Fresh image + destruction can sharply reduce cross-job residue.
Compromise persistence Compromised host may receive later jobs One-job routing limits future assignment, but the job can still compromise reachable resources.
Autoscaling guidance GitHub does not recommend persistent runners for autoscaling GitHub recommends ephemeral runners for autoscaling.
Logs Local runner logs remain on the host unless exported GitHub recommends forwarding ephemeral runner application logs externally before production use.

--ephemeral is a registration behavior, not a VM destruction mechanism. After GitHub de-registers the runner, your infrastructure automation must wipe or destroy the machine/container. Otherwise the compromised filesystem still exists even though GitHub no longer routes jobs to it.

7. Four identities beginners often collapse into “the runner token”

Credential / identity Purpose Lifetime / scope Never infer
Runner registration token Registers a machine with repository/org/enterprise runner inventory Short-lived administrative enrollment material It does not become the workflow job identity.
Runner removal token Authenticates runner removal through the config script Short-lived administrative removal material It is not a cleanup substitute for destroying the VM.
GITHUB_TOKEN Job-scoped GitHub App installation token for workflow API/repository operations Per job; effective permissions depend on workflow/repository policy A self-hosted machine does not automatically get broad GitHub access.
Machine/cloud/network identity Access granted by the host environment: IAM role, managed identity, local credential, network position Infrastructure-defined GitHub workflow permissions do not constrain an overprivileged machine identity.

The last row is why runner security is a separate control plane. Even a workflow with permissions: {} can still reach an internal unauthenticated service if the machine’s network can reach it.

8. Trust zones: separate untrusted code from sensitive networks

GitHub explicitly recommends using self-hosted runners only with private repositories because a fork of a public repository can potentially submit pull-request code that executes on the self-hosted machine. In production, go further: classify event/source trust and route it to an execution zone with only the minimum network and machine identity required.

Workload Safer default zone Why
Public/fork PR validation Standard GitHub-hosted runner Treat contributed code and event strings as untrusted; avoid internal network reach.
Trusted main-branch build needing custom compiler Ephemeral self-hosted build zone with narrow egress/internal access Custom hardware/software without turning the host into a permanent bridge.
Production deployment Separate protected deploy zone or OIDC-capable hosted execution where possible Do not mix arbitrary build code with production credentials/network reach.
Security-sensitive signing Dedicated hardened service/runner with explicit policy and provenance controls Keys and signing authority require stronger isolation than general CI.

9. Capacity, queueing, scaling, and cost are observable controls

Standard GitHub-hosted runners inherently scale without you operating a fleet. Self-hosted fleets require explicit capacity management. GitHub currently recommends Actions Runner Controller (ARC) as the Kubernetes-based reference implementation for autoscaling self-hosted runners; the Runner Scale Set Client is available for custom infrastructure outside Kubernetes. GitHub also supports autoscaling signals from workflow_job webhooks, but its documentation notes delivery timeliness can introduce delay and reliability concerns.

Larger GitHub-hosted runners sit between standard hosted and self-hosted: GitHub manages the VM, while organizations configure size, groups, networking/capacity. They are currently Team/Enterprise Cloud features and are billed per minute even for public repositories. That makes “bigger runner” a cost decision, not a free speed button.

10. Update responsibility and reproducibility

Runner drift has two forms. Hosted-image drift is managed by GitHub but can change preinstalled tools; capture versions and prefer explicit image labels when compatibility matters. Self-hosted drift is yours: OS updates, local packages, container runtime, CA certificates, build tools, and the runner application must all be governed.

If self-hosted runner auto-update is disabled using GitHub’s documented --disableupdate mode, GitHub currently requires the runner software to be updated within 30 days of a new version, and critical security updates can block job queueing sooner. A production image pipeline should therefore rebuild and roll runner images deliberately rather than allowing snowflake hosts to age indefinitely.

11. Read-only inspection first: prove hosted job and self-hosted inventory separately

Hosted job records and self-hosted runner inventory are different API resources. A standard hosted runner appears in the job record but is not a self-hosted registration in your repository.

REPO="OWNER/REPOSITORY"
RUN_ID="123456789"

gh run view "$RUN_ID" -R "$REPO" \
  --json event,headSha,status,conclusion,jobs,url \
  --jq '{event,headSha,status,conclusion,jobs:[.jobs[]|{name,conclusion,startedAt,completedAt,databaseId}]}'

gh api -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/actions/runs/$RUN_ID/jobs?filter=latest&per_page=100" \
  --jq '.jobs[] | {name,status,conclusion,runner_name,runner_group_name,labels,started_at,completed_at}'

gh api -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/actions/runners?per_page=100" \
  --jq '{total_count,runners:[.runners[]|{id,name,os,status,busy,ephemeral,version,labels:[.labels[].name]}]}'

The last query requires repository administration access because runner inventory is an administrative resource. On a personal disposable repository where you are owner, it should return zero self-hosted registrations before the optional lab.

12. Why this matters in DevOps

A production pipeline is only as isolated as the infrastructure executing it. Runner policy determines whether untrusted code can touch internal networks, whether one compromised job contaminates later jobs, whether capacity spikes become a 24-hour queue, whether patch drift breaks reproducibility, and whether a “faster” runner quietly creates recurring cost. Runner selection therefore belongs in change-control and security design, not only in build engineering.

13. Lesson summary

Standard GitHub-hosted runners minimize fleet administration and job-to-job residue. Self-hosted runners trade that simplicity for control of hardware, software, and networking—along with full lifecycle responsibility. Labels express eligibility; groups add repository access boundaries where available; state/capacity determine whether a matching runner can accept work. Ephemeral one-job runners reduce persistence and are GitHub’s recommended autoscaling pattern, but network reach and machine identity still define the job’s blast radius.

Knowledge check

A workflow has permissions: {} but runs on a self-hosted VM that can reach an unauthenticated internal admin service. Is the job effectively harmless?

Why is --ephemeral not sufficient cleanup by itself?

A job requests [self-hosted, linux, isolated-ci]. A runner has only self-hosted and linux. Will it run there?

Why prefer explicit ubuntu-24.04 over ubuntu-latest for a compatibility-sensitive test?

What does a runner group add beyond a label?

Next lesson

Next: GitHub-Hosted and Self-Hosted Runners, Labels, Groups, Scaling, and Runner Security: Guided Hands-On Workflow and Core Operations

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.