Chapter 16Lesson 03~160 minutes

GitHub-Hosted and Self-Hosted Runners, Labels, Groups, Scaling, and Runner Security: Configuration, Design Choices, and Tradeoffs

Runner choice is not a preference between “fast” and “custom.” It decides who patches the machine, who pays for idle capacity, which network the job can reach, how clean the workspace is, which repositories may route work there, and how an incident can spread. This lesson turns those consequences into a runner policy.

Runner policyTrust zonesAutoscalingCost governance

Learning objectives

  • Choose GitHub-hosted versus self-hosted runners using trust, reproducibility, network, operations, and cost criteria.
  • Choose persistent versus ephemeral self-hosted lifecycle and explain why autoscaling changes the preferred lifecycle.
  • Place repository-level runners, organization runner groups, labels, and trust zones into a coherent routing policy.
  • Compare standard hosted, larger hosted, ARC/scale-set, and bespoke self-hosted capacity strategies without making a paid feature mandatory.
  • Write a decision record that justifies runner scope, network reach, patch cadence, logging, and emergency removal.

Availability boundary: The required reasoning is free-compatible. Standard hosted public-repository runs and repository-level self-hosted runners are available without buying larger runners. Organization runner groups, private networking, enterprise policies, and larger-runner entitlements vary by organization/plan/deployment; use the current GitHub UI/docs for your account. Larger runners are currently Team/Enterprise Cloud only and billed per minute.

1. GitHub-hosted versus self-hosted: choose the owner of risk

Dimension Standard GitHub-hosted Self-hosted
Machine lifecycle GitHub provisions/disposes normal hosted VMs You provision, isolate, clean, decommission, and often autoscale.
OS/tool maintenance GitHub maintains runner images; you adapt to image changes You patch OS/tools/runtime and govern runner-app update policy.
Network reach Public GitHub-hosted network by default Whatever the machine can reach—potentially powerful and dangerous.
Hardware/software customization Bounded by offered images/specs Full control, including specialized hardware and local tools.
Residue Fresh hosted environment sharply limits cross-job residue Persistent runners can retain files/processes/credentials unless cleaned; ephemeral image destruction is safer.
Operations burden Low Fleet, image, monitoring, scaling, logs, incident response, cost.
Best default General CI and untrusted/fork validation Only when a concrete hardware/software/network requirement justifies the added trust boundary.

Self-hosted should be the answer to a requirement, not a prestige upgrade. If “we need faster CI” is the only requirement, first measure queue/runtime and evaluate matrix, caching, standard-image, or eligible larger-runner options before building a fleet.

2. Persistent versus ephemeral self-hosted: the default changes with scale

A persistent runner can be practical for a tightly controlled private repository and carefully cleaned host, but persistence couples jobs through machine history. The risk rises with repository count, contributor count, and network privilege. For autoscaling, GitHub explicitly recommends ephemeral runners and does not recommend persistent runners.

Question Persistent Ephemeral
What happens after a job? Runner remains eligible for later jobs GitHub de-registers after one job; infrastructure should destroy the instance.
Image drift Accumulates unless rebuilt/managed Image pipeline can create a known baseline per job/fleet generation.
Forensics Local history may exist, but mutable host complicates attribution Forward runner logs externally before destruction; retain immutable run/job IDs.
Capacity Static/manual unless you build scaling around it Natural fit for queue-driven provisioning/scale sets.
Compromise containment Compromise can survive into future jobs Limits future assignment, but current job still has network/machine authority.

3. Repository scope versus organization scope/groups

Repository-level self-hosted runner: dedicated to one repository. It is the smallest routing scope and is appropriate for the optional Chapter 16 private lab. Organization-level runner: can serve multiple repositories according to organization policy. Runner groups are then the key access-control layer for separating fleets such as untrusted-build, trusted-build, and deploy-prod.

Do not encode security only in label names. A label named prod is descriptive routing metadata. Restrict repository access through the runner group (where available), workflow/repository policy, protected environments, and machine/cloud permissions as well.

4. Network access: design from “deny by default” outward

Self-hosted runners can communicate with any endpoint allowed by their host network. Treat network reach as capability. Start with GitHub-required outbound HTTPS and the minimum package/test endpoints, then add internal destinations only for workloads that genuinely require them.

Zone Inbound Outbound/internal Typical job
Public/untrusted CI No unsolicited inbound Internet/package endpoints only; no internal routes Fork PR lint/test.
Trusted private build No unsolicited inbound Selected package mirrors/build dependencies Main-branch compile/test.
Deployment zone No unsolicited inbound Only deployment control plane/targets Protected deployment job.
Signing zone No unsolicited inbound Narrow signing/HSM interface only Release signing/provenance.

Network segmentation complements GitHub permissions. One cannot replace the other.

5. Four capacity models and when they fit

Model What scales Good fit Tradeoff
Standard GitHub-hosted GitHub-managed capacity General CI; public/free labs; untrusted code Less hardware/network customization.
Larger GitHub-hosted GitHub-managed larger VMs with configured max concurrency/groups Need more resources/static IP/private-network features without self-hosted fleet Team/Enterprise Cloud; billed per minute.
ARC runner scale sets Ephemeral self-hosted runners on Kubernetes Platform team already operates Kubernetes and needs elastic custom runners Kubernetes + ARC operational complexity.
Custom scale-set client / own autoscaler Your VM/container provisioning around GitHub scale-set APIs Non-Kubernetes infrastructure with strong platform engineering You own provisioning, telemetry, cleanup, and failure handling.

The workflow_job webhook can also drive autoscaling decisions, but GitHub documents delivery-timeliness concerns. Use queue data as an input to scaling, not as an assumption that every webhook arrives instantly and exactly once.

6. Cost is more than Actions minutes

For a public repository, standard hosted runner use is currently free/unlimited. Self-hosted Actions execution itself does not incur GitHub-hosted runner charges, but you pay for compute, disks, bandwidth, images, cluster/control-plane operations, patching, logging, and human time. Larger runners are billed per minute even on public repositories.

A runner that is twice as large but halves runtime may or may not reduce total cost; a self-hosted fleet with low utilization may be more expensive than hosted capacity even if the raw VM rate looks lower. Measure queue time, runtime, concurrency, utilization, failure/retry rate, and operational labor together.

7. Labels should describe stable capability, not accidental host names

Prefer labels such as linux, arm64, gpu, isolated-ci, or deploy-prod over labels tied to one machine name. A capability label makes replacement/autoscaling possible without editing every workflow.

Use a small reviewed vocabulary. If workflows request arbitrary compound labels and administrators add labels ad hoc, routing becomes invisible policy. Record who may create labels, who may change group access, and which label combinations correspond to which trust zones.

8. Image/tool update policy: pin what you control, observe what you do not

For hosted runners, choose explicit OS labels where compatibility requires it and record tool versions in release-critical logs. For self-hosted runners, prefer immutable/prebaked images and rebuild the runner rather than patching long-lived snowflakes in place. If you disable automatic runner-app updates, integrate the documented 30-day update requirement into image rollout alerts and test new runner releases before production promotion.

9. Worked decision table

Scenario Recommended default Maintainability Security/governance Reliability/compatibility Cost
Public OSS PR test Standard hosted explicit Ubuntu label Very high Strong isolation from your network Known hosted image; record versions Free/unlimited standard public path.
Private monorepo needs 32 GB RAM only Evaluate larger hosted runner first High GitHub-managed; group access if org Avoid custom fleet solely for RAM Per-minute; compare measured runtime.
Private firmware build needs USB lab hardware Self-hosted dedicated trusted zone Medium/low Strict repo/group access; no untrusted PR code Hardware compatibility drives decision Own hardware/ops cost.
High-volume private builds need custom compiler Ephemeral self-hosted scale set Medium One-job lifecycle, narrow network, external logs Prebaked image reduces drift Infra + platform ops; can scale to demand.
Production deploy to internal API Separate deployment runner/group or hosted OIDC/private-network option Medium Strongest repo/environment/network controls Keep deploy tools stable and small Pay for isolation rather than mixing with builds.

10. Runner architecture decision record template

Runner class: <standard-hosted | larger-hosted | self-hosted-ephemeral | self-hosted-persistent>
Workload trust: <public-fork | internal-PR | protected-branch | deployment>
Repository/org scope: <...>
Group and label contract: <...>
Machine identity permissions: <none/minimum>
Network destinations allowed: <...>
Secrets available: <none or named class>
Image/tool provenance and patch cadence: <...>
Runner application update policy: <...>
Autoscaling/capacity signal: <...>
Queue SLO and max concurrency: <...>
Runner/application log retention: <...>
De-registration + machine destruction procedure: <...>
Cost owner / budget signal: <...>
Emergency isolation procedure: <...>

This record makes “why this job runs there” reviewable during incident response and architecture changes.

11. Lesson summary

The runner decision is a multi-axis policy: lifecycle, access scope, labels/groups, network, machine identity, patching, logging, scaling, and cost. GitHub-hosted is the safer operational default for general CI and untrusted inputs. Self-hosted is justified by concrete hardware/software/network requirements and should trend toward ephemeral, isolated, observable execution as scale and privilege increase.

Knowledge check

A team wants self-hosted runners only because dependency installation is slow. What should they evaluate first?

Why is deploy-prod as a label not sufficient production protection?

Why can an ephemeral runner still be dangerous?

What should happen to runner logs if ephemeral machines are destroyed after jobs?

What is the main cost trap of an oversized runner?

Next lesson

Next: GitHub-Hosted and Self-Hosted Runners, Labels, Groups, Scaling, and Runner Security: Diagnostics, Failure Modes, Security, and Performance

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.