Chapter 09Lesson 04~180 minutes

Self-Hosted Runners, Runner Groups, Isolation, and Operational Security: Diagnostics, Failure Modes, and Production Practices

Self-hosted runner failures often look like ordinary CI errors while the real cause is routing, host persistence, privilege, network policy or runner maintenance. Production diagnosis must preserve both GitHub run evidence and host-side diagnostics before changing the runner, because a “fix” that deletes logs or rebuilds blindly can destroy the evidence needed to explain a compromise or repeated outage.

DiagnosticsUntrusted codeDocker socketUpdate failuresIncident response

Learning objectives

  • Diagnose queued/offline jobs by separating workflow selection from group/label matching and runner connectivity.
  • Recognize unsafe public-fork, root/admin, shared-state and Docker-socket runner designs.
  • Identify stale-runner/update failures and preserve _diag evidence before remediation.
  • Use an incident-safe sequence that isolates a potentially compromised runner before cleanup or rebuild.
  • Repair an intentionally unsafe workflow by moving untrusted code away from trusted self-hosted infrastructure.

1. Evidence-first diagnostic sequence

  1. Preserve workflow run ID/attempt, event, ref/SHA and workflow revision.
  2. Record requested runs-on group/labels and whether a run/job was actually created.
  3. Record runner name/scope/group/labels/status/version from GitHub Settings.
  4. Preserve relevant _diag/Runner_*.log and Worker_*.log plus host service logs.
  5. Record service user, host patch state and network reachability.
  6. If compromise is plausible, stop new job assignment and isolate the host before “cleaning.”
  7. Apply the smallest safe correction, or rebuild from a known image if integrity cannot be proven.

2. Failure: job remains queued

Evidence Likely layer Next check
No matching runner online Capacity/status Runner process/service and connectivity.
Runner online but wrong custom label Routing Requested versus registered labels.
Correct labels but group denies repository Authorization Runner-group repository/workflow access policy.
Runner too old Maintenance Runner version/update logs and current required release.
Runner receives assignment but not job Connectivity/process Runner diagnostics; GitHub can requeue if pickup fails.

3. Critical failure: untrusted PR code on trusted self-hosted compute

Do not “fix” this by giving the workflow fewer repository permissions while leaving untrusted code on the same host.

permissions: read-all can reduce GitHub API scope, but malicious code can still attack the persistent host, local credentials and reachable network.

# UNSAFE DESIGN — for diagnosis only; do not run
on: pull_request
jobs:
  test:
    runs-on: [self-hosted, trusted-internal]
    steps:
      - checkout untrusted PR code
      - execute repository test script

Repair the trust boundary: run untrusted PR validation on GitHub-hosted isolated runners. Allow self-hosted runners only for trusted source events/branches after the required review/control boundary.

4. Failure: runner executes as root/admin

Root/admin may make setup easier but enlarges every workflow’s local authority. Record the actual service identity and group memberships. Replace ambient administrator execution with a dedicated unprivileged account and explicit narrow elevation only where a reviewed operation requires it.

id
ps -o user,pid,ppid,cmd -C Runner.Listener 2>/dev/null || true
getent group docker 2>/dev/null || true

5. Failure: shared workspace or credentials

A later job sees a prior job’s file, modified tool or local credential. Preserve a hash/inventory of the unexpected state. Do not simply delete the file and rerun. Identify why the host permitted the state to survive, whether another job consumed it, and whether any credential exposed to the host must be revoked.

6. Failure: Docker socket exposes host control

Evidence:
  service user belongs to docker group
  /var/run/docker.sock is writable by runner job
  job can create privileged/mounted containers
Conclusion:
  containerized build does not equal host isolation

Repair by removing unnecessary daemon access, isolating the runner/daemon per job, or using a runner design whose host can be destroyed after the job. Never describe “inside Docker” as proof that untrusted code cannot reach the host.

7. Failure: updates blocked indefinitely

If a runner has --disableupdate, compare its version with the account’s current runner download instructions and release feed. GitHub currently stops queueing jobs to a manually managed runner if an available update is not applied within 30 days, and critical security updates can require immediate update before new jobs are queued.

Fix the image/update pipeline, not the symptom. Re-enable automatic update or publish a new reviewed runner image and replace the old fleet.

8. Failure: logs deleted before incident review

Deleting _diag or destroying an ephemeral host before forwarding diagnostics turns a diagnosable event into speculation. Production ephemeral fleets should stream runner/worker logs externally. In an incident, isolate first, preserve evidence second, then rebuild/revoke.

9. Incident-safe response for suspected compromise

  1. Stop or disable the affected runner/group from receiving new work.
  2. Restrict network access while preserving forensic evidence.
  3. Preserve GitHub run/job logs, runner diagnostics and relevant host audit logs.
  4. Inventory credentials/secrets/tokens that the affected jobs or host could access.
  5. Revoke/rotate exposed credentials and invalidate long-lived trust where appropriate.
  6. Rebuild the runner host from a known clean image; do not rely on workspace deletion as integrity proof.
  7. Restore routing only after root cause and preventive controls are documented.

10. Intentionally broken routing example

# Broken but non-destructive: no runner has all requested labels
on: workflow_dispatch
permissions: {}
jobs:
  probe:
    runs-on: [self-hosted, linux, x64, chapter09-typo]
    steps:
      - run: echo never-started-until-a-matching-runner-exists

Expected evidence: the workflow run/job exists but remains queued because no matching runner is available. Preserve the queued state, compare the requested label with the registered chapter09-lab label, correct only the typo, then rerun. Do not deregister/recreate the runner as a first response.

Next lesson

Checkpoint: prove a disposable runner lifecycle

Provision, route, execute, inspect residue, clean, deregister and document the controls required before any real organization use.

Knowledge check

A self-hosted job is queued for several minutes. Should you recreate the runner immediately?

Why is read-only GITHUB_TOKEN not enough to make untrusted code safe on a trusted runner?

What should happen before wiping a runner suspected of compromise?

What does a writable Docker socket imply?

A job asks for chapter09-typo but the runner has chapter09-lab. What is the least destructive fix?

Official references and version notes

Version and compatibility note

Version-sensitive self-hosted-runner behavior was rechecked against current primary GitHub documentation on 2026-09-09. The latest public actions/runner release visible at verification time is v2.337.0, released through a progressive rollout; the repository/organization “New self-hosted runner” page remains the authoritative version/download instruction for the specific account. The current reference lists x64 support on Linux/macOS/Windows, Arm64 on Linux/macOS/Windows in public preview, and Arm32 on Linux. GitHub currently lists Ubuntu 20.04+, Debian 10+, several other Linux families, Windows 10/11 and Windows Server 2016/2019/2022, and macOS 11+ as supported runner hosts. Self-hosted runners connect outbound to GitHub over HTTPS port 443 and must remain able to reach the documented Actions domains; no inbound job-listener port is required. By default the runner application self-updates, but the operating system and all other host software remain the operator’s responsibility. If automatic runner updates are disabled, the runner must be updated within 30 days of a new available release; critical security updates can block new jobs until applied. GitHub recommends ephemeral self-hosted runners for autoscaling and warns that ephemeral/JIT runner logs should be forwarded to external storage before production use. Runner groups can restrict repository/workflow access; labels select capabilities but are not an authorization boundary. The unsafe public-PR example is intentionally presented as inert text. The executable broken example only creates a queued job through a label mismatch and has no side effect.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.