Chapter 08Lesson 04~165 minutes

GitHub-Hosted Runners, Images, Labels, Hardware, and Runtime Behavior: Diagnostics, Failure Modes, and Production Practices

Runner failures are often misdiagnosed as “GitHub Actions is flaky.” This lesson uses first-failure evidence to separate image rollout, shell/path incompatibility, missing preinstalled tools, fresh-job filesystem boundaries, architecture mismatches, hardware/resource pressure, and network assumptions. The correction should target the causal layer rather than hiding it with a rerun.

DiagnosticsImage driftPortabilityJob isolationNetwork evidence

Learning objectives

  • Apply an evidence-first diagnostic sequence to hosted-runner failures.
  • Diagnose -latest migration or image-release drift without guessing.
  • Repair cross-OS shell/path failures with portable workflow design.
  • Recognize job-isolation failures caused by assuming filesystem persistence.
  • Separate runner/network/hardware evidence from application and external-service failures.

1. Preserve evidence before rerun

  1. Record run ID, attempt, source SHA and workflow revision.
  2. Open the failing job's Set up job section and record requested label, image OS/version and included-software link.
  3. Record runner.os/runner.arch, selected tool versions and failing shell.
  4. Locate the first failing command—not the final cascade.
  5. Only then compare a previous successful run's image/tool evidence.

A rerun may land on a newer image release and destroy the clean comparison if you failed to preserve the original details.

2. Failure: treating ubuntu-latest as immutable

Yesterday: requested=ubuntu-latest image_os=ubuntu24 image_version=20260901.1 python=3.x.y
Today:     requested=ubuntu-latest image_os=ubuntu24 image_version=20260908.1 python=3.x.z
Failure: tool changed behavior

The first fix is not “rerun until green.” Reproduce with the old/new tool version if possible. If the application depends on a specific tool, declare that tool through a setup action or package pin. If the OS generation changed during a -latest migration, move to a versioned label temporarily and schedule the migration deliberately.

3. Failure: hard-coded Unix paths on Windows

# Broken cross-OS assumption — diagnostic example only
- run: cat /tmp/report.txt

On Windows this is the wrong shell/path model. Prefer runner-provided paths and platform-neutral tooling:

- shell: pwsh
  run: |
    $p = Join-Path $env:RUNNER_TEMP 'report.txt'
    'portable' | Set-Content $p
    Get-Content $p

Or intentionally branch by runner.os if the operation is OS-specific. Do not pretend one shell dialect is portable.

4. Failure: expecting job A's file in job B

If job B reports “file not found,” first verify whether the file was created in another job. Separate jobs are fresh instances. The correct fix is an explicit artifact/output/rebuild design—not moving files into a more “global-looking” directory on job A's VM.

5. Failure: relying on undocumented preinstalled software

If a command disappears or changes version, inspect the run-specific included-software manifest. Then decide whether the tool is an incidental image utility or an application dependency. Application dependencies belong in workflow/package configuration with explicit versions.

6. Failure: architecture mismatch

runner_arch=ARM64
error: package/native binary only published for x64

Do not “fix” this by emulation without understanding the support requirement. Either choose an x64 runner, add an Arm64-compatible dependency/build, or constrain the matrix until the product supports Arm64. Preserve the dependency-resolution evidence.

7. Failure: assuming identical hardware/network

Public/private standard specifications differ, macOS classes differ, larger runners differ, and network paths are external dependencies. If a job is killed for memory or disk pressure, record the runner class and actual resource observations. If an API call times out, correlate runner logs with provider status/network evidence before blaming CPU or adding retries.

8. Troubleshooting shortcuts to avoid

  • Blind reruns that replace first-failure evidence.
  • Switching to a huge runner without measuring memory/CPU/disk pressure.
  • Installing “latest” packages ad hoc in the failing job without recording versions.
  • Hard-coding public hosted IP ranges as a permanent firewall contract.
  • Moving untrusted workloads to a privileged self-hosted runner because hosted networking is inconvenient.

9. Intentionally broken portability probe

strategy:
  matrix:
    os: [ubuntu-24.04, windows-2025]
runs-on: ${{ matrix.os }}
steps:
  - name: Broken — assumes Bash and /tmp
    run: |
      echo "$RUNNER_OS" > /tmp/os.txt
      cat /tmp/os.txt

Expected evidence: Ubuntu succeeds; Windows fails before proving application behavior. Repair it with explicit shell/runner temp semantics, preserve both run attempts, and explain that this was a workflow portability failure—not a Windows application failure.

Next lesson

Checkpoint: reproducibility packet

Run a deterministic probe on Ubuntu and Windows, repair one portability failure, and document the exact runtime inputs without depending on mutable defaults.

Knowledge check

What should you preserve before rerunning a hosted-runner failure?

A file is missing in job B but existed in job A. What layer is most likely wrong?

Why can a blind rerun hide an image-drift diagnosis?

A native dependency has no Arm64 build. Is the hosted runner itself necessarily broken?

When should you choose a larger runner for a failing job?

Official references and version notes

Version and compatibility note

Version-sensitive runner behavior was rechecked against current primary GitHub documentation on 2026-09-09. At verification time, ubuntu-latest maps to Ubuntu 24.04 x64, windows-latest to Windows Server 2025 x64, and macos-latest to macOS 26 Arm64; Ubuntu 26.04 and selected Arm64 images are available in preview. GitHub states that runner-image software is typically updated weekly and that -latest migrations are gradual, so a successful run must record the actual image version/toolchain rather than treating a label as a frozen machine. For standard hosted runners, every normal job receives a fresh hosted instance; steps inside one job share that instance, while separate jobs do not share its filesystem. The ubuntu-slim single-CPU option is a special container-on-shared-VM case and is not used in mandatory labs. Current public-repository standard Linux/Windows x64 runners provide 4 CPU/16 GB RAM/14 GB SSD, while private-repository standard Linux/Windows x64 runners provide 2 CPU/8 GB RAM/14 GB SSD; macOS and Arm64 classes have different specifications. Hardware figures are therefore plan/repository/image inputs, not universal constants. Larger runners are optional organization/enterprise features and are not required for course completion. Official actions/setup-python v7.0.0 is pinned in examples to 5fda3b95a4ea91299a34e894583c3862153e4b97. The intentionally broken examples are inert teaching snippets; the safe repair uses runner-provided paths and explicit shells.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.