Chapter 01Lesson 04~120 minutes

GitHub Actions Foundations, Automation Model, and CI/CD Concepts: Diagnostics, Failure Modes, and Production Practices

Troubleshooting GitHub Actions is a state-localization problem. Random reruns and YAML edits destroy evidence and can repeat side effects. This lesson engineers two safe failures—one where no run is selected and one where a second job incorrectly assumes it shares the first job’s filesystem—then repairs each failure by identifying the owning layer and applying the smallest correction.

DiagnosticsFailure modesJob boundariesFirst failureRecovery

Learning objectives

  • Use a consistent evidence-first diagnostic sequence from event/workflow selection through run, queue, runner, step, artifact/deployment, and external state.
  • Distinguish “no run was created” from “a run was created but a job/step failed.”
  • Diagnose a cross-job filesystem failure without misclassifying it as a runner or GitHub platform outage.
  • Explain why blind reruns, broad permissions, context dumps, and delete/recreate tactics are unsafe troubleshooting defaults.
  • Preserve first-failure evidence and choose a least-destructive repair or rollback.

1. Use the same diagnostic sequence every time

When a workflow behaves unexpectedly, start at the earliest state that could explain the symptom. Moving backward and forward randomly creates noise. A reliable sequence is:

Evidence-first diagnostic sequence
flowchart TD
  A[Preserve run ID / attempt / first failure] --> B[Event + ref/SHA + workflow revision]
  B --> C[Conditions + inputs + permissions]
  C --> D[Job graph + queue]
  D --> E[Runner + image + toolchain]
  E --> F[Step / action / runtime / network]
  F --> G[Outputs / artifacts / caches]
  G --> H[Environment / deployment]
  H --> I[External target + governance]
  I --> J[Smallest correction + targeted verification]

If no run exists, do not start at runner logs. If a job never left the queue, do not edit a shell command. If a deployment request succeeded but the service is unhealthy, do not assume the workflow parser is at fault. Locate ownership first.

2. Preserve first-failure evidence before changing anything

Before rerunning or editing:

  • record repository, workflow name/path, run URL/ID, attempt, event, ref, head SHA, and triggering actor;
  • record which jobs were created, skipped, queued, cancelled, failed, or succeeded;
  • save the first failing step's log and relevant setup/runner metadata;
  • note permissions, inputs, secrets availability assumptions, and any external API/resource IDs;
  • if side effects occurred, record exactly which external resource changed and its current state.

A rerun is a new attempt against a possibly changed external world. A registry tag may now exist; a deployment may be partially applied; a rate limit may have reset; a runner image may have updated. Preserve attempt 1 before creating attempt 2.

3. Failure mode A: valid workflow, no run selected

Start with the working Lesson 2 workflow. In the disposable repository only, temporarily change its push filter:

on:
  push:
    branches: [release-only-demo]
  workflow_dispatch:

Commit this change to main, then make one harmless follow-up commit to main. The follow-up push does not match release-only-demo, so you should not expect a push-triggered run for that event. The absence of a run is not a “failed job”; the workflow selection condition was not satisfied.

Question Evidence Conclusion
Did GitHub receive a push? Commit/push exists in repository history. Yes.
Did the filter match? Workflow says release-only-demo; event branch is main. No.
Was a run ID created for that workflow/event? No matching run in Actions history/API. No.
Should you inspect runner logs? No runner was assigned because no run/job existed. No.

Repair: restore the intended main filter, commit the correction, and verify a later controlled push produces a run. Preserve the non-run explanation in your note rather than calling it an Actions outage.

4. Failure mode B: jobs do not share a hosted-runner filesystem

Create a separate disposable workflow .github/workflows/ch01-broken-boundary.yml:

name: Chapter 01 - Broken Job Boundary

on:
  workflow_dispatch:

permissions: {}

jobs:
  producer:
    runs-on: ubuntu-24.04
    steps:
      - name: Create local state
        shell: bash
        run: |
          set -euo pipefail
          echo "created-by=$GITHUB_JOB" > state.txt
          test -s state.txt
          cat state.txt

  consumer:
    needs: producer
    runs-on: ubuntu-24.04
    steps:
      - name: Incorrectly expect producer filesystem
        shell: bash
        run: |
          set -euo pipefail
          echo "consumer runner: $RUNNER_OS / $RUNNER_ARCH"
          test -f state.txt
          cat state.txt

Trigger it manually. Expected state: producer succeeds; consumer starts on its own hosted job environment and fails at test -f state.txt. The dependency needs: producer controls job ordering; it does not copy the producer filesystem.

5. Interpret the broken run before fixing it

Collect this minimum evidence:

  1. Run ID and attempt.
  2. Event = workflow_dispatch.
  3. Producer conclusion = success.
  4. Consumer conclusion = failure.
  5. Producer log proves state.txt existed inside producer.
  6. Consumer log proves its own step could not find state.txt.
  7. Runner setup logs show the jobs were independently scheduled.

Do not classify the problem as “GitHub lost my file.” The file was never transferred across the job boundary. The workflow's dataflow design was wrong.

Causal diagnosis: needs establishes dependency and exposes supported job outputs; it is not a shared-disk primitive. Cross-job file transfer requires an explicit mechanism such as artifacts, which Chapter 13 teaches. For this Chapter 01 example, the smallest repair is to keep filesystem-dependent steps in the same job.

6. Apply the smallest equivalent repair

For this tiny workflow, combine the two operations into one job:

jobs:
  same_job:
    runs-on: ubuntu-24.04
    steps:
      - name: Create local state
        shell: bash
        run: echo "created" > state.txt

      - name: Read local state
        shell: bash
        run: |
          set -euo pipefail
          test -f state.txt
          cat state.txt

This repair matches the actual requirement: two steps need one ephemeral filesystem. Do not introduce artifact infrastructure merely because a tutorial says “use artifacts”; choose a cross-job transfer only when jobs genuinely need separate runners/lifecycles.

Rerun the corrected workflow as a new run after committing the fix. Compare the failing run and corrected run rather than deleting the failure.

7. Separate common failure layers

Symptom Likely layer to inspect first Evidence
No run exists Event/filter/workflow location/default-branch rules. Event/ref, workflow file at revision, filter configuration.
Job is queued Runner availability/labels/concurrency/policy. Queue state, requested runs-on, runner availability.
Step exits 127 / command not found Runner image/toolchain/shell. Runner image, PATH, tool version/setup.
API returns 403 Token/permission/policy/event trust. Requested permission, effective event restrictions, endpoint requirements.
File missing in another job Dataflow/job boundary. Producing job log, consumer workspace, explicit transfer configuration.
Deployment job green, service unhealthy External provider/target. Provider deployment/resource ID, health checks, service logs.
Rerun causes duplicate release Idempotency/side-effect design. Existing release/package/resource identity and prior attempt.

8. Troubleshooting shortcuts that create new incidents

  • Blind rerun: can repeat publication, deployment, or destructive API calls.
  • Grant write-all: can hide the real missing permission while expanding blast radius.
  • Print every context/environment variable: can expose sensitive values and creates noisy evidence.
  • Disable TLS verification: changes the trust model instead of solving identity/network configuration.
  • Switch to a privileged self-hosted runner: may bypass the symptom while exposing trusted infrastructure to workflow code.
  • Delete failed runs/logs before understanding them: destroys the first-failure record.
  • Change several YAML keys at once: makes the correction non-causal and difficult to review.

9. A green indicator has a bounded meaning

Apply this interpretation discipline in production:

  • Run created: event/workflow selection created an execution record.
  • Job started: required conditions allowed the job and a runner became available.
  • Step succeeded: that step process/action concluded successfully under its local context.
  • Job succeeded: the job's required steps reached a successful conclusion.
  • Check succeeded: GitHub recorded the check result expected by repository policy.
  • Artifact uploaded: an object was stored; verify its identity/digest/content before promotion.
  • Environment approved: authorization occurred; rollout is not thereby healthy.
  • Deployment succeeded in GitHub: deployment status was reported; confirm provider/target health independently.

10. Production incident mini-runbook

  1. Freeze cleanup and preserve run/attempt evidence.
  2. Write the exact expected state transition that failed.
  3. Locate the earliest layer where observed state differs from expected.
  4. Confirm source/ref/SHA and workflow revision before editing.
  5. Check permissions and trust context before adding credentials.
  6. Check runner/image/toolchain before blaming application code for environment drift.
  7. Inspect external target evidence if any side effect crossed GitHub's boundary.
  8. Choose one reversible correction.
  9. Rerun the smallest safe scope; do not automatically replay deployment side effects.
  10. Keep both original and corrected evidence with the causal explanation.

11. Lab cleanup

Restore the main Chapter 01 workflow's intended branch filter. Keep the broken-boundary workflow if you want the failed run as a teaching fixture, or remove that one file after preserving its run ID/log evidence. Do not delete unrelated workflows or change organization policy. The lab used no secret, deployment, package, cloud account, or self-hosted runner, so rollback is limited to repository workflow files and disposable run history.

Next lesson

Checkpoint Lab

Combine the chapter into one evidence packet: trigger the same minimal workflow by push and manual intent, predict state changes, verify them independently, and explain exactly what the green results do and do not prove.

Knowledge check

A commit is pushed but no matching workflow run exists. Where should diagnosis start?

Why does needs: producer not make state.txt appear in consumer?

What should you do before rerunning a failed deployment workflow?

A 403 disappears after changing permissions to write-all. Is that a good final fix?

Why keep the original failed run after the fix succeeds?

Official references and version notes

Version and compatibility note

Rechecked on 2026-09-09. The intentionally broken example is safe by construction: manual trigger only, permissions: {}, no secret, no external action, no publication/deployment, and a disposable repository. Preserve the first failed run before applying the same-job repair.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.