GitHub Actions Foundations, Automation Model, and CI/CD Concepts: Diagnostics, Failure Modes, and Production Practices
Troubleshooting GitHub Actions is a state-localization problem. Random reruns and YAML edits destroy evidence and can repeat side effects. This lesson engineers two safe failures—one where no run is selected and one where a second job incorrectly assumes it shares the first job’s filesystem—then repairs each failure by identifying the owning layer and applying the smallest correction.
Learning objectives
- Use a consistent evidence-first diagnostic sequence from event/workflow selection through run, queue, runner, step, artifact/deployment, and external state.
- Distinguish “no run was created” from “a run was created but a job/step failed.”
- Diagnose a cross-job filesystem failure without misclassifying it as a runner or GitHub platform outage.
- Explain why blind reruns, broad permissions, context dumps, and delete/recreate tactics are unsafe troubleshooting defaults.
- Preserve first-failure evidence and choose a least-destructive repair or rollback.
1. Use the same diagnostic sequence every time
When a workflow behaves unexpectedly, start at the earliest state that could explain the symptom. Moving backward and forward randomly creates noise. A reliable sequence is:
flowchart TD A[Preserve run ID / attempt / first failure] --> B[Event + ref/SHA + workflow revision] B --> C[Conditions + inputs + permissions] C --> D[Job graph + queue] D --> E[Runner + image + toolchain] E --> F[Step / action / runtime / network] F --> G[Outputs / artifacts / caches] G --> H[Environment / deployment] H --> I[External target + governance] I --> J[Smallest correction + targeted verification]
If no run exists, do not start at runner logs. If a job never left the queue, do not edit a shell command. If a deployment request succeeded but the service is unhealthy, do not assume the workflow parser is at fault. Locate ownership first.
2. Preserve first-failure evidence before changing anything
Before rerunning or editing:
- record repository, workflow name/path, run URL/ID, attempt, event, ref, head SHA, and triggering actor;
- record which jobs were created, skipped, queued, cancelled, failed, or succeeded;
- save the first failing step's log and relevant setup/runner metadata;
- note permissions, inputs, secrets availability assumptions, and any external API/resource IDs;
- if side effects occurred, record exactly which external resource changed and its current state.
A rerun is a new attempt against a possibly changed external world. A registry tag may now exist; a deployment may be partially applied; a rate limit may have reset; a runner image may have updated. Preserve attempt 1 before creating attempt 2.
3. Failure mode A: valid workflow, no run selected
Start with the working Lesson 2 workflow. In the disposable repository only, temporarily change its push filter:
on:
push:
branches: [release-only-demo]
workflow_dispatch:
Commit this change to main, then make one harmless
follow-up commit to main. The follow-up push does not
match release-only-demo, so you should not expect a
push-triggered run for that event. The absence of a run is not a
“failed job”; the workflow selection condition was not satisfied.
| Question | Evidence | Conclusion |
|---|---|---|
| Did GitHub receive a push? | Commit/push exists in repository history. | Yes. |
| Did the filter match? |
Workflow says release-only-demo; event branch
is main.
|
No. |
| Was a run ID created for that workflow/event? | No matching run in Actions history/API. | No. |
| Should you inspect runner logs? | No runner was assigned because no run/job existed. | No. |
Repair: restore the intended main filter, commit the
correction, and verify a later controlled push produces a run.
Preserve the non-run explanation in your note rather than calling it
an Actions outage.
4. Failure mode B: jobs do not share a hosted-runner filesystem
Create a separate disposable workflow
.github/workflows/ch01-broken-boundary.yml:
name: Chapter 01 - Broken Job Boundary
on:
workflow_dispatch:
permissions: {}
jobs:
producer:
runs-on: ubuntu-24.04
steps:
- name: Create local state
shell: bash
run: |
set -euo pipefail
echo "created-by=$GITHUB_JOB" > state.txt
test -s state.txt
cat state.txt
consumer:
needs: producer
runs-on: ubuntu-24.04
steps:
- name: Incorrectly expect producer filesystem
shell: bash
run: |
set -euo pipefail
echo "consumer runner: $RUNNER_OS / $RUNNER_ARCH"
test -f state.txt
cat state.txt
Trigger it manually. Expected state: producer succeeds;
consumer starts on its own hosted job environment and
fails at test -f state.txt. The dependency
needs: producer controls job ordering; it does not copy
the producer filesystem.
5. Interpret the broken run before fixing it
Collect this minimum evidence:
- Run ID and attempt.
- Event =
workflow_dispatch. - Producer conclusion = success.
- Consumer conclusion = failure.
-
Producer log proves
state.txtexisted inside producer. -
Consumer log proves its own step could not find
state.txt. - Runner setup logs show the jobs were independently scheduled.
Do not classify the problem as “GitHub lost my file.” The file was never transferred across the job boundary. The workflow's dataflow design was wrong.
needs establishes
dependency and exposes supported job outputs; it is not a
shared-disk primitive. Cross-job file transfer requires an explicit
mechanism such as artifacts, which Chapter 13 teaches. For this
Chapter 01 example, the smallest repair is to keep
filesystem-dependent steps in the same job.
6. Apply the smallest equivalent repair
For this tiny workflow, combine the two operations into one job:
jobs:
same_job:
runs-on: ubuntu-24.04
steps:
- name: Create local state
shell: bash
run: echo "created" > state.txt
- name: Read local state
shell: bash
run: |
set -euo pipefail
test -f state.txt
cat state.txt
This repair matches the actual requirement: two steps need one ephemeral filesystem. Do not introduce artifact infrastructure merely because a tutorial says “use artifacts”; choose a cross-job transfer only when jobs genuinely need separate runners/lifecycles.
Rerun the corrected workflow as a new run after committing the fix. Compare the failing run and corrected run rather than deleting the failure.
7. Separate common failure layers
| Symptom | Likely layer to inspect first | Evidence |
|---|---|---|
| No run exists | Event/filter/workflow location/default-branch rules. | Event/ref, workflow file at revision, filter configuration. |
| Job is queued | Runner availability/labels/concurrency/policy. |
Queue state, requested runs-on, runner
availability.
|
| Step exits 127 / command not found | Runner image/toolchain/shell. | Runner image, PATH, tool version/setup. |
| API returns 403 | Token/permission/policy/event trust. | Requested permission, effective event restrictions, endpoint requirements. |
| File missing in another job | Dataflow/job boundary. | Producing job log, consumer workspace, explicit transfer configuration. |
| Deployment job green, service unhealthy | External provider/target. | Provider deployment/resource ID, health checks, service logs. |
| Rerun causes duplicate release | Idempotency/side-effect design. | Existing release/package/resource identity and prior attempt. |
8. Troubleshooting shortcuts that create new incidents
- Blind rerun: can repeat publication, deployment, or destructive API calls.
-
Grant
write-all: can hide the real missing permission while expanding blast radius. - Print every context/environment variable: can expose sensitive values and creates noisy evidence.
- Disable TLS verification: changes the trust model instead of solving identity/network configuration.
- Switch to a privileged self-hosted runner: may bypass the symptom while exposing trusted infrastructure to workflow code.
- Delete failed runs/logs before understanding them: destroys the first-failure record.
- Change several YAML keys at once: makes the correction non-causal and difficult to review.
9. A green indicator has a bounded meaning
Apply this interpretation discipline in production:
- Run created: event/workflow selection created an execution record.
- Job started: required conditions allowed the job and a runner became available.
- Step succeeded: that step process/action concluded successfully under its local context.
- Job succeeded: the job's required steps reached a successful conclusion.
- Check succeeded: GitHub recorded the check result expected by repository policy.
- Artifact uploaded: an object was stored; verify its identity/digest/content before promotion.
- Environment approved: authorization occurred; rollout is not thereby healthy.
- Deployment succeeded in GitHub: deployment status was reported; confirm provider/target health independently.
10. Production incident mini-runbook
- Freeze cleanup and preserve run/attempt evidence.
- Write the exact expected state transition that failed.
- Locate the earliest layer where observed state differs from expected.
- Confirm source/ref/SHA and workflow revision before editing.
- Check permissions and trust context before adding credentials.
- Check runner/image/toolchain before blaming application code for environment drift.
- Inspect external target evidence if any side effect crossed GitHub's boundary.
- Choose one reversible correction.
- Rerun the smallest safe scope; do not automatically replay deployment side effects.
- Keep both original and corrected evidence with the causal explanation.
11. Lab cleanup
Restore the main Chapter 01 workflow's intended branch filter. Keep the broken-boundary workflow if you want the failed run as a teaching fixture, or remove that one file after preserving its run ID/log evidence. Do not delete unrelated workflows or change organization policy. The lab used no secret, deployment, package, cloud account, or self-hosted runner, so rollback is limited to repository workflow files and disposable run history.
Knowledge check
A commit is pushed but no matching workflow run exists. Where should diagnosis start?
At event/workflow selection: verify workflow location/revision, event type, branch/path filters, and default-branch rules before inspecting runners or shell code.
Why does needs: producer not make
state.txt appear in consumer?
needs controls job dependency/dataflow metadata;
separate hosted jobs use separate execution environments. Files
require an explicit transfer mechanism or same-job design.
What should you do before rerunning a failed deployment workflow?
Preserve the run ID/attempt, first failure, and external side-effect state; then determine whether a rerun would repeat or conflict with prior mutations.
A 403 disappears after changing permissions to
write-all. Is that a good final fix?
No. It proves authority affected the symptom but grants excessive scope. Identify the exact required permission and event/policy constraints, then narrow it.
Why keep the original failed run after the fix succeeds?
It preserves causal evidence and lets reviewers compare the failure and correction rather than relying on a rewritten history.
Official references and version notes
- Understanding GitHub Actions — current component model for workflows, events, jobs, steps, actions, and runners.
- Workflow syntax for GitHub Actions — authoritative workflow keys, permissions, jobs, runner selection, and manual dispatch syntax.
- Variables reference — definitions of GITHUB_SHA, GITHUB_REF, GITHUB_RUN_ID, GITHUB_RUN_ATTEMPT, and related default variables.
- GitHub-hosted runners reference — current hosted-runner labels, VM behavior, hardware, and image caveats.
- Billing and usage — current availability, public-repository standard-runner usage, private-repository quotas, and usage boundaries.
- Troubleshooting GitHub Actions — current troubleshooting guidance and workflow-run diagnostics.
- Events that trigger workflows — event-specific trigger, ref, SHA, and default-branch behavior.
Rechecked on 2026-09-09. The intentionally broken
example is safe by construction: manual trigger only,
permissions: {}, no secret, no external action, no
publication/deployment, and a disposable repository. Preserve the
first failed run before applying the same-job repair.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.