Workflow Files, YAML Structure, Jobs, Steps, and Actions: Diagnostics, Failure Modes, and Production Practices
A workflow failure should be localized before it is edited. Chapter 02 has two fundamentally different failure classes: structure can be rejected before runner work begins, or valid structure can reach a runner and fail during execution. This lesson preserves both forms of evidence, then diagnoses YAML/schema placement, fresh-job filesystem assumptions, unsupported keys, shell portability, and mutable action references without using broad permissions or blind reruns.
Learning objectives
- Separate YAML/schema/selection failures from runner-time step failures using run existence, job allocation, and first-failure evidence.
-
Diagnose an intentionally misplaced
stepsblock without inventing a runner problem when no job executed. -
Diagnose a two-job filesystem handoff failure and explain why
needsdoes not transfer files. - Identify OS/shell and unsupported-key failures from the owning configuration layer before changing action/tool versions.
- Repair the smallest causal layer while preserving the original workflow revision, run/attempt, logs, and action references.
1. Evidence-first diagnostic sequence
- Preserve the exact workflow revision. Save the commit SHA and the invalid/failed YAML before editing.
- Ask whether a run exists. No normal run/job evidence points toward workflow discovery, trigger, parse, or schema problems rather than runner execution.
- If a run exists, preserve run ID and attempt. Do not immediately rerun and replace first-failure context.
- Inspect job graph and queue. Confirm whether the expected job was created, skipped, queued, assigned, or completed.
- Inspect runner identity and the first failing step. Separate shell/tool/network failures from workflow structure.
- Inspect data boundaries. Verify which files were committed, checked out, generated in this job, or incorrectly assumed to arrive from another job.
- Apply the smallest correction. Do not change runner, action version, permissions, and workflow structure simultaneously.
- Rerun the smallest equivalent scope. Preserve both before/after evidence.
2. Failure A: structurally wrong job map
name: Broken structure
on:
workflow_dispatch:
permissions: {}
jobs:
verify:
runs-on: ubuntu-24.04
steps:
- run: echo "This is not under a job"
The indentation is legal enough to resemble YAML structure, but
steps is now interpreted at the wrong level of the
workflow's jobs map. GitHub's workflow schema expects each job entry
to be a job definition. Preserve the exact committed file and the
current validation message rather than teaching one hard-coded error
string.
3. Failure B: valid graph, wrong filesystem assumption
name: Broken handoff
on:
workflow_dispatch:
permissions: {}
jobs:
producer:
runs-on: ubuntu-24.04
steps:
- name: Create transient file
run: echo "producer-only" > handoff.txt
consumer:
needs: producer
runs-on: ubuntu-24.04
steps:
- name: Read producer file
run: cat handoff.txt
This workflow is structurally valid and can create a run. The
producer succeeds. The consumer is ordered after it, receives a
fresh hosted runner, and then fails because
handoff.txt was never committed or transferred. The
failure log belongs to the consumer step, not to the YAML parser.
4. Repair only the causal boundary
At this course stage, choose one of two simple repairs:
- If the consumer logic truly depends on an ephemeral file, keep producer and consumer steps in one job so the file remains on one runner.
- If the two jobs should remain isolated, make the consumer recreate/check out the inputs it legitimately owns. Chapter 13 later teaches artifacts for explicit file transfer.
Do not “repair” the example by changing to a persistent privileged self-hosted runner. That would hide the dataflow error while creating a much larger trust and state-management problem.
5. Shell portability failure: correct YAML, wrong command language
A step can be structurally valid but written for a different runner
OS. For example, Bash-specific syntax such as
[[ ... ]], set -o pipefail, or Unix paths
should not be assumed to work in a job moved to a Windows default
shell. Diagnose this from runs-on, the shell shown in
the step log, and the command error.
The least-surprising repair is to make the supported platform explicit or supply a correct shell for that runner. Do not rewrite random YAML keys when the job clearly reached a runner and the shell rejected the command.
6. YAML typing and local-parser traps
Not every YAML parser uses the same scalar rules or understands
GitHub Actions. A common example is older YAML 1.1 tooling that can
interpret the plain key on as a boolean when loading
and rewriting a workflow. That does not mean GitHub Actions wants
the workflow key renamed; it means the local serializer is the wrong
compatibility tool for this document.
Quote values that are semantically strings when ambiguity would be harmful. Tool versions are a good example:
with:
python-version: '3.13'
Likewise, distinguish GitHub expression typing from raw YAML scalar typing. A local YAML parse can prove basic document syntax, but only GitHub's current workflow schema and runtime semantics can prove that keys, contexts, scopes, and action inputs are valid for Actions. Preserve the original text before allowing automated formatters or serializers to rewrite workflow YAML.
7. Unsupported or misplaced keys: schema belongs to GitHub, not intuition
GitHub Actions has a defined workflow schema. A key supported at one scope may be invalid at another. A key from another CI product may be valid YAML yet meaningless or rejected here. Always verify current workflow syntax documentation for the exact scope rather than assuming indentation can make a key valid.
This also explains why a generic YAML linter is necessary but insufficient. It can catch malformed YAML, but it does not necessarily understand GitHub's current workflow schema, context restrictions, action metadata, or security policy.
8. Mutable action tag: successful run, weak reproducibility
- name: Checkout
uses: actions/checkout@v7
This can succeed. The problem is not immediate execution failure; it is dependency ambiguity. The tag can move, so a future run with the same workflow text may execute different action code. Production guidance in this course uses the verified full SHA and records the release label separately.
- name: Checkout
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
When diagnosing an old run, preserve what action revision actually ran before updating dependencies. “Upgrade everything” destroys causal isolation.
9. Failure localization map
flowchart TD
A[Committed workflow revision] --> B{Workflow accepted / selected?}
B -- no --> C[Trigger / YAML / schema / policy evidence]
B -- yes --> D{Expected job created?}
D -- no --> E[Condition / graph / permissions / policy]
D -- yes --> F{Runner assigned?}
F -- no --> G[Runner label / capacity / policy]
F -- yes --> H{Step executed?}
H -- no --> I[Previous step / condition / action startup]
H -- yes --> J[Shell / action / tool / network / data boundary]
J --> K[Smallest causal repair + preserved before/after run]
10. Troubleshooting anti-patterns
- Blind rerun: can repeat side effects and replaces attention to first-failure evidence.
- Switch runner OS: can make one shell error disappear while changing the environment under test.
- Grant write-all: can mask permission design bugs and increases blast radius.
- Replace full SHA with a branch/tag: makes dependency behavior less reproducible.
- Move to persistent self-hosted runner: can make missing-file assumptions appear to work because stale state survives.
- Change multiple layers at once: prevents causal attribution.
11. Guided broken-example lab
- Commit the structural failure on a disposable branch and capture GitHub's validation evidence plus commit SHA.
-
Fix only the
stepsplacement and confirm that the workflow can now create a normal run. - Replace it with the two-job handoff failure and dispatch once.
-
Record producer success, consumer failure, runner identities, run
ID/attempt, and the exact
cat handoff.txtfailure. - Repair by keeping the file-dependent commands in one job or by making the consumer independently recreate its needed input.
- Dispatch again and compare the before/after evidence without deleting the failed run.
Knowledge check
A commit contains invalid workflow structure and no expected job ever starts. Why is runner capacity the wrong first diagnosis?
The failure occurs before runner assignment. Inspect workflow selection/parsing/schema/policy at the exact committed revision first.
The producer job succeeds and consumer fails on
cat handoff.txt. What does
needs prove?
Only that the consumer depends on producer completion/result according to the job graph; it does not transfer the producer filesystem.
A workflow works on Ubuntu but its Bash step fails after moving to Windows. What evidence should you preserve?
The run/job IDs, runner OS/label, shell invocation shown in the log, and the exact command error before changing syntax or runner choice.
Why can actions/checkout@v7 be a production
problem even if the run is green?
The mutable tag weakens reproducibility and supply-chain control; future runs can execute different code with unchanged workflow text.
What is the purpose of keeping the original failed run after repair?
It preserves first-failure evidence and allows reviewers to prove that the targeted change—not an unrelated environmental change—resolved the causal layer.
Official references and version notes
- Understanding GitHub Actions — current definitions and execution relationships for workflows, jobs, steps, actions, and runners.
- Workflow syntax for GitHub Actions — authoritative workflow structure, jobs, steps, permissions, defaults, runners, and shell behavior.
-
Setting default shell and working directory
— current precedence and restrictions for workflow/job
defaults.run. - Using GitHub-hosted runners — job-to-runner isolation and filesystem sharing within a job.
- Secure use reference — current guidance to pin action dependencies to full-length commit SHAs and minimize privileges.
- Troubleshooting workflows — current workflow-run troubleshooting guidance.
- GitHub-hosted runners reference — current hosted runner filesystem/runtime boundaries and image labels.
Version-sensitive behavior was rechecked against primary GitHub
documentation and GitHub-maintained action repositories on
2026-09-09. Executable examples use
ubuntu-24.04,
actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
(upstream release v7.0.1), and where Python setup is needed
actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
(upstream release v7.0.0) with Python 3.13. At verification time
both actions declare a Node 24 runtime. Runner images, action
releases/runtimes, workflow keys, parser diagnostics, and
plan-dependent behavior can change; re-resolve current immutable
SHAs before copying these examples into long-lived production
workflows. The broken examples use no secrets, mutation
permissions, packages, deployments, self-hosted runners, or
external targets. The exact wording of GitHub validation errors is
intentionally not frozen because platform diagnostics can evolve.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.