Chapter 13Lesson 04~160 minutes

GitHub Actions Foundations: Workflows, Events, YAML, Permissions, and Execution Model: Diagnostics, Failure Modes, Security, and Performance

Actions failures often look deceptively similar in the UI: a red run, a missing run, a queued job, or a 403. The causes are very different. You will learn to preserve the event/ref/SHA evidence first, then identify whether the failure is discovery, syntax, authorization, runner, dependency, or application logic.

DiagnosticsScript injectionLogsLeast privilege

Learning objectives

  • Classify a missing/failed workflow as discovery, YAML/schema, event/ref, authorization, runner, dependency, or application failure before changing anything.
  • Diagnose an intentionally blocked write as an expected GITHUB_TOKEN permission failure rather than “fixing” it with broad access.
  • Explain why direct interpolation of pull-request/event text into shell code is a script-injection vulnerability.
  • Review workflow logs without dumping secrets or full contexts, and explain why masking is not a substitute for minimization.
  • Apply an evidence-preserving sequence that identifies repository, event, ref/SHA, workflow definition, job, runner, and API response before correction.

Safety boundary: All runnable failure examples target disposable public repositories and GitHub-hosted runners. The lesson does not register self-hosted runners, create long-lived credentials, disable organizational security policy, force-update refs, or use privileged pull_request_target execution.

1. Diagnostic sequence: preserve evidence before you “fix” the YAML

Use the same sequence every time: preserve evidence → identify repository/account/event/ref/workflow/run/job scope → inspect settings/permissions/logs/API → choose the least destructive correction → verify independently.

Question Evidence Why it comes before editing
Did GitHub discover the workflow? gh workflow list --all, Actions UI A file outside .github/workflows is not an execution failure; it is not a workflow.
Did an event create a run? gh run list --event ... A trigger/filter mismatch is different from a job failure.
Which source/workflow version? run headSha, log workflow_sha, Git refs Prevents fixing the wrong branch/revision.
Did authorization fail? Step stderr, HTTP status/body, effective permissions A 403 should not be repaired by blindly granting write-all.
Did a runner execute the job? job status, runner labels/log setup Queued/offline is different from a test failure.
Did application logic fail? specific step exit status/log Only after platform/scope evidence is known.

2. Failure mode: invalid YAML or wrong workflow location

GitHub discovers workflow files only in .github/workflows with YAML extensions. A file at .github/ci.yml may be valid YAML but is not a workflow. Conversely, a file in the correct directory can fail validation because its YAML or Actions schema is invalid.

REPO="OWNER/DISPOSABLE"

gh workflow list -R "$REPO" --all \
  --json id,name,path,state

# Inspect the committed path; this is Git evidence, not Actions evidence.
gh api \
  -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/contents/.github/workflows" \
  --jq '.[] | {name,path,sha}'

If a new file does not appear in gh workflow list, verify location and branch first. If GitHub reports “Invalid workflow file,” preserve the exact annotation/line before editing. Do not hide the original cause by replacing the whole file.

3. Failure mode: the run is real, but it ran on the wrong event/ref

A workflow can succeed perfectly on the wrong source. Filter and dispatch mistakes are operational failures even when every step is green. Inspect event, headBranch, headSha, and the workflow’s own logged workflow_ref/workflow_sha.

RUN_ID="123456789"
REPO="OWNER/DISPOSABLE"

gh run view "$RUN_ID" -R "$REPO" \
  --json event,headBranch,headSha,attempt,status,conclusion,url

gh run view "$RUN_ID" -R "$REPO" --log

Do not repair a wrong-ref run by force-pushing a branch to “make the SHA match.” Correct the trigger/ref selection and create new evidence.

4. Intentionally broken example: least privilege blocks an unnecessary write

Suppose a CI workflow declares only contents: read but a new step tries to create an Issue. That write should fail. This is a useful controlled failure because it proves the token is not overprivileged.

permissions:
  contents: read

jobs:
  validate:
    runs-on: ubuntu-latest
    steps:
      - name: Intentionally forbidden write
        env:
          GH_TOKEN: ${{ github.token }}
        run: |
          gh api --method POST \
            -H "X-GitHub-Api-Version: 2026-03-10" \
            "repos/$GITHUB_REPOSITORY/issues" \
            -f title='Chapter 13 permission probe' \
            -f body='This request should be rejected by least privilege.'

Expected evidence: the step exits non-zero and GitHub responds with a permission/forbidden error; no Issue is created. Interpret the cause: the token authenticated, but its effective authorization did not include Issues write.

Least-destructive repair: If CI does not actually need to create Issues, remove the write operation. Do not “fix” the workflow by granting write-all. If the business requirement truly requires reporting, isolate that behavior in a separate job and grant only issues: write there.

5. Security failure: untrusted context becomes shell syntax

GitHub explicitly documents script-injection risk when attacker-controlled context values are embedded directly in inline scripts. A malicious PR title can contain characters that become executable shell syntax after expression substitution.

Unsafe pattern — do not run:

# DO NOT USE: PR title is untrusted and is inserted into shell source code.
- run: |
    title="${{ github.event.pull_request.title }}"
    echo "$title"

Safer pattern: copy the value into an environment variable and let the shell read it as data.

- name: Handle PR title as data
  env:
    PR_TITLE: ${{ github.event.pull_request.title }}
  run: |
    printf '%s\n' "$PR_TITLE"

Quoting still matters. For complex structured data, prefer an action or script that parses a file/JSON input without constructing shell source code from attacker input.

6. Security failure: logs expose secrets or excessive context

Do not dump toJSON(github), env, secrets, or entire event payloads merely for convenience. The github context includes github.token during job steps; secret redaction is not guaranteed for every transformed value. Log only the fields needed to diagnose the run.

Need Log this Avoid
Identify source event, ref, exact SHA Full event payload
Identify workflow workflow ref + workflow SHA Entire github object
Identify runner OS/arch/runner label Full environment dump
Debug API failure HTTP status + sanitized message Authorization header/token
Debug secret-based integration Secret name/presence as boolean if necessary Secret value or derived encodings

7. Runner and action failures: preserve the infrastructure boundary

A job that stays queued can mean no eligible runner is available. A step that fails before your code runs can mean an action dependency was blocked by policy or unavailable. For GitHub-hosted runners, inspect the selected label and setup log. For self-hosted runners, later chapters add runner status, group policy, network, and persistent-machine diagnostics.

If repository policy requires full-SHA pinning and the workflow uses third-party/action@v1, the correct fix is not to disable policy. Pin a reviewed full SHA or choose an allowed component.

8. Reliability, rate limits, and billing only where they cause failure

Trigger loops and overly broad path/event definitions can create many runs. Public standard hosted runners are currently free, but runaway automation still consumes concurrency, API rate limit, reviewer attention, and log retention. In private repositories, hosted-runner usage also interacts with included quota/billing. Design filters and cancellation/concurrency intentionally in later chapters rather than using “free” as a reason to ignore execution volume.

9. Lesson summary

A missing workflow, a missing run, a 403, a queued job, a dependency-policy rejection, and a failing test are different failure classes. Preserve repository/event/ref/workflow/job evidence first. Least privilege, safe input handling, selective logs, and runner/dependency policy are part of correctness, not optional security decoration.

Knowledge check

A workflow file is committed at .github/ci.yml. Why does gh workflow list not show it?

A step receives HTTP 403 using GITHUB_TOKEN. Does that prove authentication failed?

Why is removing an unnecessary write step often safer than adding write-all?

A PR title is printed through an environment variable with proper quoting. What security problem did that avoid?

A workflow is green but headSha does not match the commit you intended to validate. What is the correct conclusion?

Next lesson

Next: Checkpoint Lab — GitHub Actions Foundations: Workflows, Events, YAML, Permissions, and Execution Model

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.