Chapter 16Lesson 04~195 minutes

Reusable Workflows, workflow_call, Inputs, Secrets, Outputs, and Nesting: Diagnostics, Failure Modes, and Production Practices

Reusable-workflow failures often look like syntax or permission problems but arise at different boundaries: caller validation, interface mismatch, secret propagation, token non-escalation, runner availability, mutable references or self-cancelling concurrency groups. This lesson preserves first-failure evidence and diagnoses the boundary before editing the contract.

First failureenv boundaryPermission ceilingMutable refsConcurrency

Learning objectives

  • Diagnose caller validation, called-workflow runtime and nested authorization failures as different causal layers.
  • Recognize that caller env, secrets and permissions do not propagate with identical rules.
  • Preserve the resolved workflow reference and first attempt before changing a mutable dependency.
  • Avoid self-cancelling caller/called concurrency groups and unsafe privileged cleanup patterns.
  • Repair the narrowest interface or permission defect without broadening secrets or token scope.

1. Evidence-first diagnostic sequence

Before editing YAML, preserve the run ID/attempt and first error. Confirm the event/ref/source SHA and caller workflow revision. Identify the call job and the exact reusable workflow path/ref/SHA. Confirm input values, secret mappings and caller permissions. Then inspect the called job graph, runner and first failing step. Only after that should you change the smallest responsible boundary.

1. preserve run ID / attempt / first error
2. confirm caller event, ref, SHA and workflow file
3. confirm called workflow repository/path/ref/resolved SHA
4. inspect typed inputs and declared secret names
5. inspect caller permissions and every nested reduction
6. inspect call job -> called job graph -> runner
7. inspect outputs/artifacts only after execution reached them
8. separate environment/deployment/external state
9. change the narrowest interface or implementation defect
10. rerun the smallest equivalent scope and compare evidence

2. Failure: “the caller defined env, so the called workflow should see it”

A team defines env: RUNTIME: 3.13 at the caller workflow level, then expects $RUNTIME in a called workflow. The called job observes an empty value. This is not a runner defect. Caller workflow-level env is not propagated across a reusable-workflow boundary.

# Broken assumption
name: caller
on: workflow_dispatch
env:
  RUNTIME: '3.13'
jobs:
  ci:
    uses: ./.github/workflows/reusable-ci.yml
# No `with: runtime:` interface was supplied.

Repair the dataflow, not the runner: declare a workflow_call input and pass with: runtime: '3.13'. If the value is centrally governed rather than caller-selected, use a suitable vars source.

3. Failure: calling a reusable workflow from steps

This is a structural error. Reusable workflows are job-level calls. A step-level uses target must be an action, not a reusable workflow file.

# Broken
jobs:
  build:
    runs-on: ubuntu-24.04
    steps:
      - uses: ./.github/workflows/reusable-ci.yml

# Correct boundary
jobs:
  ci:
    uses: ./.github/workflows/reusable-ci.yml

If you actually need a reusable sequence of steps inside an existing job, that is a composite-action use case, which Chapter 17 covers.

4. Failure: broad secrets: inherit for a narrow test contract

The workflow works, but review shows the called workflow can access every eligible inherited organization/repository/environment secret. Functionally green is not least privilege. Preserve the run evidence, enumerate which secret the called workflow really needs and replace inheritance with explicit mapping.

# Over-broad
jobs:
  ci:
    uses: octo-org/platform/.github/workflows/ci.yml@0123456789abcdef0123456789abcdef01234567
    secrets: inherit

# Narrower
jobs:
  ci:
    uses: octo-org/platform/.github/workflows/ci.yml@0123456789abcdef0123456789abcdef01234567
    secrets:
      package_read_token: ${{ secrets.PACKAGE_READ_TOKEN }}

The hexadecimal reference is illustrative; in a real repository it must be the reviewed commit SHA of the actual platform workflow.

5. Failure: nested workflow assumes it can elevate permissions

Suppose caller A grants only contents: read. Workflow B calls workflow C and C declares issues: write. The correct mental model is not “C asked for write, so C gets write.” The caller's permission ceiling still applies. Diagnose the intended capability: either C should not perform the write, or A must intentionally delegate the necessary narrow permission after review.

Do not “fix” the failure with write-all. Separate read-only CI from mutation where possible.

6. Failure: a mutable branch reference changed between full reruns

A production caller uses @main for a central workflow. The initial run used one central commit; a later full rerun resolves the branch again and executes newer reusable-workflow code. That can make incident comparison misleading even though the application source SHA is unchanged.

Preserve both reusable-workflow identities and rerun modes. Repair future reproducibility by pinning a reviewed commit SHA. Do not rewrite the historical evidence.

7. Failure: called workflow cancels its own caller

A called workflow sees ${{ github.workflow }} in the caller's context. If both caller and called workflow use the same concurrency group expression with cancel-in-progress: true, the called workflow can collide with and cancel the caller that invoked it.

# Dangerous when duplicated in caller and called workflow
concurrency:
  group: ${{ github.workflow }}-${{ github.ref }}
  cancel-in-progress: true

Use distinct group namespaces based on the actual critical resource, for example caller-ci-… and deploy-production-…. Concurrency is an external-side-effect invariant, not a naming convention.

8. Intentionally broken runnable example: valid type, unsupported semantic value

The most useful interface failure creates a real run so you can preserve evidence. The following caller passes a valid string type, but the central contract supports only 3.13. GitHub can create the run and enter the called workflow; the contract check then fails with an intentional exit code and exact message.

jobs:
  ci:
    permissions:
      contents: read
    uses: ./.github/workflows/reusable-ci.yml
    with:
      runtime: '3.12'   # syntactically valid string, semantically unsupported
      strict: false
    secrets:
      demo_secret: ${{ secrets.CH16_FAKE_SECRET }}

Preserve the run ID, attempt, source SHA, call-job state and message unsupported runtime: 3.12; contract supports 3.13. The repair is only runtime: '3.13'. No permission, secret or runner change is justified.

9. Causal failure map

Symptom Likely layer Evidence before repair
workflow will not validate call syntax / unsupported caller-job key / input schema workflow file and validation message; there may be no run
called job starts then rejects value contract semantics run/attempt, input value, called job log
secret required but absent caller secret mapping or nested forwarding secret name mappings, never secret value
API gets 403 permission ceiling or endpoint policy caller/called permissions and safe response headers/status
job waits/no runner runner access/capacity caller repository runner eligibility and queue state
different code on full rerun mutable reusable-workflow ref resolved called workflow SHA per attempt
caller unexpectedly cancelled concurrency collision group expression and run cancellation evidence

10. Production repair rules

Do not delete evidence, broaden all secrets, add write-all, move the call into privileged pull_request_target, or blindly rerun until green. Treat reusable workflow incidents like dependency incidents: preserve identity, prove the interface mismatch, correct one boundary, then rerun the smallest comparable scope.

Knowledge check

A called workflow sees an empty caller workflow-level env value. Which layer should you fix?

Why is a step-level reusable workflow call invalid?

What is wrong with fixing a nested 403 using permissions: write-all?

Why preserve the resolved called-workflow SHA before fixing an @main incident?

What should you change in the intentionally broken 3.12 example?

Next lesson

Prove the whole contract under failure and repair

Lesson 5 combines two callers, a nested leaf, outputs and an interface failure into one evidence packet.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked on 2026-09-09 for GitHub.com. A reusable workflow is called at the job level, not from a step. Current GitHub.com limits allow up to 10 connected workflow levels and 50 unique reusable workflows in one top-level workflow tree. Supported caller-job surfaces are name, uses, with, secrets, strategy, needs, if, concurrency and permissions. Nested GITHUB_TOKEN permissions can stay the same or become more restrictive, never more permissive. Secrets are passed only to the directly called workflow unless forwarded again. Workflow-level caller env values do not cross the boundary automatically. Same-repository ./.github/workflows/file.yml calls use the same commit as the caller; cross-repository production calls should use a reviewed full commit SHA instead of a mutable branch or tag. Re-running all jobs against a non-SHA ref resolves that ref again, while re-running failed/specific jobs uses the called workflow commit from the first attempt.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.