Chapter 29Lesson 04~205 minutes

Reusable Platform Pipelines, Golden Paths, and Organization-Wide Delivery Design: Diagnostics, Failure Modes, and Production Practices

Diagnose platform-wide breakage, secret inheritance, oversized interfaces, hidden runner trust, missing rollback versions and misleading adoption metrics without destroying first-failure evidence.

DiagnosticsBlast radiusSecretsRunner trustRecovery

Learning objectives

  • Preserve caller and platform revision evidence before repairing a platform incident.
  • Diagnose fleet breakage caused by moving refs and incompatible contracts.
  • Recognize secret inheritance, free-form interfaces and hidden runner trust as security debt.
  • Separate adoption metrics from reliability and developer-outcome metrics.
  • Apply the least destructive rollback or compatibility repair and rerun the smallest equivalent scope.

1. Evidence-first platform incident sequence

A platform incident can affect many repositories at once, so diagnosis starts by freezing identities. Preserve the first caller run ID/attempt, caller event/ref/SHA, caller workflow revision, resolved reusable-workflow SHA, evaluated inputs/permissions, runner identity, first failing step and any artifact/deployment evidence. Only then compare successful and failing platform versions.

  1. Preserve run ID/attempt, caller SHA and first-failure logs.
  2. Resolve every reusable workflow/action ref to the exact commit used.
  3. Confirm caller inputs, explicitly passed secrets and GITHUB_TOKEN permissions.
  4. Inspect job graph, queue and runner group/labels.
  5. Inspect action/runtime/network failure inside the called workflow.
  6. Inspect outputs/artifacts/caches and deployment/external state separately.
  7. Compare the failing platform release with last-known-good.
  8. Apply the smallest compatible correction or revert the caller pin.
  9. Rerun one representative caller before broader rollout.
  10. Record blast radius, exceptions and recovery evidence.

2. Failure: a platform team breaks every caller through @main

A moving branch lets a central commit change executable code without a caller repository change. The blast radius can be immediate and difficult to correlate because each caller still shows the same YAML text.

# BROKEN — do not use in production
jobs:
  ci:
    uses: acme-platform/delivery/.github/workflows/ci.yml@main
    secrets: inherit

Interpret the evidence: if callers with identical application SHAs begin failing after the platform main commit moves, resolve job_workflow_ref/workflow dependency data and compare the called implementation. Repair by publishing a tested immutable release and changing callers—or the managed channel—to a known commit SHA. Preserve the broken platform SHA for incident analysis.

3. Failure: secrets: inherit expands privilege silently

Secret inheritance makes every eligible caller secret available to the called workflow boundary even when only one credential is needed. A future repository secret can widen platform exposure without a platform code diff. That is especially dangerous when the platform also executes third-party actions or caller-controlled scripts.

Repair the contract to declare and pass specific secrets, or replace static cloud credentials with OIDC where supported. Never diagnose this by printing secrets or whole contexts. Verify by contract/configuration and by a safe presence-only check if necessary.

4. Failure: the platform becomes an unsafe mini-language

# BROKEN — do not use in production
on:
  workflow_call:
    inputs:
      runner: {type: string, required: true}
      raw-test-command: {type: string, required: true}
      raw-deploy-command: {type: string, required: true}
      bypass-policy: {type: boolean, default: false}
      provider: {type: string, required: true}
      privileged: {type: boolean, default: false}

Free-form command inputs turn the caller into a program generator and make quoting, injection, permission and portability behavior part of an undocumented language. A bypass-policy flag also places governance inside caller-controlled configuration. Repair by splitting contracts, exposing bounded intent, passing untrusted values through data channels, and moving hard authorization to independent policy/environment/runner controls.

5. Failure: a reusable workflow hides trusted runner access

Suppose v1 used GitHub-hosted runners but v2 silently changes deployment to a persistent self-hosted group with internal network access. Even if the YAML still passes, the trust boundary changed. Preserve runner labels/group evidence and any external side effect before retrying.

Repair by documenting runner requirements as a release-level contract, restricting runner-group access to intended repositories/workflows, and testing untrusted PR paths separately. Do not put public fork code on a trusted self-hosted runner just because the central workflow owns the YAML.

6. Failure: no last-known-good version exists

If callers all follow a moving branch and the platform has no release mapping, rollback becomes guesswork. Do not rebuild application artifacts or delete failed runs to “start fresh.” Identify the last successful called-workflow SHA from evidence, publish or restore a reviewed version mapping, then pin one representative caller and verify it before broader rollback.

7. Failure: “90% adoption” hides platform pain

Counting workflow files or repositories that mention the platform can look healthy while every team carries exceptions, pins an obsolete version or experiences long queues. Adoption needs a denominator and quality measures: active callers, supported-version distribution, success rate, platform-attributable failures, duration, exception age, update acceptance and rollback frequency.

8. Keep causal layers separate

Symptom Likely layer to inspect first Do not confuse with
Reusable workflow cannot be resolved visibility/access/ref/policy runner outage
Called job permission denied caller/called token permissions or secret contract workflow template copy
Job queued indefinitely runner labels/group/capacity reusable-workflow syntax
All callers fail after version update platform release/runtime/action dependency application source regression
Only deployment adapter fails environment/OIDC/provider target CI evidence generation
Required workflow not running ruleset workflow configuration/event/access ordinary status-check name alone

9. Read-only audit evidence for reusable workflow usage

Where organization audit-log API access is available, prepared_workflow_job records can identify job_workflow_ref plus caller workflow refs/SHAs. This is useful for incident scoping and adoption inventory, but audit data is not a substitute for caller-owned run evidence. The documentation notes that this event data is available through the REST API rather than the web audit interface/export path.

Without enterprise/organization audit capability, a free local alternative is to scan checked-out repositories for uses: pins and maintain a version manifest in a disposable platform lab. State clearly that this is a simulation of fleet inventory, not proof of every historical run.

10. Policy failures are not application failures

An organization may intentionally deny an action, require full-SHA action pins or require a workflow through a ruleset. If a platform upgrade is blocked by policy, preserve the denial message and policy scope. Do not work around the control with a local action copy, policy bypass or broader token. Either make the platform compatible with policy or use the formal exception path.

11. Least-destructive recovery pattern

Choose one representative caller with no production side effects. Pin it to the last-known-good platform SHA or a compatibility-repaired release, run the smallest equivalent CI path, compare outputs/check names/permissions, then expand the rollout gradually. If deployment state was changed, verify the external target independently before calling recovery complete.

12. Platform incident evidence packet

Field Retain
Caller identity repository, event, source SHA, workflow revision, run ID/attempt.
Platform identity called workflow ref + resolved SHA, release mapping.
Contract inputs, outputs, explicit secret names, effective permissions.
Compute runner group/labels/image/tool versions.
First failure job/step conclusion and log excerpt location.
Blast radius callers/version distribution and representative failures.
Correction revert or compatibility patch SHA.
Recovery representative rerun plus broader rollout evidence.
Governance exception/policy record and expiry if used.

13. Lesson summary

Platform failures are fleet dependency incidents. Preserve exact caller and platform identities, diagnose permission/runner/policy/external state separately, then recover with a known immutable version and a representative canary before broad rollout.

Next lesson

Checkpoint Lab — Reusable Platform Pipelines, Golden Paths, and Organization-Wide Delivery Design

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why is @main especially risky for a central platform workflow?

Why is secrets: inherit a poor golden-path default?

A reusable workflow suddenly queues forever after an upgrade. What state should you inspect?

Why preserve a broken platform SHA instead of deleting the release immediately?

What is a safer first recovery step than updating every caller?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.