Enterprise Policies, Allowed Actions, Runner Governance, and Auditability: Diagnostics, Failure Modes, and Production Practices
Diagnose governance failures by preserving the denied run or policy evidence, locating the controlling layer, and applying the least destructive correction instead of bypassing controls.
Learning objectives
- Use an evidence-first sequence for policy, runner and audit failures.
- Diagnose overly broad policy as well as overly restrictive policy.
- Recognize public-repository access to privileged runners as a machine-trust incident.
- Preserve original denials and weak exceptions before remediation.
- Correct the controlling layer without bypassing enterprise governance.
1. Diagnostic contract: preserve denial evidence before changing policy
A policy failure is often an intentional safety signal. The first response should be to preserve the run ID/attempt, workflow SHA, denial message, relevant policy snapshot and runner eligibility—not to broaden settings until the workflow turns green. Governance diagnostics must answer whether the workflow is wrong, the policy is wrong, the exception is missing, or the requested operation belongs on a different trust boundary.
Use the same causal order throughout this lesson: event/ref/SHA → selected workflow revision → effective policy and permissions → job graph → runner eligibility → action/runtime → artifact/deployment/external state → audit/exception evidence.
2. Evidence-first diagnostic sequence
- Preserve the original run ID, attempt number, source SHA, workflow file and first denial/log evidence.
- Confirm which enterprise, organization and repository policy layers apply and record their effective values.
- Confirm the action/reusable-workflow reference, token permission request, event/fork trust, runner group and environment/rules state.
- Inspect the job graph and queue. A job that never becomes eligible is different from a runner that starts and fails.
- Confirm runner group, labels, runner image and any private network/trust assumptions.
- Inspect the failing action or shell step only after policy and runner eligibility are understood.
- Inspect artifacts, deployments and external targets before retrying any side effect.
- Check audit/exception evidence for recent policy changes or expired waivers.
- Apply the least destructive correction and rerun the smallest equivalent scope.
3. Failure mode: policy is so broad that it grants any Marketplace action
Suppose an organization believes “verified creators only” is equivalent to a curated supply-chain policy. A new workflow uses a verified Marketplace action with broad permissions and a mutable tag. The platform permits it, so there is no denial to alert the team. This is a governance false negative: policy eligibility is broader than the organization’s intended trust boundary.
# Intentionally weak policy outcome — do not use as a production pattern.
permissions: write-all
steps:
- uses: verified-vendor/publish-action@v3
Repair by separating trust signals: select reviewed external repositories/patterns, require full action SHAs, establish least-privilege workflow templates and review external release diffs. Do not assume publisher verification evaluates the action’s permissions or code.
4. Failure mode: exception has no expiry
A migration waiver that never expires turns temporary risk into
permanent policy debt. The observable failure is not a red workflow;
it is stale governance state. Diagnose by querying the exception
registry for missing or past expires_at, confirming
whether the dependency is still used, and checking audit evidence
for extensions.
{
"exception_id": "EX-0042",
"control": "actions.allowlist",
"scope": "all repositories",
"reason": "migration"
// BUG: no owner, no exact dependency, no expiry, no removal condition
}
Least-destructive repair: narrow the resource scope, identify the risk owner, add a near-term expiry and compensating controls, then create the paved-road replacement. Do not silently rewrite history; retain the original weak exception as evidence and issue a superseding record.
5. Failure mode: a public repository can reach a privileged runner
This is a machine-boundary incident, not a YAML style problem. A public fork may submit code that executes under the pull-request workflow. If that workflow can reach a self-hosted or network-privileged runner, the attacker-controlled code may interact with local files, caches, network resources or credentials available to the machine.
Preserve the runner-group configuration, eligible repositories/workflows, queued/running job evidence and runner diagnostics. Then remove public-repository eligibility or route untrusted validation to GitHub-hosted runners. If the runner may have executed hostile code, follow the runner compromise process; merely moving the runner to another group does not prove the machine is clean.
Do not troubleshoot this by adding an approval prompt while leaving the runner broadly reachable. The control objective is to ensure untrusted workflow code cannot be scheduled onto the privileged machine in the first place.
6. Failure mode: repository YAML tries to override an enterprise deny
A team sees an error saying an action is not allowed and edits the
workflow to add broader permissions, different runner
labels or another environment. None of those changes address the
controlling action policy. The job may remain blocked before
execution because the executable dependency itself is ineligible.
The correct path is to identify the effective enterprise/organization action policy, determine whether the dependency should be reviewed and allowed, or replace it with an approved action/local implementation. Repository YAML cannot legitimately weaken the upstream deny. Preserve the original denial because it proves enforcement worked.
7. Failure mode: audit logs exist but are not retained/exported
Assume a policy changed four months ago and an incident review needs the exact administrator and old value. If the organization relied only on a short interactive window or failed to export relevant events, the missing evidence cannot be recreated from a green workflow run. A run records execution; it does not necessarily encode who changed the organization’s policy.
Repair prospectively: define required audit event categories, external retention where appropriate, access controls, integrity checks, deduplication for at-least-once streams and monitoring that detects a stopped export. The evidence gap itself belongs in the incident record.
8. Failure mode: policy blocks teams without a versioned paved road
A strict allowlist may be technically correct but operationally harmful if teams lack an approved way to build, test, publish or deploy. Symptoms include exception growth, duplicated shell scripts, long lead time and pressure to loosen global policy. The root cause is a platform-product gap, not necessarily developer resistance.
Measure denied workflows by use case, time-to-resolution, exception recurrence and platform adoption. Build versioned reusable workflows/actions that satisfy common needs with least privilege and immutable dependencies. This connects Chapter 34 back to Chapter 29: governance is sustainable when the control plane and golden path evolve together.
9. Intentionally broken example: diagnose the controlling layer
The following workflow is deliberately noncompliant. Assume the
enterprise policy allows only approved actions, requires full action
SHAs, defaults tokens to restricted mode, and the
prod-deploy runner group permits only
payments/.github/workflows/deploy.yml.
name: broken-release
on: pull_request
permissions: write-all
jobs:
deploy:
runs-on:
group: prod-deploy
steps:
- uses: vendor/release-action@v2
- run: echo "deploy"
Interpret it layer by layer. The mutable external action can be denied by action policy before execution. The runner group can deny this workflow path even if the action were allowed. The broad token request violates the platform standard and is especially inappropriate for untrusted PR code. The event itself is wrong for a privileged deployment. Fixing only the action SHA would leave several independent trust failures.
| Evidence | Interpretation | Least-destructive correction |
|---|---|---|
| “Action not allowed” before job starts | Executable dependency policy denied the workflow | Use approved dependency or complete narrow review/allow process |
| Job queued with no eligible runner / runner-group denial | Workflow is not entitled to privileged pool | Move ordinary CI to hosted runner; reserve deploy group for reviewed deploy workflow |
Workflow asks for write-all |
Token request is broader than necessary |
Set permissions:{} and grant exact keys only
where needed
|
Event is pull_request |
Untrusted code boundary conflicts with deployment trust | Split validation from privileged deployment; require trusted branch/environment path |
10. Recovery and rerun discipline
After a policy correction, do not blindly rerun every failed workflow. First verify whether any job reached a runner or external target. If the workflow was rejected before execution, a rerun after policy repair is usually lower risk. If some jobs ran before a later policy/runner/deployment failure, inspect artifacts and external side effects before repeating them.
Preserve the first run/attempt and policy snapshot. A second run under a changed policy is not equivalent evidence; it proves the new state. Your incident record should show the transition: original policy → denial or unsafe eligibility → authorized correction → new policy digest → successful verification.
11. Lesson summary
Governance failures can be too strict, too loose, incorrectly scoped or unauditable. Diagnose them at the layer that owns the state: executable dependency policy, token/fork settings, runner groups, rules/environment, platform paved road or audit retention. The checkpoint lab now asks you to write one coherent standard and prove that it accepts and rejects workflows for the right reasons.
Knowledge check
A workflow is blocked because an action is not allowed. Why
will adding contents: write not fix it?
Action eligibility and token permissions are different control layers. The workflow dependency is denied before broader token authority would matter.
Why can an overly broad policy be a failure even if every workflow is green?
It may create governance false negatives by allowing unreviewed code or trust boundaries that the organization intended to prohibit.
A public fork job ran on a privileged self-hosted runner. Is moving the runner group sufficient remediation?
No. Remove the eligibility boundary and also treat the machine as potentially compromised; inspect/replace credentials and runner state according to incident procedures.
Why preserve an expired or weak exception instead of editing it in place?
The original record is historical evidence. Create a superseding corrected record so reviewers can reconstruct the governance timeline.
A strict policy produces hundreds of repeated exceptions for the same build task. What system problem should you investigate?
A missing or inadequate paved-road capability. The platform should provide a supported versioned automation path rather than normalizing repeated waivers.
Official references and version notes
- Enforcing policies for GitHub Actions in your enterprise — Current enterprise policy options, selected-action rules and full-SHA enforcement.
- Managing GitHub Actions settings for a repository — Repository-level action policy, token defaults and fork settings, including inheritance limits.
- Managing access to self-hosted runners using groups — Runner-group organization/repository/workflow access and public-repository warnings.
- Reviewing the audit log for your organization — Audit-log search/export/API boundaries and retention guidance.
- Secure use reference — Least privilege, untrusted input, third-party action and runner security guidance.
- Compromised runners — Current risk model for runner compromise and untrusted workflow code.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.