Chapter 19Lesson 04~175 minutes

Environments, Deployment Protection Rules, Approvals, Concurrency, and Rollbacks: Diagnostics, Failure Modes, Security, and Performance

Deployment failures are especially dangerous when the workflow still looks green or when the “fix” destroys the evidence needed to understand what happened. This lesson uses controlled failures to practice an evidence-first diagnostic sequence across environment identity, concurrency, approval, artifact provenance, plan availability, and sensitive deployment operations.

DiagnosticsEnvironment typoConcurrency collisionApproval evidenceSafe recovery

Learning objectives

  • Diagnose a misspelled environment that silently creates an unprotected environment resource.
  • Detect concurrency-group collisions and choose cancellation/queue corrections that match deployment semantics.
  • Prove whether an approval was granted for the same commit/artifact actually deployed.
  • Explain why rebuild-based rollback can produce a different artifact and how to recover without history rewrite.
  • Separate permission, plan/visibility, Actions, environment, and external target failures.

Availability: All destructive examples remain synthetic or inside the disposable public lab. No real cloud credential, secret, package deletion, runner registration, policy bypass, force update, repository transfer/delete, or history rewrite is required. Private-repository plan failures are shown as diagnostic fixtures/read-only reasoning.

1. Use the same diagnostic order even during an incident

  1. Preserve evidence: run ID, attempt, head SHA, workflow revision, artifact digest, environment name, deployment ID/status, reviewer decision, logs.
  2. Identify scope: repository/host → workflow/job → environment/deployment → concurrency group → artifact → external target.
  3. Inspect permissions/settings/rules/logs/API: do not infer policy from a job name.
  4. Choose the least destructive correction: fix a name/group/permission or redeploy retained bytes before considering bypass/history mutation.
  5. Verify independently: compare deployment history, artifact digest, target health, and approval evidence after repair.

2. Broken example: prodution bypasses the intended environment configuration

Suppose production has a reviewer and environment secret, but a workflow edit says:

jobs:
  deploy:
    environment: prodution   # intentionally wrong
    runs-on: ubuntu-24.04
    steps:
      - run: echo "deploy"

Current GitHub behavior can create a referenced environment automatically. A normal workflow-created environment has no protection rules or secrets. The symptom can therefore be more dangerous than a hard failure: the job may run without waiting for the expected reviewer, while an expected environment secret is absent.

Preserve the suspicious run and inspect the environment list:

REPO="OWNER/c19-deployment-governance-lab"
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/environments?per_page=100"   --jq '.environments[] | {name,protection_rules}'

gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/environments/production"   --jq '{name,protection_rules}'
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/environments/prodution"   --jq '{name,protection_rules}' 

Repair: correct the workflow to the canonical environment name, review that change, and only then delete the typo environment in the disposable lab. In production, consider policy/lint checks over allowed environment names. Do not “fix” the problem by copying production secrets into the typo environment.

3. Broken example: a generic concurrency group cancels the wrong workflow

Concurrency groups are repository-wide and case-insensitive. If deployment and an unrelated data-maintenance workflow both use production, one can cancel or replace the other's pending/running work depending on configuration.

# Too broad in two unrelated workflows
concurrency:
  group: production
  cancel-in-progress: true

Inspect run histories and, where useful, current concurrency-group REST state. The least destructive repair is a namespaced key such as payments-production-deploy and a queue/cancellation mode chosen for that transition. Do not rerun repeatedly until something “wins”; that compounds race evidence.

4. Broken example: approval exists, but for what artifact?

A reviewer sees “production” and clicks approve without checking the candidate. Later the deploy log shows payload digest D2, while the change record expected D1. The approval UI proves a person approved a job; it does not automatically prove the person compared a digest.

Correlate four identities:

Evidence Question
Workflow headSha Which repository/workflow revision orchestrated this run?
Payload embedded source SHA Which source revision was the artifact built from?
Payload SHA-256 / attestation Which exact bytes were approved/promoted?
Environment deployment/reviewer record Who admitted the transition, and when?

Repair: reject/cancel the candidate if still waiting; otherwise perform a fresh controlled deployment of the verified artifact. Never move a tag or overwrite same-version bytes to make the historical approval “look right.”

5. Broken example: rollback rebuilds an old commit

During an incident an operator checks out yesterday's SHA and runs the build again. The new package has digest D3, while yesterday's known-good production package was D0. The source may match, but the artifact identity does not.

Preserve both digests. If retained D0 is available and compatible with current runtime/data state, redeploy it through the rollback path. If it is not retained, admit that the system no longer has byte-identical rollback capability; treat the rebuild as a new candidate requiring verification rather than mislabeling it as the old artifact.

6. “Protection rule missing” can be plan/visibility, not YAML

A private repository on GitHub Free cannot use the same environment capabilities as the mandatory public lab. Required-reviewer/wait-timer availability also changes with plan/visibility. A missing UI option or ignored policy should therefore trigger a product-boundary check before code changes.

Record repository visibility and owner type, current plan/policy documentation, and deployment (GitHub.com versus GHES version). For GHES, use the matching versioned documentation rather than assuming GitHub.com features—especially newer concurrency queue semantics—exist on every server release.

7. Security-sensitive actions: what this chapter does not normalize

Action Why sensitive Safer incident default
Bypass environment protection Erases separation of duties at the exact risk boundary. Use documented break-glass only when necessary; record actor/reason; preserve artifact identity.
Create/export deployment credential Expands credential attack surface. Prefer existing scoped environment secret or OIDC; never print credentials.
Force-update branch/tag Can invalidate source/release references. Create a new revert/repair commit or new governed deployment.
Delete deployment evidence/environment May remove context/secrets/rules and waiting jobs can fail. Preserve evidence first; cleanup only disposable resources.
Register self-hosted runner Adds persistent compute/network trust. Not needed here; use hosted runner in lab.

8. Compact runbook

REPO="OWNER/c19-deployment-governance-lab"
RUN_ID="123456789"   # synthetic example; substitute observed run ID

gh run view "$RUN_ID" -R "$REPO"   --json databaseId,event,headSha,status,conclusion,url,jobs

gh api -H "X-GitHub-Api-Version: 2026-03-10"   "repos/$REPO/environments?per_page=100"   --jq '.environments[] | {name,protection_rules}'

gh api -H "X-GitHub-Api-Version: 2026-03-10"   "repos/$REPO/deployments?per_page=100"   --jq '[.[] | {id,environment,sha,ref,created_at}]' 

Then inspect the specific deployment's statuses and compare them to workflow logs and the target system's own audit/health evidence.

Knowledge checks

A production job did not wait for review, and its environment is “prodution.” What should you inspect first?

Why can two unrelated workflows interfere even if they target different GitHub environments?

A reviewer approved run 42, but production received bytes with an unexpected digest. Is the approval sufficient evidence?

What does a different rollback rebuild digest prove?

Why should you check repository visibility before debugging a missing environment protection feature?

Summary

Deployment diagnostics should preserve state before correction. Exact environment names, concurrency groups, workflow/source/artifact identities, protection availability, deployment records, and target evidence form the causal chain. The safest correction is usually a narrow policy/name/configuration fix or a new controlled deployment—not bypass, force update, history rewrite, or evidence deletion.

Next: the checkpoint integrates two environments, explicit digest identity, protected promotion, a controlled failure, and a retained-artifact rollback into one production-style runbook.

Next lesson

Checkpoint Lab — Environments, Deployment Protection Rules, Approvals, Concurrency, and Rollbacks

Further reading

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.