Chapter 27Lesson 04~195 minutes

Infrastructure as Code, Terraform Plans, Policy Checks, and Deployment Workflows: Diagnostics, Failure Modes, and Production Practices

Diagnose IaC failures by preserving source, plan, policy, state, identity, lock and partial-apply evidence before correction or recovery.

DiagnosticsState lockingSensitive planPartial applyRecovery

Learning objectives

  • Diagnose IaC failures without erasing the original run, plan, state or lock evidence.
  • Explain why privileged PR apply, wrong workspace, raw plan exposure and blind force-unlock are causal security failures.
  • Separate YAML/event failures from provider/runtime, backend/state, policy and external-resource failures.
  • Handle partial apply and stale plans with the least destructive correction.
  • Interpret one intentionally broken workflow and repair boundaries rather than symptoms.

1. Evidence-first diagnostic sequence

IaC troubleshooting is dangerous when the first response is “rerun apply.” Infrastructure may already be partially changed even though a job is red. Preserve the run ID/attempt, event/ref/SHA, workflow revision, plan digest, policy output, tool/provider versions, backend/workspace identity, state lock/error, and provider-side resource status before making another mutation.

  1. Preserve run/attempt and first-failure logs/evidence.
  2. Confirm event/ref/SHA and workflow revision.
  3. Confirm inputs, conditions, environment and evaluated permissions.
  4. Confirm runner/toolchain and provider lock.
  5. Confirm backend/workspace and state identity/locking.
  6. Confirm saved-plan digest and policy result.
  7. Inspect provider/API or local target state independently.
  8. Apply the least destructive correction.
  9. Rerun only the smallest equivalent scope after understanding side effects.

2. Intentionally broken workflow: identify the trust violations

# BROKEN — do not use in production
on: pull_request
permissions: write-all
jobs:
  terraform:
    runs-on: self-hosted
    steps:
      - uses: actions/checkout@main
      - run: terraform init
      - run: terraform plan -out=tfplan
      - run: terraform show -json tfplan
      - run: terraform apply -auto-approve

This workflow combines almost every prohibited shortcut. Untrusted PR code executes on a self-hosted runner, write-all is granted, checkout is mutable, tool version/provider lock are uncontrolled, raw plan JSON is printed, no plan/policy/approval boundary exists, and apply calculates/executes directly against whichever backend the PR config selects. A green run would still be weak evidence.

3. Failure: applying PR code with privileged credentials

PR content can modify providers, backends, data sources, modules, provisioners and scripts. Supplying production cloud credentials or OIDC mutation authority to arbitrary PR code turns review automation into an execution boundary. Keep untrusted PRs on hosted runners with read-only/minimal permissions and move privileged planning/apply to a trusted revision or carefully validated handoff.

4. Failure: correct plan, wrong backend/workspace

A plan that was reviewed for staging must not be applied to production merely because the resource names look similar. Record backend/workspace identity in the plan evidence packet and reassert it before mutation. In cloud providers also confirm account/project/subscription and region. If those values differ, stop before apply.

5. Failure: exposing secrets through plan/state diagnostics

Terraform marks some values sensitive in human CLI output, but saved plans, state and JSON representations can still contain sensitive data. Printing terraform show -json into public logs or attaching production state for debugging can create a credential incident. Use synthetic fixtures for learning; in production retain only necessary protected evidence and rotate any secret whose value was exposed.

If a real secret appears in logs/artifacts, deletion alone is not remediation. Revoke/rotate the credential and preserve an incident record; cleanup only reduces further exposure.

6. Failure: auto-approving destructive change

A successful plan can legitimately propose deletes. Automation should make destructive actions visible to policy and review rather than hide them behind -auto-approve. The -auto-approve flag is acceptable only after an upstream noninteractive approval mechanism has already authorized the exact plan; it is not the approval mechanism itself.

7. Failure: “fixing” a lock by force-unlocking

A state lock means another operation may still own the state. Before any force-unlock, identify the exact backend, lock ID, owning run/process and provider-side operation. Blind force-unlock can create concurrent writers and state corruption. A canceled GitHub job does not prove the remote operation stopped.

For local training state, the safest repair is usually to stop the duplicate local process and retry after confirming no writer exists. For remote backends, follow the backend-specific recovery procedure and preserve lock metadata.

8. Failure: partial apply

Terraform can create some objects and fail on a later one. The red job conclusion therefore does not mean “nothing changed.” Inspect state and the provider target before retry. Correct the causal configuration/permission/quota problem, then generate a fresh plan from current state. Do not rebuild an old plan or repeatedly apply it blindly.

9. Failure: stale saved plan

Stale-plan rejection is often a safety signal. It says another state transition occurred after planning. Preserve the old plan digest and first-failure message, inspect state history/current resources, then create a new plan for review. Do not bypass the check by replacing backend state with an older copy unless a deliberate state-recovery procedure requires it.

10. Causal layer map

Symptom Likely layer First evidence
Workflow never plans event/YAML/expression selected workflow revision and event payload
Provider install differs toolchain/dependency Terraform version + lock diff/checksum
Plan differs unexpectedly backend/workspace/refresh/input target identity + state serial/remote object
Policy denies policy exact plan hash + rule version/reason
OIDC exchange denied identity/trust bounded claims + provider trust response
Apply lock error backend concurrency lock ID/owner + competing run
Apply red but resources exist partial external mutation state + provider-side resource status
Apply says stale state transition timing old plan digest + current state identity

11. Repair the broken example by restoring boundaries

The repair is not one YAML tweak. Use a GitHub-hosted runner for untrusted validation, pin checkout and setup actions by full SHA, pin Terraform and providers, commit the lock file, plan against an authorized target with only required read privileges, store protected plan evidence, evaluate policy, then perform mutation from a trusted revision behind an environment with narrowly scoped identity. Record target health/state afterward.

12. Lesson summary

IaC incident safety depends on preserving the first plan/state/resource evidence and recognizing that a failed apply can still have side effects. Fix the causal boundary; do not widen permissions, force-unlock blindly or erase evidence.

Next lesson

Checkpoint Lab — Infrastructure as Code, Terraform Plans, Policy Checks, and Deployment Workflows

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

A Terraform apply job fails halfway. Can you assume no infrastructure changed?

Why is terraform show -json tfplan dangerous in logs?

What should you do first when encountering a state lock?

A saved plan is stale after another deployment. What is the safe next step?

Why is a trusted self-hosted runner a poor place for arbitrary fork PR Terraform?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.