Infrastructure as Code, Terraform Plans, Policy Checks, and Deployment Workflows: Diagnostics, Failure Modes, and Production Practices
Diagnose IaC failures by preserving source, plan, policy, state, identity, lock and partial-apply evidence before correction or recovery.
Learning objectives
- Diagnose IaC failures without erasing the original run, plan, state or lock evidence.
- Explain why privileged PR apply, wrong workspace, raw plan exposure and blind force-unlock are causal security failures.
- Separate YAML/event failures from provider/runtime, backend/state, policy and external-resource failures.
- Handle partial apply and stale plans with the least destructive correction.
- Interpret one intentionally broken workflow and repair boundaries rather than symptoms.
1. Evidence-first diagnostic sequence
IaC troubleshooting is dangerous when the first response is “rerun apply.” Infrastructure may already be partially changed even though a job is red. Preserve the run ID/attempt, event/ref/SHA, workflow revision, plan digest, policy output, tool/provider versions, backend/workspace identity, state lock/error, and provider-side resource status before making another mutation.
- Preserve run/attempt and first-failure logs/evidence.
- Confirm event/ref/SHA and workflow revision.
- Confirm inputs, conditions, environment and evaluated permissions.
- Confirm runner/toolchain and provider lock.
- Confirm backend/workspace and state identity/locking.
- Confirm saved-plan digest and policy result.
- Inspect provider/API or local target state independently.
- Apply the least destructive correction.
- Rerun only the smallest equivalent scope after understanding side effects.
2. Intentionally broken workflow: identify the trust violations
# BROKEN — do not use in production
on: pull_request
permissions: write-all
jobs:
terraform:
runs-on: self-hosted
steps:
- uses: actions/checkout@main
- run: terraform init
- run: terraform plan -out=tfplan
- run: terraform show -json tfplan
- run: terraform apply -auto-approve
This workflow combines almost every prohibited shortcut. Untrusted
PR code executes on a self-hosted runner, write-all is
granted, checkout is mutable, tool version/provider lock are
uncontrolled, raw plan JSON is printed, no plan/policy/approval
boundary exists, and apply calculates/executes directly against
whichever backend the PR config selects. A green run would still be
weak evidence.
3. Failure: applying PR code with privileged credentials
PR content can modify providers, backends, data sources, modules, provisioners and scripts. Supplying production cloud credentials or OIDC mutation authority to arbitrary PR code turns review automation into an execution boundary. Keep untrusted PRs on hosted runners with read-only/minimal permissions and move privileged planning/apply to a trusted revision or carefully validated handoff.
4. Failure: correct plan, wrong backend/workspace
A plan that was reviewed for staging must not be applied to production merely because the resource names look similar. Record backend/workspace identity in the plan evidence packet and reassert it before mutation. In cloud providers also confirm account/project/subscription and region. If those values differ, stop before apply.
5. Failure: exposing secrets through plan/state diagnostics
Terraform marks some values sensitive in human CLI output, but saved
plans, state and JSON representations can still contain sensitive
data. Printing terraform show -json into public logs or
attaching production state for debugging can create a credential
incident. Use synthetic fixtures for learning; in production retain
only necessary protected evidence and rotate any secret whose value
was exposed.
If a real secret appears in logs/artifacts, deletion alone is not remediation. Revoke/rotate the credential and preserve an incident record; cleanup only reduces further exposure.
6. Failure: auto-approving destructive change
A successful plan can legitimately propose deletes. Automation
should make destructive actions visible to policy and review rather
than hide them behind -auto-approve. The
-auto-approve flag is acceptable only after an upstream
noninteractive approval mechanism has already authorized the exact
plan; it is not the approval mechanism itself.
7. Failure: “fixing” a lock by force-unlocking
A state lock means another operation may still own the state. Before any force-unlock, identify the exact backend, lock ID, owning run/process and provider-side operation. Blind force-unlock can create concurrent writers and state corruption. A canceled GitHub job does not prove the remote operation stopped.
For local training state, the safest repair is usually to stop the duplicate local process and retry after confirming no writer exists. For remote backends, follow the backend-specific recovery procedure and preserve lock metadata.
8. Failure: partial apply
Terraform can create some objects and fail on a later one. The red job conclusion therefore does not mean “nothing changed.” Inspect state and the provider target before retry. Correct the causal configuration/permission/quota problem, then generate a fresh plan from current state. Do not rebuild an old plan or repeatedly apply it blindly.
9. Failure: stale saved plan
Stale-plan rejection is often a safety signal. It says another state transition occurred after planning. Preserve the old plan digest and first-failure message, inspect state history/current resources, then create a new plan for review. Do not bypass the check by replacing backend state with an older copy unless a deliberate state-recovery procedure requires it.
10. Causal layer map
| Symptom | Likely layer | First evidence |
|---|---|---|
| Workflow never plans | event/YAML/expression | selected workflow revision and event payload |
| Provider install differs | toolchain/dependency | Terraform version + lock diff/checksum |
| Plan differs unexpectedly | backend/workspace/refresh/input | target identity + state serial/remote object |
| Policy denies | policy | exact plan hash + rule version/reason |
| OIDC exchange denied | identity/trust | bounded claims + provider trust response |
| Apply lock error | backend concurrency | lock ID/owner + competing run |
| Apply red but resources exist | partial external mutation | state + provider-side resource status |
| Apply says stale | state transition timing | old plan digest + current state identity |
11. Repair the broken example by restoring boundaries
The repair is not one YAML tweak. Use a GitHub-hosted runner for untrusted validation, pin checkout and setup actions by full SHA, pin Terraform and providers, commit the lock file, plan against an authorized target with only required read privileges, store protected plan evidence, evaluate policy, then perform mutation from a trusted revision behind an environment with narrowly scoped identity. Record target health/state afterward.
12. Lesson summary
IaC incident safety depends on preserving the first plan/state/resource evidence and recognizing that a failed apply can still have side effects. Fix the causal boundary; do not widen permissions, force-unlock blindly or erase evidence.
Knowledge check
A Terraform apply job fails halfway. Can you assume no infrastructure changed?
No. Inspect current state and provider-side resources before any retry.
Why is terraform show -json tfplan dangerous in
logs?
Plan JSON may contain sensitive data even when normal human output hides values.
What should you do first when encountering a state lock?
Identify the backend, lock owner/ID and competing operation; do not force-unlock blindly.
A saved plan is stale after another deployment. What is the safe next step?
Preserve the failure, inspect the new state, then generate and review a fresh plan.
Why is a trusted self-hosted runner a poor place for arbitrary fork PR Terraform?
Untrusted code may access persistent host/network state or credentials and compromise the runner/target boundary.
Official references and version notes
- Terraform CLI plan — Saved plans, planning modes, detailed exit codes and the security warning for plan files.
- Terraform CLI apply — Applying a saved plan and automation semantics.
- Terraform dependency lock file — Provider selections/checksums and why the lock file belongs with source.
- Terraform state locking — State-lock behavior and the operational risk of force-unlock.
- Terraform local backend — Local state backend used only by the disposable no-cloud lab.
- HashiCorp setup-terraform — Official setup action; current v4 line uses Node 24.
- GitHub environments — Approval/protection boundary for a guarded apply job.
- GitHub OIDC overview — Preferred short-lived cloud federation boundary for optional real-provider adapters.
- Workflow artifacts — Plan/evidence transfer is GitHub artifact state, not infrastructure state.
- OpenTofu documentation — Open-source alternative CLI with a similar plan/apply operating boundary; verify syntax/version separately before substitution.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.