Infrastructure-as-Code Pipelines, Terraform/OpenTofu Workflows, Plan Reviews, State Safety, and Drift Controls: Diagnostics, Failure Modes, Security, and Performance
Diagnose sensitive plans, code/provider mismatch, concurrent mutation, exposed state, stale credentials, lock contention, and blind drift correction without destroying evidence or broadening credentials.
Learning objectives
- Use an evidence-first diagnostic sequence across GitLab, IaC tool, backend/state and provider layers.
- Diagnose a plan containing secrets without copying it into more locations.
- Detect when apply runs different source/provider identity from plan.
- Explain concurrent-apply and lock contention causally.
- Separate drift detection from authorized reconciliation.
1. Evidence-first diagnostic sequence
Preserve evidence before retrying: pipeline ID, job IDs, source/ref/SHA, compiled configuration, first failure log, plan digest/metadata, provider lock digest, backend/state identity, lock state and external/provider observations. Then locate the failure layer.
| Layer | Question | Evidence |
|---|---|---|
| Compilation/rules | Was the expected plan/apply job created? | Merged YAML, workflow/job rules, pipeline source. |
| Runner/toolchain | Did the intended CLI/provider set run? |
Runner/executor/image, tofu version, lockfile
digest.
|
| Plan | Which exact proposed change failed/reviewed? | Plan job ID, plan digest, human summary/report. |
| Identity/backend | Could the job read/lock/write the intended state? | State name/address, authorization denial, lock owner/status. |
| Apply | Did it consume the reviewed plan? | Digest check, source SHA check, apply trace. |
| Provider/external | What target operation failed or produced unhealthy state? | Provider/API diagnostics and external read-only verification. |
| Drift/governance | Was a difference auto-fixed without authorization? | Drift pipeline, policy/manual record, later state history. |
2. Failure: the plan contains a secret
Suppose a plan artifact contains a generated password or backend token. Do not “diagnose” by downloading it to more machines or pasting it into an issue. Preserve metadata—job/artifact ID, digest, access scope and exposure window—then restrict access, rotate/revoke the affected secret if exposure is plausible, and fix the configuration so secret values are not materialized in public review output.
Remember that marking an IaC output sensitive hides
normal CLI rendering but does not guarantee the value is absent from
state or every plan representation.
3. Intentionally broken pipeline: review one thing, apply another
plan:
stage: plan
script:
- tofu init -input=false
- tofu plan -out=plan.cache
artifacts:
paths: [plan.cache]
apply:
stage: apply
when: manual
script:
- tofu init -upgrade
- tofu apply -auto-approve
This is broken in multiple layers. init -upgrade can
change provider selection. The apply command does not consume
plan.cache, so it computes a new plan. There is no
source SHA/plan digest guard. No resource group serializes mutation.
Artifact access/retention is unspecified.
Repair the cause rather than merely rerunning:
apply:
stage: apply
needs: [plan]
resource_group: production-state
when: manual
allow_failure: false
script:
- sha256sum -c plan.sha256
- grep -Fx "source_sha=$CI_COMMIT_SHA" plan-metadata.txt
- tofu init -input=false -lockfile=readonly
- tofu apply -input=false -auto-approve plan.cache
4. Failure: apply uses different code or provider version
Compare the apply job’s CI_COMMIT_SHA, repository
cleanliness, OpenTofu version and lockfile digest to the plan
metadata. If the source/provider set differs, the correct response
is a new plan and review. Do not force apply because “the change is
small.”
printf 'pipeline_source=%s sha=%s pipeline=%s job=%s\n' "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID"
tofu version
sha256sum .terraform.lock.hcl
sha256sum plan.cache
cat plan-metadata.txt
5. Failure: concurrent applies corrupt intent
A robust backend lock should prevent simultaneous state writes, but concurrent pipelines can still produce stale plans and queues. Preserve both pipeline/job IDs and each plan’s source/state assumptions. Do not disable locking to “unstick” the second pipeline.
Use a GitLab resource group keyed to the same state. If a lock remains after an interrupted job, investigate the owning operation first. Force-unlock is a disruptive recovery action and should be guarded by exact lock identity plus proof that the original writer is gone.
6. Failure: state was uploaded as an artifact
State may contain credentials, connection strings, generated secrets
and private attributes. If terraform.tfstate appears in
job artifacts, preserve the artifact/job ID and access logs where
available, restrict/delete only the identified lab artifact
according to policy, rotate exposed credentials if necessary, and
move state back to a protected backend. Never fix this by merely
adding the filename to .gitignore; Git and job
artifacts are different data flows.
7. Failure: apply cannot lock GitLab state after a saved plan
GitLab documents a specific cause: a plan created after
init -backend-config=password=$CI_JOB_TOKEN can retain
the first job’s short-lived token configuration. The apply job has a
different job token, so lock operations fail.
Evidence is a lock/authentication error in apply while the project/state/roles are otherwise correct. Fix backend configuration by using supported HTTP environment variables or the GitLab OpenTofu component so each job injects its current identity. Do not replace the short-lived job token with a broad persistent PAT by default.
8. Failure: lock contention looks like a provider problem
Differentiate “state is locked” from “provider API rejected the mutation.” Backend lock errors occur before many provider actions. Record state name, lock ID/owner metadata if available, competing job IDs and resource-group queue. If the competing job is healthy, wait. If it is dead, use the backend’s documented recovery process after review.
9. Failure: drift is auto-fixed without review
A scheduled pipeline detects that an internet-facing rule changed out-of-band and immediately applies the repository configuration. That may undo an emergency mitigation or erase forensic evidence. The bug is governance/recovery design, not merely OpenTofu syntax.
Change the scheduled job to generate a plan/evidence packet and notify the owner. Reconciliation becomes a separate authorized mutation. For an explicitly self-healing low-risk domain, document that policy and its exception/rollback path instead of relying on an implicit auto-apply.
10. Failure: plan artifact is missing because the job failed
Do not infer “no changes” from “no plan artifact.” Separate tool
exit semantics from artifact upload semantics. Preserve the failed
job trace and the provider/backend error. If a diagnostic artifact
is safe and useful, configure
artifacts:when: always for sanitized metadata—not raw
secrets/state—so the first-failure evidence survives.
11. Failure: sanitized MR report and full plan are confused
The GitLab Terraform/OpenTofu report is often a reduced review surface. A green widget or safe summary does not mean the full binary plan is safe to expose. Track the report artifact and plan artifact separately with different retention/access controls where needed.
12. Performance without correctness shortcuts
| Symptom | Unsafe shortcut | Better investigation |
|---|---|---|
| Plans are slow | Disable refresh or state locking everywhere. | Profile provider/API latency; split state by ownership; cache provider plugins safely; use refresh controls only with understood semantics. |
| Apply queue is long | Remove resource_group. |
Measure mutation duration, split independent states, reduce unnecessary applies, tune ordering policy. |
| Provider download dominates | Use mutable unreviewed mirror/image. | Pin provider lock/checksums, use reviewed mirrors/caches and immutable tool images. |
| Large state causes lock contention | Copy state into per-job artifacts. | Refactor ownership/state boundaries and use remote backend/data interfaces. |
13. Security boundaries during troubleshooting
- Do not print state, environment-variable inventories, tokens or raw backend credentials.
- Do not disable TLS verification to reach a backend or provider.
- Do not grant cloud administrator rights because a provider call is denied.
- Do not upload state/plan to public issue trackers for debugging.
- Do not force-unlock or delete state until the exact lock/state identity and recovery impact are understood.
- Do not rebuild or re-plan and pretend the new output is the original reviewed evidence.
14. Least-destructive recovery map
| Evidence | Likely layer | Smallest safe next action |
|---|---|---|
| YAML invalid / apply job absent | GitLab compilation/rules | Fix CI configuration; no state action. |
| Plan digest mismatch | Artifact/evidence | Stop; generate a new plan and review. |
| Provider lock digest changed | Toolchain | Restore/pin intended lock or intentionally upgrade and re-plan. |
| Backend lock held by live job | Concurrency/backend | Wait; preserve queue/lock evidence. |
| Backend lock held by dead job | Recovery/backend | Review exact lock ID then documented unlock procedure. |
| 403 from provider | Identity/provider | Check role/scope/trust; do not broaden to admin reflexively. |
| Code 2 in drift job | External drift | Create review/reconciliation decision; do not treat as tool failure. |
Knowledge check
Why is tofu init -upgrade in an apply job
suspicious after plan review?
It can select newer providers, so the apply toolchain may differ from the plan toolchain. Use the committed lock file with readonly initialization.
A state lock blocks the pipeline. What should you do first?
Identify the state/lock owner and competing job. Do not disable locking or force-unlock until you prove the writer is no longer active and recovery is safe.
Why can a plan artifact need credential rotation after accidental exposure?
Plans can contain sensitive values or backend/provider data. Restricting future access does not undo prior disclosure.
A drift job returns code 2. Should GitLab retry it as a failed job?
No. Code 2 means changes are proposed. Capture it as drift evidence and route to review; reserve failure handling for code 1 or other errors.
What proves apply used the reviewed plan?
At minimum the saved plan digest verification plus source SHA/tool/provider/backend identity and the apply job record that consumed that file.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12.
GitLab-managed OpenTofu/Terraform state and merge-request plan
integration are available on Free, Premium, and Ultimate. GitLab’s
legacy Terraform CI/CD templates were removed in GitLab 18.0;
current guidance is the versioned OpenTofu CI/CD component.
GitLab-managed OpenTofu/Terraform state support in the current CLI
flow was introduced in GitLab 18.3 and requires
glab 1.66+. Plan files are not encrypted by GitLab or
OpenTofu and can contain credentials or sensitive values; restrict
artifact access and never treat plan output as safe by default. The
current released OpenTofu component used for optional examples is
4.8.1; that release supports OpenTofu
1.12.5 and publishes signed rootless images. The
standalone local lab pins OpenTofu 1.12.6 (latest
stable on the verification date) and
hashicorp/local 2.9.0. GitLab’s troubleshooting note
about job-token values persisted through backend configuration is
especially important when splitting plan/apply across jobs. Use
supported runtime backend variables/component behavior.
- Infrastructure as Code with OpenTofu and GitLab — official reference.
- GitLab-managed Terraform/OpenTofu state — official reference.
- OpenTofu integration in merge requests — official reference.
- CI/CD artifacts reports — terraform — official reference.
- CI/CD YAML syntax — resource_group and artifacts access — official reference.
- Resource groups — official reference.
- GitLab IaC troubleshooting — official reference.
- GitLab OpenTofu CI/CD component — official reference.
- GitLab deprecations and removals — official reference.
- glab opentofu state — official reference.
- OpenTofu 1.12 documentation — official reference.
- OpenTofu releases — official reference.
- HashiCorp local provider — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.