Chapter 29Lesson 04~220 minutes

Infrastructure-as-Code Pipelines, Terraform/OpenTofu Workflows, Plan Reviews, State Safety, and Drift Controls: Diagnostics, Failure Modes, Security, and Performance

Diagnose sensitive plans, code/provider mismatch, concurrent mutation, exposed state, stale credentials, lock contention, and blind drift correction without destroying evidence or broadening credentials.

DiagnosticsState lockSensitive planProvider lockDrift

Learning objectives

  • Use an evidence-first diagnostic sequence across GitLab, IaC tool, backend/state and provider layers.
  • Diagnose a plan containing secrets without copying it into more locations.
  • Detect when apply runs different source/provider identity from plan.
  • Explain concurrent-apply and lock contention causally.
  • Separate drift detection from authorized reconciliation.

1. Evidence-first diagnostic sequence

Preserve evidence before retrying: pipeline ID, job IDs, source/ref/SHA, compiled configuration, first failure log, plan digest/metadata, provider lock digest, backend/state identity, lock state and external/provider observations. Then locate the failure layer.

Layer Question Evidence
Compilation/rules Was the expected plan/apply job created? Merged YAML, workflow/job rules, pipeline source.
Runner/toolchain Did the intended CLI/provider set run? Runner/executor/image, tofu version, lockfile digest.
Plan Which exact proposed change failed/reviewed? Plan job ID, plan digest, human summary/report.
Identity/backend Could the job read/lock/write the intended state? State name/address, authorization denial, lock owner/status.
Apply Did it consume the reviewed plan? Digest check, source SHA check, apply trace.
Provider/external What target operation failed or produced unhealthy state? Provider/API diagnostics and external read-only verification.
Drift/governance Was a difference auto-fixed without authorization? Drift pipeline, policy/manual record, later state history.

2. Failure: the plan contains a secret

Suppose a plan artifact contains a generated password or backend token. Do not “diagnose” by downloading it to more machines or pasting it into an issue. Preserve metadata—job/artifact ID, digest, access scope and exposure window—then restrict access, rotate/revoke the affected secret if exposure is plausible, and fix the configuration so secret values are not materialized in public review output.

Security-sensitive action: changing artifact visibility, deleting a leaked lab artifact or rotating a credential is disruptive. Scope the exact artifact/credential and preserve the audit record before remediation.

Remember that marking an IaC output sensitive hides normal CLI rendering but does not guarantee the value is absent from state or every plan representation.

3. Intentionally broken pipeline: review one thing, apply another

plan:
  stage: plan
  script:
    - tofu init -input=false
    - tofu plan -out=plan.cache
  artifacts:
    paths: [plan.cache]

apply:
  stage: apply
  when: manual
  script:
    - tofu init -upgrade
    - tofu apply -auto-approve

This is broken in multiple layers. init -upgrade can change provider selection. The apply command does not consume plan.cache, so it computes a new plan. There is no source SHA/plan digest guard. No resource group serializes mutation. Artifact access/retention is unspecified.

Repair the cause rather than merely rerunning:

apply:
  stage: apply
  needs: [plan]
  resource_group: production-state
  when: manual
  allow_failure: false
  script:
    - sha256sum -c plan.sha256
    - grep -Fx "source_sha=$CI_COMMIT_SHA" plan-metadata.txt
    - tofu init -input=false -lockfile=readonly
    - tofu apply -input=false -auto-approve plan.cache

4. Failure: apply uses different code or provider version

Compare the apply job’s CI_COMMIT_SHA, repository cleanliness, OpenTofu version and lockfile digest to the plan metadata. If the source/provider set differs, the correct response is a new plan and review. Do not force apply because “the change is small.”

printf 'pipeline_source=%s sha=%s pipeline=%s job=%s\n'   "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID"
tofu version
sha256sum .terraform.lock.hcl
sha256sum plan.cache
cat plan-metadata.txt

5. Failure: concurrent applies corrupt intent

A robust backend lock should prevent simultaneous state writes, but concurrent pipelines can still produce stale plans and queues. Preserve both pipeline/job IDs and each plan’s source/state assumptions. Do not disable locking to “unstick” the second pipeline.

Use a GitLab resource group keyed to the same state. If a lock remains after an interrupted job, investigate the owning operation first. Force-unlock is a disruptive recovery action and should be guarded by exact lock identity plus proof that the original writer is gone.

6. Failure: state was uploaded as an artifact

State may contain credentials, connection strings, generated secrets and private attributes. If terraform.tfstate appears in job artifacts, preserve the artifact/job ID and access logs where available, restrict/delete only the identified lab artifact according to policy, rotate exposed credentials if necessary, and move state back to a protected backend. Never fix this by merely adding the filename to .gitignore; Git and job artifacts are different data flows.

7. Failure: apply cannot lock GitLab state after a saved plan

GitLab documents a specific cause: a plan created after init -backend-config=password=$CI_JOB_TOKEN can retain the first job’s short-lived token configuration. The apply job has a different job token, so lock operations fail.

Evidence is a lock/authentication error in apply while the project/state/roles are otherwise correct. Fix backend configuration by using supported HTTP environment variables or the GitLab OpenTofu component so each job injects its current identity. Do not replace the short-lived job token with a broad persistent PAT by default.

8. Failure: lock contention looks like a provider problem

Differentiate “state is locked” from “provider API rejected the mutation.” Backend lock errors occur before many provider actions. Record state name, lock ID/owner metadata if available, competing job IDs and resource-group queue. If the competing job is healthy, wait. If it is dead, use the backend’s documented recovery process after review.

9. Failure: drift is auto-fixed without review

A scheduled pipeline detects that an internet-facing rule changed out-of-band and immediately applies the repository configuration. That may undo an emergency mitigation or erase forensic evidence. The bug is governance/recovery design, not merely OpenTofu syntax.

Change the scheduled job to generate a plan/evidence packet and notify the owner. Reconciliation becomes a separate authorized mutation. For an explicitly self-healing low-risk domain, document that policy and its exception/rollback path instead of relying on an implicit auto-apply.

10. Failure: plan artifact is missing because the job failed

Do not infer “no changes” from “no plan artifact.” Separate tool exit semantics from artifact upload semantics. Preserve the failed job trace and the provider/backend error. If a diagnostic artifact is safe and useful, configure artifacts:when: always for sanitized metadata—not raw secrets/state—so the first-failure evidence survives.

11. Failure: sanitized MR report and full plan are confused

The GitLab Terraform/OpenTofu report is often a reduced review surface. A green widget or safe summary does not mean the full binary plan is safe to expose. Track the report artifact and plan artifact separately with different retention/access controls where needed.

12. Performance without correctness shortcuts

Symptom Unsafe shortcut Better investigation
Plans are slow Disable refresh or state locking everywhere. Profile provider/API latency; split state by ownership; cache provider plugins safely; use refresh controls only with understood semantics.
Apply queue is long Remove resource_group. Measure mutation duration, split independent states, reduce unnecessary applies, tune ordering policy.
Provider download dominates Use mutable unreviewed mirror/image. Pin provider lock/checksums, use reviewed mirrors/caches and immutable tool images.
Large state causes lock contention Copy state into per-job artifacts. Refactor ownership/state boundaries and use remote backend/data interfaces.

13. Security boundaries during troubleshooting

  • Do not print state, environment-variable inventories, tokens or raw backend credentials.
  • Do not disable TLS verification to reach a backend or provider.
  • Do not grant cloud administrator rights because a provider call is denied.
  • Do not upload state/plan to public issue trackers for debugging.
  • Do not force-unlock or delete state until the exact lock/state identity and recovery impact are understood.
  • Do not rebuild or re-plan and pretend the new output is the original reviewed evidence.

14. Least-destructive recovery map

Evidence Likely layer Smallest safe next action
YAML invalid / apply job absent GitLab compilation/rules Fix CI configuration; no state action.
Plan digest mismatch Artifact/evidence Stop; generate a new plan and review.
Provider lock digest changed Toolchain Restore/pin intended lock or intentionally upgrade and re-plan.
Backend lock held by live job Concurrency/backend Wait; preserve queue/lock evidence.
Backend lock held by dead job Recovery/backend Review exact lock ID then documented unlock procedure.
403 from provider Identity/provider Check role/scope/trust; do not broaden to admin reflexively.
Code 2 in drift job External drift Create review/reconciliation decision; do not treat as tool failure.

Knowledge check

Why is tofu init -upgrade in an apply job suspicious after plan review?

A state lock blocks the pipeline. What should you do first?

Why can a plan artifact need credential rotation after accidental exposure?

A drift job returns code 2. Should GitLab retry it as a failed job?

What proves apply used the reviewed plan?

Next lesson

Next lesson

Lesson 5 combines these controls in a checkpoint: isolated state, plan digest, simulated review, serialized exact apply, controlled drift, evidence packet and reviewed cleanup.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. GitLab-managed OpenTofu/Terraform state and merge-request plan integration are available on Free, Premium, and Ultimate. GitLab’s legacy Terraform CI/CD templates were removed in GitLab 18.0; current guidance is the versioned OpenTofu CI/CD component. GitLab-managed OpenTofu/Terraform state support in the current CLI flow was introduced in GitLab 18.3 and requires glab 1.66+. Plan files are not encrypted by GitLab or OpenTofu and can contain credentials or sensitive values; restrict artifact access and never treat plan output as safe by default. The current released OpenTofu component used for optional examples is 4.8.1; that release supports OpenTofu 1.12.5 and publishes signed rootless images. The standalone local lab pins OpenTofu 1.12.6 (latest stable on the verification date) and hashicorp/local 2.9.0. GitLab’s troubleshooting note about job-token values persisted through backend configuration is especially important when splitting plan/apply across jobs. Use supported runtime backend variables/component behavior.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.