Chapter 21Lesson 04~185 minutes

Protected Environments, Deployment Approvals, Freeze Windows, Manual Gates, and Separation of Duties: Diagnostics, Failure Modes, Security, and Performance

Diagnose self-approval, approval-versus-rollout confusion, undocumented freeze bypass, premature secret exposure, and override gaps while preserving pipeline, authorization, deployment, and target evidence.

DiagnosticsGate bypassSecretsOverridesEvidence-first

Learning objectives

  • Use an evidence-first sequence before retrying, approving, changing policy, or altering the target.
  • Diagnose five authorization failures at the correct configuration, identity, deployment, or external-health layer.
  • Repair an intentionally broken gate without erasing the original denial/failure evidence.
  • Explain why protected variables, admin power, and blind retries can undermine a nominal approval workflow.
  • Choose the least destructive correction and rerun only the smallest safe scope.

1. Evidence-first diagnostic sequence

  1. Preserve pipeline ID, deployment job ID, deployment ID, original logs, approval/denial record, and external target observation.
  2. Confirm CI_PIPELINE_SOURCE, ref, CI_COMMIT_SHA, candidate digest, and compiled configuration.
  3. Confirm workflow/job-rule decisions, freeze state, manual-gate attributes, and non-secret inputs.
  4. Inspect job graph/queue and runner/executor metadata before blaming policy.
  5. Inspect artifact/report/cache identity and verify the deploy job received the intended bytes.
  6. Inspect protected-environment/approval state or the local policy simulation.
  7. Inspect deployment record and then the external target independently.
  8. Apply the smallest correction; do not “make it green” by weakening authorization.
Never troubleshoot authorization by printing all variables, using a broad PAT, disabling TLS, or deleting/recreating production state.

2. Failure: the pipeline triggerer self-approves unexpectedly

Symptom: approval history shows the requester as approver even though policy intended independent review.

Interpretation: first verify project approval options and eligible approver groups. Do not assume a compromised account merely from identity overlap. If self-approval was explicitly enabled, the system may be working exactly as configured.

Repair: disable self-approval for the protected environment workflow, ensure another eligible approver exists, preserve the old deployment/approval record, and test a new disposable deployment.

3. Failure: approval is interpreted as successful rollout

Symptom: an approved deployment is reported as “successful” in a change record even though no post-deploy health check exists or the target is unhealthy.

Cause: authorization and external-state evidence were collapsed. Approval only says an eligible actor permitted the attempt.

Repair: record the deployment job/deployment ID and independently query the synthetic target/service. Do not change approval policy to fix a health failure.

4. Failure: a freeze is bypassed with no evidence

Symptom: a production deployment occurs while CI_DEPLOY_FREEZE=true, but there is no exception ID or authorized reason.

Evidence: preserve the pipeline creation time, compiled rules, freeze-period definition, job actor, and deployment record. Check whether the normal deploy job ignored the variable or a distinct manual exception path was used.

Repair: move freeze logic into shared/versioned rules and require a bounded exception record. Do not teach operators to unset or overwrite CI_DEPLOY_FREEZE.

5. Failure: a protected variable is exposed before the gate

Symptom: a pre-deployment “prepare” job can access a production credential even though the later deployment is gated.

prepare_release:
  stage: prepare
  script:
    - ./ci/contact-production.sh   # BROKEN: privileged access before authorization

deploy_production:
  stage: deploy
  when: manual
  allow_failure: false
  environment:
    name: production
  script:
    - ./ci/deploy.sh

Diagnosis: the manual gate protects only the later job. Variable protection/environment scope and job placement determine where the credential is available. Masking would not repair this.

Repair: move the production access into the deployment job, scope the variable to the production environment where supported, and keep pre-gate preparation on non-privileged data.

6. Failure: administrator override has no audit narrative

Symptom: a deployment occurred through administrator authority, but the incident/change record says only “admin override.”

Interpretation: GitLab administrators can have broad protected-environment/approval authority. That is expected platform power, but governance still requires candidate identity, reason, actor, deployment ID, target effect, and verification.

Repair: create an after-action evidence packet. If emergency admin use is common, design a formal exception path so privileged recovery does not depend on undocumented institutional memory.

7. Intentionally broken example: gate exists, but policy is bypassable

stages: [prepare, deploy]

prepare:
  stage: prepare
  script:
    - echo "candidate=$CI_COMMIT_SHA" > candidate.txt
    - ./ci/write-production-marker.sh   # side effect occurs here

deploy:
  stage: deploy
  environment: production
  rules:
    - if: '$CI_DEPLOY_FREEZE'
      when: manual
      allow_failure: true              # optional, not a blocking exception
    - when: manual
  script:
    - echo "approved deployment"

Three problems are visible before running anything: the pre-gate job already performs the side effect; the freeze path is optional; and the candidate is identified only by a mutable text file with no verified digest.

Repair the architecture, not just the YAML color: move all production writes after authorization, set blocking intent explicitly, verify artifact identity, and record the exception/approval actor.

8. Failure taxonomy: fix the right layer

Observation Likely layer Do not “fix” by
Pipeline has no deployment job Rules/workflow/compiled config Changing runner or approval roles
Job waits for human Manual/approval authorization Retrying runner
Eligible approver gets denied Protected-environment membership/approval rule Broadening PAT scope
Job approved but queued Runner/capacity/tags Removing approval requirement
Deploy job gets wrong bytes Artifact/dataflow Adding approvers
Job succeeds; target unhealthy Application/provider/external target Marking approval successful again
Freeze active but job runnable normally Freeze-rule design Unsetting freeze variable

9. Performance without weakening authorization

Authorization can add latency, but the answer is not to merge roles indiscriminately. Improve lead time by making candidate evidence compact, precomputing non-privileged tests before the gate, notifying the narrow approver group, and keeping deployment scripts idempotent and fast. Measure wait-for-approval separately from runner queue and deployment duration so you optimize the actual bottleneck.

10. Least-destructive repair checklist

  • Preserve the first denial/failure and IDs.
  • Verify source SHA and candidate digest.
  • Inspect merged YAML and freeze/manual/approval state.
  • Check eligible approver/deployer identity without broadening roles.
  • Confirm production credentials appear only in the authorized job.
  • Retry only the smallest safe job/pipeline after correcting root cause.
  • Verify deployment record and external health independently.

Knowledge check

A deploy job is approved but remains pending with no runner. Should you remove the approval rule?

Why is masking a production variable insufficient when a pre-gate job can access it?

What evidence should you preserve when a freeze was bypassed?

What is wrong with allow_failure: true on a freeze exception job intended to block deployment?

A target is unhealthy after a correctly authorized, green job. Which layer should be investigated next?

Next lesson

Checkpoint lab

Stage a full simulated production promotion through freeze and approval policy, prove authorization and target health separately, exercise an emergency exception, and assemble the final evidence packet.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. The mandatory learning path uses only disposable/local simulation plus GitLab Free features. Protected environments and deployment approvals are treated as Premium/Ultimate and are simulated unless the learner already has an authorized eligible project.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.