Chapter 17Lesson 04~265 minutes

Environments, Deployments, Review Apps, Protected Environments, and Deployment Approvals: Diagnostics, Failure Modes, Security, and Performance

Preserve environment, deployment, job, artifact, runner, and target evidence; then diagnose drift, stale resources, bypass paths, and approval mistakes without weakening controls.

DiagnosticsDeployment driftStale review appBypassRollbackSecurity

Learning objectives

  • Diagnose GitLab environment/deployment state separately from external target state.
  • Detect wrong target identity and rollback rebuild drift from commit/artifact evidence.
  • Diagnose stale Review Apps and stop-action failures without deleting unrelated resources.
  • Explain how runner/token/IAM paths can bypass the intent of GitLab environment governance.
  • Interpret an intentionally failed deployment while preserving the original cause.
Availability baseline (verified 2026-08-21 against current GitLab documentation). Environments, deployment records, Review Apps/dynamic environments, environment:on_stop, environment:auto_stop_in, environment/deployment APIs, deployment rollback/redeploy, and project-level environment-scoped CI/CD variables are available on GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. Protected environments and deployment approvals are Premium/Ultimate. Group-variable environment scoping is also Premium/Ultimate, so the mandatory chapter path uses only a disposable project and synthetic project-level values. No cloud account, Kubernetes cluster, production endpoint, real credential, or paid approval feature is required.

1. Evidence-first diagnostic sequence

Deployment incidents are dangerous because an engineer under pressure can “fix” the GitLab record while the external target remains wrong. Use the same disciplined sequence throughout the academy:

  1. Preserve evidence: pipeline ID/source/SHA, deployment/job IDs, environment name/state/URL, artifact digest, runner identity, target identifier, logs, and approval/protection state.
  2. Locate scope: offering/instance → namespace/project → ref/MR → pipeline/job/runner → artifact → environment/deployment → external target/policy.
  3. Inspect controls: CI YAML, environment history, scoped variables, runner eligibility, protection/approval settings, target IAM, API responses.
  4. Correct the narrowest cause: do not weaken branch/environment protection or rotate infrastructure blindly.
  5. Verify independently: prove GitLab record and target state converge on the intended artifact identity.

2. Failure: GitLab says “production,” script writes somewhere else

Consider a broken deploy job:

deploy_production:
  script:
    - ./deploy.sh --target testing
  environment:
    name: production
    url: https://production.example.invalid

The job may succeed and GitLab may record a successful production deployment even though the command explicitly targets testing. The correct diagnosis is not “GitLab environment is wrong.” The environment metadata and external operation disagree.

Repair: make target identity a single validated input, log a non-secret target ID, and add post-deployment verification. In production, fail closed if environment name/tier and target identifier do not match the expected mapping.

3. Intentionally broken example: safe failed deployment

Use a disposable environment only:

deploy_broken:
  stage: deploy
  script:
    - echo "artifact_verified=true"
    - echo "simulated_target=review/ch17-broken"
    - echo "ERROR: synthetic deployment rejected by fixture" >&2
    - exit 42
  environment:
    name: review/ch17-broken
    url: https://review-ch17-broken.example.invalid

Expected evidence: the job exits non-zero, pipeline/job UI preserves the stderr line and exit failure, and the deployment history shows a failed attempt for the environment. Do not edit the log or rerun immediately. Record job ID, SHA, deployment status, and the synthetic error first.

Repair: change only the synthetic rejection condition, rerun a new deployment, and verify the new deployment record succeeds. The failed record remains useful evidence; do not delete history to make the dashboard look clean.

4. Failure: “rollback” rebuilds different bytes

You select an old deployment commit, but the deploy job downloads latest dependencies/base image or performs a fresh build. GitLab can correctly point the rollback deployment to the old commit while the payload differs from the original release.

Evidence to preserve: original deployment ID/job/SHA, original artifact/image digest if retained, rollback job/SHA, rebuilt digest, dependency/base-image identities. If digests differ, call the operation a rebuild/recovery—not an identity-preserving rollback.

Repair: store/promote immutable artifacts long enough for the recovery objective and design deployment jobs to accept their immutable identity.

5. Failure: Review App stays available and consumes resources

GitLab environment state can remain available because the stop job never ran, the rules for deploy and stop differ, the branch lifecycle did not trigger cleanup as expected, or the external teardown failed. Conversely, GitLab may show stopped while the cloud resource still exists.

Evidence Question
Environment state + updated time Did GitLab transition to stopped?
Stop job ID/status/log Did the cleanup job actually execute?
on_stop/auto_stop_in/rules Was cleanup wired and eligible?
Resource group Could deploy/stop overlap or UI stop requirements be violated?
External inventory/tag/target ID Does the real resource still exist?

Repair: make the stop job idempotent, align rules, use a bounded TTL, and independently sweep stale external resources with project/environment ownership tags.

6. Failure: protected-environment intent is bypassed by another path

A protected environment governs GitLab deployment authorization, but a long-lived cloud token copied into another job, an over-privileged shared runner, or an external automation system can still call the target directly. The symptom is “production changed without an authorized GitLab deployment.”

Investigate target audit logs/identity first, then GitLab pipelines/runners/variables. Do not respond by only tightening the protected environment. Revoke/rotate exposed real credentials, remove broad target permissions, prefer short-lived identity, and ensure deploy capability exists only in the intended trusted path.

7. Failure: merge approval is mistaken for deployment approval

An MR has two approvals, so an operator assumes production deployment is approved. That inference is invalid. MR approval state belongs to the change-integration workflow. Deployment approval is a separate Premium/Ultimate protected-environment state per deployment.

Current deployment-approval behavior has another operational nuance: required approvals block deployments, and approval itself does not automatically run the job. After approvals are satisfied, the deployment job must still be executed under the configured workflow. The pipeline triggerer cannot self-approve by default unless that option is enabled.

8. Failure: environment-scoped variable is missing

A job expects PROD_ENDPOINT, but its environment name is production/eu while the variable scope is exactly production. The job sees no variable.

Preserve the environment name and variable metadata (not the secret value). Compare exact/wildcard matching. Repair the scope only if the job truly belongs to that target; do not broaden to * simply to make the pipeline green.

9. Failure: stale deployment job after a newer deployment

GitLab deployment safety can prevent outdated deployment jobs. An older manual deployment may be disabled after a newer deployment has run. Job age is based on start time, which can make human intuition about commit age misleading.

Inspect environment history and deployment/job timestamps before retrying. A rollback has a specific supported path; bypassing stale-job protection is not the same as rolling back. Keep Chapter 18’s manual/delayed/freeze controls separate from this environment-state diagnosis.

10. Failure: dynamic URL metadata is stale

A deployment script discovers a real ephemeral URL at runtime, but the environment URL was hard-coded or never updated. GitLab’s link points to the wrong place even though the external deployment is healthy.

Use a supported dynamic environment URL pattern (for example a dotenv report feeding a later environment update) when the target generates its address. Sanitize URLs and never put secret query parameters/tokens in environment URLs because those become visible metadata.

11. Diagnose permission failures without granting Maintainer blindly

If a deployment returns a 403/blocked state, determine which control rejected it:

  • Project role insufficient to run the job?
  • Protected environment disallows the actor?
  • Required deployment approvals missing?
  • Protected branch/tag blocks the pipeline path?
  • Environment-scoped variable absent?
  • Runner protected/tag/scope mismatch?
  • Target IAM rejects the job identity?

Granting Maintainer or unprotecting the environment without this decomposition destroys the evidence and can create a larger incident.

12. Security-sensitive actions

High-impact changes. Treat protected-environment edits, approval bypass, environment variable/credential changes, deployment-job privilege changes, runner registration, force-push/history rewrite, environment deletion, and external resource teardown as security-sensitive. Use disposable resources in this chapter and record pre-change state plus rollback.

If a real credential appears in logs, environment URLs, artifacts, or job output, response begins with revoke/rotate, then containment/cleanup and root-cause correction. “Mask the log” is not sufficient after disclosure.

13. Cost and performance only where causally relevant

Review Apps can multiply runner work and external resources per MR. Short TTLs reduce waste but can increase redeploy latency. Serializing one shared staging environment reduces race risk but increases queue time. Measure the actual critical path, resource fan-out, and stale-environment count before adding more runner capacity or removing serialization.

14. API evidence packet

PROJECT_ID="12345678"
ENV_ID="123456"

glab api "projects/$PROJECT_ID/environments/$ENV_ID" \
  --jq '{id,name,state,external_url,tier,updated_at}'

glab api "projects/$PROJECT_ID/deployments?environment=review/ch17-broken&order_by=finished_at&sort=desc" \
  --jq '.[] | {id,status,sha,ref,deployable_job_id:.deployable.id,created_at,finished_at}'

Pair this with pipeline/job metadata and the artifact checksum. That is enough to diagnose most Chapter 17 failures without reading secrets.

15. Verification after repair

  • GitLab environment name/tier/URL corresponds to the intended target mapping.
  • Newest deployment points to expected job, pipeline, and commit SHA.
  • Deploy job verified the intended artifact digest/identity.
  • External target independently reports the same version where applicable.
  • Stop/cleanup state is consistent in GitLab and external inventory.
  • Protection/approval/variable/runner policies remain least privilege.
  • The original failed deployment remains preserved as evidence.

Knowledge check

GitLab records a successful production deployment, but deploy.sh targeted testing. What failed?

A rollback commit matches the original deployment but the artifact digest differs. Is it an identity-preserving rollback?

A Review App environment is stopped but the cloud resource still exists. What should you inspect?

A production deployment changed without a GitLab deployment record. What is the likely class of problem?

Why not broaden a missing environment-scoped variable to * immediately?

Does approving an MR approve a protected-environment deployment?

Summary

Deployment incidents become tractable when you preserve three identities independently: GitLab environment/deployment metadata, immutable artifact/commit identity, and real external target identity. Most serious failures are mismatches between those planes or trust paths that bypass intended governance. Repair the narrowest responsible layer and re-prove the full chain.

Official references

Next lesson

Checkpoint the complete deployment-state model

Lesson 5 predicts environment transitions, records a synthetic deployment, preserves a deliberate failure, recovers to a known artifact, verifies state, and removes every disposable resource.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.