Chapter 20Lesson 04~180 minutes

Review Apps, Dynamic Environments, Ephemeral Test Stacks, Route Maps, and Environment Cleanup: Diagnostics, Failure Modes, Security, and Performance

Diagnose orphaned review environments, unsafe resource identifiers, leaked production credentials, unreachable stop jobs, and cost leaks while preserving pipeline/environment/resource evidence before repair.

DiagnosticsOrphansUnsafe namesSecretsCost leaks

Learning objectives

  • Preserve pipeline/job/environment IDs, source SHA, effective configuration, resource manifest, URL metadata, and external inventory before changing anything.
  • Diagnose an orphaned environment by separating GitLab stopped/available state from actual external resource existence.
  • Recognize raw branch-name resource identifiers and shared-name collisions as identity defects before they become cleanup hazards.
  • Diagnose protected/production secret exposure and an unavailable stop job as trust/lifecycle configuration failures rather than retry problems.
  • Repair only the affected review resource and verify cleanup rather than using broad delete commands or project-wide resets.

1. Preserve evidence before cleanup: an orphan is an incident, not an invitation to wipe

When a review environment misbehaves, the fastest-looking action is often “delete it and rerun.” That destroys the evidence needed to learn whether GitLab never invoked the stop job, the provider rejected deletion, the wrong resource ID was computed, or a secret/permission boundary blocked teardown. Preserve first-failure state before repair.

Capture pipeline/job IDs, CI_PIPELINE_SOURCE, CI_COMMIT_SHA, MR IID, compiled job inclusion, environment ID/name/slug/state/URL, stop-job status, exact resource manifest, and bounded external inventory. Then repair the smallest layer that is wrong.

2. Evidence-first diagnostic sequence

  1. Preserve: pipeline/job/environment IDs, trace excerpt, manifest, provider operation/error, current external inventory.
  2. Source: prove MR/ref/SHA and pipeline source.
  3. Compiled config: verify deploy and stop jobs were included by the intended rules.
  4. Job graph: ensure stop is runnable and not stranded behind failed stages/needs.
  5. Runner/tool: prove the stop job reached a runner with the expected image/toolchain.
  6. Identity: prove non-secret resource ID, credential scope, and target account/namespace.
  7. Environment: compare GitLab state/URL with external resource existence/health.
  8. Repair: change only the failing layer; do not recreate unrelated resources.
  9. Verify: rerun the smallest safe action and independently prove final absence/health.

3. Failure: GitLab says stopped, but the resource is still running

Symptom: environment state is stopped, yet the provider inventory still lists mr-418. This is not contradictory: GitLab lifecycle metadata and provider state are separate.

Evidence: preserve the stop job ID/status, teardown request/operation ID, provider error, and current resource tags. Common causes include insufficient scoped permission, asynchronous delete not completed, wrong provider account/region, or cleanup script exiting successfully after a failed command.

Repair: fix error handling/identity and retry only teardown for mr-418. Confirm the exact resource is absent. Do not reopen/redeploy unless content deployment itself also failed.

4. Failure: raw branch name becomes an unsafe resource selector

A branch such as feature/team/a may be valid Git syntax but unsuitable as a filesystem path segment, DNS label, Kubernetes name, database schema, or cloud resource key. A more malicious string can become a traversal or selector hazard if code concatenates it without validation.

Do not run: examples that compute destructive paths directly from CI_COMMIT_REF_NAME. The safe repair is to use a constrained identifier such as mr-$CI_MERGE_REQUEST_IID, validate it against an allowlist pattern, compute the target under a dedicated root, and verify a manifest/ownership label before deletion.

For human traceability, record the original branch/ref separately as metadata; it does not need to be the provider resource key.

5. Failure: review app receives a production secret

Symptom: an MR review job can authenticate to production because a broadly scoped variable or runner-side secret is available. Masking is not a permission boundary; it only reduces accidental log exposure.

Preserve without leaking: record variable key names/scopes and protected/unprotected metadata from settings—not secret values. Record the MR trust context and whether the source is a fork. Do not enable debug trace or print environment variables.

Repair: revoke/rotate if exposure is plausible, remove production authority from review jobs, use isolated review credentials or short-lived identity, and scope variables to the review environment only where appropriate. Chapter 21 will add protected-environment authorization; that does not justify giving ordinary review code production credentials.

6. Failure: stop job cannot run after branch deletion

GitLab can request stop when a branch/MR closes, but teardown still needs a runnable job. If the stop script exists only on a deleted branch and the runner cannot fetch it, cleanup fails. Likewise, incompatible deploy/stop rules, a later blocked stage, or unmet needs can make the stop job unreachable.

Safer designs include MR pipelines whose ref remains fetchable for lifecycle work, keeping deploy and stop in a compatible stage/rules arrangement, or making teardown tooling independent of the feature branch (for example, a versioned platform component/image) so GIT_STRATEGY: none can be used deliberately. Verify current runner/ref behavior in your GitLab version.

7. Failure: auto-stop exists, but costs still grow

auto_stop_in can schedule GitLab environment stop, yet cost can still leak if the stop job fails, provider resources are not tied to the environment, deletion is asynchronous, or the provider creates secondary resources that the manifest omits. The hourly background cadence also means expiry is approximate.

Track active GitLab environments and provider resources independently. Reconcile project/MR/source labels, age, and ownership. Alert on resources older than policy, failed stop jobs, environments with no open MR, and provider resources with no matching GitLab environment.

8. Intentionally broken example: every MR points at one shared resource

# Broken on purpose: identity collision, not destructive.
review_deploy:
  stage: deploy
  script:
    - printf 'resource_id=review-shared\n'
  environment:
    name: review/shared
    url: https://review-shared.example.invalid/
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'

Two MRs can overwrite the same preview, GitLab history collapses into one environment, and cleanup cannot prove which MR owns the resource. The pipeline may be green while reviewers see another MR.

Repair: move the identity decision to configuration, not the shell: review/mr-$CI_MERGE_REQUEST_IID, resource_group: review-mr-$CI_MERGE_REQUEST_IID, provider ID mr-$CI_MERGE_REQUEST_IID, and manifest labels including source SHA. Re-run only the affected review deployment after preserving the collision evidence.

9. Failure: URL points to a stale or reassigned target

An environment URL is metadata. DNS, ingress, provider routing, or a reverse proxy may still point to old content or another resource. Compare the environment URL to DNS/routing state and then verify the target reports the expected source SHA/content digest. Do not “fix” a URL mismatch by rebuilding source unless evidence shows the artifact itself is wrong.

10. Failure layer map

Symptom Primary layer Evidence Least-destructive correction
Stopped in GitLab, resource exists Provider/teardown Stop job + provider inventory/operation. Retry exact delete after permission/error fix.
Two MRs show same content Identity/config Compiled env/resource IDs + source SHAs. Unique MR identity; redeploy affected previews.
Production secret available Identity/trust Variable scope/protection metadata; MR origin. Remove/rotate authority; isolate review credentials.
Stop job absent Compilation/rules Merged config/job inclusion. Align deploy/stop rules.
Stop job blocked Graph/stage/needs Pipeline graph/status. Make stop runnable independently.
Costs rise despite auto-stop Lifecycle/reconciliation Env ages + provider resource ages. Add exact owner reconciliation/provider TTL.

11. Repair verification

  • The intended MR/source SHA still maps to the intended environment/resource ID.
  • The stop job exists in compiled configuration and is runnable.
  • No secret value was printed while diagnosing permissions.
  • Deletion targeted one manifest-owned resource, not a broad prefix.
  • GitLab environment state and provider inventory are both captured after repair.
  • Unrelated review environments remain present and unchanged.
  • The first-failure pipeline/job/provider evidence is retained for postmortem or regression tests.

12. Performance without creating an orphan factory

Review speed comes from prebuilt artifacts/images, small isolated resources, dependency caches, parallel build/test work, and fast provider APIs—not from skipping identity/cleanup checks. Measure provisioning time separately from queue time and application readiness. If resource creation is the bottleneck, consider lighter previews or shared immutable dependencies before raising concurrency limits.

High throughput increases cleanup cardinality. Every optimization that creates more simultaneous previews should come with stronger TTL, quota, and reconciliation evidence.

13. Diagnostic summary

Review-app failures are usually identity, trust, lifecycle, or provider-state failures—not reasons for blind retries. Preserve evidence, locate the failed layer, correct the narrowest cause, and verify the exact external state afterward.

Knowledge check

GitLab shows an environment stopped but the cloud resource still exists. What is the first conclusion?

Why is masking a production secret not enough for review apps?

A stop job is missing after MR close. Which layers do you inspect before retrying?

Why is auto_stop_in not sufficient cost governance?

What is wrong with review/shared for every MR?

Next lesson

Checkpoint lab

Provision two reviews, prove unique identity, create one bounded orphan, reconcile only that resource, and produce complete lifecycle evidence.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Review-app, environment lifecycle, variable, protected-resource, route-map, and cleanup semantics are version-sensitive. Re-check the GitLab version used by your organization before copying exact lifecycle behavior into production.

  • Review apps — dynamic review environments, merge-request workflows, stop behavior, and route maps.
  • Environments — static/dynamic environments, environment states, on_stop, auto_stop_in, stale cleanup, deletion, and environment-scoped variables.
  • CI/CD YAML syntax reference — environment:name, url, on_stop, action, auto_stop_in, resource_group, and rules.
  • Predefined CI/CD variables — CI_COMMIT_REF_SLUG, CI_ENVIRONMENT_NAME, CI_ENVIRONMENT_SLUG, CI_ENVIRONMENT_URL, and merge-request variables.
  • CI/CD variables — protection, masking/hidden behavior, fork/MR exposure, and environment scope.
  • Environments API — environment metadata, stop, and stale-environment operations when API automation is appropriate.
  • Deployment safety — protected-resource and deployment-safety boundaries that also matter for review infrastructure.

Current assumptions used in this chapter: review apps and dynamic environments are available on Free, Premium, and Ultimate across GitLab.com, Self-Managed, and Dedicated. CI_COMMIT_REF_SLUG is normalized and shortened to 63 bytes; CI_ENVIRONMENT_SLUG is derived from environment:name, is truncated to 24 characters, and uppercase environment names can receive a random suffix. Route maps live in .gitlab/route-map.yml, are evaluated in declaration order, and the first matching source rule determines the public path; the merge-request widget can surface up to five mapped pages before filtering. environment:auto_stop_in accepts human-readable durations (and variables), but environment expiration is serviced by background work that runs approximately hourly, so it is not an exact timer. Deploy and stop jobs should have compatible rules/only/except; a stop job also has to be runnable when cleanup is needed. To trigger on_stop from the Environments UI, deploy and stop jobs must share a resource_group. Protected environments are Premium/Ultimate and are not required by the mandatory lab. The mandatory path uses fake .invalid URLs and guarded local filesystem resources under a temporary sandbox—no cloud account, Kubernetes cluster, production DNS, PAT, deploy token, or real secret is required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.