Review Apps, Dynamic Environments, Ephemeral Test Stacks, Route Maps, and Environment Cleanup: Diagnostics, Failure Modes, Security, and Performance
Diagnose orphaned review environments, unsafe resource identifiers, leaked production credentials, unreachable stop jobs, and cost leaks while preserving pipeline/environment/resource evidence before repair.
Learning objectives
- Preserve pipeline/job/environment IDs, source SHA, effective configuration, resource manifest, URL metadata, and external inventory before changing anything.
- Diagnose an orphaned environment by separating GitLab stopped/available state from actual external resource existence.
- Recognize raw branch-name resource identifiers and shared-name collisions as identity defects before they become cleanup hazards.
- Diagnose protected/production secret exposure and an unavailable stop job as trust/lifecycle configuration failures rather than retry problems.
- Repair only the affected review resource and verify cleanup rather than using broad delete commands or project-wide resets.
1. Preserve evidence before cleanup: an orphan is an incident, not an invitation to wipe
When a review environment misbehaves, the fastest-looking action is often “delete it and rerun.” That destroys the evidence needed to learn whether GitLab never invoked the stop job, the provider rejected deletion, the wrong resource ID was computed, or a secret/permission boundary blocked teardown. Preserve first-failure state before repair.
Capture pipeline/job IDs, CI_PIPELINE_SOURCE,
CI_COMMIT_SHA, MR IID, compiled job inclusion,
environment ID/name/slug/state/URL, stop-job status, exact resource
manifest, and bounded external inventory. Then repair the smallest
layer that is wrong.
2. Evidence-first diagnostic sequence
- Preserve: pipeline/job/environment IDs, trace excerpt, manifest, provider operation/error, current external inventory.
- Source: prove MR/ref/SHA and pipeline source.
- Compiled config: verify deploy and stop jobs were included by the intended rules.
- Job graph: ensure stop is runnable and not stranded behind failed stages/needs.
- Runner/tool: prove the stop job reached a runner with the expected image/toolchain.
- Identity: prove non-secret resource ID, credential scope, and target account/namespace.
- Environment: compare GitLab state/URL with external resource existence/health.
- Repair: change only the failing layer; do not recreate unrelated resources.
- Verify: rerun the smallest safe action and independently prove final absence/health.
3. Failure: GitLab says stopped, but the resource is still running
Symptom: environment state is stopped,
yet the provider inventory still lists mr-418. This is
not contradictory: GitLab lifecycle metadata and provider state are
separate.
Evidence: preserve the stop job ID/status, teardown request/operation ID, provider error, and current resource tags. Common causes include insufficient scoped permission, asynchronous delete not completed, wrong provider account/region, or cleanup script exiting successfully after a failed command.
Repair: fix error handling/identity and retry only
teardown for mr-418. Confirm the exact resource is
absent. Do not reopen/redeploy unless content deployment itself also
failed.
4. Failure: raw branch name becomes an unsafe resource selector
A branch such as feature/team/a may be valid Git syntax
but unsuitable as a filesystem path segment, DNS label, Kubernetes
name, database schema, or cloud resource key. A more malicious
string can become a traversal or selector hazard if code
concatenates it without validation.
CI_COMMIT_REF_NAME. The safe repair is to
use a constrained identifier such as
mr-$CI_MERGE_REQUEST_IID, validate it against an
allowlist pattern, compute the target under a dedicated root, and
verify a manifest/ownership label before deletion.
For human traceability, record the original branch/ref separately as metadata; it does not need to be the provider resource key.
5. Failure: review app receives a production secret
Symptom: an MR review job can authenticate to production because a broadly scoped variable or runner-side secret is available. Masking is not a permission boundary; it only reduces accidental log exposure.
Preserve without leaking: record variable key names/scopes and protected/unprotected metadata from settings—not secret values. Record the MR trust context and whether the source is a fork. Do not enable debug trace or print environment variables.
Repair: revoke/rotate if exposure is plausible, remove production authority from review jobs, use isolated review credentials or short-lived identity, and scope variables to the review environment only where appropriate. Chapter 21 will add protected-environment authorization; that does not justify giving ordinary review code production credentials.
6. Failure: stop job cannot run after branch deletion
GitLab can request stop when a branch/MR closes, but teardown still
needs a runnable job. If the stop script exists only on a deleted
branch and the runner cannot fetch it, cleanup fails. Likewise,
incompatible deploy/stop rules, a later blocked stage,
or unmet needs can make the stop job unreachable.
Safer designs include MR pipelines whose ref remains fetchable for
lifecycle work, keeping deploy and stop in a compatible stage/rules
arrangement, or making teardown tooling independent of the feature
branch (for example, a versioned platform component/image) so
GIT_STRATEGY: none can be used deliberately. Verify
current runner/ref behavior in your GitLab version.
7. Failure: auto-stop exists, but costs still grow
auto_stop_in can schedule GitLab environment stop, yet
cost can still leak if the stop job fails, provider resources are
not tied to the environment, deletion is asynchronous, or the
provider creates secondary resources that the manifest omits. The
hourly background cadence also means expiry is approximate.
Track active GitLab environments and provider resources independently. Reconcile project/MR/source labels, age, and ownership. Alert on resources older than policy, failed stop jobs, environments with no open MR, and provider resources with no matching GitLab environment.
8. Intentionally broken example: every MR points at one shared resource
# Broken on purpose: identity collision, not destructive.
review_deploy:
stage: deploy
script:
- printf 'resource_id=review-shared\n'
environment:
name: review/shared
url: https://review-shared.example.invalid/
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
Two MRs can overwrite the same preview, GitLab history collapses into one environment, and cleanup cannot prove which MR owns the resource. The pipeline may be green while reviewers see another MR.
Repair: move the identity decision to
configuration, not the shell:
review/mr-$CI_MERGE_REQUEST_IID,
resource_group: review-mr-$CI_MERGE_REQUEST_IID,
provider ID mr-$CI_MERGE_REQUEST_IID, and manifest
labels including source SHA. Re-run only the affected review
deployment after preserving the collision evidence.
9. Failure: URL points to a stale or reassigned target
An environment URL is metadata. DNS, ingress, provider routing, or a reverse proxy may still point to old content or another resource. Compare the environment URL to DNS/routing state and then verify the target reports the expected source SHA/content digest. Do not “fix” a URL mismatch by rebuilding source unless evidence shows the artifact itself is wrong.
10. Failure layer map
| Symptom | Primary layer | Evidence | Least-destructive correction |
|---|---|---|---|
| Stopped in GitLab, resource exists | Provider/teardown | Stop job + provider inventory/operation. | Retry exact delete after permission/error fix. |
| Two MRs show same content | Identity/config | Compiled env/resource IDs + source SHAs. | Unique MR identity; redeploy affected previews. |
| Production secret available | Identity/trust | Variable scope/protection metadata; MR origin. | Remove/rotate authority; isolate review credentials. |
| Stop job absent | Compilation/rules | Merged config/job inclusion. | Align deploy/stop rules. |
| Stop job blocked | Graph/stage/needs | Pipeline graph/status. | Make stop runnable independently. |
| Costs rise despite auto-stop | Lifecycle/reconciliation | Env ages + provider resource ages. | Add exact owner reconciliation/provider TTL. |
11. Repair verification
- The intended MR/source SHA still maps to the intended environment/resource ID.
- The stop job exists in compiled configuration and is runnable.
- No secret value was printed while diagnosing permissions.
- Deletion targeted one manifest-owned resource, not a broad prefix.
- GitLab environment state and provider inventory are both captured after repair.
- Unrelated review environments remain present and unchanged.
- The first-failure pipeline/job/provider evidence is retained for postmortem or regression tests.
12. Performance without creating an orphan factory
Review speed comes from prebuilt artifacts/images, small isolated resources, dependency caches, parallel build/test work, and fast provider APIs—not from skipping identity/cleanup checks. Measure provisioning time separately from queue time and application readiness. If resource creation is the bottleneck, consider lighter previews or shared immutable dependencies before raising concurrency limits.
High throughput increases cleanup cardinality. Every optimization that creates more simultaneous previews should come with stronger TTL, quota, and reconciliation evidence.
13. Diagnostic summary
Review-app failures are usually identity, trust, lifecycle, or provider-state failures—not reasons for blind retries. Preserve evidence, locate the failed layer, correct the narrowest cause, and verify the exact external state afterward.
Knowledge check
GitLab shows an environment stopped but the cloud resource still exists. What is the first conclusion?
GitLab lifecycle metadata and external provider state diverged; preserve teardown/provider evidence and repair the exact cleanup path.
Why is masking a production secret not enough for review apps?
Masking reduces log disclosure but does not remove authority. Unmerged code should not receive production credentials in the first place.
A stop job is missing after MR close. Which layers do you inspect before retrying?
Compiled rules/job inclusion, stage/needs reachability, repository/ref availability for cleanup tooling, then runner/provider identity.
Why is auto_stop_in not sufficient cost governance?
Stop processing can be delayed or fail, and provider resources may outlive GitLab state; independent TTL/reconciliation is needed for stronger guarantees.
What is wrong with review/shared for every MR?
It destroys unique ownership/provenance and allows MRs to overwrite one another, making both review correctness and cleanup ambiguous.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. Review-app, environment lifecycle, variable, protected-resource, route-map, and cleanup semantics are version-sensitive. Re-check the GitLab version used by your organization before copying exact lifecycle behavior into production.
- Review apps — dynamic review environments, merge-request workflows, stop behavior, and route maps.
-
Environments
— static/dynamic environments, environment states,
on_stop,auto_stop_in, stale cleanup, deletion, and environment-scoped variables. -
CI/CD YAML syntax reference
—
environment:name,url,on_stop,action,auto_stop_in,resource_group, andrules. -
Predefined CI/CD variables
—
CI_COMMIT_REF_SLUG,CI_ENVIRONMENT_NAME,CI_ENVIRONMENT_SLUG,CI_ENVIRONMENT_URL, and merge-request variables. - CI/CD variables — protection, masking/hidden behavior, fork/MR exposure, and environment scope.
- Environments API — environment metadata, stop, and stale-environment operations when API automation is appropriate.
- Deployment safety — protected-resource and deployment-safety boundaries that also matter for review infrastructure.
Current assumptions used in this chapter: review
apps and dynamic environments are available on Free, Premium, and
Ultimate across GitLab.com, Self-Managed, and Dedicated.
CI_COMMIT_REF_SLUG is normalized and shortened to 63
bytes; CI_ENVIRONMENT_SLUG is derived from
environment:name, is truncated to 24 characters, and
uppercase environment names can receive a random suffix. Route maps
live in .gitlab/route-map.yml, are evaluated in
declaration order, and the first matching source rule determines the
public path; the merge-request widget can surface up to five mapped
pages before filtering.
environment:auto_stop_in accepts human-readable
durations (and variables), but environment expiration is serviced by
background work that runs approximately hourly, so it is not an
exact timer. Deploy and stop jobs should have compatible
rules/only/except; a stop job
also has to be runnable when cleanup is needed. To trigger
on_stop from the Environments UI, deploy and stop jobs
must share a resource_group. Protected environments are
Premium/Ultimate and are not required by the mandatory lab. The
mandatory path uses fake .invalid URLs and guarded
local filesystem resources under a temporary sandbox—no cloud
account, Kubernetes cluster, production DNS, PAT, deploy token, or
real secret is required.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.