Chapter 19Lesson 04~170 minutes

Environments, Deployments, Environment Tiers, URLs, Deployment History, and Operational Traceability: Diagnostics, Failure Modes, Security, and Performance

Diagnose green deployment jobs with unhealthy targets, split history from inconsistent names, lost artifact identity, stale URLs, and rollback-by-rebuild mistakes without erasing evidence.

DiagnosticsUnhealthy targetHistory splitStale URLArtifact identity

Learning objectives

  • Preserve original pipeline/job/deployment IDs, digest, environment record, and target read-back before changing anything.
  • Diagnose a green deployment job with an unhealthy target as a deployment/provider/health problem rather than a YAML-success problem.
  • Detect history fragmentation caused by inconsistent environment names and identify stale environment URLs separately from target health.
  • Diagnose rollback-by-rebuild as loss of artifact identity and repair the process by retaining/promoting immutable evidence.
  • Apply the least destructive correction and rerun only the deployment or verification scope that actually failed.

1. Evidence-first deployment diagnostics

Deployment failures are dangerous because “green” can still be wrong. Do not immediately rerun the job or edit the environment. Preserve the first-failure evidence and the state that explains the first outcome.

Evidence-first diagnostic sequence
            flowchart TD
             A[Preserve pipeline/job/deployment IDs + digest + target evidence] --> B[Confirm source/ref/SHA + compiled config]
             B --> C[Confirm rules and environment name/tier/action]
             C --> D[Inspect job graph/runner/script]
             D --> E[Inspect artifact/report/cache identity]
             E --> F[Inspect GitLab environment/deployment record]
             F --> G[Inspect external target health/read-back]
             G --> H[Apply least destructive correction]
             H --> I[Rerun smallest safe scope]
          

2. Failure: deployment job is green but target is unhealthy

Preserve the successful job ID, deployment ID, artifact digest, target resource identity, and health/read-back result. Then ask whether the script verified the provider response or merely submitted an asynchronous request.

A green job means only that the job’s script returned success. The target might fail afterward, point at the wrong backend, reject a delayed rollout, or depend on an unhealthy external service. Add a separate action: verify job or provider-specific health gate. Do not relabel a health check as another deployment.

3. Failure: wrong environment name fragments history

Suppose one pipeline uses training/ch19 and another accidentally uses training/ch-19. GitLab creates two environments because environment names are identity, not display aliases.

# Intended
environment:
  name: training/ch19

# Broken: different identity
environment:
  name: training/ch-19

Preserve both environment IDs and deployment histories. Repair future configuration to the canonical name. Do not delete the unexpected environment until you know whether external resources or scoped variables were associated with it.

4. Failure: artifact identity is lost or rebuilt

An old deployment says source SHA A, but the deploy job runs npm build, mvn package, or another build step again. The resulting bytes are not necessarily the artifact that tests approved.

Repair by splitting build from deployment, recording a digest, retaining/promoting that payload, and making deployment verify the digest before side effects. If the old payload is already gone, record that rollback reproducibility is unavailable instead of pretending a fresh build is identical.

5. Failure: environment URL points to a stale or wrong target

A GitLab URL can remain unchanged while routing changes beneath it—or the URL itself can be stale after migration. Preserve environment ID/name, current URL, provider resource ID, DNS/backend mapping, and actual version/digest observed by the endpoint.

Fix URL metadata and routing at their causal layers. A successful HTTP response from the URL is still insufficient unless the response identifies the expected deployed version/digest.

6. Intentionally broken example: success record with stale target bytes

On a disposable branch, change the deployment script so it records success but deliberately does not copy the new artifact:

deploy:training:broken:
  stage: deploy
  image: alpine:3.22
  needs:
    - job: build:release
      artifacts: true
  script:
    - mkdir -p simulated-target/current evidence
    - echo "pretend provider accepted request" | tee evidence/deployment-receipt.txt
    - test -f simulated-target/current/app.txt || printf 'old-bytes
' > simulated-target/current/app.txt
  environment:
    name: training/ch19
    deployment_tier: testing
    url: https://training.invalid/ch19
  artifacts:
    when: always
    paths: [simulated-target/current/, evidence/]

The job can be green and GitLab can create a successful deployment record. The failure becomes visible only when a verify job compares target content/digest to the build artifact digest.

Repair: restore the exact artifact copy and digest verification. Preserve the broken deployment ID/job trace so the evidence shows why the previous green state was not operationally correct.

7. Failure: rollback rebuilds different bytes

You select an older deployment, GitLab runs the old deployment job, but its artifacts expired. Someone manually rebuilds the old commit and then deploys the new result. The Git commit is old; the payload is new.

Do not call this byte-identical rollback unless the rebuilt digest matches an independently preserved digest. The production fix is retention/promotion policy: keep immutable release payloads for at least the rollback window and verify them at deployment time.

8. Failure: deployment ordering fights Chapter 18 concurrency controls

Environment history can also be wrong when an older pipeline deploys after a newer one. Combine the stable environment identity with Chapter 18’s resource_group serialization and GitLab deployment-safety settings such as preventing outdated deployment jobs where appropriate. Ordering controls reduce races; they still do not verify target health.

9. Causal layer matrix

Symptom Most likely layer Evidence before repair Smallest safe correction
Environment missing Compilation/rules/environment name Merged config, rule result, job inclusion Fix rule/name; rerun deploy only if intended
Job pending Queue/runner/resource group Job status, tags, waiting-for-resource state Fix runner/lock ownership, not deploy script
Job green, target old Script/provider/external target Job log, deployment ID, expected vs observed digest Correct target path/provider and redeploy exact artifact
Two histories for same target Environment identity Environment IDs/names, scoped vars, deployment lists Canonicalize future name; reconcile before deletion
Rollback payload unavailable Artifact/registry retention Old digest, artifact expiry, package/image availability Restore retention/promotion design; do not fabricate evidence
URL healthy but wrong version Routing/target identity URL, response version/digest, provider backend ID Fix routing/deployment target; preserve record

10. Security and disruption boundaries

  • Never print deployment credentials/tokens while diagnosing environment variables.
  • Do not unprotect production environments or broaden runner/token scope to “see if deploy works.”
  • Do not delete environment/deployment history as troubleshooting; it is evidence.
  • Do not retry a deployment with unknown external outcome until the target is reconciled by stable identity.
  • Do not rebuild a release artifact merely to satisfy a rollback button.
  • Use fake targets in training and exact disposable resource guards for any real provider extension.

11. Performance: verify enough, but verify the right thing

Health verification adds latency, but removing it only makes pipelines faster at producing ambiguous state. Prefer small, targeted checks that prove version identity and critical health. Run deeper post-deployment tests asynchronously when they are not required to decide whether promotion succeeded, while retaining their linkage to the deployment ID/digest.

12. Diagnostic rule

Preserve first evidence, locate the failed state transition, and repair that transition only. CI_PIPELINE_SOURCE and CI_COMMIT_SHA identify the pipeline input; artifact digest identifies the payload; environment/deployment IDs identify GitLab’s operational record; target read-back identifies reality.

Knowledge check

A deployment job is green but the service serves the old digest. Which state is wrong?

Why should you preserve an accidentally created environment before deleting it?

What evidence makes rollback-by-rebuild defensible?

Why is environment URL reachability weak evidence?

What should you do before retrying a deployment with unknown external outcome?

Next lesson

Checkpoint lab

Perform two traceable deployments, inject a green-but-wrong target state, preserve history, and prepare an exact rollback dossier.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Environment/deployment keywords, environment actions, deployment-tier variable support, auto-stop behavior, rollback UI behavior, environment-scoped variables, and deployment safety are version-sensitive. Re-check the GitLab and Runner versions used by production delivery before applying the exact examples.

  • Environments — environment identity/state, tiers, URLs, stop/auto-stop behavior, environment-scoped variables, and operational views.
  • Deployments — deployment history, deployment refs, retry/rollback behavior, and auditability.
  • Deployment safety — deployment serialization, outdated-job prevention, protected environments, and rollback considerations.
  • CI/CD YAML syntax reference — environment, environment:name, url, action, auto_stop_in, deployment_tier, and related job semantics.
  • Deployments API — deployment IDs, statuses, SHA/ref, user, environment, and history queries.
  • Environments API — environment metadata/state/tier management when API access is appropriate.
  • CI/CD variables — environment scope behavior and precedence boundaries.

Current assumptions used in this chapter: environments and deployment history are available on Free, Premium, and Ultimate across GitLab.com, Self-Managed, and Dedicated. Environment states include available, stopping, and stopped. environment:action currently supports start (default, creates a deployment after job start), prepare, verify, access, and stop; the latter four model lifecycle/access work rather than ordinary deployment creation. Deployment tiers are production, staging, testing, development, and other; CI/CD-variable support for deployment_tier was added in GitLab 18.5. auto_stop_in uses natural-language durations; stop processing is background work and is not an exact real-time timer. The mandatory labs use synthetic artifacts, a stable training/ch19 environment, alpine:3.22, a reserved .invalid URL, and filesystem/artifact snapshots as a faithful target simulation—no cloud account, protected production environment, or real credential is required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.