Chapter 03Lesson 05~150 minutes

Checkpoint Lab — Jobs, Stages, Scripts, Images, Services, before_script, after_script, and Exit Behavior

The checkpoint combines the chapter into one controlled build/test pipeline. You will use a disposable service dependency, capture job/runner/image/service evidence, inject one shell failure and one service failure, preserve first-failure logs, repair the causal defects, and show that cleanup hooks and green status never substitute for independent evidence.

Checkpoint labBuild/testService failureShell failureRecovery

Learning objectives

  • Build a disposable two-stage pipeline with a synthetic build artifact and an HTTP service dependency.
  • Predict stage, job, runner/image/service, hook, and exit-status changes before execution.
  • Inject one shell failure and one service-resolution/readiness failure while preserving the original evidence for each.
  • Restore a verified pipeline and prove the repair with exact source/job/runner/service evidence.
  • Produce a compact evidence packet and state clearly what the green pipeline proves and what it does not prove.
Checkpoint contract. Use only a disposable project/branch and synthetic data. The GitLab execution path requires an authorized Docker-capable runner; if unavailable, use CI Lint plus the local Docker service simulation and record the limitation. Do not register a production runner or enable privileged mode for this exercise.

1. Goal and state model

The final pipeline has a build stage and a test stage. The build job produces a tiny synthetic artifact. The test job consumes that artifact, starts a disposable NGINX service, performs an application readiness check, validates the build marker, and writes an after_script evidence file. You will then inject one shell failure and one service failure separately.

Source

Exact lab SHA and CI_PIPELINE_SOURCE.

Graph

build_marker → stage barrier → service_test.

Runtime

Runner/executor + Alpine job image + NGINX service alias.

Hooks

Main setup/script followed by separate-shell after_script.

Evidence

Job traces + build marker + after-script artifact + assumptions note.

2. Preflight and predictions

LAB_BRANCH="ch03/checkpoint"
EVIDENCE="ch03-checkpoint-evidence"
mkdir -p "$EVIDENCE"

glab auth status
git status --short --branch
git rev-parse HEAD | tee "$EVIDENCE/head-before.txt"

test ! -e .gitlab-ci.yml || {
  echo "Existing .gitlab-ci.yml detected; use a disposable project/branch."
  exit 2
}

Write predictions before running anything:

  1. build_marker runs before service_test because of stage order.
  2. The jobs do not share a workspace; the build marker crosses the boundary only as an artifact.
  3. The NGINX service is reachable as web, not localhost.
  4. A shell failure leaves the job red even though after_script writes evidence.
  5. A wrong service alias fails at service/network resolution, not at pipeline compilation.

3. Build the known-good pipeline

stages:
  - build
  - test

default:
  image: alpine:3.22.1
  before_script:
    - printf 'pipeline=%s job=%s stage=%s sha=%s
' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_JOB_STAGE" "$CI_COMMIT_SHA"

build_marker:
  stage: build
  script:
    - printf 'artifact_sha=%s
' "$CI_COMMIT_SHA" > build-marker.txt
    - test -s build-marker.txt
  artifacts:
    expire_in: 1 day
    paths:
      - build-marker.txt

service_test:
  stage: test
  services:
    - name: nginx:1.28.0-alpine
      alias: web
  before_script:
    - test -s build-marker.txt
    - |
      i=0
      until wget -qO- http://web/ >/dev/null 2>&1; do
        i=$((i + 1))
        [ "$i" -lt 20 ] || { echo "service readiness failed"; exit 1; }
        sleep 1
      done
  script:
    - grep -F "artifact_sha=" build-marker.txt
    - wget -qO- http://web/ | head -n 1
    - printf 'test_complete
'
  after_script:
    - printf 'after_status=%s job=%s
' "$CI_JOB_STATUS" "$CI_JOB_ID" > after-script.txt
  artifacts:
    when: always
    expire_in: 1 day
    paths:
      - after-script.txt

This chapter uses the default earlier-stage artifact download only as lab glue; Chapter 10 teaches artifact transfer and retention in depth. The important point here is that explicit hosted evidence crosses a job boundary—an incidental runner workspace does not.

4. Validate before execution

glab ci lint .gitlab-ci.yml --include-jobs   | tee "$EVIDENCE/01-good-lint.txt"

glab ci config compile   | tee "$EVIDENCE/02-good-compiled.yml"

Confirm two visible jobs, expected stages, exact image/service configuration, and the service alias before creating a pipeline. If no Docker-capable runner exists, continue with the local simulation in section 9 and treat runtime steps as expected-state analysis.

5. Optional authorized execution and baseline evidence

git switch -c "$LAB_BRANCH"
git add -- .gitlab-ci.yml
git diff --cached --check
git diff --cached
git commit -m "ch03: checkpoint job execution"
LAB_SHA="$(git rev-parse HEAD)"
printf '%s
' "$LAB_SHA" | tee "$EVIDENCE/good-sha.txt"
# Optional authorized side effect:
# git push -u origin "$LAB_BRANCH"

If executed, preserve the baseline pipeline ID; both job IDs; exact SHA; runner identity/executor; image/service names; build artifact content; service_test trace; and after-script.txt. Do not proceed to failure injection until the baseline is understood.

6. Failure injection A — shell failure

In service_test.script, insert this command before the final success line:

    - echo "intentional shell failure follows"
    - sh -c 'exit 23'
    - printf 'must-not-run
' 

Create a new disposable commit/pipeline. Expected evidence: the job reaches execution, exits at code 23, the following main-script command does not run, after_script still writes its status file for a normal script failure, and the job stays failed.

Preserve before repair: failing commit SHA, pipeline/job IDs, first non-zero command, runner/image identity, and the after-script.txt artifact. Do not replace the failure with || true.

7. Repair A — remove the causal shell defect

Remove only the intentional exit command and rerun from a new commit. Verify that the main script reaches test_complete. The repair is proven by changed source + new pipeline/job evidence, not by editing or deleting the old failed trace.

8. Failure injection B — service alias mismatch

Change only the service alias from web to web-broken while the readiness check still requests http://web/:

services:
  - name: nginx:1.28.0-alpine
    alias: web-broken

Expected: configuration remains valid and the job can start, but the application readiness loop fails because the expected hostname is absent. Preserve DNS/network/service evidence. Repair only the alias back to web.

Do not “repair” with host networking, privileged mode, TLS disablement, or unbounded retries. The causal defect is the service identity mismatch.

9. Faithful local service simulation when no Docker runner is available

docker network create ch03-checkpoint-net

docker run -d --rm --name ch03-checkpoint-web   --network ch03-checkpoint-net nginx:1.28.0-alpine

docker run --rm --network ch03-checkpoint-net   alpine:3.22.1 sh -c '
    i=0
    until wget -qO- http://ch03-checkpoint-web/ >/dev/null 2>&1; do
      i=$((i+1))
      [ "$i" -lt 20 ] || exit 1
      sleep 1
    done
    wget -qO- http://ch03-checkpoint-web/ | head -n 1
  '

docker rm -f ch03-checkpoint-web 2>/dev/null || true
docker network rm ch03-checkpoint-net

For the service-failure simulation, deliberately request a nonexistent hostname and capture the failure. This validates the network/alias causal model but still does not claim GitLab Runner or artifact semantics were executed.

10. Evidence packet and acceptance checklist

  • Configuration: lint + compiled configuration for known-good state.
  • Baseline: source SHA, pipeline ID, job IDs, runner/executor/image/service identity, green traces if runtime exists.
  • Shell failure: failing SHA, job trace, exit 23 location, after-script evidence, repaired run.
  • Service failure: failing SHA, alias mismatch, readiness/DNS failure, repaired run.
  • Artifact/workspace boundary: explain why build-marker crosses jobs only through explicit artifact storage.
  • Assumptions: GitLab offering/tier, runner availability/executor, image/service tags, local-simulation use, verification date.
What a green final pipeline proves: for that exact source SHA and runtime context, the configured jobs executed successfully and the recorded service checks passed. It does not prove future runs, unrelated deployment targets, production health, or security/compliance beyond the controls actually evaluated.

11. Cleanup / rollback

docker rm -f ch03-checkpoint-web 2>/dev/null || true
docker network rm ch03-checkpoint-net 2>/dev/null || true

git status --short
git ls-remote --heads origin "$LAB_BRANCH"
# Only if the remote branch is confirmed disposable and authorized:
# git push origin --delete "$LAB_BRANCH"

Do not delete failed pipelines/logs before you have finished comparing them; they are part of the exercise’s first-failure evidence.

12. Chapter checkpoint result

You can now reason from a compiled job into actual execution: stage/DAG eligibility, runner/executor, image/workspace/services, main script hooks, shell exit behavior, post-processing, and retained artifacts. You also practiced the production habit that matters most: preserve the first causal failure, change the smallest responsible layer, and prove recovery with new immutable source/job evidence.

Next lesson

Chapter 04 — GitLab Runner Architecture, Registration, Runner Managers, Job Execution, and Lifecycle

Now that job semantics are visible from the pipeline side, Chapter 04 moves into the runner manager itself: registration/authentication, scope, lifecycle, job request flow, and operational evidence.

Knowledge check

Why does the checkpoint use two different injected failures?

Why is the old failed job trace kept after repair?

Does after-script.txt prove the failed shell command was repaired?

What is the minimum safe repair for the service failure?

What does the final green pipeline not prove?

Official references and version notes

  • CI/CD YAML syntax reference — current stages, stage, needs, image, services, before_script, after_script, allow_failure, and timeout semantics.
  • Scripts and job logs — non-zero exit behavior, multiline-command caveats, default hooks, and cancellation behavior.
  • Job execution flow — Runner source preparation, cache/artifact transfer, main execution, after_script, upload, and cleanup phases.
  • Docker executor — job images, service containers, runner workflow, shell requirements, entrypoints, and privilege risks.
  • Run jobs in Docker containers — image/service syntax, entrypoint handling, and where scripts execute.
  • Services — service aliases, networking, health checks, startup warnings, and service lifecycle.
  • Runner executors and supported shells — executor isolation and shell portability boundaries.
Version and compatibility note

Version-sensitive statements were rechecked against current primary GitLab documentation on 2026-09-11. Core jobs, stages, scripts, hooks, images, services, needs, and allow_failure are documented across GitLab Free, Premium, and Ultimate. Docker image/service examples require a Docker-capable runner or the documented local Docker simulation; the mandatory learning path does not require registering or weakening a production runner.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.