Chapter 03Lesson 04~120 minutes

Jobs, Stages, Scripts, Images, Services, before_script, after_script, and Exit Behavior: Diagnostics, Failure Modes, Security, and Performance

A red job is not a diagnosis. This lesson separates shell-status mistakes, failed pipelines whose earlier error was accidentally hidden, service-readiness races, incompatible image entrypoints/shells, and incorrect beliefs about after_script. Each failure is traced from compiled job through runner/executor evidence to the first causal command.

DiagnosticsPipe failuresService readinessEntrypointsFailure evidence

Learning objectives

  • Diagnose a swallowed non-zero exit, an unsafe pipeline/pipe assumption, a service-readiness race, and an incompatible container entrypoint or missing shell.
  • Explain why after_script is a separate shell and why its failure cannot repair or redefine the main script result.
  • Preserve exact pipeline/job IDs, SHA, runner/executor/image/service metadata, and first-failure logs before retrying.
  • Separate configuration, queue/runner, executor/image, shell/tool, service/network, and artifact/report failure layers.
  • Repair the smallest causal defect instead of adding broad retries, privileged execution, or error suppression.
Diagnostic discipline: capture pipeline/job IDs, exact SHA, runner/executor/image/service identity, and the first failing log before changing configuration or rerunning. A rerun is a new experiment with potentially different infrastructure or dependency state.

1. Evidence-first diagnostic sequence

  1. Confirm pipeline source/ref/SHA and compiled job definition.
  2. Confirm job existed and whether it is created, pending, running, failed, canceled, or allowed-to-fail.
  3. Confirm runner/executor/image/service identity.
  4. Find the first unexpected log line or non-zero command, not the last red summary.
  5. Inspect service/network readiness only if the job reached that layer.
  6. Inspect artifacts/reports only after the producer phase ran.
  7. Apply the smallest causal correction and rerun only the needed scope.
Never start by adding retries. Retries can hide deterministic configuration errors, repeat side effects, and erase a useful first-failure condition.

2. Failure mode: swallowed exit codes

GitLab documents a specific multiline-command hazard: when multiple commands are combined into one command string, failures from earlier commands can be ignored and only the final command result may be reported. Shell constructs such as || true can also intentionally suppress failure.

bad_example:
  script:
    - sh -c 'false; echo "last command succeeds"'

better_example:
  script:
    - false
    - echo "this should not run"

The repair is not “add more echo.” Keep failure-producing operations as separate script items when practical, or explicitly capture and test status inside one compound command.

3. Failure mode: a pipeline reports the last command, not the command you cared about

POSIX shell pipeline status normally follows the last command. This can make a failed producer look successful when its output is piped into a successful consumer. Bash offers pipefail; plain sh may not.

# Demonstrates why shell contract matters:
pipe_probe:
  image: alpine:3.22.1
  script:
    - sh -c 'false | cat; printf "pipeline_rc=%s\n" "$?"'

bash_pipe_probe:
  image: bash:5.2.37-alpine3.22
  script:
    - set -o pipefail
    - false | cat

Do not paste set -o pipefail into an environment that only guarantees POSIX sh. Choose the interpreter explicitly and test the semantics you depend on.

4. Failure mode: service health is not application readiness

Runner starts services before before_script and performs a port-based health check. A service can still require schema initialization, leader election, cache warm-up, or application bootstrap after a port opens.

before_script:
  - |
    i=0
    until wget -qO- http://web/health >/dev/null 2>&1; do
      i=$((i + 1))
      [ "$i" -lt 30 ] || { echo "web never became semantically ready"; exit 1; }
      sleep 1
    done

Bound readiness loops. An unbounded loop converts a clear dependency failure into a timeout. Preserve service warnings and the application-level readiness output together.

5. Failure mode: incompatible image entrypoint or missing shell

A Docker job image must be compatible with Runner’s script-delivery model. Typical evidence is a job that fails before normal script output, complains that a shell cannot be found, or starts the image’s application entrypoint instead of executing CI commands.

Evidence Likely cause Repair direction
sh: not found / no shell Minimal/distroless image lacks supported shell. Use a CI-capable image or separate runtime/build images.
Container exits before script Entrypoint does not accept shell command. Verify upstream image contract; use reviewed entrypoint override if justified.
Tool missing after image change Image contents differ from assumption. Record/pin image + verify tool version explicitly.
Security note: do not switch to privileged mode to solve an entrypoint or missing-tool problem. Privilege changes the trust boundary and can expose the runner host.

6. Failure mode: treating after_script as a status repair mechanism

broken_build:
  script:
    - printf 'starting
'
    - false
  after_script:
    - printf 'cleanup-completed
' > cleanup.txt
    - exit 0
  artifacts:
    when: always
    paths: [cleanup.txt]

The job remains failed because the main script failed. The cleanup artifact is valuable evidence that post-processing ran; it is not a green override. Conversely, if the main script succeeds and after_script fails, GitLab documents that the main job exit result remains successful by default.

7. after_script cancellation and timeout boundaries

Current GitLab behavior runs after_script when a job is canceled during before_script or script, and CI_JOB_STATUS reports canceled during that phase. after_script has its own default timeout (five minutes in current Runner documentation). For job timeouts, after_script does not run by default unless Runner script/after-script timeout settings leave time within the job limit.

after_script:
  - |
    if [ "$CI_JOB_STATUS" = "canceled" ]; then
      echo "job canceled; skip nonessential post-processing"
      exit 0
    fi
  - ./ci/generate-summary.sh

Keep post-processing bounded and non-destructive. A cleanup path that requires unlimited time is a reliability risk.

8. Failure-layer matrix

Symptom Evidence first Do not do first
No job in pipeline Compiled config/rules/source. Debug runner host.
Pending job Runner eligibility/tags/capacity. Rewrite script.
Job starts then command not found Image/tool/shell identity. Add retries.
Service DNS failure Service alias/network and Runner service logs. Disable TLS or use host networking.
Service port reachable but test fails Application readiness/semantic health. Assume Runner health check is enough.
Cleanup artifact exists but job red First main-script failure. Treat cleanup as success proof.

9. Performance without hiding correctness

Job performance is a product of queue time, image pulls, source/artifact transfer, setup, service startup, main commands, and post-processing. Measure before optimizing. A huge “all tools included” image can reduce installation time while increasing pull time; a service per job improves isolation while increasing startup cost.

Queue time

Runner fleet/capacity; not script optimization.

Image pull

Image size, registry locality, pull policy/cache.

Setup time

Runtime installs versus prebuilt image.

Service startup

Image size + readiness time.

Main execution

Actual build/test workload.

Post-processing

after_script + artifact/cache upload.

10. Intentionally broken example: preserve, classify, repair

Run this only in a disposable context:

diagnostic_probe:
  image: alpine:3.22.1
  services:
    - name: nginx:1.28.0-alpine
      alias: wrong-name
  script:
    - printf 'sha=%s job=%s
' "$CI_COMMIT_SHA" "$CI_JOB_ID"
    - wget -qO- http://web/
  after_script:
    - printf 'after_status=%s
' "$CI_JOB_STATUS" > diagnostic-after.txt
  artifacts:
    when: always
    paths: [diagnostic-after.txt]

Expected: the service exists under wrong-name, while the script asks for web. Preserve the job ID, Runner/service log warning or DNS failure, and artifact. Repair only the alias to web; do not add || true, privileged networking, or a blind retry.

Next lesson

Checkpoint Lab

Build the final controlled two-stage pipeline, inject one shell defect and one service defect, and produce a reviewable evidence packet for both failures and the repaired state.

Knowledge check

Why can sh -c "false; echo ok" be dangerous in CI?

What is wrong with using Runner’s service health check as the only database-readiness signal?

Why is privileged mode not a valid fix for a missing shell inside the job image?

Can after_script: exit 0 convert a failed main script into a successful job?

Which log should be preserved before retrying a flaky-looking job?

Official references and version notes

  • CI/CD YAML syntax reference — current stages, stage, needs, image, services, before_script, after_script, allow_failure, and timeout semantics.
  • Scripts and job logs — non-zero exit behavior, multiline-command caveats, default hooks, and cancellation behavior.
  • Job execution flow — Runner source preparation, cache/artifact transfer, main execution, after_script, upload, and cleanup phases.
  • Docker executor — job images, service containers, runner workflow, shell requirements, entrypoints, and privilege risks.
  • Run jobs in Docker containers — image/service syntax, entrypoint handling, and where scripts execute.
  • Services — service aliases, networking, health checks, startup warnings, and service lifecycle.
  • Runner executors and supported shells — executor isolation and shell portability boundaries.
Version and compatibility note

Version-sensitive statements were rechecked against current primary GitLab documentation on 2026-09-11. Core jobs, stages, scripts, hooks, images, services, needs, and allow_failure are documented across GitLab Free, Premium, and Ultimate. Docker image/service examples require a Docker-capable runner or the documented local Docker simulation; the mandatory learning path does not require registering or weakening a production runner.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.