Jobs, Stages, Scripts, Images, Services, before_script, after_script, and Exit Behavior: Diagnostics, Failure Modes, Security, and Performance
A red job is not a diagnosis. This lesson separates shell-status
mistakes, failed pipelines whose earlier error was accidentally
hidden, service-readiness races, incompatible image
entrypoints/shells, and incorrect beliefs about
after_script. Each failure is traced from compiled job
through runner/executor evidence to the first causal command.
Learning objectives
- Diagnose a swallowed non-zero exit, an unsafe pipeline/pipe assumption, a service-readiness race, and an incompatible container entrypoint or missing shell.
-
Explain why
after_scriptis a separate shell and why its failure cannot repair or redefine the main script result. - Preserve exact pipeline/job IDs, SHA, runner/executor/image/service metadata, and first-failure logs before retrying.
- Separate configuration, queue/runner, executor/image, shell/tool, service/network, and artifact/report failure layers.
- Repair the smallest causal defect instead of adding broad retries, privileged execution, or error suppression.
1. Evidence-first diagnostic sequence
- Confirm pipeline source/ref/SHA and compiled job definition.
- Confirm job existed and whether it is created, pending, running, failed, canceled, or allowed-to-fail.
- Confirm runner/executor/image/service identity.
- Find the first unexpected log line or non-zero command, not the last red summary.
- Inspect service/network readiness only if the job reached that layer.
- Inspect artifacts/reports only after the producer phase ran.
- Apply the smallest causal correction and rerun only the needed scope.
2. Failure mode: swallowed exit codes
GitLab documents a specific multiline-command hazard: when multiple
commands are combined into one command string, failures from earlier
commands can be ignored and only the final command result may be
reported. Shell constructs such as || true can also
intentionally suppress failure.
bad_example:
script:
- sh -c 'false; echo "last command succeeds"'
better_example:
script:
- false
- echo "this should not run"
The repair is not “add more echo.” Keep failure-producing operations as separate script items when practical, or explicitly capture and test status inside one compound command.
3. Failure mode: a pipeline reports the last command, not the command you cared about
POSIX shell pipeline status normally follows the last command. This
can make a failed producer look successful when its output is piped
into a successful consumer. Bash offers pipefail; plain
sh may not.
# Demonstrates why shell contract matters:
pipe_probe:
image: alpine:3.22.1
script:
- sh -c 'false | cat; printf "pipeline_rc=%s\n" "$?"'
bash_pipe_probe:
image: bash:5.2.37-alpine3.22
script:
- set -o pipefail
- false | cat
Do not paste set -o pipefail into an environment that
only guarantees POSIX sh. Choose the interpreter
explicitly and test the semantics you depend on.
4. Failure mode: service health is not application readiness
Runner starts services before before_script and
performs a port-based health check. A service can still require
schema initialization, leader election, cache warm-up, or
application bootstrap after a port opens.
before_script:
- |
i=0
until wget -qO- http://web/health >/dev/null 2>&1; do
i=$((i + 1))
[ "$i" -lt 30 ] || { echo "web never became semantically ready"; exit 1; }
sleep 1
done
Bound readiness loops. An unbounded loop converts a clear dependency failure into a timeout. Preserve service warnings and the application-level readiness output together.
5. Failure mode: incompatible image entrypoint or missing shell
A Docker job image must be compatible with Runner’s script-delivery model. Typical evidence is a job that fails before normal script output, complains that a shell cannot be found, or starts the image’s application entrypoint instead of executing CI commands.
| Evidence | Likely cause | Repair direction |
|---|---|---|
sh: not found / no shell |
Minimal/distroless image lacks supported shell. | Use a CI-capable image or separate runtime/build images. |
| Container exits before script | Entrypoint does not accept shell command. | Verify upstream image contract; use reviewed entrypoint override if justified. |
| Tool missing after image change | Image contents differ from assumption. | Record/pin image + verify tool version explicitly. |
6. Failure mode: treating after_script as a status repair mechanism
broken_build:
script:
- printf 'starting
'
- false
after_script:
- printf 'cleanup-completed
' > cleanup.txt
- exit 0
artifacts:
when: always
paths: [cleanup.txt]
The job remains failed because the main script failed. The cleanup
artifact is valuable evidence that post-processing ran; it is not a
green override. Conversely, if the main script succeeds and
after_script fails, GitLab documents that the main job
exit result remains successful by default.
7. after_script cancellation and timeout boundaries
Current GitLab behavior runs after_script when a job is
canceled during before_script or script,
and CI_JOB_STATUS reports canceled during
that phase. after_script has its own default timeout
(five minutes in current Runner documentation). For job timeouts,
after_script does not run by default unless Runner
script/after-script timeout settings leave time within the job
limit.
after_script:
- |
if [ "$CI_JOB_STATUS" = "canceled" ]; then
echo "job canceled; skip nonessential post-processing"
exit 0
fi
- ./ci/generate-summary.sh
Keep post-processing bounded and non-destructive. A cleanup path that requires unlimited time is a reliability risk.
8. Failure-layer matrix
| Symptom | Evidence first | Do not do first |
|---|---|---|
| No job in pipeline | Compiled config/rules/source. | Debug runner host. |
| Pending job | Runner eligibility/tags/capacity. | Rewrite script. |
| Job starts then command not found | Image/tool/shell identity. | Add retries. |
| Service DNS failure | Service alias/network and Runner service logs. | Disable TLS or use host networking. |
| Service port reachable but test fails | Application readiness/semantic health. | Assume Runner health check is enough. |
| Cleanup artifact exists but job red | First main-script failure. | Treat cleanup as success proof. |
9. Performance without hiding correctness
Job performance is a product of queue time, image pulls, source/artifact transfer, setup, service startup, main commands, and post-processing. Measure before optimizing. A huge “all tools included” image can reduce installation time while increasing pull time; a service per job improves isolation while increasing startup cost.
Runner fleet/capacity; not script optimization.
Image size, registry locality, pull policy/cache.
Runtime installs versus prebuilt image.
Image size + readiness time.
Actual build/test workload.
after_script + artifact/cache upload.
10. Intentionally broken example: preserve, classify, repair
Run this only in a disposable context:
diagnostic_probe:
image: alpine:3.22.1
services:
- name: nginx:1.28.0-alpine
alias: wrong-name
script:
- printf 'sha=%s job=%s
' "$CI_COMMIT_SHA" "$CI_JOB_ID"
- wget -qO- http://web/
after_script:
- printf 'after_status=%s
' "$CI_JOB_STATUS" > diagnostic-after.txt
artifacts:
when: always
paths: [diagnostic-after.txt]
Expected: the service exists under wrong-name, while
the script asks for web. Preserve the job ID,
Runner/service log warning or DNS failure, and artifact. Repair only
the alias to web; do not add || true,
privileged networking, or a blind retry.
Knowledge check
Why can sh -c "false; echo ok" be dangerous in
CI?
The compound shell command can end with a zero status even though an earlier operation failed, depending on shell/control-flow semantics. Preserve and test the status you actually care about.
What is wrong with using Runner’s service health check as the only database-readiness signal?
The Runner health check is port-oriented and does not prove schema/application-level readiness.
Why is privileged mode not a valid fix for a missing shell inside the job image?
Privilege changes the host trust boundary but does not supply the missing CI-compatible shell contract; use an appropriate image instead.
Can after_script: exit 0 convert a failed main
script into a successful job?
No. The main script result remains failed;
after_script is separate post-processing.
Which log should be preserved before retrying a flaky-looking job?
The first failing job trace plus pipeline/job/SHA/runner/image/service identity, before infrastructure or dependency state changes.
Official references and version notes
-
CI/CD YAML syntax reference
— current
stages,stage,needs,image,services,before_script,after_script,allow_failure, andtimeoutsemantics. - Scripts and job logs — non-zero exit behavior, multiline-command caveats, default hooks, and cancellation behavior.
-
Job execution flow
— Runner source preparation, cache/artifact transfer, main
execution,
after_script, upload, and cleanup phases. - Docker executor — job images, service containers, runner workflow, shell requirements, entrypoints, and privilege risks.
- Run jobs in Docker containers — image/service syntax, entrypoint handling, and where scripts execute.
- Services — service aliases, networking, health checks, startup warnings, and service lifecycle.
- Runner executors and supported shells — executor isolation and shell portability boundaries.
Version-sensitive statements were rechecked against current
primary GitLab documentation on 2026-09-11. Core
jobs, stages, scripts, hooks, images, services,
needs, and allow_failure are documented
across GitLab Free, Premium, and Ultimate. Docker image/service
examples require a Docker-capable runner or the documented local
Docker simulation; the mandatory learning path does not require
registering or weakening a production runner.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.