CI/CD YAML, Scripts, Images, Services, before_script, after_script, and Defaults: Diagnostics, Failure Modes, Security, and Performance
Diagnose YAML-versus-shell parsing surprises, moving images, service-readiness failures, after_script misconceptions, inherited defaults, and executor mismatches without hiding the original cause.
Learning objectives
- Distinguish generic YAML parsing errors from GitLab CI schema errors and shell interpretation failures.
- Diagnose a moving image tag by comparing commit/configuration identity with runtime image evidence.
- Diagnose service startup/readiness separately from DNS/alias and application-level failures.
- Explain why after_script failure can be invisible in the final job status and how to place required validation correctly.
- Trace unexpected behavior back through default/include inheritance and executor capability before changing permissions or runners.
default,
image, services, before_script,
script, and after_script—are part of core
GitLab CI/CD and are available on Free, Premium, and Ultimate across
GitLab.com, Self-Managed, and Dedicated. Runtime behavior still
depends on the runner executor. In particular, container
image/services semantics require a
compatible container-capable executor; a Shell executor runs commands
directly on the runner host and has materially different isolation and
dependency assumptions. Hosted-compute quotas and registry
availability can change, so all mandatory learning also has a CI
Lint/log-fixture or local-container fallback.
1. Diagnostic sequence for job-runtime failures
Use the same evidence-first discipline from earlier chapters:
- Preserve evidence: pipeline ID, job ID, commit SHA, runner/executor, log, image/service references, relevant merged YAML.
- Locate the layer: YAML parse → GitLab CI schema → configuration resolution → runner eligibility → environment/image pull → service startup/readiness → shell/lifecycle command → artifact/cache upload.
- Inspect the owning state: do not edit runner tags to fix a shell quote or add privileges to fix service readiness.
- Apply the least destructive correction.
- Rerun on a new commit/pipeline and compare evidence.
2. Broken example: a colon is data to the shell but syntax to YAML
The following looks like one shell command to a human, but YAML can interpret the colon-space as mapping syntax unless the whole command is quoted correctly:
broken_yaml:
script:
- curl --header "Content-Type: application/json" https://example.invalid/api
fixed_yaml:
script:
- 'curl --header "Content-Type: application/json" https://example.invalid/api'
| Failure signature | Likely layer |
|---|---|
| CI Lint/parser rejects document | YAML syntax or GitLab CI schema. |
| Job starts, shell reports command/quote error | Shell syntax/runtime. |
| Command executes but remote returns HTTP error | Application/network/auth layer. |
3. Image tag moved: same YAML, different runtime
Suppose yesterday and today both ran
image: toolchain:stable at different pipeline SHAs and
the upstream tag was republished. A new package version or
entrypoint can change behavior without a line changing in your CI
YAML.
Diagnosis requires two identities: configuration commit SHA and resolved image identity. If the repository config is unchanged but the pulled digest differs, the runtime dependency changed. Repair by selecting a reviewed immutable digest or controlled promotion tag, not by adding arbitrary retries to the script.
4. Service exists but tests race its readiness
A job can resolve the service alias and still fail because the application has not finished startup. Distinguish:
| Check | Failure meaning |
|---|---|
| Alias does not resolve | Service alias/network/executor configuration problem. |
| TCP connection refused | Container/process may not be ready or listening on expected port. |
| TCP accepts but health endpoint fails | Application is running but not ready/healthy. |
| Authentication/application error | Service is reachable; credentials/schema/protocol is the next layer. |
wait_for_service:
script:
- |
max=20
n=1
while [ "$n" -le "$max" ]; do
if wget -qO- http://web/health >/dev/null 2>&1; then
echo "ready on attempt $n"
break
fi
if [ "$n" -eq "$max" ]; then
echo "service readiness timeout" >&2
exit 1
fi
n=$((n + 1))
sleep 1
done
5. after_script failed but the job is green
This is not necessarily a GitLab bug. Current semantics give
after_script a separate shell and timeout, and its
failure does not change the job exit code when the main
script succeeded. If the after hook produces a
mandatory report, signature, or validation decision, move the
correctness gate into script. Keep post-processing in
after_script only when its failure policy is acceptable
and observable.
after_script as the only revocation mechanism for
long-lived credentials. Prefer short-lived job identity and systems
that expire/revoke credentials independently; Chapter 13 covers
secret and token lifecycle.
6. Job behavior changed after a template/default update
A local job file may be unchanged while an included template changes
default:image, tags, hooks, or another inherited
keyword. Preserve the pipeline SHA and inspect the resolved/merged
configuration associated with that pipeline. Compare the referenced
include source/ref with the previous working run.
The least destructive fix may be to pin the include, update the shared template deliberately, or add a narrowly justified job override. Copying the whole shared template into the repository is usually a poor emergency fix because it creates permanent drift.
7. image/services configured but the runner cannot implement them
If a job lands on a Shell executor, container-image assumptions are the wrong layer. The script executes on the runner host with host-installed dependencies and limited isolation. Do not “repair” this by installing arbitrary production software on the runner host during incident response. Either route the job to a compatible reviewed runner or redesign the job for the intended executor.
| Symptom | Inspect | Least destructive correction |
|---|---|---|
| Image tools absent | Runner/executor and job image metadata. | Use compatible container runner or explicit host dependency policy. |
| Service alias unavailable | Executor capability and service config. | Use compatible executor or local/external test dependency with explicit contract. |
| Different shell syntax | Runner OS/shell. | Use correct shell syntax or constrain runner platform. |
8. Intentionally broken lab: readiness assumption
Use the disposable service job from Lesson 2 but replace its bounded loop with a single immediate request. Preserve the failed job log. Then restore the loop and compare the repaired pipeline.
# Intentionally broken for a disposable lab only:
script:
- wget -qO- http://web/health
# Repair: bounded readiness check with explicit timeout and evidence.
allow_failure or an
unconditional || true.
Those change status semantics without fixing the race. Preserve the
original failed pipeline/job first, then repair readiness.
9. Security and performance only where causal
- Image trust: mutable/unreviewed images create supply-chain risk; pin and control provenance where exact identity matters.
- Executor isolation: Shell runners expose more host state; do not run untrusted code on sensitive shared hosts.
- Service secrets: use synthetic data in labs; do not bake credentials into service command lines or public YAML.
- Performance: large images and unnecessary services increase pull/startup time and hosted compute consumption.
-
Reliability: external package installs in every
before_scriptadd network dependency and upstream volatility.
Knowledge check
CI Lint succeeds but the job shell reports syntax error. Which layer owns the failure?
The runtime shell/command layer, not YAML parsing or pipeline creation.
Same YAML tag, different resolved digest: what changed?
The image registry reference resolved to different content; the runtime dependency changed even if repository configuration did not.
A service hostname resolves but the application returns “not ready.” Is this primarily DNS?
No. Network/alias resolution succeeded; diagnose application readiness/startup next.
Why is || true a poor repair for a readiness race?
It suppresses evidence and status without establishing that the required service became ready.
Why inspect merged configuration after an included template changes?
Because inherited defaults may originate outside the local file and change the job without an obvious local diff.
A Shell executor lacks tools expected from image:. What should you inspect first?
The runner/executor boundary; image-based runtime semantics are not universally implemented by every executor.
Summary
Runtime failures become diagnosable when you separate YAML parsing, GitLab schema/resolution, executor capability, image identity, service networking/readiness, shell semantics, and lifecycle status behavior. Preserve the failing job, repair the owning layer, and rerun with comparable SHA/image/runner evidence rather than weakening policy.
Official references
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — Scripts and job logs
- GitLab Docs — Deprecated CI/CD keywords
- GitLab Docs — Use CI/CD configuration from other files
- GitLab Docs — Run CI/CD jobs in Docker containers
- GitLab Docs — Services
- GitLab Docs — Docker executor
- GitLab Docs — Shell executor
- GitLab Docs — Runner executors
- GitLab Docs — Pipeline editor
- GitLab Docs — CI Lint
- GitLab Docs — Predefined CI/CD variables
- GitLab Docs — Runner security
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.