Chapter 11Lesson 04~225 minutes

CI/CD YAML, Scripts, Images, Services, before_script, after_script, and Defaults: Diagnostics, Failure Modes, Security, and Performance

Diagnose YAML-versus-shell parsing surprises, moving images, service-readiness failures, after_script misconceptions, inherited defaults, and executor mismatches without hiding the original cause.

DiagnosticsYAML quotingafter_scriptImage driftReadinessInheritance

Learning objectives

  • Distinguish generic YAML parsing errors from GitLab CI schema errors and shell interpretation failures.
  • Diagnose a moving image tag by comparing commit/configuration identity with runtime image evidence.
  • Diagnose service startup/readiness separately from DNS/alias and application-level failures.
  • Explain why after_script failure can be invisible in the final job status and how to place required validation correctly.
  • Trace unexpected behavior back through default/include inheritance and executor capability before changing permissions or runners.
Availability baseline (verified 2026-08-21). The CI/CD YAML keywords taught here—default, image, services, before_script, script, and after_script—are part of core GitLab CI/CD and are available on Free, Premium, and Ultimate across GitLab.com, Self-Managed, and Dedicated. Runtime behavior still depends on the runner executor. In particular, container image/services semantics require a compatible container-capable executor; a Shell executor runs commands directly on the runner host and has materially different isolation and dependency assumptions. Hosted-compute quotas and registry availability can change, so all mandatory learning also has a CI Lint/log-fixture or local-container fallback.

1. Diagnostic sequence for job-runtime failures

Use the same evidence-first discipline from earlier chapters:

  1. Preserve evidence: pipeline ID, job ID, commit SHA, runner/executor, log, image/service references, relevant merged YAML.
  2. Locate the layer: YAML parse → GitLab CI schema → configuration resolution → runner eligibility → environment/image pull → service startup/readiness → shell/lifecycle command → artifact/cache upload.
  3. Inspect the owning state: do not edit runner tags to fix a shell quote or add privileges to fix service readiness.
  4. Apply the least destructive correction.
  5. Rerun on a new commit/pipeline and compare evidence.

2. Broken example: a colon is data to the shell but syntax to YAML

The following looks like one shell command to a human, but YAML can interpret the colon-space as mapping syntax unless the whole command is quoted correctly:

broken_yaml:
  script:
    - curl --header "Content-Type: application/json" https://example.invalid/api
Repair: quote the complete YAML scalar, then let the shell parse the inner quotes. Validate with CI Lint before pushing.
fixed_yaml:
  script:
    - 'curl --header "Content-Type: application/json" https://example.invalid/api' 
Failure signature Likely layer
CI Lint/parser rejects document YAML syntax or GitLab CI schema.
Job starts, shell reports command/quote error Shell syntax/runtime.
Command executes but remote returns HTTP error Application/network/auth layer.

3. Image tag moved: same YAML, different runtime

Suppose yesterday and today both ran image: toolchain:stable at different pipeline SHAs and the upstream tag was republished. A new package version or entrypoint can change behavior without a line changing in your CI YAML.

Diagnosis requires two identities: configuration commit SHA and resolved image identity. If the repository config is unchanged but the pulled digest differs, the runtime dependency changed. Repair by selecting a reviewed immutable digest or controlled promotion tag, not by adding arbitrary retries to the script.

4. Service exists but tests race its readiness

A job can resolve the service alias and still fail because the application has not finished startup. Distinguish:

Check Failure meaning
Alias does not resolve Service alias/network/executor configuration problem.
TCP connection refused Container/process may not be ready or listening on expected port.
TCP accepts but health endpoint fails Application is running but not ready/healthy.
Authentication/application error Service is reachable; credentials/schema/protocol is the next layer.
wait_for_service:
  script:
    - |
      max=20
      n=1
      while [ "$n" -le "$max" ]; do
        if wget -qO- http://web/health >/dev/null 2>&1; then
          echo "ready on attempt $n"
          break
        fi
        if [ "$n" -eq "$max" ]; then
          echo "service readiness timeout" >&2
          exit 1
        fi
        n=$((n + 1))
        sleep 1
      done

5. after_script failed but the job is green

This is not necessarily a GitLab bug. Current semantics give after_script a separate shell and timeout, and its failure does not change the job exit code when the main script succeeded. If the after hook produces a mandatory report, signature, or validation decision, move the correctness gate into script. Keep post-processing in after_script only when its failure policy is acceptable and observable.

Security-sensitive cleanup: do not rely on after_script as the only revocation mechanism for long-lived credentials. Prefer short-lived job identity and systems that expire/revoke credentials independently; Chapter 13 covers secret and token lifecycle.

6. Job behavior changed after a template/default update

A local job file may be unchanged while an included template changes default:image, tags, hooks, or another inherited keyword. Preserve the pipeline SHA and inspect the resolved/merged configuration associated with that pipeline. Compare the referenced include source/ref with the previous working run.

The least destructive fix may be to pin the include, update the shared template deliberately, or add a narrowly justified job override. Copying the whole shared template into the repository is usually a poor emergency fix because it creates permanent drift.

7. image/services configured but the runner cannot implement them

If a job lands on a Shell executor, container-image assumptions are the wrong layer. The script executes on the runner host with host-installed dependencies and limited isolation. Do not “repair” this by installing arbitrary production software on the runner host during incident response. Either route the job to a compatible reviewed runner or redesign the job for the intended executor.

Symptom Inspect Least destructive correction
Image tools absent Runner/executor and job image metadata. Use compatible container runner or explicit host dependency policy.
Service alias unavailable Executor capability and service config. Use compatible executor or local/external test dependency with explicit contract.
Different shell syntax Runner OS/shell. Use correct shell syntax or constrain runner platform.

8. Intentionally broken lab: readiness assumption

Use the disposable service job from Lesson 2 but replace its bounded loop with a single immediate request. Preserve the failed job log. Then restore the loop and compare the repaired pipeline.

# Intentionally broken for a disposable lab only:
script:
  - wget -qO- http://web/health

# Repair: bounded readiness check with explicit timeout and evidence.
Do not hide the failure with allow_failure or an unconditional || true. Those change status semantics without fixing the race. Preserve the original failed pipeline/job first, then repair readiness.

9. Security and performance only where causal

  • Image trust: mutable/unreviewed images create supply-chain risk; pin and control provenance where exact identity matters.
  • Executor isolation: Shell runners expose more host state; do not run untrusted code on sensitive shared hosts.
  • Service secrets: use synthetic data in labs; do not bake credentials into service command lines or public YAML.
  • Performance: large images and unnecessary services increase pull/startup time and hosted compute consumption.
  • Reliability: external package installs in every before_script add network dependency and upstream volatility.

Knowledge check

CI Lint succeeds but the job shell reports syntax error. Which layer owns the failure?

Same YAML tag, different resolved digest: what changed?

A service hostname resolves but the application returns “not ready.” Is this primarily DNS?

Why is || true a poor repair for a readiness race?

Why inspect merged configuration after an included template changes?

A Shell executor lacks tools expected from image:. What should you inspect first?

Summary

Runtime failures become diagnosable when you separate YAML parsing, GitLab schema/resolution, executor capability, image identity, service networking/readiness, shell semantics, and lifecycle status behavior. Preserve the failing job, repair the owning layer, and rerun with comparable SHA/image/runner evidence rather than weakening policy.

Official references

Next lesson

Checkpoint: predict the exact runtime, break it safely, repair it

Lesson 5 integrates defaults, lifecycle hooks, image/service dependencies, log ordering, failure evidence, pinning/recording, and cleanup into one reproducible two-job exercise.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.