Chapter 03Lesson 03~105 minutes

Jobs, Stages, Scripts, Images, Services, before_script, after_script, and Exit Behavior: Configuration, Design Choices, and Tradeoffs

Job syntax is easy to copy and hard to operate when its design assumptions are hidden. This lesson compares stages with DAG edges, global/default hooks with job-local hooks, rich images with explicit setup, service containers with external dependencies, and shell portability with convenience—always tying the choice to runner trust, reproducibility, latency, and evidence.

Design tradeoffsneeds vs stagesImage strategyService boundariesPortability

Learning objectives

  • Choose between stage barriers and needs edges according to real data dependencies and feedback latency.
  • Decide when default hooks reduce duplication and when job-local setup/cleanup is easier to review.
  • Compare a prebuilt tool image with runtime dependency installation using reproducibility, network dependence, image size, and patching as criteria.
  • Choose service containers only for job-local network dependencies and distinguish them from long-lived external systems.
  • Write shell commands that are explicit about interpreter assumptions, pipeline exit behavior, quoting, and expected failure policy.
Design principle: optimize for explicit dependencies and reproducible execution, not the shortest YAML. A clever one-line job that depends on hidden runner state is usually more expensive to operate than a slightly longer job whose image, inputs, failure policy, and evidence are obvious.

1. Stages versus needs: barrier scheduling versus explicit dependency edges

Stages are simple: every successful job in an earlier stage forms a barrier before the next stage. This is easy to read and is a good default for small pipelines. needs expresses specific dependency edges and can start downstream work earlier, reducing critical-path latency.

Choice Strength Tradeoff Use when
Stages only Simple mental model and broad barrier. Independent jobs may wait unnecessarily. Small pipelines where stage-level ordering matches real dependencies.
needs DAG Earlier feedback and explicit job dependencies. More graph reasoning; artifact flow changes with needs. Latency matters and jobs have precise dependencies.

Chapter 09 covers DAG engineering in depth. In Chapter 03, the operational rule is enough: do not add needs merely for speed until you can name the exact data/control dependency it replaces.

2. Default hooks versus job-local hooks

default:
  before_script:
    - ./ci/common-preflight.sh
  after_script:
    - ./ci/common-summary.sh

special_job:
  before_script:
    - ./ci/special-preflight.sh
  script:
    - ./ci/run-special.sh
  after_script: []

Defaults reduce duplication, but job-level hook configuration replaces the default for that keyword rather than merging line by line. A default hook is appropriate when every inheriting job truly shares the same precondition. Job-local hooks are safer when setup has different privilege, network, or toolchain requirements.

Reviewability test: a reviewer opening one job should be able to discover its effective setup and cleanup without hunting through unrelated files or assuming a hidden runner convention.

3. Rich prebuilt image versus runtime installation

Strategy Advantages Risks / cost
Versioned prebuilt image Fast jobs, repeatable toolchain, less network installation during job. Image build/patch ownership; registry trust; image size.
Minimal image + install tools Simple base, flexible for experiments. Network/package-repo dependency, slower runs, version drift unless pinned.
Runner-host tools No image pull; can be fast on shell executor. Hidden host state, weaker isolation, portability and patching burden.

For production, prefer an image whose human-readable version and immutable digest can be traced to a reviewed source. A tag such as alpine:3.22.1 is easier to understand than latest, but tags remain registry references and can be mutable; digest pinning gives stronger identity when operationally practical.

4. Entrypoint design is part of the image contract

Docker executor expects the image to accept the Runner-generated shell script. GitLab documents that the image needs a supported shell in PATH and grep. An image whose ENTRYPOINT insists on starting a server or exits without accepting sh/bash can break job startup.

tool_job:
  image:
    name: example.invalid/tool-image:1.4.2
    entrypoint: [""]
  script:
    - tool --version

The empty entrypoint is a deliberate compatibility override, not a universal fix. First verify the upstream image contract and trust. Do not blindly neutralize security or initialization logic embedded by the image maintainer.

5. Service container versus external dependency

Dependency Best fit Main risk
Job-local database/cache/HTTP server GitLab services on a container-capable executor. Startup/readiness, image trust, resource usage.
Long-lived shared test environment External system with explicit credentials/network policy. Contention, state leakage, cleanup, availability.
Production system Deployment/integration boundary, not a “service” shortcut. Real side effects and blast radius.

A service should be disposable with the job. If multiple pipelines depend on durable mutable shared state, call that an external integration and manage its identity/authorization separately.

6. Shell portability versus convenience

Runner supports several shells. Unix jobs commonly use Bash or sh; Windows paths commonly use PowerShell variants. Syntax that is valid in Bash may fail in POSIX sh, and PowerShell has different error and quoting rules. Decide the shell contract deliberately.

portable_sh_job:
  image: alpine:3.22.1
  script:
    - set -eu
    - value="synthetic"
    - printf '%s
' "$value"

bash_specific_job:
  image: bash:5.2.37-alpine3.22
  script:
    - set -euo pipefail
    - printf 'bash=%s
' "$BASH_VERSION"

If a script requires Bash features such as arrays or pipefail, use an image/environment that explicitly provides Bash and record that assumption instead of hoping the runner fallback happens to be Bash.

7. Failure policy: fail, allow, or classify by exit code

A failed job is valuable control information. allow_failure: true changes pipeline interpretation and should correspond to a documented policy, not a desire to hide instability. GitLab also supports allow_failure:exit_codes when only specific exits are non-blocking.

optional_probe:
  script:
    - ./ci/probe.sh
  allow_failure:
    exit_codes:
      - 42

This says “exit 42 is a recognized non-blocking condition”; it does not say all probe failures are harmless. Preserve the actual exit code and evidence.

8. Worked design scenario

A team has a 40-second lint job, a 6-minute build, and a database-backed integration test. The lint does not need build output; integration does. A reasonable progression is:

  1. Keep simple stages while the pipeline is small.
  2. Move lint to needs: [] only when the team wants earlier feedback and understands DAG semantics.
  3. Package compiler/runtime dependencies in a versioned build image instead of installing them from the internet on every job.
  4. Use a disposable database service for integration tests, with explicit readiness checks and fake lab credentials.
  5. Keep deployment outside this job until artifact identity and deployment authorization are designed explicitly.
Maintainability

One clear runtime contract per job; avoid hidden host dependencies.

Least privilege

Test service has no production credentials or shared network reach.

Latency

Only independent work bypasses stage barriers.

Auditability

Image/service identities and exit policies are visible in CI config/evidence.

9. Anti-patterns and safer replacements

Anti-pattern Why it hurts Safer pattern
image: latest everywhere Runtime can change without CI diff. Version + digest where practical; record upstream release.
Install all tools in before_script Slow, network-dependent, drift-prone. Prebuilt reviewed tool image for stable requirements.
Shared mutable external DB for every branch Cross-pipeline state leakage. Disposable job-local service or isolated per-pipeline namespace.
|| true after flaky commands Converts unknown failures into green jobs. Capture exit code; classify allowed failures explicitly.
Privileged runner to “make Docker work” Expands host compromise blast radius. Use approved isolated runner pattern; privilege only when justified.
Next lesson

Diagnostics, Failure Modes, Security, and Performance

Intentionally break shell, service, and image assumptions; preserve first-failure evidence; and repair the actual execution layer rather than masking symptoms.

Knowledge check

When is stage-only scheduling usually preferable to needs?

Why is a versioned image generally better than installing a tool on every run?

Should a production database be modeled as a GitLab service container?

What must be true before using Bash-only shell features?

What does allow_failure:exit_codes add over blanket allow_failure: true?

Official references and version notes

  • CI/CD YAML syntax reference — current stages, stage, needs, image, services, before_script, after_script, allow_failure, and timeout semantics.
  • Scripts and job logs — non-zero exit behavior, multiline-command caveats, default hooks, and cancellation behavior.
  • Job execution flow — Runner source preparation, cache/artifact transfer, main execution, after_script, upload, and cleanup phases.
  • Docker executor — job images, service containers, runner workflow, shell requirements, entrypoints, and privilege risks.
  • Run jobs in Docker containers — image/service syntax, entrypoint handling, and where scripts execute.
  • Services — service aliases, networking, health checks, startup warnings, and service lifecycle.
  • Runner executors and supported shells — executor isolation and shell portability boundaries.
Version and compatibility note

Version-sensitive statements were rechecked against current primary GitLab documentation on 2026-09-11. Core jobs, stages, scripts, hooks, images, services, needs, and allow_failure are documented across GitLab Free, Premium, and Ultimate. Docker image/service examples require a Docker-capable runner or the documented local Docker simulation; the mandatory learning path does not require registering or weakening a production runner.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.