Stages, needs DAGs, Dependency Graphs, Early Execution, and Pipeline Critical-Path Design: Concepts, Architecture, and Mental Model
Stages are a safe default, but stage barriers can make independent work wait for unrelated work. This lesson builds a precise model of stage ordering, explicit needs edges, runnable frontiers, runner queues, artifact/data dependencies, fan-in, and the pipeline critical path so optimization never outruns correctness.
Learning objectives
- Explain the difference between a stage barrier and an explicit needs dependency edge.
- Identify the runnable frontier, queue delay, execution time, fan-out/fan-in points, and the pipeline critical path.
- Separate control dependencies from file/data dependencies and understand how needs changes artifact transfer.
- Predict when needs: [] or same-stage needs can start work earlier without violating correctness.
- Preserve source/SHA, compiled graph, job IDs, timestamps, runner context, and producer/consumer evidence before optimization.
1. The practical problem: unrelated work should not become a dependency
Chapter 08 answered which jobs exist. Chapter 09 answers a different question: when is each included job allowed to start? A traditional stage pipeline groups jobs into coarse barriers. Every successful job in one stage must finish before the next stage opens, even when a later job depends on only one producer. That model is simple and safe, but it can hide unnecessary waiting.
A DAG replaces some barrier assumptions with explicit dependency edges. The goal is not “make everything parallel.” The goal is to encode the real correctness graph so GitLab can expose the next runnable jobs as soon as their prerequisites are complete.
2. Mental model: graph, runnable frontier, queue, execution, fan-in, critical path
Start with jobs as nodes. A dependency edge means the consumer must wait for the producer. GitLab evaluates the compiled configuration and creates a graph. At any moment, jobs whose required predecessors are satisfied form the runnable frontier. Runnable does not mean running: a job may still sit pending while waiting for a matching runner. Execution then consumes runner capacity. Downstream fan-in nodes wait for every required predecessor before they enter the frontier.
The critical path is the longest dependency path that determines the earliest possible pipeline completion under the observed graph and capacity. A long independent job can still dominate completion even if it is not in a traditional “main” stage.
flowchart TD
A[Compiled jobs] --> B{Ordering model}
B -->|Stage barriers| C[Whole prior stage must pass]
B -->|needs edges| D[Only explicit prerequisites must pass]
C --> E[Runnable frontier]
D --> E
E --> F[Runner queue]
F --> G[Parallel execution]
G --> H[Fan-in gates]
H --> I[Pipeline completion]
G --> J[Job timing evidence]
J --> K[Observed critical path]
3. Keep graph state, runner state, and data state separate
| State | Owned by | Evidence | Common confusion |
|---|---|---|---|
| Source/ref/SHA | Pipeline request + Git repository | CI_PIPELINE_SOURCE, ref, immutable SHA |
Graph shape does not change source identity. |
| Compiled graph | GitLab CI compiler | Full config, stages, job list, needs edges | Repository YAML text alone may omit includes/expansion. |
| Runnable/queued state | GitLab scheduler | Job status, created/start timestamps, queued duration | A dependency-ready job may still wait for a runner. |
| Runner/executor state | Runner manager/executor | Runner ID/version/executor; trace | Runner capacity is not a dependency edge. |
| Artifact/data state | Producer job + GitLab artifact storage | Producer job ID, artifact path/digest, consumer download | Control dependency does not automatically mean arbitrary external data exists. |
| External system state | External service/provider | Target-specific health/record | Pipeline green is not external health. |
4. Stage barriers are coarse but easy to reason about
Without needs, GitLab runs jobs in the same stage in
parallel and waits for the entire successful stage before opening
the next one. If any ordinary job in an earlier stage fails, later
stages do not start. The model is intentionally conservative.
stages: [build, test, package]
build_app:
stage: build
script: ./ci/build-app.sh
build_docs:
stage: build
script: ./ci/build-docs.sh
unit_test:
stage: test
script: ./ci/test-app.sh
package:
stage: package
script: ./ci/package.sh
Even if unit_test needs only build_app, it
waits for build_docs because the
build stage is the barrier. That wait may be appropriate if
the stage intentionally represents a policy gate; it may be waste if
the jobs are independent.
5. needs replaces a barrier with explicit prerequisites
With needs, the consumer may start as soon as the
listed jobs complete successfully, without waiting for unrelated
jobs in earlier stages. GitLab allows jobs in different stages to
run concurrently and also permits needs edges between
jobs in the same stage.
unit_test:
stage: test
needs:
- job: build_app
artifacts: true
script: ./ci/test-app.sh
unit_test can now start when
build_app finishes even if build_docs is
still running. This is a real dependency graph: the edge is both an
ordering statement and, because artifacts defaults to
true for a listed need, an artifact-transfer statement for that
producer.
needs. Self-Managed and Dedicated
default to 50, but administrators can change that limit. A giant
fan-in is often a design smell even before the limit is reached.
6. needs: [] means “no predecessor barrier”
An empty needs array says that the job has no job
dependencies and can enter the runnable frontier immediately when
the pipeline is created. It is useful for source-only linting or
fast metadata checks that truly do not require build outputs.
lint_source:
stage: test
needs: []
script:
- ./ci/lint-source.sh
needs: [] as a speed hack.
If the command reads generated files, test binaries, downloaded
dependencies, or setup artifacts from another job, the dependency is
not empty.
7. Control edges and artifact edges must agree
Stage-only jobs normally fetch artifacts from all successful jobs in
previous stages unless artifact behavior is narrowed. A job that
uses needs no longer gets that broad previous-stage
artifact set. It can fetch artifacts only from jobs listed in its
needs, controlled by
artifacts: true|false. This is why DAG refactoring can
reveal a hidden data dependency that a stage barrier previously
supplied implicitly.
build_app:
stage: build
script:
- mkdir -p out
- printf 'sha=%s\n' "$CI_COMMIT_SHA" > out/app.txt
artifacts:
paths: [out/app.txt]
unit_test:
stage: test
needs:
- job: build_app
artifacts: true
script:
- test -s out/app.txt
- grep -F "sha=$CI_COMMIT_SHA" out/app.txt
The consumer proves two things: the file arrived, and it was produced for the same immutable SHA. The second check prevents “some artifact exists” from being mistaken for “the intended artifact is present.”
8. Critical-path evidence requires timestamps and queue context
Do not infer performance from stage names. Record job
created_at, started_at,
finished_at, duration, and queued duration from the
UI/API where available. A DAG may expose work earlier while total
time remains unchanged because jobs queue behind limited runner
capacity. Conversely, runner expansion does not fix a truly
serialized dependency chain.
| Observation | What it means | Next question |
|---|---|---|
| Consumer starts soon after producer finishes | Dependency edge is releasing work early | Was queue duration small or did runner scarcity mask the gain? |
| Job is ready but remains pending | Graph is not the current bottleneck | Do matching runners have capacity/tags? |
| Many jobs start together and all slow down | Fan-out may contend for CPU/network/cache | Is parallelism exceeding useful capacity? |
| Pipeline completion follows one long chain | That chain is likely critical | Can any edge on the chain be safely removed or work split? |
9. Optional producers require an explicit correctness decision
Rules can omit a producer from a pipeline. If a consumer has an
ordinary required needs entry for that absent job,
pipeline creation can fail because GitLab validates the graph before
execution. Use optional: true only when the consumer
remains correct without the producer.
docs_build:
stage: build
rules:
- if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
script: ./ci/build-docs.sh
summary:
stage: test
needs:
- job: docs_build
optional: true
script: ./ci/summarize.sh
If summary actually requires generated documentation,
marking the edge optional is wrong; align the rules so both jobs
exist together instead. optional is a statement about
data/control correctness, not an error-suppression switch.
10. Read-only inspection sequence before graph changes
- Record pipeline ID, source, ref, and immutable SHA.
- Inspect full/merged configuration and CI Lint simulation for the exact revision.
-
Open the pipeline graph and list stages plus every
needsrelationship. - For each consumer, write its actual inputs: repository files, variables, artifacts, services, external state.
- Record job IDs and timing fields; separate dependency wait from queued duration.
- Record producer artifact identity/path before changing transfer behavior.
- Only then propose an edge removal/addition and predict the observable start-time change.
11. Mental-model summary
Simple coarse ordering: all successful jobs in a prior stage complete before the next stage opens.
Explicit same-pipeline prerequisite that can release a consumer before unrelated earlier-stage jobs finish.
Jobs whose required graph predecessors are satisfied; runner assignment can still delay them.
With needs, only listed producers are eligible artifact sources; artifacts true is the default.
Longest effective dependency path to completion, influenced by execution and queue/capacity realities.
Every removed wait must be justified by verified absence of a control/data dependency.
Knowledge check
A test job waits for an unrelated 15-minute docs build because both builds share one stage. What is the potential optimization layer?
The job dependency graph. If the test truly depends only on its own build, a needs edge can release it earlier without changing runner configuration.
Does a job becoming runnable prove it starts immediately?
No. It can remain pending in the runner queue. Preserve queued duration and runner capacity evidence.
What artifact behavior changes when a job starts using needs?
It no longer implicitly downloads all previous-stage artifacts; it can download only artifacts from listed needs, controlled per edge.
When is optional: true safe?
Only when the consumer is semantically correct if the producer is absent. It is not a generic fix for graph-validation errors.
What is more important than a shorter critical path?
A correct critical path whose consumers receive all required inputs and whose external side effects remain safe and verifiable.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. GitLab CI/CD DAG and artifact-transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.
-
Make jobs start earlier with
needs— stage barriers, DAG execution, immediate jobs, and practical examples. -
CI/CD YAML syntax reference
— authoritative
stages,needs,needs:artifacts,needs:optional,needs:project, andneeds:pipeline:jobsemantics and limits. -
Pipeline editor
— visualization of jobs, stages, and
needsrelationships plus full configuration inspection. -
CI Lint
— syntax/logic validation and pipeline simulation that can expose
invalid
needsrelationships before execution. -
Job artifacts
— default previous-stage artifact fetching and how
needs:artifactschanges data transfer. - Troubleshooting job artifacts — missing/expired/inaccessible artifact failures.
-
Jobs API
— job IDs, stage/status,
created_at,started_at,finished_at, duration, queued duration, and runner metadata for timing evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.