Stages, needs DAGs, Dependency Graphs, Early Execution, and Pipeline Critical-Path Design: Diagnostics, Failure Modes, Security, and Performance
Fast pipelines can be wrong pipelines. This lesson diagnoses premature consumers, accidentally serialized critical paths, invalid graphs, conditionally missing producers, artifact mistakes, and queue pressure while preserving the original pipeline/job/timing evidence before changing the graph.
Learning objectives
- Diagnose early consumers that lack required data, over-serialized graphs, invalid/cyclic graphs, and missing conditional producers.
- Separate configuration/graph failures from queue/runner, script, artifact, deployment, API, and policy failures.
- Recognize when a fast graph is logically wrong even though every individual job can succeed.
- Measure queue duration and critical-path delay before adding more parallelism or runner capacity.
- Apply the smallest safe graph correction and rerun only the necessary disposable scope.
1. Evidence-first diagnostic sequence
When a DAG behaves incorrectly, preserve the first failure before retrying. Use the same layer order every time:
- Pipeline ID, source, ref, SHA, first-failure timestamp.
- Compiled/full configuration and CI Lint result.
-
workflow/job-rule membership: did the producer and consumer both exist? -
Graph: stages,
needs, optional flags, artifact flags. - Job status/timing: created, started, finished, queued duration.
- Runner/executor/image identity.
- Script/tool/network trace.
- Producer artifact/report metadata and consumer download evidence.
- External target/policy state if the job has side effects.
Only after the causal layer is known should you rerun the smallest safe scope.
2. Failure mode: consumer starts before required data exists
This is the classic unsafe DAG refactor. A developer removes a stage barrier and gives the consumer a dependency on a fast linter because “that releases it early,” but the consumer actually reads a build artifact.
# Intentionally broken.
lint:
stage: test
needs: []
script: ./ci/lint.sh
build_app:
stage: build
script:
- mkdir -p out
- echo "$CI_COMMIT_SHA" > out/app.sha
artifacts:
paths: [out/app.sha]
integration_test:
stage: test
needs:
- lint
script:
- test -s out/app.sha
- grep -F "$CI_COMMIT_SHA" out/app.sha
The graph is syntactically plausible but logically wrong.
integration_test can start after
lint while build_app is still running. The
expected first failure is test -s out/app.sha. Repair
the causal edge:
integration_test:
stage: test
needs:
- job: build_app
artifacts: true
script:
- test -s out/app.sha
- grep -F "$CI_COMMIT_SHA" out/app.sha
sleep 30. Time delay is
not a dependency contract.
3. Failure mode: accidental serialized critical path
Extra dependencies can be as damaging as missing ones. If
linux_test needs both linux_build and an
unrelated slow windows_build, the DAG preserves the
stage-era wait under a new syntax.
linux_test:
stage: test
needs:
- linux_build
- windows_build # Remove if no control/data dependency exists.
script: ./ci/test-linux.sh
Compare start timestamps. If linux_build finishes at
10:00:05, windows_build at 10:00:40, and
linux_test becomes runnable only after 10:00:40, the
graph—not runner capacity—is imposing the wait. Remove the edge only
after proving the Linux test consumes no Windows output/policy
state.
4. Failure mode: impossible or cyclic graph
A DAG must be acyclic. An impossible dependency graph should be caught at configuration validation/pipeline creation, before runner work. Use CI Lint rather than pushing repeated broken commits.
# Intentionally invalid dependency cycle for CI Lint practice.
job_a:
stage: test
needs: [job_b]
script: echo A
job_b:
stage: test
needs: [job_a]
script: echo B
Preserve the linter/configuration error as evidence. No amount of runner troubleshooting can fix a graph that GitLab cannot construct.
5. Failure mode: required need points to a job omitted by rules
Suppose docs_build exists only on the default branch,
but package always exists and requires it:
docs_build:
stage: build
rules:
- if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
script: ./ci/build-docs.sh
package:
stage: package
needs:
- docs_build
script: ./ci/package.sh
On another branch, the producer can be absent while the required edge remains. GitLab validates needs before execution, so the pipeline can fail to create. There are two legitimate repairs:
- If package is invalid without docs, align package rules so it is absent whenever docs_build is absent.
-
If docs are genuinely optional, mark that edge
optional: trueand make the package script robust to their absence.
Choose based on correctness, not which YAML makes the error disappear.
6. Failure mode: a faster pipeline uses the wrong dependency
Imagine package needs the build artifact but does
not require unit_test to pass. Pointing it
only at build_app makes packaging start earlier, but it
changes the quality gate. The pipeline may finish faster while
publishing or retaining a package that was never validated.
| Question | Evidence |
|---|---|
| What was the old gate? | Stage order or explicit test requirement |
| What is the new gate? | List of needs edges on package |
| Did artifact availability change? | Producer job ID/path/SHA |
| Did authorization/quality policy change? | Required check/job status before package start |
| Is the speedup acceptable? | Only if the removed gate was intentionally unnecessary |
7. Failure mode: graph is correct, artifact dataflow is not
A consumer can wait for the right producer yet still fail if
artifact transfer is disabled, expired, overwritten, or misnamed.
With needs, inspect each edge’s artifact flag and
producer artifact metadata.
verify_binary:
stage: test
needs:
- job: build_binary
artifacts: false # Broken if verify_binary reads dist/app.bin.
script:
- test -s dist/app.bin
Repair artifacts: false to true only if the file is
meant to cross the job boundary. If data belongs in the repository
or an external immutable registry, fix the correct data layer
instead of overloading job artifacts.
8. Performance failure: correct DAG, no latency improvement
A graph can be perfectly optimized while every newly runnable job
waits pending. Compare dependency release time with
started_at and queued_duration. If queue
dominates, inspect matching runner capacity, tags, autoscaling,
executor startup, and workload contention from Chapters 04–05.
| Symptom | Likely layer | Do not do |
|---|---|---|
| Many ready jobs pending | Runner capacity/routing | Delete needs edges blindly |
| Low queue, one long chain | Graph/critical path | Add runners as universal fix |
| High artifact download time | Dataflow/storage | Add fan-out without measuring I/O |
| Parallel jobs slow together | Resource contention | Assume concurrency is free |
9. Security-sensitive graph changes
Graph edits can change when privileged work becomes runnable. Treat these as security-sensitive if they affect jobs that can publish packages, deploy, use protected variables, access internal networks, or run on privileged runners. Preserve the old graph and required checks before changing order.
Never use a broad PAT, print tokens, disable TLS, move untrusted jobs onto privileged runners, or bypass policy simply to test whether a DAG is “faster.” The mandatory examples need none of those actions.
10. Smallest-safe-repair playbook
Add the actual producer need with artifact transfer; verify SHA-bound content.
Remove only an edge proven unnecessary by control/data/policy analysis.
Align rules or mark truly optional edge optional.
Fix configuration in CI Lint before any runner investigation.
Measure runner capacity/tags/queued duration; graph may already be correct.
Restore quality/security dependency even if it lengthens the critical path.
Knowledge check
Why is sleep 30 not a valid fix for a premature consumer?
It guesses timing rather than encoding a dependency; runner/load changes can make the race return.
A consumer needs a conditionally omitted producer. What are the two legitimate repair families?
Align the consumer’s rules so both exist together, or use optional true only if the consumer is correct without the producer.
What evidence distinguishes graph serialization from runner queue delay?
Producer completion times, consumer readiness/start time, and queued duration. An unnecessary edge delays readiness; runner scarcity delays start after readiness.
Can a pipeline be faster and still be a regression?
Yes. Removing a required data, quality, security, or authorization dependency can reduce elapsed time while violating correctness.
Where should a cyclic needs graph be diagnosed?
At configuration validation/pipeline creation using full configuration and CI Lint, not at the runner layer.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. GitLab CI/CD DAG and artifact-transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.
-
Make jobs start earlier with
needs— stage barriers, DAG execution, immediate jobs, and practical examples. -
CI/CD YAML syntax reference
— authoritative
stages,needs,needs:artifacts,needs:optional,needs:project, andneeds:pipeline:jobsemantics and limits. -
Pipeline editor
— visualization of jobs, stages, and
needsrelationships plus full configuration inspection. -
CI Lint
— syntax/logic validation and pipeline simulation that can expose
invalid
needsrelationships before execution. -
Job artifacts
— default previous-stage artifact fetching and how
needs:artifactschanges data transfer. - Troubleshooting job artifacts — missing/expired/inaccessible artifact failures.
-
Jobs API
— job IDs, stage/status,
created_at,started_at,finished_at, duration, queued duration, and runner metadata for timing evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.