Chapter 09Lesson 04~170 minutes

Stages, needs DAGs, Dependency Graphs, Early Execution, and Pipeline Critical-Path Design: Diagnostics, Failure Modes, Security, and Performance

Fast pipelines can be wrong pipelines. This lesson diagnoses premature consumers, accidentally serialized critical paths, invalid graphs, conditionally missing producers, artifact mistakes, and queue pressure while preserving the original pipeline/job/timing evidence before changing the graph.

DiagnosticsFailure modesQueue pressureArtifactsRecovery

Learning objectives

  • Diagnose early consumers that lack required data, over-serialized graphs, invalid/cyclic graphs, and missing conditional producers.
  • Separate configuration/graph failures from queue/runner, script, artifact, deployment, API, and policy failures.
  • Recognize when a fast graph is logically wrong even though every individual job can succeed.
  • Measure queue duration and critical-path delay before adding more parallelism or runner capacity.
  • Apply the smallest safe graph correction and rerun only the necessary disposable scope.

1. Evidence-first diagnostic sequence

When a DAG behaves incorrectly, preserve the first failure before retrying. Use the same layer order every time:

  1. Pipeline ID, source, ref, SHA, first-failure timestamp.
  2. Compiled/full configuration and CI Lint result.
  3. workflow/job-rule membership: did the producer and consumer both exist?
  4. Graph: stages, needs, optional flags, artifact flags.
  5. Job status/timing: created, started, finished, queued duration.
  6. Runner/executor/image identity.
  7. Script/tool/network trace.
  8. Producer artifact/report metadata and consumer download evidence.
  9. External target/policy state if the job has side effects.

Only after the causal layer is known should you rerun the smallest safe scope.

2. Failure mode: consumer starts before required data exists

This is the classic unsafe DAG refactor. A developer removes a stage barrier and gives the consumer a dependency on a fast linter because “that releases it early,” but the consumer actually reads a build artifact.

# Intentionally broken.
lint:
  stage: test
  needs: []
  script: ./ci/lint.sh

build_app:
  stage: build
  script:
    - mkdir -p out
    - echo "$CI_COMMIT_SHA" > out/app.sha
  artifacts:
    paths: [out/app.sha]

integration_test:
  stage: test
  needs:
    - lint
  script:
    - test -s out/app.sha
    - grep -F "$CI_COMMIT_SHA" out/app.sha

The graph is syntactically plausible but logically wrong. integration_test can start after lint while build_app is still running. The expected first failure is test -s out/app.sha. Repair the causal edge:

integration_test:
  stage: test
  needs:
    - job: build_app
      artifacts: true
  script:
    - test -s out/app.sha
    - grep -F "$CI_COMMIT_SHA" out/app.sha
Do not “repair” this by adding sleep 30. Time delay is not a dependency contract.

3. Failure mode: accidental serialized critical path

Extra dependencies can be as damaging as missing ones. If linux_test needs both linux_build and an unrelated slow windows_build, the DAG preserves the stage-era wait under a new syntax.

linux_test:
  stage: test
  needs:
    - linux_build
    - windows_build   # Remove if no control/data dependency exists.
  script: ./ci/test-linux.sh

Compare start timestamps. If linux_build finishes at 10:00:05, windows_build at 10:00:40, and linux_test becomes runnable only after 10:00:40, the graph—not runner capacity—is imposing the wait. Remove the edge only after proving the Linux test consumes no Windows output/policy state.

4. Failure mode: impossible or cyclic graph

A DAG must be acyclic. An impossible dependency graph should be caught at configuration validation/pipeline creation, before runner work. Use CI Lint rather than pushing repeated broken commits.

# Intentionally invalid dependency cycle for CI Lint practice.
job_a:
  stage: test
  needs: [job_b]
  script: echo A

job_b:
  stage: test
  needs: [job_a]
  script: echo B

Preserve the linter/configuration error as evidence. No amount of runner troubleshooting can fix a graph that GitLab cannot construct.

5. Failure mode: required need points to a job omitted by rules

Suppose docs_build exists only on the default branch, but package always exists and requires it:

docs_build:
  stage: build
  rules:
    - if: '$CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
  script: ./ci/build-docs.sh

package:
  stage: package
  needs:
    - docs_build
  script: ./ci/package.sh

On another branch, the producer can be absent while the required edge remains. GitLab validates needs before execution, so the pipeline can fail to create. There are two legitimate repairs:

  • If package is invalid without docs, align package rules so it is absent whenever docs_build is absent.
  • If docs are genuinely optional, mark that edge optional: true and make the package script robust to their absence.

Choose based on correctness, not which YAML makes the error disappear.

6. Failure mode: a faster pipeline uses the wrong dependency

Imagine package needs the build artifact but does not require unit_test to pass. Pointing it only at build_app makes packaging start earlier, but it changes the quality gate. The pipeline may finish faster while publishing or retaining a package that was never validated.

Question Evidence
What was the old gate? Stage order or explicit test requirement
What is the new gate? List of needs edges on package
Did artifact availability change? Producer job ID/path/SHA
Did authorization/quality policy change? Required check/job status before package start
Is the speedup acceptable? Only if the removed gate was intentionally unnecessary
Critical-path optimization must preserve policy dependencies as well as file dependencies.

7. Failure mode: graph is correct, artifact dataflow is not

A consumer can wait for the right producer yet still fail if artifact transfer is disabled, expired, overwritten, or misnamed. With needs, inspect each edge’s artifact flag and producer artifact metadata.

verify_binary:
  stage: test
  needs:
    - job: build_binary
      artifacts: false   # Broken if verify_binary reads dist/app.bin.
  script:
    - test -s dist/app.bin

Repair artifacts: false to true only if the file is meant to cross the job boundary. If data belongs in the repository or an external immutable registry, fix the correct data layer instead of overloading job artifacts.

8. Performance failure: correct DAG, no latency improvement

A graph can be perfectly optimized while every newly runnable job waits pending. Compare dependency release time with started_at and queued_duration. If queue dominates, inspect matching runner capacity, tags, autoscaling, executor startup, and workload contention from Chapters 04–05.

Symptom Likely layer Do not do
Many ready jobs pending Runner capacity/routing Delete needs edges blindly
Low queue, one long chain Graph/critical path Add runners as universal fix
High artifact download time Dataflow/storage Add fan-out without measuring I/O
Parallel jobs slow together Resource contention Assume concurrency is free

9. Security-sensitive graph changes

Graph edits can change when privileged work becomes runnable. Treat these as security-sensitive if they affect jobs that can publish packages, deploy, use protected variables, access internal networks, or run on privileged runners. Preserve the old graph and required checks before changing order.

Never use a broad PAT, print tokens, disable TLS, move untrusted jobs onto privileged runners, or bypass policy simply to test whether a DAG is “faster.” The mandatory examples need none of those actions.

10. Smallest-safe-repair playbook

Missing data

Add the actual producer need with artifact transfer; verify SHA-bound content.

Extra wait

Remove only an edge proven unnecessary by control/data/policy analysis.

Missing conditional job

Align rules or mark truly optional edge optional.

Invalid graph

Fix configuration in CI Lint before any runner investigation.

Queue bottleneck

Measure runner capacity/tags/queued duration; graph may already be correct.

Wrong gate

Restore quality/security dependency even if it lengthens the critical path.

Knowledge check

Why is sleep 30 not a valid fix for a premature consumer?

A consumer needs a conditionally omitted producer. What are the two legitimate repair families?

What evidence distinguishes graph serialization from runner queue delay?

Can a pipeline be faster and still be a regression?

Where should a cyclic needs graph be diagnosed?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-11. GitLab CI/CD DAG and artifact-transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.

  • Make jobs start earlier with needs — stage barriers, DAG execution, immediate jobs, and practical examples.
  • CI/CD YAML syntax reference — authoritative stages, needs, needs:artifacts, needs:optional, needs:project, and needs:pipeline:job semantics and limits.
  • Pipeline editor — visualization of jobs, stages, and needs relationships plus full configuration inspection.
  • CI Lint — syntax/logic validation and pipeline simulation that can expose invalid needs relationships before execution.
  • Job artifacts — default previous-stage artifact fetching and how needs:artifacts changes data transfer.
  • Troubleshooting job artifacts — missing/expired/inaccessible artifact failures.
  • Jobs API — job IDs, stage/status, created_at, started_at, finished_at, duration, queued duration, and runner metadata for timing evidence.
Next lesson

Checkpoint Lab — Stages, needs DAGs, Dependency Graphs, Early Execution, and Pipeline Critical-Path Design

Refactor a complete staged pipeline into a measured DAG, inject one dependency fault, repair it, and prove the final graph is both faster and correct.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.