Chapter 15Lesson 01~245 minutes

DAG Pipelines, needs, Parallel Jobs, Matrices, Resource Groups, and Concurrency: Concepts, Architecture, and Mental Model

Model pipeline scheduling as explicit dependency edges, finite runner capacity, and shared-resource constraints rather than as stage order alone.

DAGneedsParallelismMatricesResource groupsCritical path

Learning objectives

  • Explain why stage ordering is a barrier model and how needs replaces selected barriers with explicit DAG edges.
  • Distinguish numeric parallel, parallel:matrix, runner concurrency, and real simultaneous execution.
  • Explain how resource_group serializes access to a named resource across pipelines in one project.
  • Reason about artifact availability, optional dependencies, interruptibility, and the critical path.
  • Inspect a pipeline graph and job timestamps before attempting optimization.
Availability baseline (verified 2026-08-21 against current GitLab documentation). The mandatory mechanisms in this chapter—needs, parallel, parallel:matrix, resource_group, and interruptible—are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. A job may list at most 50 needs entries. Numeric parallel accepts 1–200 instances, and parallel:matrix can create at most 200 permutations. Those are configuration limits, not guarantees of simultaneous execution: runner capacity, tags, protection, resource groups, and external-service limits can keep jobs pending. Hosted-runner quota/billing can change, so every required exercise includes a CI-Lint/fixture path that does not require paid compute.

1. The problem: stages express safety, but often too much waiting

Chapter 14 established that runners supply finite execution capacity. Now consider a pipeline with two independent build paths: an API build takes three minutes, a UI build takes twelve, and the API test needs only the API build. In a pure stage pipeline, the API test waits for the slow UI build because the entire build stage must finish. That wait is not a true dependency; it is an artifact of the stage barrier.

Pipeline optimization starts by separating dependency from presentation order. Stages answer “which broad phase normally comes before which?” A directed acyclic graph (DAG) answers “which exact jobs must finish before this exact job is eligible?”

2. Mental model: a pipeline has dependency edges and capacity constraints

From stage barriers to a DAG
flowchart LR
  A[api_build 3m] -- needs --> B[api_test 7m]
  U[ui_build 12m] -- needs --> V[ui_test 4m]
  B -- needs --> P[package 3m]
  V -- needs --> P
  P -- resource_group --> R[(shared release resource)]

The arrows are data/order dependencies created with needs. They do not create runner capacity. Jobs with satisfied dependencies become runnable; a matching runner still has to accept them. The resource-group edge is different: it is a mutual-exclusion constraint around an external/shared resource, not a dependency on another job.

A useful scheduling model is:

Runnable job = configuration includes the job + every required needs predecessor is satisfied + resource lock is available + an eligible runner slot exists. Running concurrently is therefore narrower than “the graph allows concurrency.”

3. Stage ordering versus needs

Model Eligibility rule Strength Failure mode
Stages only A later-stage job waits for all successful jobs in earlier stages. Simple mental model and broad safety barrier. Independent work waits unnecessarily.
needs DAG A job can start when its listed predecessors finish, even if other jobs in earlier stages are still running. Expresses actual dependencies and shortens idle waits. Missing/incorrect edges can start a consumer too early.
needs: [] Job is eligible as soon as the pipeline is created. Fast lint/static checks. Can compete for runner slots with critical work.

GitLab currently allows up to 50 jobs in a job's needs list. A graph with hundreds of edges is usually a design smell even before that limit: optimize around stable dependency boundaries instead of making every job know every other job.

4. Critical path: the path you cannot parallelize away

The critical path is the longest dependency chain that determines the earliest possible completion time if runner capacity is unlimited. In the example above, the API path takes 3 + 7 = 10 minutes before packaging, while the UI path takes 12 + 4 = 16. Packaging needs both and adds three minutes, so the theoretical minimum is 19 minutes. Adding more runners cannot reduce that below 19 unless you shorten or change a dependency on the critical path.

Contrast that with the stage-barrier approximation: build stage 12 minutes + test stage 7 minutes + package 3 minutes = 22 minutes. The DAG removes three minutes of false waiting; it does not make the work itself faster.

5. parallel creates job instances; runners determine simultaneity

Numeric parallel expands one job definition into multiple instances. GitLab names them job_name 1/N through N/N and gives each instance CI_NODE_INDEX and CI_NODE_TOTAL. Current syntax accepts 1–200 instances.

sharded_test:
  stage: test
  parallel: 4
  script:
    - printf 'node=%s total=%s\n' "$CI_NODE_INDEX" "$CI_NODE_TOTAL"

Four instances do not imply four jobs run simultaneously. With one eligible runner slot, three instances can remain pending. With four slots but an external test database that permits only one writer, true four-way concurrency may be incorrect even though GitLab can schedule it.

6. parallel:matrix expands combinations into named jobs

A matrix is useful when each instance needs explicit configuration values rather than only a shard number. GitLab expands combinations and exposes each matrix identifier as a CI/CD variable.

compatibility:
  stage: test
  parallel:
    matrix:
      - OS: [linux, windows]
        RUNTIME: [v1, v2]
  script:
    - printf 'os=%s runtime=%s\n' "$OS" "$RUNTIME"

This creates four job instances. Current GitLab permits at most 200 matrix permutations. Keep values concise: matrix values become part of job names, and duplicate value combinations can generate the same job name and overwrite each other. A matrix should represent a meaningful compatibility dimension, not a convenient way to multiply compute.

7. Depending on a matrix: all instances versus selected instances

If a normal needs entry points to a parallelized job, it refers to all instances of that parallel job and, by default, can download artifacts from all of them. If those instances publish artifacts with identical file names, downloads can overwrite one another. For selective dependencies use needs:parallel:matrix.

Current GitLab also supports matrix expressions such as $[[ matrix.OS ]] for one-to-one mappings. They were introduced in GitLab 18.6; current CI expressions documentation still treats the matrix context as a fast-moving feature, so this chapter teaches explicit/manual mapping first and uses matrix expressions as an optional readability improvement.

8. DAG conversion changes artifact flow

In a stage-only pipeline, a later-stage job normally receives artifacts from jobs in previous stages. Once a job uses needs, it can start before all previous stages finish, so GitLab cannot safely assume every previous artifact exists. The consumer downloads artifacts only from jobs listed in its needs entries; artifacts: false can suppress a particular download.

package:
  stage: package
  needs:
    - job: api_test
      artifacts: true
    - job: ui_test
      artifacts: false
  script:
    - test -f api-result.txt

Do not combine needs and dependencies casually in the same job. A DAG refactor is not complete until you re-prove the consumer's file inputs.

9. Optional dependencies protect pipeline creation when jobs are conditional

If rules can remove a producer from a pipeline, a mandatory needs edge to that producer can make the entire pipeline fail to create. optional: true means “wait for this job if it exists; otherwise proceed when the remaining needs are satisfied.”

optional_scan:
  rules:
    - if: '$RUN_SCAN == "true"'
  script: echo scan

report:
  needs:
    - job: optional_scan
      optional: true
  script: echo report

Do not read optional as “ignore failure.” It governs whether the dependency may be absent from the pipeline graph. Also note the current syntax restriction: needs:optional and needs:parallel:matrix cannot be combined on the same dependency.

10. resource_group: a semaphore around shared state

Pipelines run concurrently by default. A deployment, firmware programmer, shared integration appliance, or mutable test tenant may require mutual exclusion. A resource_group key creates a project-scoped lock: only one job using that key can hold it at a time; other jobs wait.

deploy_demo:
  stage: deploy
  script:
    - ./deploy-to-disposable-target.sh
  resource_group: chapter15-demo

The key models the resource identity. A key that is too broad serializes unrelated work; a key that is too narrow allows concurrent mutation of the same resource. GitLab supports process modes including unordered (default), oldest_first, newest_first, and newest_ready_first; changing an existing group's process mode is an API operation.

11. interruptible controls cancellation safety, not dependency order

When redundant-pipeline auto-cancel is enabled, interruptible: true tells GitLab that a running job may be canceled for a newer pipeline according to the project's auto-cancel mode. It is well suited to repeatable builds/tests. Long-lived mutation such as a deployment should normally remain non-interruptible unless it is explicitly designed to resume or roll back safely.

Parallelism and interruptibility interact operationally: a wide test matrix can consume many runner slots; making safe test jobs interruptible reduces wasted capacity when a newer commit supersedes them. This is a capacity optimization, not a substitute for correct rules.

12. Read-only inspection before optimization

Open Build → Pipelines, select a pipeline, and switch between the stage view and Job dependencies when available. Capture pipeline ID/SHA/source, each job ID/name/stage/status, created_at, started_at, finished_at, and runner identity. The Jobs API is useful for a machine-readable timeline:

PROJECT_ID="12345678"   # disposable project only
PIPELINE_ID="123456789"

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
  --jq '.[] | {id,name,stage,status,created_at,started_at,finished_at,runner:(.runner.description // null)}'

This evidence distinguishes three different waits: dependency wait, resource-group wait, and runner-capacity wait. Without timestamps and graph edges, “the DAG is slow” is not a diagnosis.

13. DevOps connection: optimize constraints, not YAML aesthetics

A production DAG is successful when it shortens feedback while preserving causality and bounded resource use. The optimization unit is the delivery constraint: dependency, runner slot, shared environment, artifact, network service, or human gate. A visually complex graph with the same critical path and higher compute consumption is worse, not more advanced.

14. Common misconceptions

  • “More jobs means a faster pipeline.” Only if the jobs are independent and capacity exists.
  • “needs means download every earlier artifact.” No. With needs, artifact flow is explicit to needed jobs.
  • “A resource group orders pipelines oldest first.” Not by default; unordered is the default process mode.
  • “optional: true ignores a failed producer.” It handles an absent producer, not a failed required job.
  • “Parallel matrix dimensions are free metadata.” Their Cartesian product creates real jobs and consumes runner/instance capacity.

Knowledge check

What does a needs edge change compared with stages?

A matrix expands to eight jobs but only two runner slots exist. What is the maximum number that can execute simultaneously?

Why can a pipeline fail to be created when needs points to a job controlled by rules?

What does resource_group: production protect?

After adding needs, a package job can no longer find an artifact from an unrelated previous-stage job. Why?

Why is a shorter YAML file not proof of a better DAG?

Summary

Stages create broad barriers; needs creates explicit DAG dependencies. parallel and matrices create multiple job instances but runner capacity determines real concurrency. resource_group serializes shared mutable resources, and interruptible can reclaim capacity from superseded safe work. Every optimization must re-prove artifact flow, conditional-job existence, critical path, and external-resource safety.

Official references

Next lesson

Measure a stage pipeline, convert it to a DAG, and prove the difference

Lesson 2 builds a disposable pipeline, captures timestamps, expands a small matrix, serializes a synthetic resource, and calculates runner-slot demand from observed jobs.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.