DAG Pipelines, needs, Parallel Jobs, Matrices, Resource Groups, and Concurrency: Concepts, Architecture, and Mental Model
Model pipeline scheduling as explicit dependency edges, finite runner capacity, and shared-resource constraints rather than as stage order alone.
Learning objectives
-
Explain why stage ordering is a barrier model and how
needsreplaces selected barriers with explicit DAG edges. -
Distinguish numeric
parallel,parallel:matrix, runner concurrency, and real simultaneous execution. -
Explain how
resource_groupserializes access to a named resource across pipelines in one project. - Reason about artifact availability, optional dependencies, interruptibility, and the critical path.
- Inspect a pipeline graph and job timestamps before attempting optimization.
needs,
parallel, parallel:matrix,
resource_group, and interruptible—are
available in GitLab Free/Premium/Ultimate across GitLab.com,
Self-Managed, and Dedicated. A job may list at most 50
needs entries. Numeric parallel accepts
1–200 instances, and parallel:matrix can create at most
200 permutations. Those are configuration limits, not guarantees of
simultaneous execution: runner capacity, tags, protection, resource
groups, and external-service limits can keep jobs pending.
Hosted-runner quota/billing can change, so every required exercise
includes a CI-Lint/fixture path that does not require paid compute.
1. The problem: stages express safety, but often too much waiting
Chapter 14 established that runners supply finite execution
capacity. Now consider a pipeline with two independent build paths:
an API build takes three minutes, a UI build takes twelve, and the
API test needs only the API build. In a pure stage pipeline, the API
test waits for the slow UI build because the entire
build stage must finish. That wait is not a true
dependency; it is an artifact of the stage barrier.
Pipeline optimization starts by separating dependency from presentation order. Stages answer “which broad phase normally comes before which?” A directed acyclic graph (DAG) answers “which exact jobs must finish before this exact job is eligible?”
2. Mental model: a pipeline has dependency edges and capacity constraints
flowchart LR A[api_build 3m] -- needs --> B[api_test 7m] U[ui_build 12m] -- needs --> V[ui_test 4m] B -- needs --> P[package 3m] V -- needs --> P P -- resource_group --> R[(shared release resource)]
The arrows are data/order dependencies created with
needs. They do not create runner capacity. Jobs with
satisfied dependencies become runnable; a matching runner still
has to accept them. The resource-group edge is different: it is a
mutual-exclusion constraint around an external/shared resource,
not a dependency on another job.
A useful scheduling model is:
needs predecessor is satisfied +
resource lock is available + an eligible runner slot exists.
Running concurrently is therefore narrower than
“the graph allows concurrency.”
3. Stage ordering versus needs
| Model | Eligibility rule | Strength | Failure mode |
|---|---|---|---|
| Stages only | A later-stage job waits for all successful jobs in earlier stages. | Simple mental model and broad safety barrier. | Independent work waits unnecessarily. |
needs DAG
|
A job can start when its listed predecessors finish, even if other jobs in earlier stages are still running. | Expresses actual dependencies and shortens idle waits. | Missing/incorrect edges can start a consumer too early. |
needs: []
|
Job is eligible as soon as the pipeline is created. | Fast lint/static checks. | Can compete for runner slots with critical work. |
GitLab currently allows up to 50 jobs in a job's
needs list. A graph with hundreds of edges is usually a
design smell even before that limit: optimize around stable
dependency boundaries instead of making every job know every other
job.
4. Critical path: the path you cannot parallelize away
The critical path is the longest dependency chain
that determines the earliest possible completion time if runner
capacity is unlimited. In the example above, the API path takes
3 + 7 = 10 minutes before packaging, while the UI path
takes 12 + 4 = 16. Packaging needs both and adds three
minutes, so the theoretical minimum is 19 minutes. Adding more
runners cannot reduce that below 19 unless you shorten or change a
dependency on the critical path.
Contrast that with the stage-barrier approximation: build stage 12 minutes + test stage 7 minutes + package 3 minutes = 22 minutes. The DAG removes three minutes of false waiting; it does not make the work itself faster.
5. parallel creates job instances; runners determine
simultaneity
Numeric parallel expands one job definition into
multiple instances. GitLab names them
job_name 1/N through N/N and gives each
instance CI_NODE_INDEX and CI_NODE_TOTAL.
Current syntax accepts 1–200 instances.
sharded_test:
stage: test
parallel: 4
script:
- printf 'node=%s total=%s\n' "$CI_NODE_INDEX" "$CI_NODE_TOTAL"
Four instances do not imply four jobs run simultaneously. With one eligible runner slot, three instances can remain pending. With four slots but an external test database that permits only one writer, true four-way concurrency may be incorrect even though GitLab can schedule it.
6. parallel:matrix expands combinations into named jobs
A matrix is useful when each instance needs explicit configuration values rather than only a shard number. GitLab expands combinations and exposes each matrix identifier as a CI/CD variable.
compatibility:
stage: test
parallel:
matrix:
- OS: [linux, windows]
RUNTIME: [v1, v2]
script:
- printf 'os=%s runtime=%s\n' "$OS" "$RUNTIME"
This creates four job instances. Current GitLab permits at most 200 matrix permutations. Keep values concise: matrix values become part of job names, and duplicate value combinations can generate the same job name and overwrite each other. A matrix should represent a meaningful compatibility dimension, not a convenient way to multiply compute.
7. Depending on a matrix: all instances versus selected instances
If a normal needs entry points to a parallelized job,
it refers to all instances of that parallel job and, by default, can
download artifacts from all of them. If those instances publish
artifacts with identical file names, downloads can overwrite one
another. For selective dependencies use
needs:parallel:matrix.
Current GitLab also supports matrix expressions such as
$[[ matrix.OS ]] for one-to-one mappings. They were
introduced in GitLab 18.6; current CI expressions documentation
still treats the matrix context as a fast-moving feature, so this
chapter teaches explicit/manual mapping first and uses matrix
expressions as an optional readability improvement.
8. DAG conversion changes artifact flow
In a stage-only pipeline, a later-stage job normally receives
artifacts from jobs in previous stages. Once a job uses
needs, it can start before all previous stages finish,
so GitLab cannot safely assume every previous artifact exists. The
consumer downloads artifacts only from jobs listed in its
needs entries; artifacts: false can
suppress a particular download.
package:
stage: package
needs:
- job: api_test
artifacts: true
- job: ui_test
artifacts: false
script:
- test -f api-result.txt
Do not combine needs and
dependencies casually in the same job. A DAG refactor
is not complete until you re-prove the consumer's file inputs.
9. Optional dependencies protect pipeline creation when jobs are conditional
If rules can remove a producer from a pipeline, a
mandatory needs edge to that producer can make the
entire pipeline fail to create. optional: true means
“wait for this job if it exists; otherwise proceed when the
remaining needs are satisfied.”
optional_scan:
rules:
- if: '$RUN_SCAN == "true"'
script: echo scan
report:
needs:
- job: optional_scan
optional: true
script: echo report
Do not read optional as “ignore failure.” It governs
whether the dependency may be absent from the pipeline graph. Also
note the current syntax restriction: needs:optional and
needs:parallel:matrix cannot be combined on the same
dependency.
10. resource_group: a semaphore around shared state
Pipelines run concurrently by default. A deployment, firmware
programmer, shared integration appliance, or mutable test tenant may
require mutual exclusion. A resource_group key creates
a project-scoped lock: only one job using that key can hold it at a
time; other jobs wait.
deploy_demo:
stage: deploy
script:
- ./deploy-to-disposable-target.sh
resource_group: chapter15-demo
The key models the resource identity. A key that is too
broad serializes unrelated work; a key that is too narrow allows
concurrent mutation of the same resource. GitLab supports process
modes including unordered (default),
oldest_first, newest_first, and
newest_ready_first; changing an existing group's
process mode is an API operation.
11. interruptible controls cancellation safety, not
dependency order
When redundant-pipeline auto-cancel is enabled,
interruptible: true tells GitLab that a running job may
be canceled for a newer pipeline according to the project's
auto-cancel mode. It is well suited to repeatable builds/tests.
Long-lived mutation such as a deployment should normally remain
non-interruptible unless it is explicitly designed to resume or roll
back safely.
Parallelism and interruptibility interact operationally: a wide test
matrix can consume many runner slots; making safe test jobs
interruptible reduces wasted capacity when a newer commit supersedes
them. This is a capacity optimization, not a substitute for correct
rules.
12. Read-only inspection before optimization
Open Build → Pipelines, select a pipeline, and
switch between the stage view and
Job dependencies when available. Capture pipeline
ID/SHA/source, each job ID/name/stage/status,
created_at, started_at,
finished_at, and runner identity. The Jobs API is
useful for a machine-readable timeline:
PROJECT_ID="12345678" # disposable project only
PIPELINE_ID="123456789"
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
--jq '.[] | {id,name,stage,status,created_at,started_at,finished_at,runner:(.runner.description // null)}'
This evidence distinguishes three different waits: dependency wait, resource-group wait, and runner-capacity wait. Without timestamps and graph edges, “the DAG is slow” is not a diagnosis.
13. DevOps connection: optimize constraints, not YAML aesthetics
A production DAG is successful when it shortens feedback while preserving causality and bounded resource use. The optimization unit is the delivery constraint: dependency, runner slot, shared environment, artifact, network service, or human gate. A visually complex graph with the same critical path and higher compute consumption is worse, not more advanced.
14. Common misconceptions
- “More jobs means a faster pipeline.” Only if the jobs are independent and capacity exists.
-
“
needsmeans download every earlier artifact.” No. Withneeds, artifact flow is explicit to needed jobs. -
“A resource group orders pipelines oldest first.”
Not by default;
unorderedis the default process mode. -
“
optional: trueignores a failed producer.” It handles an absent producer, not a failed required job. - “Parallel matrix dimensions are free metadata.” Their Cartesian product creates real jobs and consumes runner/instance capacity.
Knowledge check
What does a needs edge change compared with
stages?
It changes job eligibility from waiting for whole earlier stages to waiting for explicitly listed predecessor jobs.
A matrix expands to eight jobs but only two runner slots exist. What is the maximum number that can execute simultaneously?
At most two from runner capacity, and possibly fewer if tags, protection, or resource-group constraints apply.
Why can a pipeline fail to be created when
needs points to a job controlled by
rules?
The producer may be absent from the graph. Make the edge
optional: true when absence is valid, or align the
jobs' rules.
What does
resource_group: production protect?
It creates mutual exclusion for jobs using that key in the same project so they do not concurrently mutate the named resource.
After adding needs, a package job can no longer
find an artifact from an unrelated previous-stage job.
Why?
A job using needs only downloads artifacts from needed jobs; the refactor changed the data-flow contract.
Why is a shorter YAML file not proof of a better DAG?
Performance depends on the critical path, runner capacity, external constraints, and correctness—not configuration length.
Summary
Stages create broad barriers; needs creates explicit
DAG dependencies. parallel and matrices create multiple
job instances but runner capacity determines real concurrency.
resource_group serializes shared mutable resources, and
interruptible can reclaim capacity from superseded safe
work. Every optimization must re-prove artifact flow,
conditional-job existence, critical path, and external-resource
safety.
Official references
- GitLab Docs — Make jobs start earlier with needs
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — needs keyword
- GitLab Docs — needs:artifacts
- GitLab Docs — needs:optional
- GitLab Docs — needs:parallel:matrix
- GitLab Docs — parallel
- GitLab Docs — parallel:matrix
- GitLab Docs — Matrix expressions
- GitLab Docs — Resource groups
- GitLab Docs — resource_group keyword
- GitLab Docs — interruptible keyword
- GitLab Docs — Auto-cancel redundant pipelines
- GitLab Docs — Pipeline efficiency
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — Resource groups API
- GitLab Docs — Runner advanced configuration
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.