DAG Pipelines, needs, Parallel Jobs, Matrices, Resource Groups, and Concurrency: Configuration, Design Choices, and Tradeoffs
Turn DAG mechanics into maintainable policy choices for dependency precision, matrix breadth, runner capacity, serialization, and cancellation.
Learning objectives
- Choose between stage barriers and DAG edges based on dependency volatility and auditability.
- Size parallel/matrix work against runner capacity and external-system limits rather than syntax maxima.
- Select resource-group key granularity and process mode deliberately.
-
Use
interruptibleand auto-cancel only for work that is safe to terminate. - Justify an optimization across maintainability, reliability, security, performance, and cost.
needs,
parallel, parallel:matrix,
resource_group, and interruptible—are
available in GitLab Free/Premium/Ultimate across GitLab.com,
Self-Managed, and Dedicated. A job may list at most 50
needs entries. Numeric parallel accepts
1–200 instances, and parallel:matrix can create at most
200 permutations. Those are configuration limits, not guarantees of
simultaneous execution: runner capacity, tags, protection, resource
groups, and external-service limits can keep jobs pending.
Hosted-runner quota/billing can change, so every required exercise
includes a CI-Lint/fixture path that does not require paid compute.
1. Design principle: make constraints explicit, but no more explicit than necessary
A simple stage pipeline is not “less DevOps” than a DAG. It is often the best representation when jobs truly share a broad barrier or when the pipeline is small enough that a few seconds of waiting cost less than a dense dependency graph. DAG precision becomes valuable when false barriers materially delay feedback or when different product paths can proceed independently.
The design question is therefore not “Can we use
needs?” but “Which dependency is stable enough to
encode, and what operational benefit does the edge buy?”
2. Stage model versus DAG
| Question | Prefer stages | Prefer DAG / needs |
|---|---|---|
| Dependency shape | Most jobs genuinely require the whole prior phase. | Jobs form independent chains/fan-out/fan-in paths. |
| Change frequency | Pipeline is small and changes rarely. | Long-running paths evolve independently. |
| Performance value | Queue/startup dominates; barrier removal saves little. | False waits sit on the critical path. |
| Cognitive load | Beginner or small team benefits from one broad order. | Team can maintain explicit edges and artifact contracts. |
| Failure isolation | Broad “build then test” gate is intentional. | A failing unrelated path should not block useful feedback elsewhere. |
3. Parallelism versus runner capacity and cost
GitLab allows up to 200 numeric parallel instances and 200 matrix permutations, but those maxima are guardrails, not targets. Increasing width can raise hosted-runner consumption, self-managed fleet size, cache pressure, registry pulls, network egress, test-account concurrency, and API rate pressure.
Use an evidence loop: measure serial duration → identify divisible work → estimate ideal speedup → cap by runner/external slots → run a small experiment → compare saved lead time with added compute and operational complexity. If eight shards save 15 seconds while doubling flakiness and queue contention, two shards may be the better design.
4. The non-parallel remainder limits speedup
Even without formal performance theory, one fact matters: serial setup, packaging, or deployment remains on the critical path. If a 10-minute pipeline contains seven minutes of parallelizable tests but three minutes of serial setup/package work, infinite test shards cannot reduce the pipeline below those three minutes plus overhead. Optimize the longest constrained path, not the easiest YAML to split.
5. Matrix breadth: coverage has a cost surface
A matrix should encode dimensions whose cross-product is actually meaningful. Testing three operating systems × four runtime versions × two database versions creates 24 jobs. If only specific supported combinations matter, list those combinations instead of the full Cartesian product.
| Matrix choice | Benefit | Risk | Policy |
|---|---|---|---|
| Full Cartesian product | Maximum compatibility coverage. | Duplicate/low-value work and high compute. | Use only when every combination is supported and consequential. |
| Curated combinations | Controls cost and noise. | Can miss unexpected interactions. | Tie combinations to a documented support matrix. |
| One dimension per MR; broad nightly | Fast developer feedback plus periodic breadth. | Different pipeline sources have different evidence. | Make source/schedule semantics explicit. |
6. Matrix dependencies: all-to-all versus one-to-one
A downstream job that merely names a matrix producer in
needs waits for all instances. That can accidentally
turn a matrix back into a barrier. Use
needs:parallel:matrix when each downstream instance
should wait only for its corresponding producer. Explicit mappings
are stable and easy to audit; current matrix expressions can reduce
repetition but are compile-time string substitution and should not
be treated as a general programming language.
7. Resource-group key granularity
| Key design | Effect | Typical symptom |
|---|---|---|
production for every service |
Safe but broad serialization. | Long deployment queue despite independent targets. |
production/$SERVICE |
One deployment at a time per service. | Higher throughput while preserving per-service exclusivity. |
device/$DEVICE_ID |
One job per physical fixture/device. | Good fit for scarce hardware. |
| Unique key per pipeline/job | No practical serialization. | Race conditions return because jobs never contend on one key. |
The key should name the mutable thing that cannot safely handle concurrent writers. Do not use a resource group merely to “slow CI down.”
8. Queue order is a deployment policy
unordered maximizes simplicity but does not guarantee
pipeline order. oldest_first preserves chronological
order at some throughput cost. newest_first and
newest_ready_first prioritize newer pipelines and
require idempotent jobs because older deployment work may execute
later or be skipped by your surrounding policy. Select a mode based
on release semantics, not aesthetics.
9. Interruptibility: reclaim stale capacity without corrupting state
| Job type | Usually interruptible? | Reason |
|---|---|---|
| Lint/unit test | Yes | Repeatable and safe to restart for a newer commit. |
| Deterministic build before publication | Often | Output is tied to SHA and can be recreated. |
| Database migration | Usually no | Cancellation can leave partial external mutation. |
| Deployment with transactional/rollback design | Depends | Only if the deployment tool explicitly supports safe interruption. |
Under conservative auto-cancel behavior, once a non-interruptible job has started, the pipeline may no longer be considered safely cancelable. Treat this as part of concurrency design: stale matrices should not occupy capacity, but mutation jobs should not be terminated merely to make dashboards look efficient.
10. DAG edges are also data contracts
Every needs edge should answer two questions: “Do I
need completion?” and “Do I need artifacts?” If a test only needs a
build's completion but downloads a 2 GB artifact unnecessarily, the
DAG may become slower and more expensive. Use
artifacts: false when the edge is ordering-only.
Conversely, if a consumer needs a file, prove the producer publishes
it and the edge downloads it.
11. Security: concurrency multiplies side effects
Parallel jobs amplify whatever permissions and network reach each job possesses. Four safe read-only tests are different from four jobs sharing a mutable cloud account, database, package namespace, or runner host workspace. Before widening a matrix, inventory external writes, credentials, rate-limited APIs, cache keys, and protected-variable availability. Concurrency is a blast-radius multiplier when job boundaries are weak.
12. Worked scenario: payments monorepo
A team has API, UI, and docs components. API tests need the API build; UI tests need the UI build; docs lint is independent; production deployment mutates one environment. The team has four runner slots.
| Choice | Decision | Justification |
|---|---|---|
| Build/test ordering | DAG per component | Removes false waits without connecting unrelated components. |
| Compatibility tests | 2×2 curated matrix | Four jobs fit runner capacity and cover supported combinations. |
| Docs lint | needs: [], interruptible |
Fast feedback; safe to cancel for newer commits. |
| Production deploy |
resource_group: production, non-interruptible
|
One writer to production; avoid mid-deploy cancellation. |
| Artifact edges | Only component-specific artifacts | Avoid broad downloads and preserve provenance. |
This design is not maximally parallel. It is intentionally bounded by four runner slots and one production writer.
13. Design review questions
- Which edges express true data/order dependencies, and which only reproduce old stage boundaries?
- What is the current critical path, and which proposed edge actually shortens it?
- How many jobs can become runnable at once? How many runners and external slots are safe?
- Which jobs write shared state and therefore need a resource key?
- Which jobs are safe to cancel when superseded?
-
After adding
needs, which artifacts stop downloading automatically?
Knowledge check
When is a stage-only design preferable to a DAG?
When broad phase barriers reflect real dependencies and the simplicity benefit outweighs negligible waiting.
Why should you not size a matrix to the syntax maximum?
Real capacity, cost, external service limits, and coverage value—not the 200-permutation limit—should determine width.
A deployment resource-group key is unique per pipeline. What is wrong?
Jobs no longer contend on one shared key, so the key does not serialize access to the actual shared resource.
Why is interruptible: true risky on a migration
job?
Auto-cancel could stop the job after partial external mutation, leaving inconsistent state.
What extra contract does a needs edge carry when
artifacts are required?
The producer must publish the expected artifact and the consumer must explicitly receive it through the needs artifact semantics.
Summary
Choose DAG complexity only where explicit dependencies create measurable value. Bound matrix width by useful coverage and actual capacity. Name resource groups after the mutable resource, select queue order according to deployment semantics, and mark only safely repeatable work as interruptible. Every concurrency decision is simultaneously a performance, reliability, security, and cost decision.
Official references
- GitLab Docs — Make jobs start earlier with needs
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — needs keyword
- GitLab Docs — needs:artifacts
- GitLab Docs — needs:optional
- GitLab Docs — needs:parallel:matrix
- GitLab Docs — parallel
- GitLab Docs — parallel:matrix
- GitLab Docs — Matrix expressions
- GitLab Docs — Resource groups
- GitLab Docs — resource_group keyword
- GitLab Docs — interruptible keyword
- GitLab Docs — Auto-cancel redundant pipelines
- GitLab Docs — Pipeline efficiency
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — Resource groups API
- GitLab Docs — Runner advanced configuration
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.