Chapter 15Lesson 03~235 minutes

DAG Pipelines, needs, Parallel Jobs, Matrices, Resource Groups, and Concurrency: Configuration, Design Choices, and Tradeoffs

Turn DAG mechanics into maintainable policy choices for dependency precision, matrix breadth, runner capacity, serialization, and cancellation.

DesignRunner capacityCostSerializationinterruptibleTradeoffs

Learning objectives

  • Choose between stage barriers and DAG edges based on dependency volatility and auditability.
  • Size parallel/matrix work against runner capacity and external-system limits rather than syntax maxima.
  • Select resource-group key granularity and process mode deliberately.
  • Use interruptible and auto-cancel only for work that is safe to terminate.
  • Justify an optimization across maintainability, reliability, security, performance, and cost.
Availability baseline (verified 2026-08-21 against current GitLab documentation). The mandatory mechanisms in this chapter—needs, parallel, parallel:matrix, resource_group, and interruptible—are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. A job may list at most 50 needs entries. Numeric parallel accepts 1–200 instances, and parallel:matrix can create at most 200 permutations. Those are configuration limits, not guarantees of simultaneous execution: runner capacity, tags, protection, resource groups, and external-service limits can keep jobs pending. Hosted-runner quota/billing can change, so every required exercise includes a CI-Lint/fixture path that does not require paid compute.

1. Design principle: make constraints explicit, but no more explicit than necessary

A simple stage pipeline is not “less DevOps” than a DAG. It is often the best representation when jobs truly share a broad barrier or when the pipeline is small enough that a few seconds of waiting cost less than a dense dependency graph. DAG precision becomes valuable when false barriers materially delay feedback or when different product paths can proceed independently.

The design question is therefore not “Can we use needs?” but “Which dependency is stable enough to encode, and what operational benefit does the edge buy?”

2. Stage model versus DAG

Question Prefer stages Prefer DAG / needs
Dependency shape Most jobs genuinely require the whole prior phase. Jobs form independent chains/fan-out/fan-in paths.
Change frequency Pipeline is small and changes rarely. Long-running paths evolve independently.
Performance value Queue/startup dominates; barrier removal saves little. False waits sit on the critical path.
Cognitive load Beginner or small team benefits from one broad order. Team can maintain explicit edges and artifact contracts.
Failure isolation Broad “build then test” gate is intentional. A failing unrelated path should not block useful feedback elsewhere.

3. Parallelism versus runner capacity and cost

GitLab allows up to 200 numeric parallel instances and 200 matrix permutations, but those maxima are guardrails, not targets. Increasing width can raise hosted-runner consumption, self-managed fleet size, cache pressure, registry pulls, network egress, test-account concurrency, and API rate pressure.

Use an evidence loop: measure serial duration → identify divisible work → estimate ideal speedup → cap by runner/external slots → run a small experiment → compare saved lead time with added compute and operational complexity. If eight shards save 15 seconds while doubling flakiness and queue contention, two shards may be the better design.

4. The non-parallel remainder limits speedup

Even without formal performance theory, one fact matters: serial setup, packaging, or deployment remains on the critical path. If a 10-minute pipeline contains seven minutes of parallelizable tests but three minutes of serial setup/package work, infinite test shards cannot reduce the pipeline below those three minutes plus overhead. Optimize the longest constrained path, not the easiest YAML to split.

5. Matrix breadth: coverage has a cost surface

A matrix should encode dimensions whose cross-product is actually meaningful. Testing three operating systems × four runtime versions × two database versions creates 24 jobs. If only specific supported combinations matter, list those combinations instead of the full Cartesian product.

Matrix choice Benefit Risk Policy
Full Cartesian product Maximum compatibility coverage. Duplicate/low-value work and high compute. Use only when every combination is supported and consequential.
Curated combinations Controls cost and noise. Can miss unexpected interactions. Tie combinations to a documented support matrix.
One dimension per MR; broad nightly Fast developer feedback plus periodic breadth. Different pipeline sources have different evidence. Make source/schedule semantics explicit.

6. Matrix dependencies: all-to-all versus one-to-one

A downstream job that merely names a matrix producer in needs waits for all instances. That can accidentally turn a matrix back into a barrier. Use needs:parallel:matrix when each downstream instance should wait only for its corresponding producer. Explicit mappings are stable and easy to audit; current matrix expressions can reduce repetition but are compile-time string substitution and should not be treated as a general programming language.

7. Resource-group key granularity

Key design Effect Typical symptom
production for every service Safe but broad serialization. Long deployment queue despite independent targets.
production/$SERVICE One deployment at a time per service. Higher throughput while preserving per-service exclusivity.
device/$DEVICE_ID One job per physical fixture/device. Good fit for scarce hardware.
Unique key per pipeline/job No practical serialization. Race conditions return because jobs never contend on one key.

The key should name the mutable thing that cannot safely handle concurrent writers. Do not use a resource group merely to “slow CI down.”

8. Queue order is a deployment policy

unordered maximizes simplicity but does not guarantee pipeline order. oldest_first preserves chronological order at some throughput cost. newest_first and newest_ready_first prioritize newer pipelines and require idempotent jobs because older deployment work may execute later or be skipped by your surrounding policy. Select a mode based on release semantics, not aesthetics.

Deadlock warning. Resource groups can create deadlock-like waits when a parent pipeline holds or waits on a resource while synchronously waiting for a child pipeline that needs the same resource. Model parent/child waiting direction before combining resource locks with downstream pipeline strategies.

9. Interruptibility: reclaim stale capacity without corrupting state

Job type Usually interruptible? Reason
Lint/unit test Yes Repeatable and safe to restart for a newer commit.
Deterministic build before publication Often Output is tied to SHA and can be recreated.
Database migration Usually no Cancellation can leave partial external mutation.
Deployment with transactional/rollback design Depends Only if the deployment tool explicitly supports safe interruption.

Under conservative auto-cancel behavior, once a non-interruptible job has started, the pipeline may no longer be considered safely cancelable. Treat this as part of concurrency design: stale matrices should not occupy capacity, but mutation jobs should not be terminated merely to make dashboards look efficient.

10. DAG edges are also data contracts

Every needs edge should answer two questions: “Do I need completion?” and “Do I need artifacts?” If a test only needs a build's completion but downloads a 2 GB artifact unnecessarily, the DAG may become slower and more expensive. Use artifacts: false when the edge is ordering-only. Conversely, if a consumer needs a file, prove the producer publishes it and the edge downloads it.

11. Security: concurrency multiplies side effects

Parallel jobs amplify whatever permissions and network reach each job possesses. Four safe read-only tests are different from four jobs sharing a mutable cloud account, database, package namespace, or runner host workspace. Before widening a matrix, inventory external writes, credentials, rate-limited APIs, cache keys, and protected-variable availability. Concurrency is a blast-radius multiplier when job boundaries are weak.

12. Worked scenario: payments monorepo

A team has API, UI, and docs components. API tests need the API build; UI tests need the UI build; docs lint is independent; production deployment mutates one environment. The team has four runner slots.

Choice Decision Justification
Build/test ordering DAG per component Removes false waits without connecting unrelated components.
Compatibility tests 2×2 curated matrix Four jobs fit runner capacity and cover supported combinations.
Docs lint needs: [], interruptible Fast feedback; safe to cancel for newer commits.
Production deploy resource_group: production, non-interruptible One writer to production; avoid mid-deploy cancellation.
Artifact edges Only component-specific artifacts Avoid broad downloads and preserve provenance.

This design is not maximally parallel. It is intentionally bounded by four runner slots and one production writer.

13. Design review questions

  • Which edges express true data/order dependencies, and which only reproduce old stage boundaries?
  • What is the current critical path, and which proposed edge actually shortens it?
  • How many jobs can become runnable at once? How many runners and external slots are safe?
  • Which jobs write shared state and therefore need a resource key?
  • Which jobs are safe to cancel when superseded?
  • After adding needs, which artifacts stop downloading automatically?

Knowledge check

When is a stage-only design preferable to a DAG?

Why should you not size a matrix to the syntax maximum?

A deployment resource-group key is unique per pipeline. What is wrong?

Why is interruptible: true risky on a migration job?

What extra contract does a needs edge carry when artifacts are required?

Summary

Choose DAG complexity only where explicit dependencies create measurable value. Bound matrix width by useful coverage and actual capacity. Name resource groups after the mutable resource, select queue order according to deployment semantics, and mark only safely repeatable work as interruptible. Every concurrency decision is simultaneously a performance, reliability, security, and cost decision.

Official references

Next lesson

Debug the graph when “more parallel” becomes slower or unsafe

Lesson 4 engineers missing producers, artifact regressions, runner saturation, bad resource keys, and a critical-path regression, then repairs each with the least destructive control.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.