DAG Pipelines, needs, Parallel Jobs, Matrices, Resource Groups, and Concurrency: Diagnostics, Failure Modes, Security, and Performance
Preserve evidence and distinguish graph-creation, capacity, artifact, and resource-lock failures before applying the smallest correction.
Learning objectives
- Apply a repeatable evidence-first diagnostic sequence to DAG/concurrency failures.
-
Repair a missing conditional producer with aligned rules or
needs:optional. - Distinguish runner saturation from dependency or resource-group waiting.
- Recover artifact flow broken by a DAG conversion.
- Detect a “faster-looking” graph whose critical path or external contention actually worsened.
needs,
parallel, parallel:matrix,
resource_group, and interruptible—are
available in GitLab Free/Premium/Ultimate across GitLab.com,
Self-Managed, and Dedicated. A job may list at most 50
needs entries. Numeric parallel accepts
1–200 instances, and parallel:matrix can create at most
200 permutations. Those are configuration limits, not guarantees of
simultaneous execution: runner capacity, tags, protection, resource
groups, and external-service limits can keep jobs pending.
Hosted-runner quota/billing can change, so every required exercise
includes a CI-Lint/fixture path that does not require paid compute.
1. Diagnostic sequence: preserve → scope → graph → capacity → data → resource → verify
- Preserve evidence: pipeline ID, SHA, source, YAML revision, failed creation message, job timestamps/statuses, and relevant runner/resource metadata.
- Scope: identify project/ref/MR/pipeline/job and whether the problem is creation-time or execution-time.
-
Graph: list required
needsedges and which conditional jobs actually exist. - Capacity: inspect eligible runner slots, tags, protected state, and pending duration.
- Data: list artifacts each consumer expects and which needed producers provide them.
-
Resource: identify
resource_groupkeys and current holder/queue behavior. - Correct minimally: change the smallest edge/rule/key/capacity assumption, then rerun and compare evidence.
2. Failure: needs points to a job excluded by
rules
Intentionally broken configuration:
security_scan:
rules:
- if: '$RUN_SECURITY_SCAN == "true"'
script: echo scan
package:
needs: [security_scan]
script: echo package
With RUN_SECURITY_SCAN unset/false,
security_scan is not added. GitLab validates DAG
relationships during pipeline creation, so the pipeline can fail
with a message equivalent to:
'package' job needs 'security_scan' job, but it was not added to
the pipeline.
Repair A: if package may proceed without the scan,
make the dependency optional. Repair B: if package
must never exist without the scan, give both jobs aligned rules. Do
not add allow_failure; that addresses runtime failure,
not graph absence.
package:
needs:
- job: security_scan
optional: true
script: echo package
3. Failure: parallel jobs overload runner capacity
Symptoms: many matrix instances are pending; started
jobs run normally; no dependency error exists. Inspect job
queued_duration, runner tags, runner status, and fleet
concurrency. A 20-job matrix with two runner slots is a queue
generator, not a 20-way speedup.
Least-destructive fixes depend on intent: reduce matrix breadth, run broader coverage on a different pipeline source/schedule, add capacity only if justified, or optimize job duration. Do not remove tests merely to clear the queue without documenting coverage loss.
4. Failure: runners are available but the external test system collapses
All matrix jobs start, but tests fail with timeouts/rate-limit
errors because they share one test tenant or API. This is not a
GitLab runner problem. Preserve the external error and identify the
shared resource. Possible corrections include a narrower matrix,
per-instance isolated test resources, client-side rate control, or
resource_group around jobs that truly require mutual
exclusion.
Do not serialize the entire pipeline if only one external operation needs serialization.
5. Failure: resource-group key is too broad
API and UI deployments target independent environments but both use
resource_group: production. They queue behind one
another despite being safe to run concurrently. Rename keys to the
actual resources, for example production/api and
production/ui, only after confirming there is no hidden
shared database/migration lock.
6. Failure: resource-group key is too narrow
Each deployment uses
resource_group: production/$CI_PIPELINE_ID. Every
pipeline gets a different key, so two pipelines can deploy
production concurrently. The YAML looks sophisticated but mutual
exclusion is gone. Replace it with a stable key that identifies the
shared production target.
7. Failure: queue order is mistaken for dependency order
With the default unordered process mode, waiting
deployment jobs are mutually exclusive but not guaranteed
oldest-first. If release correctness depends on order, inspect the
resource group's process_mode through the API and
choose an explicit mode. For newest-oriented modes, make deployment
jobs idempotent and define how obsolete pipelines are handled.
8. Failure: DAG conversion breaks an artifact assumption
Before conversion, a test in a later stage implicitly downloaded
artifacts from earlier jobs. After adding
needs: [build_api], it no longer sees
schema.json produced by generate_schema.
The original pipeline depended on an implicit data path that the new
DAG did not encode.
Repair by adding the real producer to needs with
artifacts enabled, or by changing the design so one canonical
producer creates the artifact. Do not copy the file through cache
just to make the symptom disappear; cache is not authoritative
job-to-job transport.
contract_test:
needs:
- job: build_api
artifacts: true
- job: generate_schema
artifacts: true
script:
- test -f schema.json
9. Failure: artifacts from parallel instances overwrite one another
If a consumer needs a parallelized producer without selecting a
subset, it can download artifacts from all instances. When each
instance publishes the same path, later downloads can overwrite
earlier ones. Give artifacts unique names/paths derived from matrix
identifiers or use needs:parallel:matrix to select only
the instance the consumer requires.
10. Failure: more parallel jobs lengthen the critical path
A team splits a 6-minute test into six 2-minute shards, but each shard independently performs a 90-second dependency setup and all compete for two runners. Theoretical job work grows from 6 minutes to 21 minutes of runner time. Queueing plus repeated setup can make pipeline completion slower.
Measure end-to-end start/finish and runner consumption, not only individual shard duration. Consider a shared build artifact, fewer shards, or moving setup to a predecessor job.
11. Failure: stale pipelines consume every runner slot
New commits arrive faster than a wide matrix completes. If those
tests are deterministic and safe to stop, mark them
interruptible: true and verify project auto-cancel
behavior. Keep deployment/migration jobs non-interruptible unless
their tooling guarantees safe cancellation.
The evidence is a reduction in running/pending jobs from superseded pipeline IDs—not simply a “canceled” badge.
12. Deadlock-like resource waits
A classic resource-group mistake appears with parent/child
pipelines: a parent deployment job or trigger waits synchronously
for a child pipeline while both need the same resource key under an
ordering mode such as oldest_first. Neither side can
make progress as designed. Draw the wait-for graph: which job
holds/waits for the resource, which pipeline waits for which
downstream result, and whether the lock should live only at the
actual mutation point.
13. Intentionally broken example: preserve the creation error
Use CI Lint against the missing-producer example. Save the
validation response/message before fixing it. Then add
optional: true only if the business logic allows
absence, validate again, and compare responses. This proves the
original cause rather than obscuring it with retries or unrelated
runner changes.
14. Security and cost are causal here
Wide concurrency can multiply simultaneous use of job tokens, protected variables, cloud identities, external API calls, and shared caches. If one job is over-privileged, 50 parallel instances create 50 concurrent opportunities to misuse that authority. Least privilege, protected-ref rules, runner isolation, and bounded matrices therefore matter directly to concurrency safety.
Likewise, resource-group serialization can protect scarce/critical targets but can also create long queues that waste upstream compute if pipelines continue building work that will never deploy. Combine concurrency controls with deliberate cancellation and pipeline-source rules, not generic warning banners.
15. Verification checklist
-
The pipeline creates successfully and every non-optional
needstarget exists. - Pending jobs are classified as dependency wait, runner wait, or resource-group wait.
- Each consumer receives only the artifacts it needs.
- Matrix width matches runner/external capacity assumptions.
- Resource-group keys map one-to-one to real mutually exclusive resources.
- Critical-path timing improves or remains intentionally unchanged after the fix.
Knowledge check
A pipeline fails before any job starts with a missing
needs target. Should you add runner
capacity?
No. This is a configuration/graph-creation failure; inspect rules and use aligned rules or optional needs where absence is valid.
Eight matrix jobs are pending while two run successfully. What is the likely first capacity check?
Eligible runner slots and tags/protection; the graph may allow more concurrency than the fleet can execute.
Two independent services queue on the same resource group. What might be wrong?
The resource-group key may be too broad and serialize independent resources.
A DAG consumer lost an artifact after adding
needs. What changed?
Artifact downloads became limited to needed jobs, so the missing producer/data edge must be modeled explicitly.
Why can more shards make a pipeline slower?
Repeated setup, runner contention, scheduling overhead, and external contention can increase total work and the effective critical path.
What is wrong with fixing a missing producer by setting
allow_failure?
The producer is absent at pipeline creation; allow_failure applies to a job that exists and runs/fails.
Summary
DAG diagnostics begin before runners: validate that required graph nodes exist. Then separate dependency wait from capacity wait and resource locking, re-prove artifact flow, and measure the true critical path. The safest correction is the narrowest one that restores the intended dependency/resource contract.
Official references
- GitLab Docs — Make jobs start earlier with needs
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — needs keyword
- GitLab Docs — needs:artifacts
- GitLab Docs — needs:optional
- GitLab Docs — needs:parallel:matrix
- GitLab Docs — parallel
- GitLab Docs — parallel:matrix
- GitLab Docs — Matrix expressions
- GitLab Docs — Resource groups
- GitLab Docs — resource_group keyword
- GitLab Docs — interruptible keyword
- GitLab Docs — Auto-cancel redundant pipelines
- GitLab Docs — Pipeline efficiency
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — Resource groups API
- GitLab Docs — Runner advanced configuration
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.