Chapter 15Lesson 04~255 minutes

DAG Pipelines, needs, Parallel Jobs, Matrices, Resource Groups, and Concurrency: Diagnostics, Failure Modes, Security, and Performance

Preserve evidence and distinguish graph-creation, capacity, artifact, and resource-lock failures before applying the smallest correction.

Diagnosticsneeds:optionalArtifactsQueuesRace conditionsPerformance

Learning objectives

  • Apply a repeatable evidence-first diagnostic sequence to DAG/concurrency failures.
  • Repair a missing conditional producer with aligned rules or needs:optional.
  • Distinguish runner saturation from dependency or resource-group waiting.
  • Recover artifact flow broken by a DAG conversion.
  • Detect a “faster-looking” graph whose critical path or external contention actually worsened.
Availability baseline (verified 2026-08-21 against current GitLab documentation). The mandatory mechanisms in this chapter—needs, parallel, parallel:matrix, resource_group, and interruptible—are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. A job may list at most 50 needs entries. Numeric parallel accepts 1–200 instances, and parallel:matrix can create at most 200 permutations. Those are configuration limits, not guarantees of simultaneous execution: runner capacity, tags, protection, resource groups, and external-service limits can keep jobs pending. Hosted-runner quota/billing can change, so every required exercise includes a CI-Lint/fixture path that does not require paid compute.

1. Diagnostic sequence: preserve → scope → graph → capacity → data → resource → verify

  1. Preserve evidence: pipeline ID, SHA, source, YAML revision, failed creation message, job timestamps/statuses, and relevant runner/resource metadata.
  2. Scope: identify project/ref/MR/pipeline/job and whether the problem is creation-time or execution-time.
  3. Graph: list required needs edges and which conditional jobs actually exist.
  4. Capacity: inspect eligible runner slots, tags, protected state, and pending duration.
  5. Data: list artifacts each consumer expects and which needed producers provide them.
  6. Resource: identify resource_group keys and current holder/queue behavior.
  7. Correct minimally: change the smallest edge/rule/key/capacity assumption, then rerun and compare evidence.

2. Failure: needs points to a job excluded by rules

Intentionally broken configuration:

security_scan:
  rules:
    - if: '$RUN_SECURITY_SCAN == "true"'
  script: echo scan

package:
  needs: [security_scan]
  script: echo package

With RUN_SECURITY_SCAN unset/false, security_scan is not added. GitLab validates DAG relationships during pipeline creation, so the pipeline can fail with a message equivalent to: 'package' job needs 'security_scan' job, but it was not added to the pipeline.

Repair A: if package may proceed without the scan, make the dependency optional. Repair B: if package must never exist without the scan, give both jobs aligned rules. Do not add allow_failure; that addresses runtime failure, not graph absence.

package:
  needs:
    - job: security_scan
      optional: true
  script: echo package

3. Failure: parallel jobs overload runner capacity

Symptoms: many matrix instances are pending; started jobs run normally; no dependency error exists. Inspect job queued_duration, runner tags, runner status, and fleet concurrency. A 20-job matrix with two runner slots is a queue generator, not a 20-way speedup.

Least-destructive fixes depend on intent: reduce matrix breadth, run broader coverage on a different pipeline source/schedule, add capacity only if justified, or optimize job duration. Do not remove tests merely to clear the queue without documenting coverage loss.

4. Failure: runners are available but the external test system collapses

All matrix jobs start, but tests fail with timeouts/rate-limit errors because they share one test tenant or API. This is not a GitLab runner problem. Preserve the external error and identify the shared resource. Possible corrections include a narrower matrix, per-instance isolated test resources, client-side rate control, or resource_group around jobs that truly require mutual exclusion.

Do not serialize the entire pipeline if only one external operation needs serialization.

5. Failure: resource-group key is too broad

API and UI deployments target independent environments but both use resource_group: production. They queue behind one another despite being safe to run concurrently. Rename keys to the actual resources, for example production/api and production/ui, only after confirming there is no hidden shared database/migration lock.

6. Failure: resource-group key is too narrow

Each deployment uses resource_group: production/$CI_PIPELINE_ID. Every pipeline gets a different key, so two pipelines can deploy production concurrently. The YAML looks sophisticated but mutual exclusion is gone. Replace it with a stable key that identifies the shared production target.

7. Failure: queue order is mistaken for dependency order

With the default unordered process mode, waiting deployment jobs are mutually exclusive but not guaranteed oldest-first. If release correctness depends on order, inspect the resource group's process_mode through the API and choose an explicit mode. For newest-oriented modes, make deployment jobs idempotent and define how obsolete pipelines are handled.

8. Failure: DAG conversion breaks an artifact assumption

Before conversion, a test in a later stage implicitly downloaded artifacts from earlier jobs. After adding needs: [build_api], it no longer sees schema.json produced by generate_schema. The original pipeline depended on an implicit data path that the new DAG did not encode.

Repair by adding the real producer to needs with artifacts enabled, or by changing the design so one canonical producer creates the artifact. Do not copy the file through cache just to make the symptom disappear; cache is not authoritative job-to-job transport.

contract_test:
  needs:
    - job: build_api
      artifacts: true
    - job: generate_schema
      artifacts: true
  script:
    - test -f schema.json

9. Failure: artifacts from parallel instances overwrite one another

If a consumer needs a parallelized producer without selecting a subset, it can download artifacts from all instances. When each instance publishes the same path, later downloads can overwrite earlier ones. Give artifacts unique names/paths derived from matrix identifiers or use needs:parallel:matrix to select only the instance the consumer requires.

10. Failure: more parallel jobs lengthen the critical path

A team splits a 6-minute test into six 2-minute shards, but each shard independently performs a 90-second dependency setup and all compete for two runners. Theoretical job work grows from 6 minutes to 21 minutes of runner time. Queueing plus repeated setup can make pipeline completion slower.

Measure end-to-end start/finish and runner consumption, not only individual shard duration. Consider a shared build artifact, fewer shards, or moving setup to a predecessor job.

11. Failure: stale pipelines consume every runner slot

New commits arrive faster than a wide matrix completes. If those tests are deterministic and safe to stop, mark them interruptible: true and verify project auto-cancel behavior. Keep deployment/migration jobs non-interruptible unless their tooling guarantees safe cancellation.

The evidence is a reduction in running/pending jobs from superseded pipeline IDs—not simply a “canceled” badge.

12. Deadlock-like resource waits

A classic resource-group mistake appears with parent/child pipelines: a parent deployment job or trigger waits synchronously for a child pipeline while both need the same resource key under an ordering mode such as oldest_first. Neither side can make progress as designed. Draw the wait-for graph: which job holds/waits for the resource, which pipeline waits for which downstream result, and whether the lock should live only at the actual mutation point.

13. Intentionally broken example: preserve the creation error

Use CI Lint against the missing-producer example. Save the validation response/message before fixing it. Then add optional: true only if the business logic allows absence, validate again, and compare responses. This proves the original cause rather than obscuring it with retries or unrelated runner changes.

14. Security and cost are causal here

Wide concurrency can multiply simultaneous use of job tokens, protected variables, cloud identities, external API calls, and shared caches. If one job is over-privileged, 50 parallel instances create 50 concurrent opportunities to misuse that authority. Least privilege, protected-ref rules, runner isolation, and bounded matrices therefore matter directly to concurrency safety.

Likewise, resource-group serialization can protect scarce/critical targets but can also create long queues that waste upstream compute if pipelines continue building work that will never deploy. Combine concurrency controls with deliberate cancellation and pipeline-source rules, not generic warning banners.

15. Verification checklist

  • The pipeline creates successfully and every non-optional needs target exists.
  • Pending jobs are classified as dependency wait, runner wait, or resource-group wait.
  • Each consumer receives only the artifacts it needs.
  • Matrix width matches runner/external capacity assumptions.
  • Resource-group keys map one-to-one to real mutually exclusive resources.
  • Critical-path timing improves or remains intentionally unchanged after the fix.

Knowledge check

A pipeline fails before any job starts with a missing needs target. Should you add runner capacity?

Eight matrix jobs are pending while two run successfully. What is the likely first capacity check?

Two independent services queue on the same resource group. What might be wrong?

A DAG consumer lost an artifact after adding needs. What changed?

Why can more shards make a pipeline slower?

What is wrong with fixing a missing producer by setting allow_failure?

Summary

DAG diagnostics begin before runners: validate that required graph nodes exist. Then separate dependency wait from capacity wait and resource locking, re-prove artifact flow, and measure the true critical path. The safest correction is the narrowest one that restores the intended dependency/resource contract.

Official references

Next lesson

Prove an optimization end to end

Lesson 5 combines a serialized baseline, DAG conversion, a small matrix, one resource lock, predictions, timestamps, failure diagnosis, evidence capture, and cleanup into the chapter checkpoint.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.