Chapter 15Lesson 02~270 minutes

DAG Pipelines, needs, Parallel Jobs, Matrices, Resource Groups, and Concurrency: Guided Hands-On Workflow and Core Operations

Measure a serialized pipeline, replace false barriers with needs, inspect matrix expansion, prove a resource lock, and calculate critical path from evidence.

Hands-onneedsparallel:matrixresource_groupTimestampsCI Lint

Learning objectives

  • Establish a serialized baseline and predict its theoretical duration before changing dependencies.
  • Convert the baseline to a needs DAG and verify start times from the API/UI.
  • Expand a safe 2×2 matrix and distinguish generated jobs from actual simultaneous execution.
  • Use a synthetic resource_group to prove mutual exclusion without touching a real environment.
  • Calculate critical path and runner-slot demand from observed evidence.
Availability baseline (verified 2026-08-21 against current GitLab documentation). The mandatory mechanisms in this chapter—needs, parallel, parallel:matrix, resource_group, and interruptible—are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. A job may list at most 50 needs entries. Numeric parallel accepts 1–200 instances, and parallel:matrix can create at most 200 permutations. Those are configuration limits, not guarantees of simultaneous execution: runner capacity, tags, protection, resource groups, and external-service limits can keep jobs pending. Hosted-runner quota/billing can change, so every required exercise includes a CI-Lint/fixture path that does not require paid compute.

1. Disposable scenario and preflight

Use a throwaway GitLab Free project such as gitlab-ch15-dag-lab or a clearly disposable branch in an existing lab project. Do not use a production repository because this lesson intentionally changes pipeline scheduling.

Preflight item Required state Why
Project/ref Disposable project or branch ch15/dag-lab. Prevents experiments from changing real merge/release behavior.
Role Developer can commit/run jobs; Maintainer is useful only for project settings/API inspection. No admin privileges are required.
Runner At least one eligible runner for live timing, or use CI Lint + supplied fixture. Hosted compute may be unavailable/quota constrained.
External systems None. Jobs use sleep, printf, and tiny artifacts only. No network service, deployment, or secret is required.
Evidence Record commit SHA, pipeline ID, job IDs/timestamps, runner IDs/descriptions if visible. Binds performance claims to one configuration revision.

2. Baseline: deliberately stage-serialized

Create the following configuration. The sleep times are intentionally short in a live lab; the comments also show the conceptual “minutes” used for critical-path arithmetic. Keep the live sleeps under a few seconds to minimize compute consumption.

stages: [build, test, package]

api_build:
  stage: build
  script:
    - sleep 3
    - printf 'api-build=%s\n' "$CI_COMMIT_SHA" > api.txt
  artifacts:
    paths: [api.txt]

ui_build:
  stage: build
  script:
    - sleep 12
    - printf 'ui-build=%s\n' "$CI_COMMIT_SHA" > ui.txt
  artifacts:
    paths: [ui.txt]

api_test:
  stage: test
  script:
    - sleep 7
    - test -f api.txt

ui_test:
  stage: test
  script:
    - sleep 4
    - test -f ui.txt

package:
  stage: package
  script:
    - sleep 3
    - printf 'package=%s\n' "$CI_COMMIT_SHA"
Prediction before running. With unlimited runner slots, stage barriers produce a conceptual lower bound of max(3,12) + max(7,4) + 3 = 22 time units. In a real hosted run, queue/startup overhead adds to that number.

3. Validate before spending compute

Use Build → Pipeline editor → Validate (CI Lint) and inspect the merged configuration. If no runner is available, this validation plus the supplied timeline fixture later in the lesson is the mandatory completion path.

git switch -c ch15/dag-lab
git add .gitlab-ci.yml
git diff --cached --check
git diff --cached
git commit -m "ch15: add serialized baseline"
BASE_SHA="$(git rev-parse HEAD)"
printf 'baseline_sha=%s\n' "$BASE_SHA"
git push -u origin ch15/dag-lab

Do not benchmark by rerunning the same job repeatedly until one looks fast. Record one representative pipeline and note queue time separately from script time.

4. Inspect the baseline timeline

PROJECT_ID="12345678"
PIPELINE_ID="123456789"

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
  --jq '.[] | {id,name,stage,status,created_at,started_at,finished_at,queued_duration,duration}'

Expected relationship: both build jobs become eligible together. Even if api_build finishes early, api_test remains blocked until ui_build completes because the build stage has not cleared. If only one runner slot exists, execution is even more serialized; record that capacity fact rather than blaming stages alone.

5. Convert to explicit needs

Now describe actual dependencies. Keep stages for readability, but let DAG edges bypass unrelated stage work.

api_test:
  stage: test
  needs:
    - job: api_build
      artifacts: true
  script:
    - sleep 7
    - test -f api.txt

ui_test:
  stage: test
  needs:
    - job: ui_build
      artifacts: true
  script:
    - sleep 4
    - test -f ui.txt

package:
  stage: package
  needs:
    - api_test
    - ui_test
  script:
    - sleep 3
    - printf 'package=%s\n' "$CI_COMMIT_SHA"

The package job does not need the build artifacts directly, only completion of both test jobs. If package really consumes files from builds, model those artifact dependencies explicitly instead of relying on old stage behavior.

New prediction. With enough runner slots the two chains are 3+7=10 and 12+4=16; package waits for the longer chain and adds 3, giving a conceptual lower bound of 19. The DAG removes three units of false stage waiting.

6. Prove causality with timestamps, not the graph picture alone

Commit the DAG version, run one pipeline, and capture the same API fields. Verify that api_test.started_at is after api_build.finished_at but can be before ui_build.finished_at. That single comparison proves the stage barrier was bypassed.

git add .gitlab-ci.yml
git diff --cached --check
git diff --cached
git commit -m "ch15: replace false stage barriers with needs"
DAG_SHA="$(git rev-parse HEAD)"
git push
printf 'dag_sha=%s\n' "$DAG_SHA"

If runner capacity is one, those timestamps might not overlap even though the DAG allows overlap. In that case the conclusion is “the dependency graph is less restrictive, but runner capacity is the current bottleneck.”

7. No-runner evidence fixture

If compute is unavailable, use this deterministic fixture to practice the same reasoning:

[
  {"name":"api_build","start":0,"finish":3},
  {"name":"ui_build","start":0,"finish":12},
  {"name":"api_test","start":3,"finish":10},
  {"name":"ui_test","start":12,"finish":16},
  {"name":"package","start":16,"finish":19}
]

Draw edges from the configuration, recompute the 19-unit critical path, and mark the maximum runnable frontier: two jobs at pipeline start, then two independent chains, then one package job. This proves the scheduling model without pretending fixture time is a live GitLab benchmark.

8. Add a harmless 2×2 matrix

Add an independent compatibility job. It intentionally performs only deterministic output.

compatibility:
  stage: test
  needs: []
  parallel:
    matrix:
      - OS: [linux, windows]
        MODE: [unit, smoke]
  script:
    - printf 'os=%s mode=%s sha=%s\n' "$OS" "$MODE" "$CI_COMMIT_SHA"

CI Lint should expand four instances. In the UI/API, inspect names and the values printed by each instance. Maximum potential concurrency increases by four, but real concurrency increases only if runner slots exist. If an external system could tolerate two concurrent tests at most, a four-wide matrix would need additional application-level or resource-group constraints.

9. Serialize a synthetic shared resource

Create two jobs that represent writes to the same disposable resource. They use the same resource-group key and only sleep/print—no environment deployment occurs.

mutate_fixture_a:
  stage: package
  needs: []
  resource_group: chapter15-shared-fixture
  script:
    - printf 'A start=%s\n' "$(date -u +%FT%TZ)"
    - sleep 5
    - printf 'A finish=%s\n' "$(date -u +%FT%TZ)"

mutate_fixture_b:
  stage: package
  needs: []
  resource_group: chapter15-shared-fixture
  script:
    - printf 'B start=%s\n' "$(date -u +%FT%TZ)"
    - sleep 5
    - printf 'B finish=%s\n' "$(date -u +%FT%TZ)"

If both jobs exist and runner capacity is at least two, only one should hold the resource at a time. The other waits for the resource even though its DAG dependencies are already satisfied. This is a controlled demonstration of a queue caused by resource serialization rather than runner scarcity.

10. Optional: inspect and change resource-group process mode

The default process mode is unordered. For a real deployment pipeline you may prefer oldest_first or a newest-oriented mode, depending on idempotence and rollback semantics. Do not change a production resource group during training.

PROJECT_ID="12345678"
RG="chapter15-shared-fixture"

glab api "projects/$PROJECT_ID/resource_groups/$RG" \
  --jq '{key,process_mode}'

# Disposable project only; optional demonstration:
glab api --method PUT "projects/$PROJECT_ID/resource_groups/$RG" \
  -f process_mode=oldest_first \
  --jq '{key,process_mode}'

Changing process mode alters which waiting job GitLab selects next; it does not allow two jobs to hold the same resource simultaneously.

11. Calculate runner-slot demand

For each moment in the DAG, count jobs whose dependency and resource constraints are satisfied. That is the runnable frontier. The observed simultaneous running count is bounded by:

running concurrency ≤ min(runnable frontier, eligible runner slots, resource-group allowance, external-system safe concurrency).

In this lab the matrix can add four immediately runnable jobs. If two build jobs, four matrix jobs, and two resource jobs are all eligible at once, the graph may expose eight runnable jobs. That does not mean you should provision eight runners; compare time saved against compute and downstream capacity.

12. Challenge: choose the smallest control

A team says, “Our integration tests must not overlap because they write to one shared emulator, but compile jobs should remain parallel.” Choose one control and justify it.

The correct surface is a shared resource_group on the emulator-mutating jobs, not removing needs, shrinking the entire runner fleet to one slot, or forcing all tests into one stage. The narrow control preserves parallelism everywhere else.

13. Cleanup and rollback

Capture the final YAML, SHA, pipeline/job IDs, and timing table first. Then either restore the previous .gitlab-ci.yml or remove the disposable branch/project. If you changed a resource-group process mode in the disposable project, restore it to the original value or delete the disposable project. Do not delete real artifacts, branches, or runners for this lesson.

git switch main
git branch --show-current
# Delete only the known disposable remote branch after confirming its name:
git ls-remote --heads origin refs/heads/ch15/dag-lab
# Then, if the output is exactly the disposable ref:
git push origin --delete ch15/dag-lab

Knowledge check

In the baseline, why did api_test wait after api_build finished?

After DAG conversion, what timestamp relationship proves the optimization?

What does a 2×2 matrix create?

Two resource-group jobs are pending despite spare runner slots. What should you inspect?

Why can a DAG show potential concurrency that does not appear in live timestamps?

What is the safest training response if hosted compute is unavailable?

Summary

You measured a deliberately serialized pipeline, replaced false stage barriers with needs, proved the start-time change, expanded a small matrix, and serialized a synthetic shared resource. The key operational habit is to separate dependency, capacity, and resource-lock evidence before claiming an optimization.

Official references

Next lesson

Decide how much concurrency is worth operating

Lesson 3 turns the mechanics into policy choices: stage simplicity versus DAG precision, matrix breadth versus capacity/cost, serialization granularity, and safe cancellation.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.