DAG Pipelines, needs, Parallel Jobs, Matrices, Resource Groups, and Concurrency: Guided Hands-On Workflow and Core Operations
Measure a serialized pipeline, replace false barriers with needs, inspect matrix expansion, prove a resource lock, and calculate critical path from evidence.
Learning objectives
- Establish a serialized baseline and predict its theoretical duration before changing dependencies.
-
Convert the baseline to a
needsDAG and verify start times from the API/UI. - Expand a safe 2×2 matrix and distinguish generated jobs from actual simultaneous execution.
-
Use a synthetic
resource_groupto prove mutual exclusion without touching a real environment. - Calculate critical path and runner-slot demand from observed evidence.
needs,
parallel, parallel:matrix,
resource_group, and interruptible—are
available in GitLab Free/Premium/Ultimate across GitLab.com,
Self-Managed, and Dedicated. A job may list at most 50
needs entries. Numeric parallel accepts
1–200 instances, and parallel:matrix can create at most
200 permutations. Those are configuration limits, not guarantees of
simultaneous execution: runner capacity, tags, protection, resource
groups, and external-service limits can keep jobs pending.
Hosted-runner quota/billing can change, so every required exercise
includes a CI-Lint/fixture path that does not require paid compute.
1. Disposable scenario and preflight
Use a throwaway GitLab Free project such as
gitlab-ch15-dag-lab or a clearly disposable branch in
an existing lab project. Do not use a production repository because
this lesson intentionally changes pipeline scheduling.
| Preflight item | Required state | Why |
|---|---|---|
| Project/ref |
Disposable project or branch ch15/dag-lab.
|
Prevents experiments from changing real merge/release behavior. |
| Role | Developer can commit/run jobs; Maintainer is useful only for project settings/API inspection. | No admin privileges are required. |
| Runner | At least one eligible runner for live timing, or use CI Lint + supplied fixture. | Hosted compute may be unavailable/quota constrained. |
| External systems |
None. Jobs use sleep, printf, and
tiny artifacts only.
|
No network service, deployment, or secret is required. |
| Evidence | Record commit SHA, pipeline ID, job IDs/timestamps, runner IDs/descriptions if visible. | Binds performance claims to one configuration revision. |
2. Baseline: deliberately stage-serialized
Create the following configuration. The sleep times are intentionally short in a live lab; the comments also show the conceptual “minutes” used for critical-path arithmetic. Keep the live sleeps under a few seconds to minimize compute consumption.
stages: [build, test, package]
api_build:
stage: build
script:
- sleep 3
- printf 'api-build=%s\n' "$CI_COMMIT_SHA" > api.txt
artifacts:
paths: [api.txt]
ui_build:
stage: build
script:
- sleep 12
- printf 'ui-build=%s\n' "$CI_COMMIT_SHA" > ui.txt
artifacts:
paths: [ui.txt]
api_test:
stage: test
script:
- sleep 7
- test -f api.txt
ui_test:
stage: test
script:
- sleep 4
- test -f ui.txt
package:
stage: package
script:
- sleep 3
- printf 'package=%s\n' "$CI_COMMIT_SHA"
max(3,12) + max(7,4) + 3 = 22 time units. In a real
hosted run, queue/startup overhead adds to that number.
3. Validate before spending compute
Use Build → Pipeline editor → Validate (CI Lint) and inspect the merged configuration. If no runner is available, this validation plus the supplied timeline fixture later in the lesson is the mandatory completion path.
git switch -c ch15/dag-lab
git add .gitlab-ci.yml
git diff --cached --check
git diff --cached
git commit -m "ch15: add serialized baseline"
BASE_SHA="$(git rev-parse HEAD)"
printf 'baseline_sha=%s\n' "$BASE_SHA"
git push -u origin ch15/dag-lab
Do not benchmark by rerunning the same job repeatedly until one looks fast. Record one representative pipeline and note queue time separately from script time.
4. Inspect the baseline timeline
PROJECT_ID="12345678"
PIPELINE_ID="123456789"
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
--jq '.[] | {id,name,stage,status,created_at,started_at,finished_at,queued_duration,duration}'
Expected relationship: both build jobs become eligible together.
Even if api_build finishes early,
api_test remains blocked until
ui_build completes because the build stage
has not cleared. If only one runner slot exists, execution is even
more serialized; record that capacity fact rather than blaming
stages alone.
5. Convert to explicit needs
Now describe actual dependencies. Keep stages for readability, but let DAG edges bypass unrelated stage work.
api_test:
stage: test
needs:
- job: api_build
artifacts: true
script:
- sleep 7
- test -f api.txt
ui_test:
stage: test
needs:
- job: ui_build
artifacts: true
script:
- sleep 4
- test -f ui.txt
package:
stage: package
needs:
- api_test
- ui_test
script:
- sleep 3
- printf 'package=%s\n' "$CI_COMMIT_SHA"
The package job does not need the build artifacts directly, only completion of both test jobs. If package really consumes files from builds, model those artifact dependencies explicitly instead of relying on old stage behavior.
3+7=10 and 12+4=16; package
waits for the longer chain and adds 3, giving a conceptual lower
bound of 19. The DAG removes three units of false stage
waiting.
6. Prove causality with timestamps, not the graph picture alone
Commit the DAG version, run one pipeline, and capture the same API
fields. Verify that api_test.started_at is after
api_build.finished_at but can be before
ui_build.finished_at. That single comparison proves the
stage barrier was bypassed.
git add .gitlab-ci.yml
git diff --cached --check
git diff --cached
git commit -m "ch15: replace false stage barriers with needs"
DAG_SHA="$(git rev-parse HEAD)"
git push
printf 'dag_sha=%s\n' "$DAG_SHA"
If runner capacity is one, those timestamps might not overlap even though the DAG allows overlap. In that case the conclusion is “the dependency graph is less restrictive, but runner capacity is the current bottleneck.”
7. No-runner evidence fixture
If compute is unavailable, use this deterministic fixture to practice the same reasoning:
[
{"name":"api_build","start":0,"finish":3},
{"name":"ui_build","start":0,"finish":12},
{"name":"api_test","start":3,"finish":10},
{"name":"ui_test","start":12,"finish":16},
{"name":"package","start":16,"finish":19}
]
Draw edges from the configuration, recompute the 19-unit critical path, and mark the maximum runnable frontier: two jobs at pipeline start, then two independent chains, then one package job. This proves the scheduling model without pretending fixture time is a live GitLab benchmark.
8. Add a harmless 2×2 matrix
Add an independent compatibility job. It intentionally performs only deterministic output.
compatibility:
stage: test
needs: []
parallel:
matrix:
- OS: [linux, windows]
MODE: [unit, smoke]
script:
- printf 'os=%s mode=%s sha=%s\n' "$OS" "$MODE" "$CI_COMMIT_SHA"
CI Lint should expand four instances. In the UI/API, inspect names and the values printed by each instance. Maximum potential concurrency increases by four, but real concurrency increases only if runner slots exist. If an external system could tolerate two concurrent tests at most, a four-wide matrix would need additional application-level or resource-group constraints.
9. Serialize a synthetic shared resource
Create two jobs that represent writes to the same disposable resource. They use the same resource-group key and only sleep/print—no environment deployment occurs.
mutate_fixture_a:
stage: package
needs: []
resource_group: chapter15-shared-fixture
script:
- printf 'A start=%s\n' "$(date -u +%FT%TZ)"
- sleep 5
- printf 'A finish=%s\n' "$(date -u +%FT%TZ)"
mutate_fixture_b:
stage: package
needs: []
resource_group: chapter15-shared-fixture
script:
- printf 'B start=%s\n' "$(date -u +%FT%TZ)"
- sleep 5
- printf 'B finish=%s\n' "$(date -u +%FT%TZ)"
If both jobs exist and runner capacity is at least two, only one should hold the resource at a time. The other waits for the resource even though its DAG dependencies are already satisfied. This is a controlled demonstration of a queue caused by resource serialization rather than runner scarcity.
10. Optional: inspect and change resource-group process mode
The default process mode is unordered. For a real
deployment pipeline you may prefer oldest_first or a
newest-oriented mode, depending on idempotence and rollback
semantics. Do not change a production resource group during
training.
PROJECT_ID="12345678"
RG="chapter15-shared-fixture"
glab api "projects/$PROJECT_ID/resource_groups/$RG" \
--jq '{key,process_mode}'
# Disposable project only; optional demonstration:
glab api --method PUT "projects/$PROJECT_ID/resource_groups/$RG" \
-f process_mode=oldest_first \
--jq '{key,process_mode}'
Changing process mode alters which waiting job GitLab selects next; it does not allow two jobs to hold the same resource simultaneously.
11. Calculate runner-slot demand
For each moment in the DAG, count jobs whose dependency and resource constraints are satisfied. That is the runnable frontier. The observed simultaneous running count is bounded by:
In this lab the matrix can add four immediately runnable jobs. If two build jobs, four matrix jobs, and two resource jobs are all eligible at once, the graph may expose eight runnable jobs. That does not mean you should provision eight runners; compare time saved against compute and downstream capacity.
12. Challenge: choose the smallest control
A team says, “Our integration tests must not overlap because they write to one shared emulator, but compile jobs should remain parallel.” Choose one control and justify it.
The correct surface is a shared resource_group on the
emulator-mutating jobs, not removing needs, shrinking
the entire runner fleet to one slot, or forcing all tests into one
stage. The narrow control preserves parallelism everywhere else.
13. Cleanup and rollback
Capture the final YAML, SHA, pipeline/job IDs, and timing table
first. Then either restore the previous
.gitlab-ci.yml or remove the disposable branch/project.
If you changed a resource-group process mode in the disposable
project, restore it to the original value or delete the disposable
project. Do not delete real artifacts, branches, or runners for this
lesson.
git switch main
git branch --show-current
# Delete only the known disposable remote branch after confirming its name:
git ls-remote --heads origin refs/heads/ch15/dag-lab
# Then, if the output is exactly the disposable ref:
git push origin --delete ch15/dag-lab
Knowledge check
In the baseline, why did api_test wait after
api_build finished?
The stage barrier required all build-stage jobs, including the unrelated UI build, to finish.
After DAG conversion, what timestamp relationship proves the optimization?
api_test can start after api_build finishes while ui_build is still running.
What does a 2×2 matrix create?
Four distinct job instances with matrix values exposed as CI/CD variables.
Two resource-group jobs are pending despite spare runner slots. What should you inspect?
Whether one already holds the shared resource-group key and the group process mode/queue.
Why can a DAG show potential concurrency that does not appear in live timestamps?
Eligible runner capacity, tag/protection constraints, resource locks, or external limits can cap actual execution.
What is the safest training response if hosted compute is unavailable?
Use CI Lint/expanded configuration and the deterministic timing fixture; do not purchase compute merely to complete the chapter.
Summary
You measured a deliberately serialized pipeline, replaced false
stage barriers with needs, proved the start-time
change, expanded a small matrix, and serialized a synthetic shared
resource. The key operational habit is to separate dependency,
capacity, and resource-lock evidence before claiming an
optimization.
Official references
- GitLab Docs — Make jobs start earlier with needs
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — needs keyword
- GitLab Docs — needs:artifacts
- GitLab Docs — needs:optional
- GitLab Docs — needs:parallel:matrix
- GitLab Docs — parallel
- GitLab Docs — parallel:matrix
- GitLab Docs — Matrix expressions
- GitLab Docs — Resource groups
- GitLab Docs — resource_group keyword
- GitLab Docs — interruptible keyword
- GitLab Docs — Auto-cancel redundant pipelines
- GitLab Docs — Pipeline efficiency
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — Resource groups API
- GitLab Docs — Runner advanced configuration
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.