Chapter 15Lesson 05~290 minutes

Checkpoint Lab — DAG Pipelines, needs, Parallel Jobs, Matrices, Resource Groups, and Concurrency

Optimize and verify one disposable delivery graph end to end, including predictions, timing evidence, a controlled failure, and cleanup.

CheckpointDAGMatrixResource groupEvidenceCleanup

Checkpoint objectives

  • Write predictions for critical path, maximum runnable concurrency, and artifact availability before execution.
  • Optimize a serialized pipeline into a measured DAG with only justified edges.
  • Add one 2×2 matrix and one synthetic serialized resource without touching production systems.
  • Diagnose one intentional missing-dependency or queue condition from preserved evidence.
  • Remove unnecessary complexity and clean up all synthetic refs/configuration.
Availability baseline (verified 2026-08-21 against current GitLab documentation). The mandatory mechanisms in this chapter—needs, parallel, parallel:matrix, resource_group, and interruptible—are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. A job may list at most 50 needs entries. Numeric parallel accepts 1–200 instances, and parallel:matrix can create at most 200 permutations. Those are configuration limits, not guarantees of simultaneous execution: runner capacity, tags, protection, resource groups, and external-service limits can keep jobs pending. Hosted-runner quota/billing can change, so every required exercise includes a CI-Lint/fixture path that does not require paid compute.

1. Mission and production-model connection

The checkpoint models a common production problem: a pipeline was written as broad stages, then grew until independent work blocked one another. Your task is not to maximize parallel jobs. Your task is to encode the smallest correct dependency graph, bound concurrency, serialize one shared resource, and prove the change from evidence.

This chapter adds a scheduling/concurrency layer to the GitLab operating model built so far: Chapter 12 controlled which jobs exist, Chapter 13 controlled which configuration/secrets they can receive, Chapter 14 controlled which runners may execute them, and Chapter 15 controls which jobs become runnable together and which shared resources remain exclusive.

2. Preflight and assumptions

Item Mandatory assumption
Offering/tier GitLab Free is sufficient on GitLab.com, Self-Managed, or Dedicated.
Project Disposable project or branch only.
Role Developer for code/pipelines; no administrator role required.
Compute Tiny live jobs if eligible compute exists; otherwise CI Lint + deterministic fixture path.
Secrets None. Do not add CI/CD variables or print environment dumps.
External systems None. The serialized “resource” is synthetic.
Evidence SHA, pipeline/job IDs, graph, timestamps/status, and selected safe logs.

3. Define the metadata/scheduling policy before YAML

Write these rules into your lab notes:

  1. API tests depend only on API build; UI tests depend only on UI build.
  2. Package depends on both test paths.
  3. Compatibility coverage is exactly four combinations: two OS labels × two modes.
  4. Only the two synthetic mutation jobs share resource_group: chapter15-checkpoint.
  5. Read-only tests may be interruptible; synthetic mutation jobs are left non-interruptible by default.
  6. Artifacts must cross only declared edges.

4. Predictions — record before running

Use conceptual durations API build 3, UI build 12, API test 7, UI test 4, package 3.

Prediction Expected result
Stage-only lower bound 22 units: 12 build + 7 test + 3 package.
DAG lower bound 19 units: max(3+7, 12+4) + 3.
Initial DAG runnable jobs API build + UI build; plus immediate matrix/resource jobs if configured with needs: [].
Matrix expansion 4 compatibility instances.
Resource concurrency At most 1 job holding chapter15-checkpoint.
Package artifacts Only artifacts explicitly provided by needed producers; do not assume every earlier-stage artifact.

Also predict two operational changes: (1) api_test can become eligible before ui_build finishes, and (2) resource-group jobs can remain waiting even when a runner slot is available.

5. Build and validate the serialized baseline

stages: [build, test, package]

api_build:
  stage: build
  script: ["sleep 3", "printf 'api\n' > api.txt"]
  artifacts: { paths: [api.txt] }

ui_build:
  stage: build
  script: ["sleep 12", "printf 'ui\n' > ui.txt"]
  artifacts: { paths: [ui.txt] }

api_test:
  stage: test
  script: ["sleep 7", "test -f api.txt"]

ui_test:
  stage: test
  script: ["sleep 4", "test -f ui.txt"]

package:
  stage: package
  script: ["sleep 3", "printf 'package\n'"]

Run CI Lint first. If compute is available, push to ch15/checkpoint and capture one baseline pipeline. If not, proceed using the stage semantics and fixture; the checkpoint remains complete without purchasing compute.

6. Optimize only the false barriers

api_test:
  stage: test
  needs:
    - job: api_build
      artifacts: true
  script: ["sleep 7", "test -f api.txt"]

ui_test:
  stage: test
  needs:
    - job: ui_build
      artifacts: true
  script: ["sleep 4", "test -f ui.txt"]

package:
  stage: package
  needs: [api_test, ui_test]
  script: ["sleep 3", "printf 'package\n'"]

Do not add edges from package to both builds unless package actually consumes their artifacts. An edge with no dependency purpose is unnecessary complexity.

7. Add the required small matrix

compatibility:
  stage: test
  needs: []
  interruptible: true
  parallel:
    matrix:
      - OS: [linux, windows]
        MODE: [unit, smoke]
  script:
    - printf 'compatibility os=%s mode=%s sha=%s\n' "$OS" "$MODE" "$CI_COMMIT_SHA"

Before running, count four generated jobs and state your runner-slot assumption. If only two slots are available, predict a maximum of two simultaneous matrix instances even though four are runnable.

8. Add one serialized synthetic resource

fixture_write_a:
  stage: package
  needs: []
  resource_group: chapter15-checkpoint
  script:
    - printf 'A-start=%s\n' "$(date -u +%FT%TZ)"
    - sleep 5
    - printf 'A-end=%s\n' "$(date -u +%FT%TZ)"

fixture_write_b:
  stage: package
  needs: []
  resource_group: chapter15-checkpoint
  script:
    - printf 'B-start=%s\n' "$(date -u +%FT%TZ)"
    - sleep 5
    - printf 'B-end=%s\n' "$(date -u +%FT%TZ)"

These jobs mutate nothing. The exercise is to prove lock behavior: their execution intervals must not overlap if both run.

9. Commit, run, and bind evidence to one SHA

git switch -c ch15/checkpoint
git add .gitlab-ci.yml
git diff --cached --check
git diff --cached
git commit -m "ch15: checkpoint DAG and concurrency"
CHECKPOINT_SHA="$(git rev-parse HEAD)"
git push -u origin ch15/checkpoint
printf 'checkpoint_sha=%s\n' "$CHECKPOINT_SHA"

Record the created pipeline ID and confirm its SHA matches CHECKPOINT_SHA. Do not mix timestamps from reruns of a different SHA into the same comparison.

10. Observe graph, timestamps, and runner state

PROJECT_ID="12345678"
PIPELINE_ID="123456789"

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
  --jq '.[] | {id,name,stage,status,queued_duration,duration,started_at,finished_at,runner:(.runner.description // null)}'

Verify independently:

  • api_test starts only after api_build but need not wait for ui_build.
  • ui_test starts only after ui_build.
  • Package waits for both test jobs.
  • Four matrix jobs exist; actual overlap does not exceed available runner capacity.
  • The two fixture-write intervals do not overlap.

11. Calculate observed critical path

Create a table from timestamps. Subtract runner queue time from script duration only when you are explicitly analyzing compute time; for user-visible lead time, queue time is real and should remain included. Identify the path whose final predecessor completes last before package begins.

If live results do not beat the stage baseline, explain why: insufficient runner slots, startup overhead, unrelated immediate jobs competing for capacity, or the critical path itself. A valid checkpoint diagnosis can conclude “the DAG is correct but capacity prevents speedup.”

12. Required failure: missing conditional producer

Add this temporary broken pair and validate it without committing if you prefer:

optional_fixture:
  rules:
    - if: '$ENABLE_FIXTURE == "true"'
  script: echo optional

consumer_fixture:
  needs: [optional_fixture]
  script: echo consumer

With the condition false, preserve the CI Lint/pipeline-creation error. Decide the intended policy. For this checkpoint the producer is genuinely optional, so repair it with:

consumer_fixture:
  needs:
    - job: optional_fixture
      optional: true
  script: echo consumer

Validate again. Do not use allow_failure or retry because they do not solve absence from the graph.

13. Remove unnecessary complexity

Review every needs and resource-group entry. For each, write one sentence naming the dependency/resource it protects. Delete any edge that merely restates a stage and buys no measurable scheduling/data benefit. Keep the 2×2 matrix only if those four combinations represent a real support policy; otherwise the production lesson is to narrow it.

14. Evidence packet

  • Baseline and optimized commit SHAs.
  • Pipeline IDs and pipeline sources.
  • CI Lint result for final YAML.
  • Job timeline table with IDs, start/finish, queue/duration, and safe runner description.
  • Screenshot/export of Job dependencies graph if useful.
  • Critical-path calculation and runner-slot assumption.
  • Preserved missing-producer error and repaired validation.
  • Resource-group key/process mode and proof of non-overlap.

Do not include job tokens, full environment dumps, runner configuration, or secrets in the evidence packet.

15. Cleanup / rollback

First confirm every commit you need is reachable from your evidence or local branch. Then remove only the synthetic CI changes/ref. Resource-group objects are project metadata created by jobs; deleting the disposable project is the cleanest full cleanup. If keeping the project, restore the original CI file and leave no live deployment/environment integration behind.

git switch main
git ls-remote --heads origin refs/heads/ch15/checkpoint
# Delete only after the exact disposable ref is confirmed:
git push origin --delete ch15/checkpoint

git branch -D ch15/checkpoint  # local disposable branch only
Do not generalize this cleanup command. Never delete an unknown branch, project, artifact, runner, or resource group merely because its name resembles the lab. Confirm scope and exact identity first.

16. Final verification checklist

  • Exactly the intended DAG jobs exist for the checkpoint pipeline.
  • Your predictions are compared against independent UI/API evidence.
  • Artifact flow is explicit and no consumer relies on an accidental previous-stage download.
  • Matrix width and observed concurrency are documented separately.
  • Resource-group jobs do not overlap.
  • The intentionally broken missing producer has a preserved cause and deterministic repair.
  • The disposable branch/configuration is removed or the project is clearly retained as disposable.

Knowledge check

What two predictions must be written before running this checkpoint?

The graph permits six runnable jobs but only three execute. Is the DAG necessarily wrong?

Why is package not given needs edges to every build job?

What evidence proves the resource group worked?

The missing-producer example fails during creation. What is the correct repair here?

What would invalidate the safety of the evidence packet?

Checkpoint summary

You can now treat GitLab pipeline concurrency as an operating model rather than a collection of YAML tricks: explicit dependency edges, finite runner capacity, bounded matrix expansion, serialized mutable resources, cancellation safety, artifact contracts, and measured critical paths. Chapter 16 builds directly on this by focusing on the data that moves across those job edges—artifacts, reports, caches, retention, and integrity.

Official references

Next chapter

Artifacts, reports, cache, dependencies, retention, and job-to-job data flow

Chapter 16 turns the DAG edges you just designed into explicit data contracts: what is authoritative, what is disposable, who can download it, and how long it should survive.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.