Chapter 09Lesson 05~205 minutes

Checkpoint Lab — Stages, needs DAGs, Dependency Graphs, Early Execution, and Pipeline Critical-Path Design

The checkpoint starts from a deliberately slow but correct staged pipeline, refactors it into a correct DAG, measures the observed improvement, injects one dependency failure, repairs it from preserved evidence, and produces an evidence packet proving that early execution did not sacrifice required inputs.

Checkpoint labCritical pathDAG refactorEvidence packetVerification

Learning objectives

  • Predict stage-only and DAG timing before running either pipeline and state the required producer-consumer data edges.
  • Quantify observed improvement using preserved job timestamps/queued durations rather than intuition.
  • Inject a routing/data dependency fault, preserve first-failure evidence, and repair the causal edge.
  • Verify every early-running consumer has its required artifact or repository input and that optional work remains truly optional.
  • Produce a compact evidence packet and bridge dependency design into Chapter 10 artifact/report/retention dataflow.

1. Checkpoint scenario and operating boundary

You maintain a synthetic two-path application: an app binary path and a documentation path. The starting pipeline is correct but slow because stage barriers serialize unrelated work. Your task is to measure it, design a DAG, quantify the improvement, inject one causal dependency failure, preserve the failure evidence, and restore a verified pipeline.

Everything remains inside one disposable GitLab project/branch. No production deployment, package publication, cloud/Kubernetes/IaC, privileged runner, real credential, or paid feature is required.

2. Assumptions and preflight

Item Checkpoint requirement
Branch glci/ch09-checkpoint in a disposable project
GitLab Primary docs checked 2026-09-11; record actual GitLab version when known
Runner Ordinary Linux-capable runner; record ID/description/version/executor where visible
Tools POSIX shell, date, sleep, test, grep, sha256sum optional
Data Synthetic text only; no secrets or proprietary source
Evidence Two successful pipeline IDs (stage + DAG), one failed pipeline ID, job IDs/timings, graphs, artifact identity

3. Write predictions before execution

Record at least these predictions in evidence/predictions.md before running:

  1. In the stage-only pipeline, app_test will wait for the slower docs_build even though it needs only app_build.
  2. In the DAG pipeline, app_test will become runnable immediately after app_build; docs_test will follow docs_build.
  3. final_bundle will not become runnable until both validation paths succeed and will receive both producer artifacts for the same SHA.
  4. After the injected fault, app_test will fail because its required app artifact was not transferred; the fix will be the data dependency edge, not a retry or delay.

4. Setup synthetic scripts

Create the helper from Lesson 2 and this simple evidence-friendly build logic. The sleeps make the two paths visibly different.

mkdir -p ci evidence
cat > ci/stamp.sh <<'EOF'
#!/bin/sh
set -eu
mkdir -p evidence
printf '%s job=%s event=%s sha=%s\n'   "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$CI_JOB_NAME" "$1" "$CI_COMMIT_SHA"   >> "evidence/${CI_JOB_NAME}.log"
EOF
chmod +x ci/stamp.sh

5. Phase A — run the deliberately slow staged pipeline

Use this exact baseline:

stages: [build, test, bundle]

default:
  before_script:
    - chmod +x ci/stamp.sh

app_build:
  stage: build
  script:
    - ci/stamp.sh start
    - sleep 5
    - mkdir -p out
    - printf 'app_sha=%s\n' "$CI_COMMIT_SHA" > out/app.txt
    - ci/stamp.sh finish
  artifacts:
    paths: [out/app.txt, evidence/]
    expire_in: 1 day

docs_build:
  stage: build
  script:
    - ci/stamp.sh start
    - sleep 12
    - mkdir -p out
    - printf 'docs_sha=%s\n' "$CI_COMMIT_SHA" > out/docs.txt
    - ci/stamp.sh finish
  artifacts:
    paths: [out/docs.txt, evidence/]
    expire_in: 1 day

app_test:
  stage: test
  script:
    - ci/stamp.sh start
    - grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
    - sleep 8
    - ci/stamp.sh finish
  artifacts:
    paths: [evidence/]
    expire_in: 1 day

docs_test:
  stage: test
  script:
    - ci/stamp.sh start
    - grep -F "docs_sha=$CI_COMMIT_SHA" out/docs.txt
    - sleep 2
    - ci/stamp.sh finish
  artifacts:
    paths: [evidence/]
    expire_in: 1 day

final_bundle:
  stage: bundle
  script:
    - ci/stamp.sh start
    - grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
    - grep -F "docs_sha=$CI_COMMIT_SHA" out/docs.txt
    - sleep 4
    - ci/stamp.sh finish

Validate first, then run. Preserve the pipeline ID and timing evidence. The idealized lower bound is 24 seconds, while observed time also includes GitLab/runner overhead.

6. Phase B — refactor to the correct DAG

Replace only the consumer definitions with explicit needs:

app_test:
  stage: test
  needs:
    - job: app_build
      artifacts: true
  script:
    - ci/stamp.sh start
    - grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
    - sleep 8
    - ci/stamp.sh finish
  artifacts:
    paths: [evidence/]
    expire_in: 1 day

docs_test:
  stage: test
  needs:
    - job: docs_build
      artifacts: true
  script:
    - ci/stamp.sh start
    - grep -F "docs_sha=$CI_COMMIT_SHA" out/docs.txt
    - sleep 2
    - ci/stamp.sh finish
  artifacts:
    paths: [evidence/]
    expire_in: 1 day

final_bundle:
  stage: bundle
  needs:
    - job: app_test
      artifacts: false
    - job: docs_test
      artifacts: false
    - job: app_build
      artifacts: true
    - job: docs_build
      artifacts: true
  script:
    - ci/stamp.sh start
    - grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
    - grep -F "docs_sha=$CI_COMMIT_SHA" out/docs.txt
    - sleep 4
    - ci/stamp.sh finish

Again validate, then run. Preserve the second successful pipeline ID. The idealized DAG lower bound is approximately max(5+8, 12+2)+4 = 18 seconds.

7. Quantify observed improvement without overclaiming

Create a comparison table from actual GitLab metadata:

Metric Stage pipeline DAG pipeline Interpretation
Pipeline ID record record Different immutable execution records
Source/ref/SHA record record Prefer same source class and comparable revision
app_test start after app_build finish record delta record delta Shows barrier removal
Queued duration app_test record record Separates graph gain from runner wait
Overall observed duration record record Report actual, not idealized sleep math
Idealized dependency lower bound 24 s 18 s Model only; not measured scheduler performance

Compute observed percentage improvement only from observed durations:

improvement_percent = 100 * (stage_duration - dag_duration) / stage_duration
If the DAG is not faster because of queueing, that is a valid result. Document runner capacity as the next bottleneck instead of manipulating evidence.

8. Phase C — inject one causal artifact/dependency failure

Deliberately break the app_test edge by disabling artifact transfer while keeping the ordering dependency:

app_test:
  stage: test
  needs:
    - job: app_build
      artifacts: false   # Deliberate fault for this disposable lab.
  script:
    - ci/stamp.sh start
    - grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
    - sleep 8

Validate: the YAML/graph can still be valid. Run it and preserve the failed pipeline ID, failed job ID, trace, and producer artifact metadata. The expected causal result is that app_test starts after app_build but cannot find/read the required out/app.txt because transfer was disabled.

Do not immediately retry. First prove the producer succeeded, artifact exists, consumer need has artifacts: false, and failure is dataflow—not runner/network/secret state.

9. Repair the causal edge and verify the smallest scope

Restore artifacts: true. Commit the correction and run one fresh pipeline. Verify:

  • app_build artifact contains the current SHA.
  • app_test starts only after app_build and reads the SHA-bound artifact.
  • docs path remains independent.
  • final_bundle waits for both tests and receives both build artifacts.
  • No new runner, token, secret, environment, release, registry, or external side effect was introduced.

10. Required evidence packet

Evidence Minimum content
Identity Project/path, branch, pipeline source, immutable SHA
Configuration Stage baseline + DAG revision identity; CI Lint/full configuration result
Graph Stage graph and needs graph; explicit producer-consumer edges
Successful runs Stage pipeline ID + DAG pipeline ID; job IDs/statuses/timings/queued durations
Failure run Failed pipeline/job ID, first-failure trace, broken artifacts flag
Data Producer artifact path plus SHA-binding verification in consumers
Runner Runner ID/description/version/executor when visible; note if shared/unknown
Analysis Theoretical 24s/18s model clearly separated from observed timings
Assumptions GitLab docs verification date 2026-09-11; no paid/admin/cloud dependency

11. Final verification checklist

  • Exactly one dependency reason is documented for every needs edge: control gate, artifact/data transfer, or both.
  • No early consumer relies on a generated file from a job absent from its required needs.
  • Artifact content is verified against the exact CI_COMMIT_SHA.
  • Observed start/finish/queue data supports the performance statement.
  • No optional need is used to hide a required missing producer.
  • No privileged runner, real secret, cloud target, release/package publication, or production resource is touched.

12. Cleanup/rollback

  1. Retain only the non-secret evidence packet needed for learning/audit.
  2. Delete glci/ch09-checkpoint after verifying the exact branch name and ensuring no unrelated work is present.
  3. If the project exists solely for this lab, delete only that exact disposable project after recording its identity.
  4. Do not delete historical pipeline/job evidence while troubleshooting is still active.

13. What Chapter 09 adds to a production GitLab CI/CD model

You can now distinguish a stage barrier from a true dependency, prove when a job becomes runnable, separate graph delay from runner queue delay, and preserve data/quality/security gates while shortening the critical path. The graph is no longer “just execution order”; it is an auditable statement of what each job requires.

Chapter 10 deepens the data side of those edges: artifacts, reports, retention, dependencies, needs:artifacts, and cross-job data flow. That is the natural next step because a correct DAG must not only order jobs correctly—it must move the intended evidence between them.

Next lesson

Next: Artifacts, Reports, Retention, Dependencies, needs:artifacts, and Cross-Job Data Flow: Concepts, Architecture, and Mental Model

Continue with the next lesson in the course sequence and carry forward the evidence-first GitLab CI/CD operating model.

Knowledge check

What three pipeline IDs should the checkpoint preserve?

Why is artifacts:false a useful injected failure here?

What must be separated when claiming a DAG performance improvement?

If final_bundle receives both artifacts but starts before app_test passes, what class of dependency is missing?

What is the bridge to Chapter 10?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-11. GitLab CI/CD DAG and artifact-transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.

  • Make jobs start earlier with needs — stage barriers, DAG execution, immediate jobs, and practical examples.
  • CI/CD YAML syntax reference — authoritative stages, needs, needs:artifacts, needs:optional, needs:project, and needs:pipeline:job semantics and limits.
  • Pipeline editor — visualization of jobs, stages, and needs relationships plus full configuration inspection.
  • CI Lint — syntax/logic validation and pipeline simulation that can expose invalid needs relationships before execution.
  • Job artifacts — default previous-stage artifact fetching and how needs:artifacts changes data transfer.
  • Troubleshooting job artifacts — missing/expired/inaccessible artifact failures.
  • Jobs API — job IDs, stage/status, created_at, started_at, finished_at, duration, queued duration, and runner metadata for timing evidence.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.