Checkpoint Lab — Stages, needs DAGs, Dependency Graphs, Early Execution, and Pipeline Critical-Path Design
The checkpoint starts from a deliberately slow but correct staged pipeline, refactors it into a correct DAG, measures the observed improvement, injects one dependency failure, repairs it from preserved evidence, and produces an evidence packet proving that early execution did not sacrifice required inputs.
Learning objectives
- Predict stage-only and DAG timing before running either pipeline and state the required producer-consumer data edges.
- Quantify observed improvement using preserved job timestamps/queued durations rather than intuition.
- Inject a routing/data dependency fault, preserve first-failure evidence, and repair the causal edge.
- Verify every early-running consumer has its required artifact or repository input and that optional work remains truly optional.
- Produce a compact evidence packet and bridge dependency design into Chapter 10 artifact/report/retention dataflow.
1. Checkpoint scenario and operating boundary
You maintain a synthetic two-path application: an app binary path and a documentation path. The starting pipeline is correct but slow because stage barriers serialize unrelated work. Your task is to measure it, design a DAG, quantify the improvement, inject one causal dependency failure, preserve the failure evidence, and restore a verified pipeline.
Everything remains inside one disposable GitLab project/branch. No production deployment, package publication, cloud/Kubernetes/IaC, privileged runner, real credential, or paid feature is required.
2. Assumptions and preflight
| Item | Checkpoint requirement |
|---|---|
| Branch |
glci/ch09-checkpoint in a disposable project
|
| GitLab | Primary docs checked 2026-09-11; record actual GitLab version when known |
| Runner | Ordinary Linux-capable runner; record ID/description/version/executor where visible |
| Tools | POSIX shell, date, sleep, test, grep, sha256sum optional |
| Data | Synthetic text only; no secrets or proprietary source |
| Evidence | Two successful pipeline IDs (stage + DAG), one failed pipeline ID, job IDs/timings, graphs, artifact identity |
3. Write predictions before execution
Record at least these predictions in
evidence/predictions.md before running:
-
In the stage-only pipeline,
app_testwill wait for the slowerdocs_buildeven though it needs onlyapp_build. -
In the DAG pipeline,
app_testwill become runnable immediately afterapp_build;docs_testwill followdocs_build. -
final_bundlewill not become runnable until both validation paths succeed and will receive both producer artifacts for the same SHA. -
After the injected fault,
app_testwill fail because its required app artifact was not transferred; the fix will be the data dependency edge, not a retry or delay.
4. Setup synthetic scripts
Create the helper from Lesson 2 and this simple evidence-friendly build logic. The sleeps make the two paths visibly different.
mkdir -p ci evidence
cat > ci/stamp.sh <<'EOF'
#!/bin/sh
set -eu
mkdir -p evidence
printf '%s job=%s event=%s sha=%s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$CI_JOB_NAME" "$1" "$CI_COMMIT_SHA" >> "evidence/${CI_JOB_NAME}.log"
EOF
chmod +x ci/stamp.sh
5. Phase A — run the deliberately slow staged pipeline
Use this exact baseline:
stages: [build, test, bundle]
default:
before_script:
- chmod +x ci/stamp.sh
app_build:
stage: build
script:
- ci/stamp.sh start
- sleep 5
- mkdir -p out
- printf 'app_sha=%s\n' "$CI_COMMIT_SHA" > out/app.txt
- ci/stamp.sh finish
artifacts:
paths: [out/app.txt, evidence/]
expire_in: 1 day
docs_build:
stage: build
script:
- ci/stamp.sh start
- sleep 12
- mkdir -p out
- printf 'docs_sha=%s\n' "$CI_COMMIT_SHA" > out/docs.txt
- ci/stamp.sh finish
artifacts:
paths: [out/docs.txt, evidence/]
expire_in: 1 day
app_test:
stage: test
script:
- ci/stamp.sh start
- grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
- sleep 8
- ci/stamp.sh finish
artifacts:
paths: [evidence/]
expire_in: 1 day
docs_test:
stage: test
script:
- ci/stamp.sh start
- grep -F "docs_sha=$CI_COMMIT_SHA" out/docs.txt
- sleep 2
- ci/stamp.sh finish
artifacts:
paths: [evidence/]
expire_in: 1 day
final_bundle:
stage: bundle
script:
- ci/stamp.sh start
- grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
- grep -F "docs_sha=$CI_COMMIT_SHA" out/docs.txt
- sleep 4
- ci/stamp.sh finish
Validate first, then run. Preserve the pipeline ID and timing evidence. The idealized lower bound is 24 seconds, while observed time also includes GitLab/runner overhead.
6. Phase B — refactor to the correct DAG
Replace only the consumer definitions with explicit needs:
app_test:
stage: test
needs:
- job: app_build
artifacts: true
script:
- ci/stamp.sh start
- grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
- sleep 8
- ci/stamp.sh finish
artifacts:
paths: [evidence/]
expire_in: 1 day
docs_test:
stage: test
needs:
- job: docs_build
artifacts: true
script:
- ci/stamp.sh start
- grep -F "docs_sha=$CI_COMMIT_SHA" out/docs.txt
- sleep 2
- ci/stamp.sh finish
artifacts:
paths: [evidence/]
expire_in: 1 day
final_bundle:
stage: bundle
needs:
- job: app_test
artifacts: false
- job: docs_test
artifacts: false
- job: app_build
artifacts: true
- job: docs_build
artifacts: true
script:
- ci/stamp.sh start
- grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
- grep -F "docs_sha=$CI_COMMIT_SHA" out/docs.txt
- sleep 4
- ci/stamp.sh finish
Again validate, then run. Preserve the second successful pipeline ID. The idealized DAG lower bound is approximately max(5+8, 12+2)+4 = 18 seconds.
7. Quantify observed improvement without overclaiming
Create a comparison table from actual GitLab metadata:
| Metric | Stage pipeline | DAG pipeline | Interpretation |
|---|---|---|---|
| Pipeline ID | record | record | Different immutable execution records |
| Source/ref/SHA | record | record | Prefer same source class and comparable revision |
| app_test start after app_build finish | record delta | record delta | Shows barrier removal |
| Queued duration app_test | record | record | Separates graph gain from runner wait |
| Overall observed duration | record | record | Report actual, not idealized sleep math |
| Idealized dependency lower bound | 24 s | 18 s | Model only; not measured scheduler performance |
Compute observed percentage improvement only from observed durations:
improvement_percent = 100 * (stage_duration - dag_duration) / stage_duration
8. Phase C — inject one causal artifact/dependency failure
Deliberately break the app_test edge by disabling
artifact transfer while keeping the ordering dependency:
app_test:
stage: test
needs:
- job: app_build
artifacts: false # Deliberate fault for this disposable lab.
script:
- ci/stamp.sh start
- grep -F "app_sha=$CI_COMMIT_SHA" out/app.txt
- sleep 8
Validate: the YAML/graph can still be valid. Run it and preserve the
failed pipeline ID, failed job ID, trace, and producer artifact
metadata. The expected causal result is that
app_test starts after app_build but cannot
find/read the required out/app.txt because transfer was
disabled.
artifacts: false, and failure is dataflow—not
runner/network/secret state.
9. Repair the causal edge and verify the smallest scope
Restore artifacts: true. Commit the correction and run
one fresh pipeline. Verify:
- app_build artifact contains the current SHA.
- app_test starts only after app_build and reads the SHA-bound artifact.
- docs path remains independent.
- final_bundle waits for both tests and receives both build artifacts.
- No new runner, token, secret, environment, release, registry, or external side effect was introduced.
10. Required evidence packet
| Evidence | Minimum content |
|---|---|
| Identity | Project/path, branch, pipeline source, immutable SHA |
| Configuration | Stage baseline + DAG revision identity; CI Lint/full configuration result |
| Graph | Stage graph and needs graph; explicit producer-consumer edges |
| Successful runs | Stage pipeline ID + DAG pipeline ID; job IDs/statuses/timings/queued durations |
| Failure run | Failed pipeline/job ID, first-failure trace, broken artifacts flag |
| Data | Producer artifact path plus SHA-binding verification in consumers |
| Runner | Runner ID/description/version/executor when visible; note if shared/unknown |
| Analysis | Theoretical 24s/18s model clearly separated from observed timings |
| Assumptions | GitLab docs verification date 2026-09-11; no paid/admin/cloud dependency |
11. Final verification checklist
-
Exactly one dependency reason is documented for every
needsedge: control gate, artifact/data transfer, or both. - No early consumer relies on a generated file from a job absent from its required needs.
-
Artifact content is verified against the exact
CI_COMMIT_SHA. - Observed start/finish/queue data supports the performance statement.
- No optional need is used to hide a required missing producer.
- No privileged runner, real secret, cloud target, release/package publication, or production resource is touched.
12. Cleanup/rollback
- Retain only the non-secret evidence packet needed for learning/audit.
-
Delete
glci/ch09-checkpointafter verifying the exact branch name and ensuring no unrelated work is present. - If the project exists solely for this lab, delete only that exact disposable project after recording its identity.
- Do not delete historical pipeline/job evidence while troubleshooting is still active.
13. What Chapter 09 adds to a production GitLab CI/CD model
You can now distinguish a stage barrier from a true dependency, prove when a job becomes runnable, separate graph delay from runner queue delay, and preserve data/quality/security gates while shortening the critical path. The graph is no longer “just execution order”; it is an auditable statement of what each job requires.
Chapter 10 deepens the data side of those edges: artifacts, reports,
retention, dependencies, needs:artifacts,
and cross-job data flow. That is the natural next step because a
correct DAG must not only order jobs correctly—it must move the
intended evidence between them.
Knowledge check
What three pipeline IDs should the checkpoint preserve?
At minimum: the successful staged baseline, the successful DAG comparison, and the deliberately failed artifact-transfer run.
Why is artifacts:false a useful injected failure here?
It preserves the ordering edge but breaks the data-transfer contract, letting you distinguish control dependency from artifact dependency.
What must be separated when claiming a DAG performance improvement?
The theoretical dependency lower bound, observed overall duration, and runner queued duration/other overhead.
If final_bundle receives both artifacts but starts before app_test passes, what class of dependency is missing?
A control/quality-gate dependency on app_test, even though the file data is available.
What is the bridge to Chapter 10?
The graph now defines who can run when; Chapter 10 defines how artifacts/reports are retained and transferred across those job boundaries.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. GitLab CI/CD DAG and artifact-transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.
-
Make jobs start earlier with
needs— stage barriers, DAG execution, immediate jobs, and practical examples. -
CI/CD YAML syntax reference
— authoritative
stages,needs,needs:artifacts,needs:optional,needs:project, andneeds:pipeline:jobsemantics and limits. -
Pipeline editor
— visualization of jobs, stages, and
needsrelationships plus full configuration inspection. -
CI Lint
— syntax/logic validation and pipeline simulation that can expose
invalid
needsrelationships before execution. -
Job artifacts
— default previous-stage artifact fetching and how
needs:artifactschanges data transfer. - Troubleshooting job artifacts — missing/expired/inaccessible artifact failures.
-
Jobs API
— job IDs, stage/status,
created_at,started_at,finished_at, duration, queued duration, and runner metadata for timing evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.