Checkpoint Lab — Pipeline Performance, Interruptible Jobs, Retry, Failure Handling, Debugging, and Cost Control
Benchmark an intentionally inefficient pipeline, apply two evidence-based optimizations, classify transient and deterministic failures differently, and prove required gates and artifacts still hold.
Learning objectives
- Benchmark a deliberately inefficient pipeline with exact SHA and IDs.
- Apply two optimizations while preserving required gates and artifacts.
- Inject transient-class and deterministic failures with different policies.
- Quantify feedback latency and compute/resource consumption separately.
- Verify cleanup, evidence integrity, and the bridge to package distribution.
interruptible, workflow:auto_cancel,
retry, allow_failure, CI Lint, job/pipeline
APIs, job traces, artifacts, and runner metadata available to the
learner—are usable on Free/Premium/Ultimate across GitLab.com,
Self-Managed, and Dedicated. Hosted/instance-runner compute quotas and
cost factors are installation- and namespace-specific. The labs
therefore use tiny jobs and include a no-runner fixture/calculation
path. No cloud account, Premium/Ultimate control, privileged runner,
or purchased compute is required.
1. Checkpoint scenario and acceptance criteria
You own a disposable merge pipeline whose feedback is slower than necessary. Your task is to benchmark it, apply two evidence-based optimizations, classify two injected failures differently, and prove that the latest pipeline still produces the required artifact and fails when its required quality gate fails.
Acceptance criteria: exact SHA/pipeline IDs recorded; baseline and optimized timestamps captured; critical path predicted before each run; one safe job made interruptible; one classified retry fixture; one deterministic required failure that is not retried/ignored; required artifact identity preserved; no real secrets; cleanup verified.
2. Preflight and predictions
Use GitLab 19.3 behavior as the reference baseline. Required path is Free-compatible. Use an ordinary disposable project/branch with Developer+ push rights and an eligible runner if available. If no compute is available, use CI Lint plus the synthetic timing table in Step 6.
Before modifying the pipeline, write these predictions:
| Prediction | Record before execution |
|---|---|
| Pipeline graph | Which jobs can overlap and which dependencies block them. |
| Maximum useful concurrency | How many jobs could safely run at once, independent of actual runner slots. |
| Interruptibility | Which job can be abandoned with zero unsafe external state. |
| Required invariant | Which job must fail the pipeline if the invariant is violated. |
| Artifact identity | Which producer creates the authoritative file and which consumer verifies its SHA content. |
3. Baseline pipeline
stages: [build, test, report]
workflow:
auto_cancel:
on_new_commit: none
build_artifact:
stage: build
script:
- mkdir -p out
- printf "artifact_sha=%s\n" "$CI_COMMIT_SHA" > out/release.txt
- sleep 6
artifacts:
paths: [out/release.txt]
expire_in: 1 day
docs_gate:
stage: test
script:
- test -f README.md
- sleep 8
required_gate:
stage: test
dependencies: [build_artifact]
script:
- test -s out/release.txt
- grep -F "$CI_COMMIT_SHA" out/release.txt
- sleep 6
summary:
stage: report
script:
- printf "pipeline=%s sha=%s\n" "$CI_PIPELINE_ID" "$CI_COMMIT_SHA"
Predict a stage-serialized critical path before you run. The documentation check has no artifact dependency but still waits for build because it is in the next stage.
4. Validate before spending runner time
glab ci lint --dry-run --include-jobs
glab ci config compile > evidence-ch21-baseline-compiled.yml
If CI Lint is invalid, stop. Configuration failure is not a performance benchmark.
5. Execute and capture baseline evidence
git switch -c ch21/checkpoint
git add .gitlab-ci.yml
git commit -m "lab: ch21 baseline"
git push -u origin ch21/checkpoint
glab ci status --wait
PROJECT_ID="12345678"
PIPELINE_ID="12345"
mkdir -p evidence/ch21-checkpoint
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID" > evidence/ch21-checkpoint/baseline-pipeline.json
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs?per_page=100" > evidence/ch21-checkpoint/baseline-jobs.json
git rev-parse HEAD > evidence/ch21-checkpoint/baseline-sha.txt
Do not retry a failed baseline before capturing its job/log
evidence. Record pipeline and job
queued_duration alongside execution duration so runner
wait is not mistaken for slow code. A benchmark is only meaningful
if the state is preserved.
6. No-runner fixture for the benchmark math
If live compute is unavailable, use this fixture as evidence supplied by the exercise, not as a claim about your project:
BASELINE
created=12:00:00 started=12:00:03 finished=12:00:24
build_artifact: queue=3s duration=6s start=12:00:03
docs_gate: queue=0s duration=8s start=12:00:09
required_gate: queue=0s duration=6s start=12:00:09
summary: queue=0s duration=1s start=12:00:17
OPTIMIZED (predicted shape, verify live when possible)
docs_gate may start near pipeline start
required_gate starts after build_artifact
summary starts after required_gate rather than waiting for docs_gate
7. Apply two optimizations without weakening gates
Optimization A — dependency graph: let independent docs work start immediately, make the required gate explicitly consume the artifact, and let summary depend only on the required gate.
Optimization B — obsolete work: mark only docs work interruptible and enable interruptible auto-cancel for newer commits.
workflow:
auto_cancel:
on_new_commit: interruptible
docs_gate:
stage: test
needs: []
interruptible: true
script:
- test -f README.md
- sleep 8
required_gate:
stage: test
needs:
- job: build_artifact
artifacts: true
interruptible: false
script:
- test -s out/release.txt
- grep -F "$CI_COMMIT_SHA" out/release.txt
- sleep 6
summary:
stage: report
needs:
- job: required_gate
artifacts: false
script:
- printf "pipeline=%s sha=%s\n" "$CI_PIPELINE_ID" "$CI_COMMIT_SHA"
Re-run CI Lint and compare the compiled graph before pushing. Required artifact checking and required failure policy remain intact.
8. Execute optimized pipeline and compare
git add .gitlab-ci.yml
git commit -m "lab: optimize ch21 scheduling"
git push
glab ci status --wait
# Replace with the new pipeline ID.
PIPELINE_ID="12346"
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID" > evidence/ch21-checkpoint/optimized-pipeline.json
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs?per_page=100" > evidence/ch21-checkpoint/optimized-jobs.json
git rev-parse HEAD > evidence/ch21-checkpoint/optimized-sha.txt
Compute both wall-clock interval and sum of job durations. If the same jobs ran, the clean-run compute may be nearly unchanged even though feedback latency improved. Record that as a valid result.
9. Inject a synthetic transient-class failure
Add an optional fixture that emits exit 75 and is retried exactly once:
transient_classification_probe:
stage: test
needs: []
when: manual
allow_failure: true
script:
- echo "synthetic transient class; no network or secret involved"
- exit 75
retry:
max: 1
exit_codes: 75
Trigger it after linting. Expected: GitLab processes one retry because the selected exit code matches. It still fails because this fixture is deliberately deterministic. Your written conclusion must say “retry policy selected the intended class,” not “retry fixed the failure.”
10. Inject a deterministic required failure and fail fast
Temporarily change only the required gate:
required_gate:
# keep the existing needs/artifact contract
script:
- echo "intentional deterministic checkpoint failure"
- exit 2
retry: 0
allow_failure: false
Expected: one attempt, failed required job, failed pipeline.
Preserve the log/API response. Do not add
allow_failure or retry to make the checkpoint green.
Restore the real artifact-verification script, lint again, and run a
final successful pipeline.
11. Prove or simulate safe supersession
If quota permits, create two harmless commits while
docs_gate from the older pipeline is still running.
Verify that the obsolete interruptible docs job may be canceled
while required non-interruptible work is not casually abandoned. If
timing makes this difficult, use the current YAML and API state as a
policy simulation and explain the expected behavior; do not burn
compute merely to force overlap.
12. Quantify latency and compute without inventing billing
Create a small report with:
| Metric | Baseline | Optimized | Interpretation |
|---|---|---|---|
| Wall-clock created→finished | record | record | Developer feedback latency. |
| Pipeline queued duration | record | record | Scheduling before execution. |
| Critical-path job chain | record | record | Dependency bottleneck. |
| Sum of job durations | record | record | Raw runner work before cost factors. |
| Estimated compute minutes | calculate with current factor | calculate with current factor | Runner-duration consumption; use Usage quotas for authoritative charged usage. |
| Artifact bytes | record | record | Evidence that optimization did not drop authoritative output. |
Do not claim a currency saving unless you have current billing data for the exact runner/subscription. The chapter’s cost control is resource accounting, not a price quote.
13. Verification checklist
- The final pipeline SHA matches the artifact content.
-
required_gateis present, non-advisory, and consumes the build artifact through an explicit dependency. -
docs_gateis the only job made interruptible in this checkpoint. - The deterministic failure was attempted once and failed the pipeline.
- The transient-classification fixture was retried only for exit 75 and remains clearly labeled optional.
- Baseline and optimized API evidence use the same metric definitions.
- No debug trace, real credential, external deployment, registry write, or privileged runner was introduced.
14. Cleanup and evidence retention
Restore the final good YAML, preserve only the evidence you intentionally want, and remove the disposable branch when safe:
git switch main
git ls-remote --heads origin ch21/checkpoint
# If and only if this is your disposable branch:
# git push origin --delete ch21/checkpoint
git show-ref --verify --quiet refs/heads/ch21/checkpoint && git branch -D ch21/checkpoint
sha256sum evidence/ch21-checkpoint/*.json evidence/ch21-checkpoint/*.txt > evidence/ch21-checkpoint/evidence.sha256
Deleting a branch does not erase historical pipeline/job records automatically. If you choose to erase job logs/artifacts, remember that erasure is destructive evidence deletion; do it only for synthetic data or required security cleanup.
15. What Chapter 21 adds to the production operating model
You now have a performance-governance loop: define required invariants, measure wall-clock/queue/critical-path/compute separately, eliminate unnecessary dependencies, cancel only safely obsolete work, retry only classified transient faults, protect failure evidence, and verify that optimization did not weaken quality or artifact identity. Chapter 22 moves from ephemeral job artifacts to governed package distribution and registry identities.
Knowledge check
Which two optimizations did the checkpoint apply?
It shortened dependency barriers with explicit
needs and enabled cancellation only for obsolete
stateless docs work with interruptible/auto-cancel.
Why should the deterministic exit-2 failure not be retried?
Its cause is intentionally deterministic. Retrying would only repeat the same defect and consume more time/compute.
What did the exit-75 fixture prove if it still failed after retry?
That GitLab selected and retried the configured failure class. It deliberately does not pretend a deterministic fixture became transient.
How do you prove optimization did not weaken artifact correctness?
The final required gate must still fetch the authoritative build artifact and verify that its recorded SHA equals the pipeline commit SHA.
Why is sum of job durations different from wall-clock latency?
Jobs can overlap, and queueing/orchestration add elapsed time. Summed runner work measures resource use; wall-clock measures feedback delay.
What should happen if CI_DEBUG_TRACE accidentally
exposes a real credential?
Revoke/rotate the credential first, restrict exposure, then clean up the log/evidence. Do not treat log deletion as the primary containment action.
Summary
You produced a measurable before/after optimization, differentiated transient retry policy from deterministic failure, preserved required gates and artifact identity, and accounted for latency and compute without fabricating pricing. That is the production standard: faster feedback through evidence and scheduling correctness, not through bypassing controls.
Official references
Primary sources used for the current GitLab 19.3 behavior taught in this lesson:
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — CI/CD pipelines
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — Debugging CI/CD pipelines
- GitLab Docs — Validate CI/CD configuration
- GitLab Docs — Compute minutes
- GitLab Docs — Configure runners
- GitLab CLI — glab ci
- GitLab CLI — glab ci trace
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.