Chapter 21Lesson 05~340 minutes

Checkpoint Lab — Pipeline Performance, Interruptible Jobs, Retry, Failure Handling, Debugging, and Cost Control

Benchmark an intentionally inefficient pipeline, apply two evidence-based optimizations, classify transient and deterministic failures differently, and prove required gates and artifacts still hold.

CheckpointBenchmarkOptimizationFailure injectionEvidenceCleanup

Learning objectives

  • Benchmark a deliberately inefficient pipeline with exact SHA and IDs.
  • Apply two optimizations while preserving required gates and artifacts.
  • Inject transient-class and deterministic failures with different policies.
  • Quantify feedback latency and compute/resource consumption separately.
  • Verify cleanup, evidence integrity, and the bridge to package distribution.
Availability baseline (verified 2026-08-22 against current GitLab 19.3 documentation). The CI/CD keywords and evidence surfaces used in the mandatory path—interruptible, workflow:auto_cancel, retry, allow_failure, CI Lint, job/pipeline APIs, job traces, artifacts, and runner metadata available to the learner—are usable on Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. Hosted/instance-runner compute quotas and cost factors are installation- and namespace-specific. The labs therefore use tiny jobs and include a no-runner fixture/calculation path. No cloud account, Premium/Ultimate control, privileged runner, or purchased compute is required.

1. Checkpoint scenario and acceptance criteria

You own a disposable merge pipeline whose feedback is slower than necessary. Your task is to benchmark it, apply two evidence-based optimizations, classify two injected failures differently, and prove that the latest pipeline still produces the required artifact and fails when its required quality gate fails.

Acceptance criteria: exact SHA/pipeline IDs recorded; baseline and optimized timestamps captured; critical path predicted before each run; one safe job made interruptible; one classified retry fixture; one deterministic required failure that is not retried/ignored; required artifact identity preserved; no real secrets; cleanup verified.

2. Preflight and predictions

Use GitLab 19.3 behavior as the reference baseline. Required path is Free-compatible. Use an ordinary disposable project/branch with Developer+ push rights and an eligible runner if available. If no compute is available, use CI Lint plus the synthetic timing table in Step 6.

Before modifying the pipeline, write these predictions:

Prediction Record before execution
Pipeline graph Which jobs can overlap and which dependencies block them.
Maximum useful concurrency How many jobs could safely run at once, independent of actual runner slots.
Interruptibility Which job can be abandoned with zero unsafe external state.
Required invariant Which job must fail the pipeline if the invariant is violated.
Artifact identity Which producer creates the authoritative file and which consumer verifies its SHA content.

3. Baseline pipeline

stages: [build, test, report]

workflow:
  auto_cancel:
    on_new_commit: none

build_artifact:
  stage: build
  script:
    - mkdir -p out
    - printf "artifact_sha=%s\n" "$CI_COMMIT_SHA" > out/release.txt
    - sleep 6
  artifacts:
    paths: [out/release.txt]
    expire_in: 1 day

docs_gate:
  stage: test
  script:
    - test -f README.md
    - sleep 8

required_gate:
  stage: test
  dependencies: [build_artifact]
  script:
    - test -s out/release.txt
    - grep -F "$CI_COMMIT_SHA" out/release.txt
    - sleep 6

summary:
  stage: report
  script:
    - printf "pipeline=%s sha=%s\n" "$CI_PIPELINE_ID" "$CI_COMMIT_SHA"

Predict a stage-serialized critical path before you run. The documentation check has no artifact dependency but still waits for build because it is in the next stage.

4. Validate before spending runner time

glab ci lint --dry-run --include-jobs
glab ci config compile > evidence-ch21-baseline-compiled.yml

If CI Lint is invalid, stop. Configuration failure is not a performance benchmark.

5. Execute and capture baseline evidence

git switch -c ch21/checkpoint
git add .gitlab-ci.yml
git commit -m "lab: ch21 baseline"
git push -u origin ch21/checkpoint
glab ci status --wait

PROJECT_ID="12345678"
PIPELINE_ID="12345"
mkdir -p evidence/ch21-checkpoint
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID" > evidence/ch21-checkpoint/baseline-pipeline.json
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs?per_page=100" > evidence/ch21-checkpoint/baseline-jobs.json
git rev-parse HEAD > evidence/ch21-checkpoint/baseline-sha.txt

Do not retry a failed baseline before capturing its job/log evidence. Record pipeline and job queued_duration alongside execution duration so runner wait is not mistaken for slow code. A benchmark is only meaningful if the state is preserved.

6. No-runner fixture for the benchmark math

If live compute is unavailable, use this fixture as evidence supplied by the exercise, not as a claim about your project:

BASELINE
created=12:00:00 started=12:00:03 finished=12:00:24
build_artifact: queue=3s duration=6s start=12:00:03
docs_gate:      queue=0s duration=8s start=12:00:09
required_gate:  queue=0s duration=6s start=12:00:09
summary:        queue=0s duration=1s start=12:00:17

OPTIMIZED (predicted shape, verify live when possible)
docs_gate may start near pipeline start
required_gate starts after build_artifact
summary starts after required_gate rather than waiting for docs_gate

7. Apply two optimizations without weakening gates

Optimization A — dependency graph: let independent docs work start immediately, make the required gate explicitly consume the artifact, and let summary depend only on the required gate.

Optimization B — obsolete work: mark only docs work interruptible and enable interruptible auto-cancel for newer commits.

workflow:
  auto_cancel:
    on_new_commit: interruptible

docs_gate:
  stage: test
  needs: []
  interruptible: true
  script:
    - test -f README.md
    - sleep 8

required_gate:
  stage: test
  needs:
    - job: build_artifact
      artifacts: true
  interruptible: false
  script:
    - test -s out/release.txt
    - grep -F "$CI_COMMIT_SHA" out/release.txt
    - sleep 6

summary:
  stage: report
  needs:
    - job: required_gate
      artifacts: false
  script:
    - printf "pipeline=%s sha=%s\n" "$CI_PIPELINE_ID" "$CI_COMMIT_SHA"

Re-run CI Lint and compare the compiled graph before pushing. Required artifact checking and required failure policy remain intact.

8. Execute optimized pipeline and compare

git add .gitlab-ci.yml
git commit -m "lab: optimize ch21 scheduling"
git push
glab ci status --wait

# Replace with the new pipeline ID.
PIPELINE_ID="12346"
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID" > evidence/ch21-checkpoint/optimized-pipeline.json
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs?per_page=100" > evidence/ch21-checkpoint/optimized-jobs.json
git rev-parse HEAD > evidence/ch21-checkpoint/optimized-sha.txt

Compute both wall-clock interval and sum of job durations. If the same jobs ran, the clean-run compute may be nearly unchanged even though feedback latency improved. Record that as a valid result.

9. Inject a synthetic transient-class failure

Add an optional fixture that emits exit 75 and is retried exactly once:

transient_classification_probe:
  stage: test
  needs: []
  when: manual
  allow_failure: true
  script:
    - echo "synthetic transient class; no network or secret involved"
    - exit 75
  retry:
    max: 1
    exit_codes: 75

Trigger it after linting. Expected: GitLab processes one retry because the selected exit code matches. It still fails because this fixture is deliberately deterministic. Your written conclusion must say “retry policy selected the intended class,” not “retry fixed the failure.”

10. Inject a deterministic required failure and fail fast

Temporarily change only the required gate:

required_gate:
  # keep the existing needs/artifact contract
  script:
    - echo "intentional deterministic checkpoint failure"
    - exit 2
  retry: 0
  allow_failure: false

Expected: one attempt, failed required job, failed pipeline. Preserve the log/API response. Do not add allow_failure or retry to make the checkpoint green. Restore the real artifact-verification script, lint again, and run a final successful pipeline.

11. Prove or simulate safe supersession

If quota permits, create two harmless commits while docs_gate from the older pipeline is still running. Verify that the obsolete interruptible docs job may be canceled while required non-interruptible work is not casually abandoned. If timing makes this difficult, use the current YAML and API state as a policy simulation and explain the expected behavior; do not burn compute merely to force overlap.

12. Quantify latency and compute without inventing billing

Create a small report with:

Metric Baseline Optimized Interpretation
Wall-clock created→finished record record Developer feedback latency.
Pipeline queued duration record record Scheduling before execution.
Critical-path job chain record record Dependency bottleneck.
Sum of job durations record record Raw runner work before cost factors.
Estimated compute minutes calculate with current factor calculate with current factor Runner-duration consumption; use Usage quotas for authoritative charged usage.
Artifact bytes record record Evidence that optimization did not drop authoritative output.

Do not claim a currency saving unless you have current billing data for the exact runner/subscription. The chapter’s cost control is resource accounting, not a price quote.

13. Verification checklist

  • The final pipeline SHA matches the artifact content.
  • required_gate is present, non-advisory, and consumes the build artifact through an explicit dependency.
  • docs_gate is the only job made interruptible in this checkpoint.
  • The deterministic failure was attempted once and failed the pipeline.
  • The transient-classification fixture was retried only for exit 75 and remains clearly labeled optional.
  • Baseline and optimized API evidence use the same metric definitions.
  • No debug trace, real credential, external deployment, registry write, or privileged runner was introduced.

14. Cleanup and evidence retention

Restore the final good YAML, preserve only the evidence you intentionally want, and remove the disposable branch when safe:

git switch main
git ls-remote --heads origin ch21/checkpoint
# If and only if this is your disposable branch:
# git push origin --delete ch21/checkpoint
git show-ref --verify --quiet refs/heads/ch21/checkpoint && git branch -D ch21/checkpoint

sha256sum evidence/ch21-checkpoint/*.json evidence/ch21-checkpoint/*.txt   > evidence/ch21-checkpoint/evidence.sha256

Deleting a branch does not erase historical pipeline/job records automatically. If you choose to erase job logs/artifacts, remember that erasure is destructive evidence deletion; do it only for synthetic data or required security cleanup.

15. What Chapter 21 adds to the production operating model

You now have a performance-governance loop: define required invariants, measure wall-clock/queue/critical-path/compute separately, eliminate unnecessary dependencies, cancel only safely obsolete work, retry only classified transient faults, protect failure evidence, and verify that optimization did not weaken quality or artifact identity. Chapter 22 moves from ephemeral job artifacts to governed package distribution and registry identities.

Knowledge check

Which two optimizations did the checkpoint apply?

Why should the deterministic exit-2 failure not be retried?

What did the exit-75 fixture prove if it still failed after retry?

How do you prove optimization did not weaken artifact correctness?

Why is sum of job durations different from wall-clock latency?

What should happen if CI_DEBUG_TRACE accidentally exposes a real credential?

Summary

You produced a measurable before/after optimization, differentiated transient retry policy from deterministic failure, preserved required gates and artifact identity, and accounted for latency and compute without fabricating pricing. That is the production standard: faster feedback through evidence and scheduling correctness, not through bypassing controls.

Official references

Primary sources used for the current GitLab 19.3 behavior taught in this lesson:

Next chapter

Chapter 22 — Package Registry, Dependency Proxy, Package Formats, Permissions, and Artifact Distribution

The next chapter promotes delivery outputs into governed package identities with namespace, authentication, versioning, retention, provenance, and consumer impact.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.