Chapter 21Lesson 02~325 minutes

Pipeline Performance, Interruptible Jobs, Retry, Failure Handling, Debugging, and Cost Control: Guided Hands-On Workflow and Core Operations

Measure a disposable pipeline, inspect job and runner evidence, apply safe interruptibility and a narrowly scoped retry fixture, then quantify latency and compute changes without weakening required checks.

Hands-onMetricsinterruptibleRetry policyglabRunner evidence

Learning objectives

  • Capture a structured baseline from pipeline/job timestamps and runner metadata.
  • Shorten a real critical path without deleting required work.
  • Mark only stateless obsolete work interruptible and explain supersession.
  • Demonstrate narrowly scoped retry selection with a synthetic fixture.
  • Compare latency, queueing, and compute using consistent evidence.
Availability baseline (verified 2026-08-22 against current GitLab 19.3 documentation). The CI/CD keywords and evidence surfaces used in the mandatory path—interruptible, workflow:auto_cancel, retry, allow_failure, CI Lint, job/pipeline APIs, job traces, artifacts, and runner metadata available to the learner—are usable on Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. Hosted/instance-runner compute quotas and cost factors are installation- and namespace-specific. The labs therefore use tiny jobs and include a no-runner fixture/calculation path. No cloud account, Premium/Ultimate control, privileged runner, or purchased compute is required.

1. Lab objective: improve a measured pipeline, not a hypothetical one

This lab uses a disposable branch/project and only synthetic files. It deliberately separates a clean baseline from two optimization questions: “can independent work start earlier?” and “can obsolete work be canceled safely?” It also uses a synthetic exit-code fixture to demonstrate narrow retry selection without depending on a flaky public network service.

2. Preflight, role, runner, and no-runner fallback

Required: GitLab Free-compatible project, Developer or higher to push the disposable branch, basic Git, and an eligible runner if you want live timing. Keep every job tiny. If hosted compute is unavailable or quota is exhausted, stop after CI Lint and use the supplied timing fixture below; the reasoning exercise is still complete.

git status --short
git remote -v
glab auth status
glab ci lint --help | sed -n "1,80p"

Record the exact project path and default branch. Do not create runner credentials, variables, or external services for this chapter.

3. Create a deliberately serialized baseline

The baseline waits for build before both test jobs even though docs_check does not consume build output. The report stage waits for the entire test stage. This is intentionally correct but inefficient.

stages: [build, test, report]

workflow:
  auto_cancel:
    on_new_commit: none

build:
  stage: build
  script:
    - mkdir -p out
    - printf "sha=%s\n" "$CI_COMMIT_SHA" > out/build.txt
    - sleep 6
  artifacts:
    paths: [out/build.txt]
    expire_in: 1 day

docs_check:
  stage: test
  script:
    - printf "docs=%s\n" "$CI_COMMIT_SHORT_SHA"
    - sleep 8

required_test:
  stage: test
  dependencies: [build]
  script:
    - test -s out/build.txt
    - grep -F "$CI_COMMIT_SHA" out/build.txt
    - sleep 6

required_report:
  stage: report
  script:
    - printf "required gates complete for %s\n" "$CI_COMMIT_SHA"

4. Predict first, then lint

Before running anything, predict the earliest possible start of each job assuming an idle runner fleet. Then validate the effective configuration:

glab ci lint --dry-run --include-jobs
glab ci config compile > ch21-baseline-compiled.yml

Expected dependency logic: build starts first; both test-stage jobs wait for build; required_report waits until both test jobs finish. CI Lint proves job existence and graph semantics but cannot prove real runner queue time.

5. Run the baseline and preserve identity

Create a disposable branch, push once, and capture the resulting pipeline ID. If your project has branch/MR rules from prior chapters, adapt the branch name only—do not bypass governance.

git switch -c ch21/performance-lab
git add .gitlab-ci.yml
git commit -m "lab: add ch21 performance baseline"
git push -u origin ch21/performance-lab

glab ci status --wait
glab ci list --ref ch21/performance-lab --per-page 5 --output json

Save the pipeline ID and SHA together. Never compare timings from different code identities without recording that difference.

6. Collect pipeline and job timing evidence

PROJECT_ID="12345678"
PIPELINE_ID="12345"
mkdir -p evidence/ch21

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID" \
  > evidence/ch21/baseline-pipeline.json

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs?per_page=100" \
  > evidence/ch21/baseline-jobs.json

jq "{id,sha,ref,status,created_at,started_at,finished_at,duration,queued_duration}" \
  evidence/ch21/baseline-pipeline.json

jq "map({id,name,status,duration,queued_duration,started_at,finished_at,runner:(.runner.description // null)})" \
  evidence/ch21/baseline-jobs.json

Record three different conclusions: (1) wall-clock interval from timestamps, (2) pipeline duration/queue fields, and (3) sum of job running durations. They answer different questions.

7. No-runner timing fixture

If you cannot run the pipeline, use this fictional but internally consistent evidence:

pipeline: created 10:00:00, started 10:00:04, finished 10:00:24
job build:         queue=4s, start=10:00:04, duration=6s
job docs_check:    queue=0s, start=10:00:10, duration=8s
job required_test: queue=0s, start=10:00:10, duration=6s
job required_report: start=10:00:18, duration=1s
# Remaining wall time illustrates orchestration/runner timing; use actual API evidence in a live lab.

Your optimization decision should still follow from dependency structure rather than from a fabricated billing number.

8. Optimization 1: shorten the critical path without deleting a gate

docs_check does not need build output, so let it start immediately. required_test needs the build artifact explicitly. required_report needs the required test result, not every test-stage job.

docs_check:
  stage: test
  needs: []
  script:
    - printf "docs=%s\n" "$CI_COMMIT_SHORT_SHA"
    - sleep 8

required_test:
  stage: test
  needs:
    - job: build
      artifacts: true
  script:
    - test -s out/build.txt
    - grep -F "$CI_COMMIT_SHA" out/build.txt
    - sleep 6

required_report:
  stage: report
  needs:
    - job: required_test
      artifacts: false
  script:
    - printf "required gates complete for %s\n" "$CI_COMMIT_SHA"

Nothing became optional. The graph simply stopped waiting on unrelated work. Re-lint and predict the new critical path before pushing.

9. Optimization 2: cancel only obsolete stateless work

Mark only docs_check interruptible and enable the workflow policy that cancels interruptible jobs on a newer commit:

workflow:
  auto_cancel:
    on_new_commit: interruptible

docs_check:
  interruptible: true
  needs: []
  script:
    - printf "docs=%s\n" "$CI_COMMIT_SHORT_SHA"
    - sleep 8

Prediction: if a newer commit starts while the obsolete docs_check is running, GitLab may cancel that job. The required test/report remain non-interruptible. Do not label a deployment, database migration, release publication, or stateful mutation interruptible merely to save minutes.

10. Demonstrate narrow retry classification safely

The following optional manual probe intentionally emits exit code 75. We call it a synthetic transient-classification fixture: it proves that GitLab selects the configured exit code, but it is intentionally deterministic so you do not depend on a flaky internet endpoint.

transient_retry_probe:
  stage: test
  when: manual
  allow_failure: true
  script:
    - echo "fixture: imagine exit 75 means temporary upstream unavailable"
    - exit 75
  retry:
    max: 1
    exit_codes: 75

When triggered, GitLab retries once and the fixture fails again because the synthetic cause is deliberately permanent. That is useful evidence: retry mechanics cannot turn a deterministic condition into success. In production, use an exit-code wrapper only when it maps a real known transient condition consistently.

11. Observe supersession and retry without hiding evidence

For the optimized branch, push two harmless commits close together only if your runner quota allows it. Keep the first job trace and pipeline status. Use glab or the API to prove which job was canceled and why:

glab ci status --live
glab ci get --output json --with-job-details
# For a chosen job ID:
glab ci trace 67890

If a job is canceled, do not immediately retry it. First confirm it belonged to the obsolete SHA. If the required job failed, preserve its failure reason and log before changing YAML.

12. Compare before/after latency and compute

Use the same fields for the optimized pipeline and create a small comparison record:

jq "[.[] | {name,duration,queued_duration,status}]" evidence/ch21/baseline-jobs.json > evidence/ch21/baseline-summary.json
# Repeat for optimized-jobs.json, then compare.

# Compute-minute estimate for a single runner class:
# sum(job_duration_seconds) / 60 * current_cost_factor
# Use the current Usage quotas/runner dashboard for authoritative charged usage.

A DAG change may reduce wall-clock latency without reducing total compute if the same jobs still run. Interruptibility reduces wasted compute only when obsolete pipelines actually overlap. Keep those conclusions separate.

13. Challenge: identify the right control

A required test takes 45 seconds but spends 90 seconds queued. A documentation job runs 60 seconds on every superseded branch commit. Which control addresses each problem?

Reason before clicking: the test is primarily a capacity/eligibility queueing problem; making it interruptible does not make the current required pipeline start sooner. The obsolete documentation job is a good candidate for safe interruptibility if it is stateless. Adding retry to either problem is unrelated.

14. Verification and cleanup

Verify that the optimized pipeline still contains the required test and report, that the build artifact is consumed by the exact required job, and that no quality gate changed to allow_failure. Then remove the disposable branch only after recording the evidence you want to keep.

git switch main
git ls-remote --heads origin ch21/performance-lab
# If and only if it is your disposable lab branch:
# git push origin --delete ch21/performance-lab
git show-ref --verify --quiet refs/heads/ch21/performance-lab && git branch -D ch21/performance-lab

Knowledge check

Which optimization reduced dependency latency in the lab?

Why is the exit-75 retry probe allow_failure?

A required job is slow because it waits for an eligible runner. What evidence identifies that?

Why can the optimized clean pipeline use roughly the same compute as the baseline?

What must remain true after marking a job interruptible?

Summary

You measured before changing YAML, separated queueing from execution, shortened a real critical path, marked only stateless obsolete work interruptible, demonstrated narrow retry selection without relying on network randomness, and compared wall-clock latency with compute consumption. The next lesson turns those observations into policy tradeoffs.

Official references

Primary sources used for the current GitLab 19.3 behavior taught in this lesson:

Next lesson

Design choices — speed, reliability, capacity, and cost

You will decide when retries are justified, which work may be canceled, when additional runners help, and when cache optimization creates more risk than value.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.