Pipeline Performance, Interruptible Jobs, Retry, Failure Handling, Debugging, and Cost Control: Guided Hands-On Workflow and Core Operations
Measure a disposable pipeline, inspect job and runner evidence, apply safe interruptibility and a narrowly scoped retry fixture, then quantify latency and compute changes without weakening required checks.
Learning objectives
- Capture a structured baseline from pipeline/job timestamps and runner metadata.
- Shorten a real critical path without deleting required work.
- Mark only stateless obsolete work interruptible and explain supersession.
- Demonstrate narrowly scoped retry selection with a synthetic fixture.
- Compare latency, queueing, and compute using consistent evidence.
interruptible, workflow:auto_cancel,
retry, allow_failure, CI Lint, job/pipeline
APIs, job traces, artifacts, and runner metadata available to the
learner—are usable on Free/Premium/Ultimate across GitLab.com,
Self-Managed, and Dedicated. Hosted/instance-runner compute quotas and
cost factors are installation- and namespace-specific. The labs
therefore use tiny jobs and include a no-runner fixture/calculation
path. No cloud account, Premium/Ultimate control, privileged runner,
or purchased compute is required.
1. Lab objective: improve a measured pipeline, not a hypothetical one
This lab uses a disposable branch/project and only synthetic files. It deliberately separates a clean baseline from two optimization questions: “can independent work start earlier?” and “can obsolete work be canceled safely?” It also uses a synthetic exit-code fixture to demonstrate narrow retry selection without depending on a flaky public network service.
2. Preflight, role, runner, and no-runner fallback
Required: GitLab Free-compatible project, Developer or higher to push the disposable branch, basic Git, and an eligible runner if you want live timing. Keep every job tiny. If hosted compute is unavailable or quota is exhausted, stop after CI Lint and use the supplied timing fixture below; the reasoning exercise is still complete.
git status --short
git remote -v
glab auth status
glab ci lint --help | sed -n "1,80p"
Record the exact project path and default branch. Do not create runner credentials, variables, or external services for this chapter.
3. Create a deliberately serialized baseline
The baseline waits for build before both test jobs even
though docs_check does not consume build output. The
report stage waits for the entire test stage. This is intentionally
correct but inefficient.
stages: [build, test, report]
workflow:
auto_cancel:
on_new_commit: none
build:
stage: build
script:
- mkdir -p out
- printf "sha=%s\n" "$CI_COMMIT_SHA" > out/build.txt
- sleep 6
artifacts:
paths: [out/build.txt]
expire_in: 1 day
docs_check:
stage: test
script:
- printf "docs=%s\n" "$CI_COMMIT_SHORT_SHA"
- sleep 8
required_test:
stage: test
dependencies: [build]
script:
- test -s out/build.txt
- grep -F "$CI_COMMIT_SHA" out/build.txt
- sleep 6
required_report:
stage: report
script:
- printf "required gates complete for %s\n" "$CI_COMMIT_SHA"
4. Predict first, then lint
Before running anything, predict the earliest possible start of each job assuming an idle runner fleet. Then validate the effective configuration:
glab ci lint --dry-run --include-jobs
glab ci config compile > ch21-baseline-compiled.yml
Expected dependency logic: build starts first; both
test-stage jobs wait for build; required_report waits
until both test jobs finish. CI Lint proves job existence and graph
semantics but cannot prove real runner queue time.
5. Run the baseline and preserve identity
Create a disposable branch, push once, and capture the resulting pipeline ID. If your project has branch/MR rules from prior chapters, adapt the branch name only—do not bypass governance.
git switch -c ch21/performance-lab
git add .gitlab-ci.yml
git commit -m "lab: add ch21 performance baseline"
git push -u origin ch21/performance-lab
glab ci status --wait
glab ci list --ref ch21/performance-lab --per-page 5 --output json
Save the pipeline ID and SHA together. Never compare timings from different code identities without recording that difference.
6. Collect pipeline and job timing evidence
PROJECT_ID="12345678"
PIPELINE_ID="12345"
mkdir -p evidence/ch21
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID" \
> evidence/ch21/baseline-pipeline.json
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs?per_page=100" \
> evidence/ch21/baseline-jobs.json
jq "{id,sha,ref,status,created_at,started_at,finished_at,duration,queued_duration}" \
evidence/ch21/baseline-pipeline.json
jq "map({id,name,status,duration,queued_duration,started_at,finished_at,runner:(.runner.description // null)})" \
evidence/ch21/baseline-jobs.json
Record three different conclusions: (1) wall-clock interval from timestamps, (2) pipeline duration/queue fields, and (3) sum of job running durations. They answer different questions.
7. No-runner timing fixture
If you cannot run the pipeline, use this fictional but internally consistent evidence:
pipeline: created 10:00:00, started 10:00:04, finished 10:00:24
job build: queue=4s, start=10:00:04, duration=6s
job docs_check: queue=0s, start=10:00:10, duration=8s
job required_test: queue=0s, start=10:00:10, duration=6s
job required_report: start=10:00:18, duration=1s
# Remaining wall time illustrates orchestration/runner timing; use actual API evidence in a live lab.
Your optimization decision should still follow from dependency structure rather than from a fabricated billing number.
8. Optimization 1: shorten the critical path without deleting a gate
docs_check does not need build output, so let it start
immediately. required_test needs the build artifact
explicitly. required_report needs the required test
result, not every test-stage job.
docs_check:
stage: test
needs: []
script:
- printf "docs=%s\n" "$CI_COMMIT_SHORT_SHA"
- sleep 8
required_test:
stage: test
needs:
- job: build
artifacts: true
script:
- test -s out/build.txt
- grep -F "$CI_COMMIT_SHA" out/build.txt
- sleep 6
required_report:
stage: report
needs:
- job: required_test
artifacts: false
script:
- printf "required gates complete for %s\n" "$CI_COMMIT_SHA"
Nothing became optional. The graph simply stopped waiting on unrelated work. Re-lint and predict the new critical path before pushing.
9. Optimization 2: cancel only obsolete stateless work
Mark only docs_check interruptible and enable the
workflow policy that cancels interruptible jobs on a newer commit:
workflow:
auto_cancel:
on_new_commit: interruptible
docs_check:
interruptible: true
needs: []
script:
- printf "docs=%s\n" "$CI_COMMIT_SHORT_SHA"
- sleep 8
Prediction: if a newer commit starts while the
obsolete docs_check is running, GitLab may cancel that
job. The required test/report remain non-interruptible. Do not label
a deployment, database migration, release publication, or stateful
mutation interruptible merely to save minutes.
10. Demonstrate narrow retry classification safely
The following optional manual probe intentionally emits exit code 75. We call it a synthetic transient-classification fixture: it proves that GitLab selects the configured exit code, but it is intentionally deterministic so you do not depend on a flaky internet endpoint.
transient_retry_probe:
stage: test
when: manual
allow_failure: true
script:
- echo "fixture: imagine exit 75 means temporary upstream unavailable"
- exit 75
retry:
max: 1
exit_codes: 75
When triggered, GitLab retries once and the fixture fails again because the synthetic cause is deliberately permanent. That is useful evidence: retry mechanics cannot turn a deterministic condition into success. In production, use an exit-code wrapper only when it maps a real known transient condition consistently.
11. Observe supersession and retry without hiding evidence
For the optimized branch, push two harmless commits close together
only if your runner quota allows it. Keep the first job trace and
pipeline status. Use glab or the API to prove which job
was canceled and why:
glab ci status --live
glab ci get --output json --with-job-details
# For a chosen job ID:
glab ci trace 67890
If a job is canceled, do not immediately retry it. First confirm it belonged to the obsolete SHA. If the required job failed, preserve its failure reason and log before changing YAML.
12. Compare before/after latency and compute
Use the same fields for the optimized pipeline and create a small comparison record:
jq "[.[] | {name,duration,queued_duration,status}]" evidence/ch21/baseline-jobs.json > evidence/ch21/baseline-summary.json
# Repeat for optimized-jobs.json, then compare.
# Compute-minute estimate for a single runner class:
# sum(job_duration_seconds) / 60 * current_cost_factor
# Use the current Usage quotas/runner dashboard for authoritative charged usage.
A DAG change may reduce wall-clock latency without reducing total compute if the same jobs still run. Interruptibility reduces wasted compute only when obsolete pipelines actually overlap. Keep those conclusions separate.
13. Challenge: identify the right control
A required test takes 45 seconds but spends 90 seconds queued. A documentation job runs 60 seconds on every superseded branch commit. Which control addresses each problem?
Reason before clicking: the test is primarily a capacity/eligibility queueing problem; making it interruptible does not make the current required pipeline start sooner. The obsolete documentation job is a good candidate for safe interruptibility if it is stateless. Adding retry to either problem is unrelated.
14. Verification and cleanup
Verify that the optimized pipeline still contains the required test
and report, that the build artifact is consumed by the exact
required job, and that no quality gate changed to
allow_failure. Then remove the disposable branch only
after recording the evidence you want to keep.
git switch main
git ls-remote --heads origin ch21/performance-lab
# If and only if it is your disposable lab branch:
# git push origin --delete ch21/performance-lab
git show-ref --verify --quiet refs/heads/ch21/performance-lab && git branch -D ch21/performance-lab
Knowledge check
Which optimization reduced dependency latency in the lab?
Giving independent work needs: [] and expressing
only the real needs edges so unrelated stage
barriers no longer controlled the critical path.
Why is the exit-75 retry probe
allow_failure?
It is an optional synthetic mechanics probe, not a production quality gate. The required pipeline gates remain required.
A required job is slow because it waits for an eligible runner. What evidence identifies that?
Its queued_duration, runner/tag eligibility, and
runner availability/status—not merely script duration.
Why can the optimized clean pipeline use roughly the same compute as the baseline?
The DAG can shorten elapsed time by overlapping the same jobs. Compute is still the sum of their runner durations unless work is removed or obsolete runs are canceled.
What must remain true after marking a job interruptible?
Canceling it mid-run must leave no unsafe partial external state and must not remove a required invariant from the latest pipeline.
Summary
You measured before changing YAML, separated queueing from execution, shortened a real critical path, marked only stateless obsolete work interruptible, demonstrated narrow retry selection without relying on network randomness, and compared wall-clock latency with compute consumption. The next lesson turns those observations into policy tradeoffs.
Official references
Primary sources used for the current GitLab 19.3 behavior taught in this lesson:
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — CI/CD pipelines
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — Debugging CI/CD pipelines
- GitLab Docs — Validate CI/CD configuration
- GitLab Docs — Compute minutes
- GitLab Docs — Configure runners
- GitLab CLI — glab ci
- GitLab CLI — glab ci trace
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.