Chapter 21Lesson 01~290 minutes

Pipeline Performance, Interruptible Jobs, Retry, Failure Handling, Debugging, and Cost Control: Concepts, Architecture, and Mental Model

Build an evidence-first model of GitLab pipeline latency, queueing, runner utilization, critical paths, interruptibility, retry semantics, failure policy, debugging evidence, and compute cost.

Mental modelLatencyQueue timeCritical pathRetryCompute

Learning objectives

  • Separate pipeline wall-clock latency, queue time, running duration, critical path, and compute usage.
  • Explain interruptible together with redundant-pipeline cancellation policy.
  • Classify retryable transient faults separately from deterministic failures.
  • Keep retry, required/advisory status, and exit-code policy conceptually distinct.
  • Use CI Lint, APIs, job logs, and runner evidence safely before tuning.
Availability baseline (verified 2026-08-22 against current GitLab 19.3 documentation). The CI/CD keywords and evidence surfaces used in the mandatory path—interruptible, workflow:auto_cancel, retry, allow_failure, CI Lint, job/pipeline APIs, job traces, artifacts, and runner metadata available to the learner—are usable on Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. Hosted/instance-runner compute quotas and cost factors are installation- and namespace-specific. The labs therefore use tiny jobs and include a no-runner fixture/calculation path. No cloud account, Premium/Ultimate control, privileged runner, or purchased compute is required.

1. The practical problem: a green pipeline can still be operationally unhealthy

By Chapter 20 you can compose reusable CI/CD safely. The next production problem is execution quality. A pipeline can be correct and still be too slow, spend long periods waiting for runners, repeat deterministic failures, consume avoidable compute, or encourage developers to bypass checks because feedback arrives too late.

Performance work must begin with evidence. “The pipeline took 12 minutes” is incomplete until you know how much of that was queueing, which jobs formed the critical path, which work ran in parallel, which attempts were retries, and which runner class executed each job.

2. Mental model: one change creates several clocks and several costs

Evidence flow for pipeline performance
flowchart TD
  C[Commit / ref] --> P[Pipeline record]
  P --> J1[Job A]
  P --> J2[Job B]
  J1 --> Q1[Queue / eligible runner]
  J2 --> Q2[Queue / eligible runner]
  Q1 --> R1[Runner execution]
  Q2 --> R2[Runner execution]
  R1 --> A[Artifacts / cache / network]
  R2 --> A
  P --> E[API + UI + glab evidence]
  J1 --> E
  J2 --> E

The pipeline record gives pipeline-level timestamps and queue/duration fields. Each job has its own created, started, finished, duration, queued-duration, runner, status, and failure evidence. Runner execution consumes compute; artifacts, caches, images, and dependencies add storage and network work. Optimization is safe only when these identities stay connected to the exact commit and pipeline.

3. Separate wall-clock latency, queue time, running duration, and compute

Use four different measurements rather than one vague “pipeline time.”

Measure What it answers Important GitLab nuance
Wall-clock feedback latency How long the developer waited from pipeline creation to finish. Derive from created_at and finished_at when you need a human-facing elapsed interval.
Pipeline queued_duration How long the pipeline waited before execution began. Returned by the Pipelines API; investigate runner eligibility/capacity when it dominates.
Pipeline duration How much running-time span GitLab attributes to the pipeline. GitLab excludes pending queue time and ignores the initial attempt of a retried/re-run job when calculating pipeline duration.
Job duration / queued_duration Which individual job executes or waits longest. Returned by Jobs API; compare jobs before adding runners or changing YAML.
Compute usage How much runner execution the namespace consumed. For quota-tracked instance runners, usage is based on job running duration multiplied by the current runner/project cost factor. Parallel jobs can consume more compute than wall-clock pipeline duration.

4. Critical path: the jobs that actually control feedback latency

The critical path is the longest chain of scheduling and dependency constraints from pipeline creation to the required result. A ten-minute job outside the merge-critical path may be less urgent than a two-minute job that blocks five downstream jobs.

Stages, needs, resource groups, runner scarcity, delayed/manual jobs, and external dependencies all change the observed path. Draw the graph, then confirm it with timestamps. More parallel jobs do not automatically shorten the critical path if the runner fleet or external service cannot execute them concurrently.

5. interruptible means “safe to abandon,” not “low priority”

interruptible: true tells GitLab that a running job may be canceled when redundant-pipeline auto-cancellation applies. The default is false. The project/workflow cancellation policy controls whether that marker actually causes cancellation.

workflow:
  auto_cancel:
    on_new_commit: interruptible

static-analysis:
  script: ./run-static-analysis.sh
  interruptible: true

deploy-production:
  script: ./deploy-known-artifact.sh
  interruptible: false

A pure compile, lint, or test of an obsolete commit is often safe to abandon. A stateful migration, release publication, production deployment, or other non-idempotent operation usually is not. Under the default conservative mode, once a non-interruptible job has started, the pipeline is protected from redundant-pipeline cancellation.

6. Retry is a fault-classification policy

GitLab automatic retry is deliberately bounded: retry:max can be 0, 1, or 2. The dangerous configuration is a broad retry of every failure, because deterministic defects simply run several times and consume more compute before reporting the same defect.

integration-probe:
  script: ./probe-dependency.sh
  retry:
    max: 1
    when:
      - runner_external_dependency_failure
      - runner_interrupted

worker:
  script: ./worker-test.sh
  retry:
    max: 1
    exit_codes: 75

GitLab 19.1 made failure reasons more granular. New configuration should prefer specific current reasons such as runner_external_dependency_failure, runner_interrupted, or runner_configuration_error rather than relying on deprecated broad timeout/stuck categories. An exit-code contract is appropriate only if your wrapper assigns that code consistently to a known retryable condition.

7. allow_failure changes governance, not reliability

A retry says “try the same required work again under a narrowly understood transient condition.” allow_failure: true instead says “this failed job does not make the pipeline fail.” Those are fundamentally different policies.

advisory-benchmark:
  script: ./benchmark.sh
  allow_failure: true

required-quality-gate:
  script: ./verify-contract.sh
  allow_failure: false

GitLab also supports allow_failure:exit_codes for explicitly tolerated result classes. Treat this as a release-policy decision: a required security, correctness, migration, or deployment check should not become advisory merely to make the pipeline green.

8. Debugging evidence: config → graph → job → runner → dependency

Start with the cheapest, safest evidence. CI Lint checks syntax and can simulate pipeline creation. The expanded configuration proves what includes, extends, defaults, and rules produced. Pipeline and Jobs APIs expose timestamps, status, runner identity, failure reasons, and artifacts. glab ci trace streams the job log.

glab ci lint --dry-run --include-jobs
glab ci status --compact
glab ci get --pipeline-id 12345 --output json --with-job-details
glab ci trace 67890
Do not turn on CI_DEBUG_TRACE casually. GitLab warns that debug logging can expose every variable available to the job and uploads that output to the job log. Prefer ordinary logs and API metadata first; if deep trace is unavoidable, use a disposable secret-free job, restrict log visibility, and remove sensitive debug evidence afterward.

9. Cost model: optimize resource consumption without inventing prices

For quota-tracked instance runners, GitLab documents compute usage conceptually as job duration / 60 × cost factor. The exact quota and cost factors depend on offering, subscription, runner type, and project context, so production policy should read the current Usage quotas/runner usage surfaces instead of embedding a price table in YAML.

Also measure bytes and transfer work: large images, repeatedly restored caches, and oversized artifacts can increase latency even when compute time looks small. A cache is beneficial only when the time and network saved by reuse exceed the restore/upload cost and the integrity risk is acceptable.

10. Read-only inspection before optimization

For a real project, capture the pipeline/ref/SHA first, then collect structured evidence without changing configuration:

PROJECT_ID="12345678"
PIPELINE_ID="12345"

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID" \
  | jq "{id,sha,ref,status,created_at,started_at,finished_at,duration,queued_duration}"

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs?per_page=100" \
  | jq "map({id,name,stage,status,created_at,started_at,finished_at,duration,queued_duration,runner:(.runner.description // null),failure_reason})"

That snapshot becomes the baseline. Do not tune runner count, retries, cache, stages, or failure policy until you can name the observed bottleneck.

Summary

Pipeline performance has multiple clocks and multiple trust boundaries. Measure wall-clock feedback, queueing, critical path, job running time, retries, runner identity, artifact/cache transfer, and compute separately. Mark work interruptible only when cancellation is safe; retry only a classified transient failure; keep required gates required; and debug from structured evidence before increasing verbosity or capacity.

Knowledge check

Why can total compute usage exceed pipeline wall-clock duration?

What does interruptible: true guarantee by itself?

Why is retry: 2 dangerous on a deterministic unit-test defect?

What is the difference between retry and allow_failure?

Which evidence should you inspect before adding more runners?

Official references

Primary sources used for the current GitLab 19.3 behavior taught in this lesson:

Next lesson

Guided workflow — measure before optimizing

You will build a tiny disposable baseline, collect timestamp and runner evidence, apply safe cancellation and retry policy, and quantify what actually changed.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.