Chapter 35Lesson 01~220 minutes

Pipeline Performance, Cost Control, Caching Strategy, Queue Reduction, Selective Execution, and Optimization: Concepts, Architecture, and Mental Model

Optimize GitLab CI/CD from measured baselines: critical path, queue time, cache and artifact I/O, runner usage, selective execution, correctness evidence, and cost assumptions.

PerformanceCritical pathQueueComputeCorrectness

Learning objectives

  • Distinguish end-to-end pipeline latency, job execution time, queue time and compute usage.
  • Find the critical path instead of optimizing the longest-looking job in isolation.
  • Explain when cache/artifact transfer and parallelism help or hurt.
  • Treat selective execution as a correctness policy, not merely a cost switch.
  • Define a baseline/effect experiment that preserves source/configuration/runner comparability.

1. The practical problem: fast-looking jobs can still produce a slow, expensive pipeline

Chapter 34 established an observability baseline. Optimization begins only after that baseline exists. A pipeline can be slow because jobs wait in a queue, because the DAG serializes independent work, because every job downloads large artifacts, because a cache miss is more expensive than rebuilding, or because too many parallel shards saturate the runner pool. These causes require different corrections.

Cost and latency also diverge. GitLab compute usage for instance runners is based on job execution, not created/pending queue time. A pipeline can have terrible developer latency from queueing while consuming little additional compute. Conversely, aggressive parallelism can reduce wall-clock time while consuming more total compute minutes.

Optimization rule: never change pipeline structure until you have a comparable baseline that preserves pipeline source, source SHA, compiled configuration, runner context, selected tests and external side effects.

2. Mental model: baseline → critical path → hypothesis → bounded change → comparable rerun → verify

Start with evidence, not a YAML trick. Measure pipeline wall time, per-job queue/run time, cache/artifact transfer, runner usage and correctness. Identify the critical path: the dependency chain whose delays determine completion. Form one hypothesis, change one layer, rerun a comparable workload, then verify latency, compute/cost and correctness together.

Measured optimization loop
            flowchart TD
             A[Baseline IDs + SHA + graph] --> B[Queue / critical path / I-O analysis]
             B --> C[One optimization hypothesis]
             C --> D[Bounded config or runner change]
             D --> E[Comparable rerun]
             E --> F[Latency + cost + correctness evidence]
             F -->|better and safe| G[Adopt + monitor]
             F -->|regression / uncertainty| H[Rollback / reject]
          

3. Freeze the state you are about to optimize

State layer Record Why
Source/revision CI_PIPELINE_SOURCE, ref, exact CI_COMMIT_SHA Keeps before/after runs comparable.
Compiled config/rules merged YAML, rule result, job graph Proves whether fewer jobs came from deliberate selection.
Pipeline/job timing pipeline/job IDs, queue, run, start/finish Separates wall time from capacity and script cost.
Runner/executor runner pool, executor, version, limits Explains queue pressure and cost/isolation boundaries.
Cache/artifact I/O key, hit/miss, bytes transferred, artifact dependencies Shows whether transfer saves or adds time.
Correctness evidence tests selected, pass/fail, reports, coverage/invariants Prevents speedups that quietly reduce assurance.
External/deployment environment/deployment and target health Prevents optimizing CI while target readiness regresses.
Cost assumptions compute minutes, runner unit cost, storage/egress assumptions Makes cost conclusions explicit rather than guessed.

4. Four different performance numbers

Metric Meaning Common mistake
Pipeline wall time Elapsed time from pipeline creation to terminal state Treating total job seconds as wall time
Job execution duration Time runner executes a job Ignoring queue before runner assignment
Queued duration Time waiting for eligible capacity Trying to speed test code to fix a capacity bottleneck
Compute usage Runner job duration × cost factor Assuming queue time directly consumes compute quota

5. Critical path beats “optimize the biggest job”

Suppose build runs 180 s, unit tests 140 s, integration 220 s and package 40 s. If unit and integration both depend on build and package depends only on integration, the wall-time critical path is build → integration → package. A 60-second improvement to unit tests may save compute but not change completion time at all.

build=180
unit=140
integration=220
package=40
critical=build+integration+package
print("critical_path_s=", critical)

6. Queue is a separate optimization surface

Runner capacity, tags, executor eligibility, global concurrent, per-runner limit and request_concurrency can dominate wait time. Current Runner guidance warns that long polling with too-low concurrency/request settings can create severe processing delays.

Adding more matrix jobs while capacity is saturated can increase p95 queue time even though each shard is shorter. Optimize the system, not a single job.

7. Cache is a speculative optimization, not a dependency guarantee

GitLab caches are designed for reusable dependencies and may live locally or in distributed object storage. They can miss. They can also cost more to compress, upload, download and extract than the work they avoid. Key caches from dependency-lock content where practical, measure bytes and hit rate, and make jobs correct when the cache is absent.

Security boundary: protected/non-protected cache separation exists for a reason. Do not merge trust domains just to improve hit rate.

8. Artifacts are dataflow, and dataflow has transfer cost

Later-stage jobs normally fetch artifacts from earlier stages unless you constrain dependencies. With needs, artifact downloads come only from jobs listed in needs; an edge can set artifacts: false. This can shorten both graph wait and transfer time, but only when the downstream job truly does not need those bytes.

9. Selective execution saves resources only if the selection is correct

rules:changes:compare_to can skip jobs when relevant files did not change. In a monorepo this can be powerful. But a “frontend-only” change may still require shared-contract or security tests. Define the dependency model and a safe fallback before skipping anything.

10. More parallelism can reduce latency and increase cost—or increase latency too

Parallel jobs reduce a critical path when capacity exists and setup/merge overhead is modest. If runner capacity is constrained, extra shards become queued work and may worsen tail latency. Total compute often rises because setup and duplicated dependency work occur per shard.

11. Read-only baseline worksheet

Before changing YAML, capture an evidence row per job:

pipeline_id,job_id,name,queue_s,run_s,cache_bytes,artifact_bytes,status,selected_tests
3500,35001,build,8,180,0,42000000,success,build
3500,35002,unit,12,140,18000000,2000000,success,unit:all
3500,35003,integration,160,220,18000000,3000000,success,integration:all
3500,35004,package,4,40,0,45000000,success,package

This already shows two possible causes: integration has both the longest execution and dominant queue. You need graph and runner evidence to know which change attacks wall time.

12. Common wrong approaches

Shortcut Why it fails Safer pattern
“Turn on cache everywhere” Transfer/compression can exceed saved work Measure hit rate and transfer time
“Shard every test” Can saturate runners and duplicate setup Shard only critical-path work with capacity
“Skip tests by path” May miss hidden/shared dependencies Codify dependency model and safety fallback
“Use a larger shared runner” Can change trust/isolation boundary Preserve Chapter 31 security model
“Optimize average only” p95/tail can regress Track median and p95 plus queue/run separation

Knowledge check

Why can queue time worsen while compute usage stays similar?

What is the critical path?

Why must cache misses remain correct?

What can needs: artifacts: false optimize?

Why is rules:changes a correctness decision?

13. Summary

Optimization is an experiment: freeze identity and correctness evidence, distinguish queue/run/wall/compute, identify the critical path, change one bounded layer, rerun comparably, and adopt only improvements that preserve assurance and trust boundaries.

Continue

Guided measured optimization workflow

Lesson 2 benchmarks a disposable model and changes DAG, cache, artifacts, rules and sharding one variable at a time.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-13. Examples use GitLab/GitLab Runner 19.3.2 semantics as the timestamped baseline where a concrete version matters. Project CI/CD analytics, needs, caches, artifacts and job rules are available across Free/Premium/Ultimate. Current project analytics exposes median and p95 pipeline duration. Job execution on instance runners contributes compute usage; created/pending queue time does not, so latency and compute cost are related but distinct metrics. Runner flow is bounded by global concurrent, per-runner limit, and job-request request_concurrency; long-polling misconfiguration can create queue delays. Caches are an optimization and are not guaranteed to exist. With needs, jobs fetch artifacts only from listed dependencies, and artifacts: false can avoid transfers. rules:changes:compare_to can skip unaffected work, but only after correctness requirements are made explicit. The mandatory labs use synthetic local data and Python standard-library tooling; no paid analytics, cloud account, privileged runner or production workload is required.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.