Chapter 35Lesson 02~265 minutes

Pipeline Performance, Cost Control, Caching Strategy, Queue Reduction, Selective Execution, and Optimization: Guided Hands-On Workflow and Core Operations

Benchmark a disposable pipeline model, change needs, cache, artifact scope and rules one variable at a time, and prove which changes improve end-to-end latency without hiding correctness costs.

Hands-onneedsCacheArtifactsRules

Learning objectives

  • Benchmark a synthetic pipeline using deterministic local data.
  • Change DAG, cache, artifact and selective-execution assumptions one at a time.
  • Compute queue/run/wall/compute-like metrics before and after.
  • Identify an optimization that improves a local job but not end-to-end latency.
  • Produce a challenge decision based on the correct causal layer.

1. Disposable lab setup

This mandatory workflow is a faithful local simulation. It does not need a GitLab token, paid analytics, Docker daemon or cloud runner. The model uses fixed synthetic job traces so every learner can reproduce the reasoning.

mkdir -p ch35-lab/input ch35-lab/evidence
cd ch35-lab
python --version

2. Timestamped assumptions

  • Docs verified: 2026-09-13.
  • GitLab/Runner baseline: 19.3.2.
  • Project CI/CD analytics is available on Free; per-job limited-availability analytics is not required.
  • Compute-like cost in this lab is total execution seconds, not queue seconds.
  • All data is synthetic; no production runner or cache backend is touched.

3. Create the baseline model

cat > input/pipeline.json <<'JSON'
{
  "pipeline_id": 3500,
  "pipeline_source": "merge_request_event",
  "sha": "3535353535353535353535353535353535353535",
  "capacity": 2,
  "jobs": [
    {"name":"build","needs":[],"run":90,"queue":5,"cache_io":0,"artifact_out":30,"required":true},
    {"name":"unit","needs":["build"],"run":110,"queue":8,"cache_io":12,"artifact_out":2,"required":true},
    {"name":"integration","needs":["build"],"run":180,"queue":75,"cache_io":12,"artifact_out":3,"required":true},
    {"name":"lint","needs":[],"run":50,"queue":4,"cache_io":8,"artifact_out":0,"required":true},
    {"name":"package","needs":["integration"],"run":35,"queue":3,"cache_io":0,"artifact_out":35,"required":true}
  ]
}
JSON

4. Compute a simple comparable baseline

The script intentionally separates wall-like critical path, queue exposure and total execution. It is a teaching model, not GitLab scheduler emulation.

cat > analyze.py <<'PYLAB'
import json,sys
p=json.load(open(sys.argv[1]))
jobs={j['name']:j for j in p['jobs']}
memo={}
def finish(name):
    if name in memo: return memo[name]
    j=jobs[name]
    start=max([finish(n) for n in j['needs']] or [0])+j['queue']
    memo[name]=start+j['run']
    return memo[name]
wall=max(finish(n) for n in jobs)
run=sum(j['run'] for j in jobs.values())
queue=sum(j['queue'] for j in jobs.values())
io=sum(j['cache_io']+j['artifact_out'] for j in jobs.values())
print(json.dumps({'wall_s':wall,'total_run_s':run,'sum_queue_s':queue,'io_units':io,'critical_finish':memo},indent=2))
PYLAB
python analyze.py input/pipeline.json | tee evidence/baseline.json

5. Optimization A: use DAG needs to remove stage-only waiting

In real GitLab, needs allows a job to start as soon as its dependencies finish rather than waiting for an entire previous stage. The baseline model already expresses a DAG, so model a deliberately stage-serialized package job and then restore the direct dependency.

python - <<'PYLAB'
import json
p=json.load(open('input/pipeline.json'))
for j in p['jobs']:
    if j['name']=='package': j['needs']=['unit','integration']
json.dump(p,open('input/stagebarrier.json','w'),indent=2)
PYLAB
python analyze.py input/stagebarrier.json | tee evidence/stagebarrier.json
python analyze.py input/pipeline.json | tee evidence/needs-dag.json

6. Optimization B: reduce artifact scope

A downstream job should not fetch artifacts it never reads. In GitLab, needs narrows artifact sources; artifacts: false disables transfer from a specific needed job. The lab represents that as lower I/O units, not magical runtime savings.

python - <<'PYLAB'
import json
p=json.load(open('input/pipeline.json'))
for j in p['jobs']:
    if j['name']=='package': j['artifact_out']=8
json.dump(p,open('input/smaller-artifacts.json','w'),indent=2)
PYLAB
python analyze.py input/smaller-artifacts.json | tee evidence/smaller-artifacts.json

7. Optimization C: test whether cache transfer is worth it

Assume unit tests save 10 seconds of setup when cache hits but spend 12 units transferring/extracting it. If the cache backend is remote and slow, the net may be negative. Build a scenario with a cache miss that increases I/O without reducing run time.

python - <<'PYLAB'
import json
p=json.load(open('input/pipeline.json'))
for j in p['jobs']:
    if j['name']=='unit':
        j['cache_io']=30
        j['run']=110
json.dump(p,open('input/cache-regression.json','w'),indent=2)
PYLAB
python analyze.py input/cache-regression.json | tee evidence/cache-regression.json

8. Optimization D: selective execution with a correctness manifest

For a docs-only change, imagine integration tests are truly unaffected. Do not simply remove the job; record the selection rule and required invariant. In GitLab a real job can use rules:changes:compare_to.

cat > evidence/selection-policy.txt <<'TXT'
Scenario: docs-only MR
Selected: lint, docs validation
Skipped: integration
Safety condition: no shared schema/build/runtime files changed
Fallback: if dependency mapping is uncertain, run integration
TXT
integration:
  script: ./test-integration.sh
  rules:
    - changes:
        compare_to: refs/heads/main
        paths:
          - src/**/*
          - schema/**/*
          - integration/**/*
    - when: never

9. Optimization E that may not improve wall time: add shards under saturated capacity

Split the 180-second integration job into two 100-second shards, but assume each waits 120 seconds because the two-runner pool is already busy. The local job runtime improves, yet wall time can worsen.

python - <<'PYLAB'
import json
p=json.load(open('input/pipeline.json'))
p['jobs']=[j for j in p['jobs'] if j['name'] not in ('integration','package')]
p['jobs'] += [
 {'name':'integration_a','needs':['build'],'run':100,'queue':120,'cache_io':12,'artifact_out':2,'required':True},
 {'name':'integration_b','needs':['build'],'run':100,'queue':120,'cache_io':12,'artifact_out':2,'required':True},
 {'name':'package','needs':['integration_a','integration_b'],'run':35,'queue':3,'cache_io':0,'artifact_out':35,'required':True}
]
json.dump(p,open('input/sharded-saturated.json','w'),indent=2)
PYLAB
python analyze.py input/sharded-saturated.json | tee evidence/sharded-saturated.json

10. A coherent GitLab YAML pattern

This example shows the mechanisms without claiming they always improve performance:

stages: [build, test, package]

build:
  stage: build
  script: ./build.sh
  artifacts:
    paths: [dist/app.bin]
    expire_in: 1 day

unit:
  stage: test
  needs:
    - job: build
      artifacts: false
  cache:
    key:
      files: [requirements.lock]
    paths: [.cache/pip/]
    policy: pull
  script: ./test-unit.sh

integration:
  stage: test
  needs:
    - job: build
      artifacts: true
  script: ./test-integration.sh
  rules:
    - changes:
        compare_to: refs/heads/main
        paths: [src/**/*, schema/**/*, integration/**/*]
    - when: never

package:
  stage: package
  needs:
    - job: build
      artifacts: true
    - job: integration
      artifacts: false
  script: ./package.sh

11. Challenge: choose the layer

Your sharded test jobs each fall from 180 s to 100 s, but p95 queue rises from 20 s to 150 s and pipeline completion worsens. Which layer owns the next experiment? A strong answer chooses runner capacity/eligibility or shard count—not test code—and preserves the security/isolation requirements from Chapter 31.

Knowledge check

Why change one optimization variable at a time?

What does needs optimize?

When is a cache a regression?

Why record a selection policy for skipped tests?

Why can sharding worsen wall time?

12. Cleanup

The lab changes only local synthetic files:

cd ..
test -d ch35-lab && rm -rf ch35-lab

13. Summary

You benchmarked first, then modeled DAG, artifact, cache, selective-execution and parallelism changes separately. The key result is methodological: a faster local job is not enough; end-to-end latency, compute/cost, correctness and runner trust must all remain acceptable.

Continue

Optimization design choices and tradeoffs

Lesson 3 balances parallelism, cache, selective execution, runner sizing, cost and correctness.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-13. Examples use GitLab/GitLab Runner 19.3.2 semantics as the timestamped baseline where a concrete version matters. Project CI/CD analytics, needs, caches, artifacts and job rules are available across Free/Premium/Ultimate. Current project analytics exposes median and p95 pipeline duration. Job execution on instance runners contributes compute usage; created/pending queue time does not, so latency and compute cost are related but distinct metrics. Runner flow is bounded by global concurrent, per-runner limit, and job-request request_concurrency; long-polling misconfiguration can create queue delays. Caches are an optimization and are not guaranteed to exist. With needs, jobs fetch artifacts only from listed dependencies, and artifacts: false can avoid transfers. rules:changes:compare_to can skip unaffected work, but only after correctness requirements are made explicit. The mandatory labs use synthetic local data and Python standard-library tooling; no paid analytics, cloud account, privileged runner or production workload is required.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.