Checkpoint Lab — Pipeline Performance, Cost Control, Caching Strategy, Queue Reduction, Selective Execution, and Optimization
Reduce a disposable pipeline's verified end-to-end latency, preserve coverage, and document one rejected optimization with an evidence packet and rollback plan.
Learning objectives
- Measure a disposable baseline and predict optimization effects before changing it.
- Reduce verified end-to-end latency while preserving the required test set.
- Compare wall latency, queue, total execution and I/O evidence.
- Reject one tempting optimization with explicit reasoning.
- Produce an evidence packet and exact cleanup/rollback plan.
1. Checkpoint mission
You inherit a synthetic merge-request pipeline whose integration path is slow. Your task is to reduce verified end-to-end latency using evidence, keep every required test, and document one rejected optimization. No real runner, cache backend or production project is touched.
Pass condition: a lower modeled wall time is not enough. The final evidence must preserve the exact source identity, required jobs/tests, runner trust assumption, and a cost/compute-like comparison.
2. Preflight and assumptions
- Docs verified: 2026-09-13.
- GitLab/Runner semantics baseline: 19.3.2.
- Python 3.11+ standard library.
- Synthetic project only; no credentials or remote mutations.
-
All required checks are listed in
required_checks.json; you may not silently delete one.
mkdir -p ch35-checkpoint/{input,evidence}
cd ch35-checkpoint
3. Predict before changing
Record at least two state predictions:
cat > evidence/predictions.txt <<'TXT'
1. Removing a stage-only wait with an explicit needs edge will reduce modeled wall time without removing tests.
2. A huge cache will be rejected if its transfer overhead exceeds saved setup time.
3. Increasing integration shards beyond available capacity will increase queue exposure.
4. Final evidence will contain the same required check set as baseline.
TXT
4. Create baseline data and required checks
cat > input/baseline.json <<'JSON'
{
"pipeline_id": 3550,
"pipeline_source": "merge_request_event",
"sha": "3535353535353535353535353535353535353550",
"runner_pool": "trusted-linux",
"jobs": [
{"name":"build","needs":[],"queue":5,"run":95,"io":28},
{"name":"unit","needs":["build"],"queue":7,"run":105,"io":12},
{"name":"integration","needs":["build","unit"],"queue":12,"run":190,"io":14},
{"name":"lint","needs":[],"queue":3,"run":55,"io":3},
{"name":"package","needs":["integration"],"queue":4,"run":40,"io":36}
]
}
JSON
cat > input/required_checks.json <<'JSON'
["build","unit","integration","lint","package"]
JSON
5. Analyze baseline and candidate
Create one analyzer and keep it with the evidence packet:
cat > optimize.py <<'PYLAB'
import json,sys
p=json.load(open(sys.argv[1]))
jobs={j['name']:j for j in p['jobs']}; memo={}
def finish(n):
if n in memo:return memo[n]
j=jobs[n]
start=max([finish(x) for x in j['needs']] or [0])+j['queue']
memo[n]=start+j['run']
return memo[n]
required=set(json.load(open('input/required_checks.json')))
actual=set(jobs)
out={'pipeline_id':p['pipeline_id'],'pipeline_source':p['pipeline_source'],'sha':p['sha'],'runner_pool':p['runner_pool'],'wall_s':max(finish(n) for n in jobs),'total_run_s':sum(j['run'] for j in jobs.values()),'sum_queue_s':sum(j['queue'] for j in jobs.values()),'io_units':sum(j['io'] for j in jobs.values()),'required_checks_preserved':required==actual,'critical_finish':memo}
print(json.dumps(out,indent=2))
PYLAB
python optimize.py input/baseline.json | tee evidence/baseline.json
6. Candidate: remove unnecessary unit → integration serialization
Assume unit and integration both consume the build output but integration does not require unit’s result. The baseline accidentally serialized them. Change only the dependency graph; do not remove a job.
python - <<'PYLAB'
import json
p=json.load(open('input/baseline.json'))
for j in p['jobs']:
if j['name']=='integration': j['needs']=['build']
p['pipeline_id']=3551
json.dump(p,open('input/candidate.json','w'),indent=2)
PYLAB
python optimize.py input/candidate.json | tee evidence/candidate.json
7. Verify correctness coverage did not shrink
The analyzer must report
required_checks_preserved: true. In a real pipeline,
also compare JUnit/report coverage and any path-based rule
decisions. A graph optimization is acceptable only if the dependency
semantics are true.
python - <<'PYLAB'
import json
b=json.load(open('evidence/baseline.json')); c=json.load(open('evidence/candidate.json'))
print('wall_delta_s=', c['wall_s']-b['wall_s'])
print('run_delta_s=', c['total_run_s']-b['total_run_s'])
print('queue_delta_s=', c['sum_queue_s']-b['sum_queue_s'])
print('coverage_ok=', c['required_checks_preserved'])
PYLAB
8. Reject one tempting optimization
Now model a “cheap” global cache: it saves 15 seconds of unit setup but adds 60 I/O units to both unit and integration. Because this exercise has no measured conversion from I/O units to seconds, it cannot claim end-to-end improvement. Reject it pending a measured transfer-time experiment.
cat > evidence/rejected-optimization.txt <<'TXT'
Rejected: add a large shared dependency cache to unit + integration.
Reason: no measured cache hit/restore/save timing; synthetic estimate adds substantial transfer I/O for only a small execution saving.
Security note: do not merge protected/non-protected cache trust domains to chase hit rate.
Required follow-up: measure cache bytes, hit rate, restore/save seconds and miss behavior on the same runner class.
TXT
9. Preserve runner capacity and security assumptions
The candidate keeps runner_pool=trusted-linux. Do not
claim a cost saving by migrating protected work to a cheaper
untrusted shared runner. If future measurements show queue
dominance, tune capacity/request flow within the same trust model
first.
10. Map the accepted change to GitLab YAML
The structural change is simply making both tests depend on build, while package still depends on integration:
build:
stage: build
script: ./build.sh
artifacts:
paths: [dist/]
unit:
stage: test
needs:
- job: build
artifacts: false
script: ./test-unit.sh
integration:
stage: test
needs:
- job: build
artifacts: true
script: ./test-integration.sh
package:
stage: package
needs:
- job: build
artifacts: true
- job: integration
artifacts: false
script: ./package.sh
11. Build the evidence packet
Preserve source/config/runner/correctness and before/after metrics together.
cp input/required_checks.json evidence/
cp optimize.py evidence/
cat > evidence/assumptions.txt <<'TXT'
Docs verified 2026-09-13; GitLab/Runner semantics baseline 19.3.2.
Synthetic local model; no real scheduler/cache/storage network measured.
Candidate changes dependency graph only; required job set is unchanged.
Runner trust pool remains trusted-linux.
Compute-like comparison uses total execution seconds; real GitLab compute uses duration × cost factor.
TXT
sha256sum evidence/* > evidence/SHA256SUMS
cat evidence/SHA256SUMS
12. Verification checklist
- Baseline and candidate preserve pipeline source and exact SHA.
- Candidate has a distinct synthetic pipeline ID.
- Required check set is identical.
- Wall-time delta is computed from the same analyzer.
- Total execution and queue sums are reported separately.
- Runner trust class is unchanged.
- Rejected optimization includes a measurable follow-up condition.
- Evidence packet has digests.
13. Rollback and cleanup
If production p95 or correctness regresses, restore the old dependency graph. For this disposable lab, remove only the generated directory:
cd ..
test -d ch35-checkpoint && rm -rf ch35-checkpoint
14. Bridge to Chapter 36
Chapter 35 optimized one pipeline using measured evidence. Chapter 36 moves outward to Governance, Compliance Pipelines, Scan Execution Policies, Pipeline Execution Policies, and Organization-Wide Controls: how an organization enforces required pipeline behavior without turning governance into opaque or unsafe centralization.
Knowledge check
What made the accepted optimization safe?
It changed only the dependency graph, preserved every required job, kept runner trust unchanged, and compared the same source/workload.
Why was the large cache rejected?
There was no measured transfer-time evidence showing it improved end-to-end latency, and it increased I/O assumptions.
Why keep total execution separate from wall time?
Wall time can fall through concurrency while total compute-like work stays the same or rises.
What would invalidate the needs change?
If integration actually depends on unit output/result, the dependency graph would be semantically wrong.
What is the rollback trigger?
A verified regression in p95 latency, correctness coverage, security/isolation, cost bounds or external delivery behavior.
15. Checkpoint summary
You improved a measured critical path without removing required assurance, preserved exact identity and runner trust, rejected an unproven cache “optimization,” and produced reproducible evidence. That is performance engineering rather than YAML folklore.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-13.
Examples use GitLab/GitLab Runner 19.3.2 semantics as the
timestamped baseline where a concrete version matters. Project CI/CD
analytics, needs, caches, artifacts and job rules are
available across Free/Premium/Ultimate. Current project analytics
exposes median and p95 pipeline duration. Job execution on instance
runners contributes compute usage; created/pending queue time does
not, so latency and compute cost are related but distinct metrics.
Runner flow is bounded by global concurrent, per-runner
limit, and job-request
request_concurrency; long-polling misconfiguration can
create queue delays. Caches are an optimization and are not
guaranteed to exist. With needs, jobs fetch artifacts
only from listed dependencies, and artifacts: false can
avoid transfers. rules:changes:compare_to can skip
unaffected work, but only after correctness requirements are made
explicit. The mandatory labs use synthetic local data and Python
standard-library tooling; no paid analytics, cloud account,
privileged runner or production workload is required.
- Compute minutes — official reference.
- Instance runner compute usage — official reference.
- CI/CD analytics — official reference.
- Runner advanced configuration — official reference.
- Caching in GitLab CI/CD — official reference.
- Caching examples — official reference.
- Job artifacts — official reference.
- CI/CD YAML reference — official reference.
- needs DAGs — official reference.
- Job rules — official reference.
- Pipeline settings and auto-cancel — official reference.
- Runner monitoring — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.