Chapter 35Lesson 04~240 minutes

Pipeline Performance, Cost Control, Caching Strategy, Queue Reduction, Selective Execution, and Optimization: Diagnostics, Failure Modes, Security, and Performance

Diagnose optimization failures such as p95 regressions, skipped tests, cache-transfer overhead, queue amplification and unsafe shared-runner cost cutting while preserving first-failure evidence.

Diagnosticsp95Queue saturationCache I/OSecurity

Learning objectives

  • Diagnose tail-latency regressions hidden by averages.
  • Recognize skipped-test optimizations that invalidate correctness evidence.
  • Separate cache/network overhead from execution bottlenecks.
  • Identify parallelism-driven queue amplification.
  • Reject cost reductions that weaken runner isolation.

1. Evidence-first diagnostic sequence

Preserve the original pipeline/job IDs and first-failure evidence. Confirm pipeline source/ref/SHA and compiled YAML; confirm rule results and selected jobs/tests; inspect graph and queue; confirm runner/executor/image/toolchain; inspect script and transfer phases; inspect artifacts/cache metadata and external target state; then change only the smallest proven layer.

Do not rerun first. A rerun changes queue, cache, runner, external-service and timing state, which can erase the condition you need to understand.

2. Failure mode: average improves while tail latency worsens

Before optimization: median 7m, p95 10m. After adding shards: median 5m, p95 18m. The average team experience may look better, but rare queue saturation has become severe. Inspect queue p95 by runner pool and job family; do not declare success from median alone.

3. Failure mode: required tests were skipped

A path rule cuts compute by 40%, but a shared schema change does not trigger integration tests. The pipeline is faster because assurance disappeared. The failing layer is repository rule/dependency modeling, not the test runner.

# INTENTIONALLY BROKEN: too narrow for a shared schema dependency
integration:
  script: ./test-integration.sh
  rules:
    - changes:
        paths: [service-a/**/*]
    - when: never

Repair by including shared triggers or falling back to the full suite when dependency mapping is uncertain.

4. Failure mode: cache dominates network

A 1.2 GB cache saves 20 seconds of package resolution but takes 65 seconds to download and extract on cold autoscaled workers. Measure cache restore/save sections and bytes. Narrow the path, key, or remove it. Do not raise runner size to fix object-storage transfer unless evidence shows CPU compression is the bottleneck.

5. Failure mode: matrix parallelism increases queue time

Eight shards each run 45 seconds, but only two eligible runners exist. Queue p95 rises beyond the original two-shard pipeline. The local job metric improved; the system metric regressed. Reduce shard count or address bounded capacity/request flow.

6. Failure mode: cost reduced by using an insecure shared runner

A protected release job is moved from an ephemeral protected runner to a cheap shared Shell runner. Compute cost falls, but untrusted jobs can share host state and credentials. This violates the trust model; reject the optimization even if latency/cost metrics look better.

7. Intentionally broken “optimization bundle”

This combines multiple unreviewed changes, so any result is hard to attribute and some are unsafe:

# INTENTIONALLY BROKEN — do not copy.
default:
  cache:
    key: global
    paths: [./]

fast_tests:
  parallel: 12
  script: ./test.sh --changed-only
  rules:
    - changes: [src/**/*]

release:
  tags: [cheap-shared]
  script: ./release.sh

Problems: giant global cache, uncontrolled cache content, high shard count without capacity evidence, ambiguous “changed-only” coverage, and release trust boundary regression. Repair each layer separately and compare against the same baseline.

8. Runner configuration can mimic application slowness

If jobs remain pending, inspect eligible runner count and Runner concurrent/limit/request_concurrency. Current GitLab Runner warns about long-polling scenarios such as global concurrency below runner count or high-volume runners capped at request concurrency 1.

9. Blind retries distort optimization data

Retries can warm caches, land on a different runner, or occur after load subsides. Preserve the first attempt and label retries as separate samples. Do not combine them without explaining sampling bias.

10. Do not optimize CI by weakening deployment verification

Removing rollout checks can make the pipeline “finish” earlier while moving failure detection outside the pipeline. Deployment submission and external health are different states; performance work must keep the required health evidence.

11. Least-destructive corrections

Evidence Likely layer Bounded correction
p95 queue spikes after more shards runner capacity / parallelism reduce shards or bounded capacity/request-flow test
cache restore > work saved cache/I-O narrow/remove cache; remeasure
shared file bypasses integration rule rules/correctness broaden dependency map or full-suite fallback
median improves, p95 worsens tail capacity variability optimize p95 cause, not average
cost win requires weaker runner security/governance reject; find savings inside same trust boundary

Knowledge check

Why preserve the first attempt before retry?

What proves a selective-execution optimization is safe?

What if cache restore is slower than dependency install?

Why can 12 shards be slower than 4?

When should a cost optimization be rejected even if it saves money?

12. Summary

Optimization failures are causal-layer failures: tail queue, rule coverage, transfer overhead, capacity saturation or trust regression. Preserve evidence, isolate the layer and correct only that layer.

Continue

Checkpoint measured optimization lab

Lesson 5 reduces verified latency while preserving required checks and documents a rejected optimization.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-13. Examples use GitLab/GitLab Runner 19.3.2 semantics as the timestamped baseline where a concrete version matters. Project CI/CD analytics, needs, caches, artifacts and job rules are available across Free/Premium/Ultimate. Current project analytics exposes median and p95 pipeline duration. Job execution on instance runners contributes compute usage; created/pending queue time does not, so latency and compute cost are related but distinct metrics. Runner flow is bounded by global concurrent, per-runner limit, and job-request request_concurrency; long-polling misconfiguration can create queue delays. Caches are an optimization and are not guaranteed to exist. With needs, jobs fetch artifacts only from listed dependencies, and artifacts: false can avoid transfers. rules:changes:compare_to can skip unaffected work, but only after correctness requirements are made explicit. The mandatory labs use synthetic local data and Python standard-library tooling; no paid analytics, cloud account, privileged runner or production workload is required.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.