Chapter 32Lesson 04~235 minutes

Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Diagnostics, Failure Modes, and Production Practices

Diagnose noisy measurements, oversized matrices, stale caches, runner-capacity bottlenecks, version drift and false savings without deleting evidence or weakening tests.

DiagnosticsCache correctnessQueueingMatrix explosionRecovery

Learning objectives

  • Use an evidence-first sequence to distinguish code/runtime bottlenecks from queue, runner, cache and external-service bottlenecks.
  • Diagnose misleading one-run benchmarks and compare equivalent workloads statistically.
  • Recognize matrix explosion, cache-key overbreadth, self-hosted capacity starvation and image/tool drift.
  • Preserve run/attempt and first-failure evidence before reruns, cancellation or cache deletion.
  • Apply the least destructive correction and rerun the smallest equivalent scope without cutting required coverage.

1. Evidence-first diagnostic sequence

Performance troubleshooting follows the same incident-safe pattern as earlier chapters. Preserve the run ID/attempt and first-failure evidence; confirm event/ref/SHA and workflow revision; confirm evaluated conditions and permissions; inspect the job graph; inspect queue/runner eligibility; record runner image/toolchain; inspect step timings; inspect cache/artifact I/O; inspect external dependencies; then change the smallest causal variable and rerun an equivalent workload.

This order prevents a common mistake: changing cache keys or matrix size before proving whether the wait is in runner capacity, application compute, package network, artifact upload or an external test environment.

2. Failure mode: optimizing from one noisy run

A single baseline run takes 9 minutes, the next optimized run takes 6, and the team claims a 33% speedup. But the baseline root job waited two minutes for a runner and the optimized job started immediately. The workflow change may have done almost nothing. Preserve several equivalent runs and compare medians/percentiles or at minimum queue/start context alongside execution duration.

Repair: repeat the same SHA/workload multiple times, annotate cold/warm cache state, and separate dispatch/queue signals from execution. Never discard the “outlier” merely because it makes the optimization look worse; investigate why it differs.

3. Failure mode: the fastest pipeline is the one that stopped testing

A team notices integration tests dominate the critical path and changes a path filter so most pull requests never run them. Wall-clock time improves immediately—but so does the probability of shipping an undetected cross-component regression. This is not optimization; it is a coverage-policy change.

# BROKEN OPTIMIZATION — do not use as a speed shortcut.
jobs:
  integration:
    if: ${{ false }}  # “temporary” test removal makes the benchmark invalid.
    runs-on: ubuntu-24.04
    steps:
      - run: ./run-required-integration-tests.sh

Repair: preserve the required test gate, then optimize its internal work: shard proven-independent tests, reduce redundant setup, use precise dependency caches, improve fixtures or schedule additional non-blocking depth only when the quality policy explicitly permits it.

4. Failure mode: matrix explosion

The job matrix expands OS × runtime × database × feature flags into 144 cells. Most cells complete in under a minute, but each repeats runner startup and setup. Queue time rises, per-job private billing rounds up, and meaningful failures are harder to triage.

Repair: classify dimensions by risk. Keep representative required cells on pull requests, move exhaustive compatibility to scheduled/release workflows only if governance permits, and use include for meaningful combinations rather than a Cartesian product. Stay well below the platform’s 256-cell ceiling unless coverage genuinely requires it.

5. Failure mode: broad cache key makes a fast wrong build

# BROKEN: lockfile/runtime identity is absent.
- uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
  with:
    path: ~/.cache/pip
    key: pip-${{ runner.os }}

This key can restore content created under different Python or dependency-lock states. The immediate symptom may be “great cache hit rate,” not a visible failure. The causal defect is cache identity, not network speed.

Repair by adding architecture, runtime/toolchain and lockfile digest to the primary key and by always reconciling dependencies against the lockfile. If poisoning or stale executable state is suspected, preserve cache metadata first, then bypass/re-key; delete only exact disposable cache entries after evidence capture.

6. Failure mode: blaming GitHub for an owned runner queue

Jobs targeting [self-hosted, linux, x64, gpu] wait 25 minutes. GitHub’s service is healthy, but only one eligible GPU runner exists and it is occupied. The queue is determined by your label/group routing and fleet capacity.

Repair: inspect runner-group eligibility, online/busy state, autoscaling controller health and demand patterns. Add bounded capacity or reschedule work if justified. Do not remove trust-routing labels merely to make the job start on an unintended runner.

7. Failure mode: latest images/tools turn drift into “performance regression”

A workflow uses ubuntu-latest, installs the latest compiler and floats action majors. A week later it slows down. Without exact runner image/tool versions, the team cannot tell whether source code, package resolution, action runtime or image composition changed.

Repair: use versioned runner labels where appropriate, pin executable actions by full SHA, pin language/tool versions, record them in evidence, and change one dependency at a time. -latest is a moving platform choice, not a stable benchmark environment.

8. Failure mode: external service latency is hidden inside “test time”

A test job spends six minutes waiting on a sandbox API. Adding a larger runner does nothing because CPU is idle. Use step-level timing and service-side correlation IDs where available to separate runner compute from network/provider latency. The least destructive fix may be a faithful local fixture or a better integration environment—not more runner cores.

9. Preserve first-failure evidence before rerun

Retries can erase the story by replacing attention with a later green attempt. Keep run ID, attempt, source SHA, job/step timestamps, runner identity, cache-hit state, artifact IDs and external correlation data before rerunning. A rerun should answer a hypothesis such as “same source on a fresh runner without cache reproduces the slowdown,” not simply “try again.”

10. Read-only diagnostic commands

RUN_ID=123456789
GH_REPO=OWNER/repo

gh run view "$RUN_ID" --json databaseId,attempt,event,headSha,status,conclusion,createdAt,updatedAt,url

gh api \
  -H 'Accept: application/vnd.github+json' \
  -H 'X-GitHub-Api-Version: 2026-03-10' \
  "repos/$GH_REPO/actions/runs/$RUN_ID/jobs?per_page=100" \
  --jq '.jobs[] | [.id,.name,.started_at,.completed_at,.conclusion,(.labels|join(","))] | @tsv'

gh cache list --limit 100

11. Symptom → evidence → causal layer → correction

Symptom Evidence that discriminates Likely layer Least destructive next action
High average queue, short jobs Performance Metrics + runner labels Capacity/concurrency Inspect eligible capacity before editing tests/cache.
Long dependency install, low cache hits lock hash + keys + cache metadata Cache/input identity Fix precise key or determine cache not worth it.
Fast matrix wall time, huge processing minutes job durations + cell count Graph granularity Merge tiny cells / cap parallelism.
Self-hosted wait only runner group + online/busy state Owned runner fleet Scale/reroute within trust boundary.
Performance changes after no source diff image/action/tool versions Platform/tool drift Freeze/record versions and reproduce.
Green faster run after test removal job graph + required-check coverage Correctness policy Restore missing tests; optimize internally.

12. Shortcuts this chapter rejects

  • Do not grant broader token permissions to make a cache/artifact call “work faster.”
  • Do not print contexts or secrets to investigate timing.
  • Do not delete caches/artifacts/runs before preserving failure evidence.
  • Do not move untrusted pull-request code onto persistent self-hosted runners to avoid hosted queue time.
  • Do not disable TLS, skip integrity checks or float dependencies/actions for convenience.
  • Do not replace required tests with retries or optimistic path filters.

13. Lesson summary

Performance incidents are evidence problems before they are tuning problems. Separate workload correctness, job graph, runner capacity, setup, cache, I/O and external waits; preserve the first attempt; then change one bounded cause. The checkpoint now asks you to prove two optimizations and reject one tempting change that would only make the metric look better by weakening the pipeline.

Next lesson

Checkpoint Lab — Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

A new workflow is 30% faster, but it no longer runs integration tests. Is that a valid optimization comparison?

Why can a broad cache key look successful?

Self-hosted jobs queue while hosted jobs do not. What should you inspect first?

Why pin runner/tool/action versions for benchmarks?

What should be preserved before a diagnostic rerun?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.