Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Diagnostics, Failure Modes, and Production Practices
Diagnose noisy measurements, oversized matrices, stale caches, runner-capacity bottlenecks, version drift and false savings without deleting evidence or weakening tests.
Learning objectives
- Use an evidence-first sequence to distinguish code/runtime bottlenecks from queue, runner, cache and external-service bottlenecks.
- Diagnose misleading one-run benchmarks and compare equivalent workloads statistically.
- Recognize matrix explosion, cache-key overbreadth, self-hosted capacity starvation and image/tool drift.
- Preserve run/attempt and first-failure evidence before reruns, cancellation or cache deletion.
- Apply the least destructive correction and rerun the smallest equivalent scope without cutting required coverage.
1. Evidence-first diagnostic sequence
Performance troubleshooting follows the same incident-safe pattern as earlier chapters. Preserve the run ID/attempt and first-failure evidence; confirm event/ref/SHA and workflow revision; confirm evaluated conditions and permissions; inspect the job graph; inspect queue/runner eligibility; record runner image/toolchain; inspect step timings; inspect cache/artifact I/O; inspect external dependencies; then change the smallest causal variable and rerun an equivalent workload.
This order prevents a common mistake: changing cache keys or matrix size before proving whether the wait is in runner capacity, application compute, package network, artifact upload or an external test environment.
2. Failure mode: optimizing from one noisy run
A single baseline run takes 9 minutes, the next optimized run takes 6, and the team claims a 33% speedup. But the baseline root job waited two minutes for a runner and the optimized job started immediately. The workflow change may have done almost nothing. Preserve several equivalent runs and compare medians/percentiles or at minimum queue/start context alongside execution duration.
Repair: repeat the same SHA/workload multiple times, annotate cold/warm cache state, and separate dispatch/queue signals from execution. Never discard the “outlier” merely because it makes the optimization look worse; investigate why it differs.
3. Failure mode: the fastest pipeline is the one that stopped testing
A team notices integration tests dominate the critical path and changes a path filter so most pull requests never run them. Wall-clock time improves immediately—but so does the probability of shipping an undetected cross-component regression. This is not optimization; it is a coverage-policy change.
# BROKEN OPTIMIZATION — do not use as a speed shortcut.
jobs:
integration:
if: ${{ false }} # “temporary” test removal makes the benchmark invalid.
runs-on: ubuntu-24.04
steps:
- run: ./run-required-integration-tests.sh
Repair: preserve the required test gate, then optimize its internal work: shard proven-independent tests, reduce redundant setup, use precise dependency caches, improve fixtures or schedule additional non-blocking depth only when the quality policy explicitly permits it.
4. Failure mode: matrix explosion
The job matrix expands OS × runtime × database × feature flags into 144 cells. Most cells complete in under a minute, but each repeats runner startup and setup. Queue time rises, per-job private billing rounds up, and meaningful failures are harder to triage.
Repair: classify dimensions by risk. Keep representative required
cells on pull requests, move exhaustive compatibility to
scheduled/release workflows only if governance permits, and use
include for meaningful combinations rather than a
Cartesian product. Stay well below the platform’s 256-cell ceiling
unless coverage genuinely requires it.
5. Failure mode: broad cache key makes a fast wrong build
# BROKEN: lockfile/runtime identity is absent.
- uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
with:
path: ~/.cache/pip
key: pip-${{ runner.os }}
This key can restore content created under different Python or dependency-lock states. The immediate symptom may be “great cache hit rate,” not a visible failure. The causal defect is cache identity, not network speed.
Repair by adding architecture, runtime/toolchain and lockfile digest to the primary key and by always reconciling dependencies against the lockfile. If poisoning or stale executable state is suspected, preserve cache metadata first, then bypass/re-key; delete only exact disposable cache entries after evidence capture.
6. Failure mode: blaming GitHub for an owned runner queue
Jobs targeting [self-hosted, linux, x64, gpu] wait 25
minutes. GitHub’s service is healthy, but only one eligible GPU
runner exists and it is occupied. The queue is determined by your
label/group routing and fleet capacity.
Repair: inspect runner-group eligibility, online/busy state, autoscaling controller health and demand patterns. Add bounded capacity or reschedule work if justified. Do not remove trust-routing labels merely to make the job start on an unintended runner.
7. Failure mode: latest images/tools turn drift into “performance regression”
A workflow uses ubuntu-latest, installs the latest
compiler and floats action majors. A week later it slows down.
Without exact runner image/tool versions, the team cannot tell
whether source code, package resolution, action runtime or image
composition changed.
Repair: use versioned runner labels where appropriate, pin
executable actions by full SHA, pin language/tool versions, record
them in evidence, and change one dependency at a time.
-latest is a moving platform choice, not a stable
benchmark environment.
8. Failure mode: external service latency is hidden inside “test time”
A test job spends six minutes waiting on a sandbox API. Adding a larger runner does nothing because CPU is idle. Use step-level timing and service-side correlation IDs where available to separate runner compute from network/provider latency. The least destructive fix may be a faithful local fixture or a better integration environment—not more runner cores.
9. Preserve first-failure evidence before rerun
Retries can erase the story by replacing attention with a later green attempt. Keep run ID, attempt, source SHA, job/step timestamps, runner identity, cache-hit state, artifact IDs and external correlation data before rerunning. A rerun should answer a hypothesis such as “same source on a fresh runner without cache reproduces the slowdown,” not simply “try again.”
10. Read-only diagnostic commands
RUN_ID=123456789
GH_REPO=OWNER/repo
gh run view "$RUN_ID" --json databaseId,attempt,event,headSha,status,conclusion,createdAt,updatedAt,url
gh api \
-H 'Accept: application/vnd.github+json' \
-H 'X-GitHub-Api-Version: 2026-03-10' \
"repos/$GH_REPO/actions/runs/$RUN_ID/jobs?per_page=100" \
--jq '.jobs[] | [.id,.name,.started_at,.completed_at,.conclusion,(.labels|join(","))] | @tsv'
gh cache list --limit 100
11. Symptom → evidence → causal layer → correction
| Symptom | Evidence that discriminates | Likely layer | Least destructive next action |
|---|---|---|---|
| High average queue, short jobs | Performance Metrics + runner labels | Capacity/concurrency | Inspect eligible capacity before editing tests/cache. |
| Long dependency install, low cache hits | lock hash + keys + cache metadata | Cache/input identity | Fix precise key or determine cache not worth it. |
| Fast matrix wall time, huge processing minutes | job durations + cell count | Graph granularity | Merge tiny cells / cap parallelism. |
| Self-hosted wait only | runner group + online/busy state | Owned runner fleet | Scale/reroute within trust boundary. |
| Performance changes after no source diff | image/action/tool versions | Platform/tool drift | Freeze/record versions and reproduce. |
| Green faster run after test removal | job graph + required-check coverage | Correctness policy | Restore missing tests; optimize internally. |
12. Shortcuts this chapter rejects
- Do not grant broader token permissions to make a cache/artifact call “work faster.”
- Do not print contexts or secrets to investigate timing.
- Do not delete caches/artifacts/runs before preserving failure evidence.
- Do not move untrusted pull-request code onto persistent self-hosted runners to avoid hosted queue time.
- Do not disable TLS, skip integrity checks or float dependencies/actions for convenience.
- Do not replace required tests with retries or optimistic path filters.
13. Lesson summary
Performance incidents are evidence problems before they are tuning problems. Separate workload correctness, job graph, runner capacity, setup, cache, I/O and external waits; preserve the first attempt; then change one bounded cause. The checkpoint now asks you to prove two optimizations and reject one tempting change that would only make the metric look better by weakening the pipeline.
Knowledge check
A new workflow is 30% faster, but it no longer runs integration tests. Is that a valid optimization comparison?
No. The workload/quality policy changed. Restore equivalent coverage before comparing performance.
Why can a broad cache key look successful?
It can produce a high hit rate and short downloads while restoring incompatible/stale state; hit rate is not correctness.
Self-hosted jobs queue while hosted jobs do not. What should you inspect first?
Eligible self-hosted runner group/labels, online/busy state and fleet/autoscaler capacity—not GitHub-hosted pricing or cache keys.
Why pin runner/tool/action versions for benchmarks?
Otherwise platform/tool drift becomes confounded with the workflow/source change and you cannot attribute the performance difference causally.
What should be preserved before a diagnostic rerun?
Run ID/attempt, SHA/workflow revision, job/step timestamps, runner identity, cache/artifact metadata and relevant external correlation evidence.
Official references and version notes
- GitHub Actions metrics — Current usage and performance metrics, including run time, queue time and failure-rate views.
- Viewing Actions metrics — Repository and organization Actions Usage/Performance Metrics and aggregation windows.
- Actions limits — Current matrix, concurrency, queue and job-duration limits; limits are explicitly subject to change.
- GitHub-hosted runners reference — Current public/private standard runner hardware, labels and isolation characteristics.
- Actions runner pricing — Current per-minute hosted runner prices and per-job minute rounding.
- GitHub Actions billing — Current plan allowances, free public standard-runner use, storage pricing and billing behavior.
- Concurrency — Concurrency groups, cancellation behavior and current queue:max semantics.
- Dependency caching reference — Cache identity, restore behavior, storage/eviction and rate-limit behavior.
- REST: workflow jobs — Job IDs, runner labels, started/completed timestamps and step timing evidence.
- REST: workflow runs — Run metadata, attempts, source SHA, status/conclusion and usage-related inspection.
- actions/checkout v7.0.1 — Full commit SHA used by executable lab examples.
- actions/setup-python v7.0.0 — Full commit SHA used for Python 3.13 setup in the lab.
- actions/cache v6.1.0 — Full commit SHA used for the explicit dependency-download cache.
- actions/upload-artifact v7.0.1 — Full commit SHA used only for tiny bounded benchmark evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.