Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Core Concepts and Mental Model
Model GitHub Actions performance as event demand flowing through queues, runner capacity, setup, compute, cache/artifact I/O and billing so optimization starts from evidence rather than intuition.
Learning objectives
- Explain queue time, execution time, critical path, processing time and billable usage as different measurements.
- Separate event demand, job graph, runner capacity, setup/dependency time, compute time, cache/artifact I/O and external waits.
- Inspect run/job/runner/cache evidence before proposing a performance change.
- Explain why more parallelism, more caching and larger runners can each make cost or latency worse.
- Build a cost model without confusing GitHub billing with self-hosted infrastructure total cost of ownership.
1. The practical problem: a faster workflow is not automatically a better workflow
Earlier chapters made GitHub Actions reproducible, secure, observable and controllable. Performance work is the next temptation: split every test into a matrix, cache everything, cancel old work and buy a bigger runner. Those changes can reduce latency, but they can also increase startup overhead, create stale-state risk, multiply rounded billable minutes, hide required tests or move a queue bottleneck from GitHub to your own runner fleet.
The operating rule for this chapter is therefore measure → explain → change one bounded variable → measure again → preserve correctness. You optimize the critical path and waste, not whichever number happens to look largest in one run.
2. Mental model: event demand becomes queue, work, I/O and cost
Every workflow begins with event volume: pushes, pull requests, schedules, dispatches and dependent workflows create runs. The workflow revision expands those runs into jobs and matrices. Jobs wait for dependencies and eligible runner capacity, then spend time on setup, dependency restoration, actual compute/test work and artifact/cache I/O. The longest dependency chain determines user-visible wall-clock latency, while the sum of runner processing time drives usage and, where applicable, billing.
flowchart TD A[Event volume] --> B[Workflow + matrix expansion] B --> C[Queued eligible jobs] C --> D[Runner availability] D --> E[Setup + dependency restore] E --> F[Build / test / analysis compute] F --> G[Artifact + cache I/O] G --> H[Job completion] H --> I[Workflow critical path] B --> J[Total runner processing time] H --> J J --> K[Plan allowance / billable usage]
The diagram has two endings because latency and usage are not the same objective. Four five-minute jobs that run perfectly in parallel may finish in roughly five minutes of compute wall time but consume roughly twenty processing minutes before per-job rounding and setup overhead. That can be a good trade when feedback latency matters; it is not a free optimization.
3. Record the performance state before changing it
| State | Evidence to capture | Optimization question |
|---|---|---|
| Event/revision | event, run ID/attempt, source SHA, workflow SHA | Are we comparing the same workload and configuration? |
| Job graph | job names, needs, matrix cells, skipped/cancelled state | What is the actual critical path and parallelizable work? |
| Queue/capacity | runner label/type, performance-metrics queue time, start timestamps | Is waiting caused by eligible runner capacity rather than code? |
| Execution | job/step started_at and completed_at, exit code | Which setup/compute step dominates? |
| Cache | primary key, hit state, restored path, bytes, retention | Does reuse save more time than restore/save costs? |
| Artifacts | artifact count, size, upload/download time, retention | Is evidence I/O on the critical path or storage budget? |
| Waste | reruns, cancellations, superseded commits, failed retries | Can unnecessary work be prevented safely? |
| Billing assumptions | visibility, plan, runner SKU, included quota, rounding | Is a modeled dollar saving actually relevant to this repository? |
4. Queue time, run time, critical path and processing time are different
Queue time is waiting for an eligible runner after
a job becomes runnable. GitHub Actions Performance Metrics exposes
average queue-time signals directly. The workflow-job REST payload
exposes started_at and completed_at, but
not a universal per-job queued_at field. Subtracting
workflow created_at from a job start can be a useful
dispatch-to-start signal for a root job, but it is not authoritative
queue time for a dependent job because dependency waiting is mixed
in.
Job duration is elapsed time from job start to completion. Critical path is the longest dependency chain that determines workflow completion. Total processing time is the sum of runner execution across jobs. Billable time depends on repository/runner type and current billing rules; on private repositories, GitHub-hosted job minutes are rounded up per job. Keep these quantities named correctly in every comparison.
5. Runner hardware is part of the benchmark input
As verified on September 10, 2026, standard
ubuntu-24.04 differs by repository visibility: public
repositories currently receive 4 vCPU and 16 GB RAM, while private
repositories receive 2 vCPU and 8 GB RAM. Both have 14 GB SSD in the
current reference. A timing comparison between a public lab and a
private production repository is therefore not an apples-to-apples
benchmark even when the workflow YAML is identical.
GitHub also offers ubuntu-slim, a 1-core containerized
runner intended for lightweight automation with a 15-minute job
limit and reduced privileges/tooling. Larger runners are
Team/Enterprise features and are billed even for public
repositories. Record the exact label and observed runner/image
metadata rather than writing “Linux runner” in your evidence.
6. Parallelism has platform limits before it has design limits
Current GitHub.com limits include a maximum of 256 generated jobs for one matrix. Standard hosted concurrency is plan-dependent: the current reference lists 20 total concurrent jobs on Free, 40 on Pro, 60 on Team and 500 on Enterprise, with separate macOS caps. Larger-runner concurrency is a different capacity model. These are operational ceilings, not optimization targets.
A matrix of 200 tiny jobs can be technically valid and still be a poor design because runner startup, checkout, tool setup, cache traffic and per-job billing rounding dominate the work. Start with the workload graph: split only work that is independent and expensive enough to amortize job startup.
7. Cache economics: reuse only when correctness survives a miss
A dependency cache is a best-effort acceleration layer. The primary key must bind reuse to inputs that determine cache validity—typically OS/architecture, tool/runtime version and a lockfile digest. A cache hit can reduce download or compilation work; it must not become proof that dependencies are installed correctly or that tests can be skipped.
As of September 10, 2026, repositories include 10 GB of Actions cache storage by default. The default retention is 7 days; eligible settings can raise retention and storage limits, and storage above the included 10 GB can be billed. The performance question is therefore three-dimensional: time saved, correctness risk and storage/network cost.
8. Artifact I/O is evidence cost, not merely “free storage”
Artifacts are durable run outputs; caches are replaceable acceleration state. Uploading a 3 GB test tree from every matrix cell can dominate runtime and storage without improving diagnosis. Prefer small, meaningful evidence: failing logs, coverage reports, manifests and exact binaries that downstream jobs truly consume. Record artifact size and retention alongside upload time.
Current included artifact storage depends on plan and is shared with GitHub Packages. Extra shared artifact/package storage is billed separately from Actions cache storage. Optimization should shorten retention or reduce redundant payloads only after evidence obligations are understood.
9. Current cost model: model the repository you actually own
Standard GitHub-hosted runners are currently free and unlimited for public repositories. Private repositories receive plan-specific included minutes and then use per-minute pricing; as of September 10, 2026, baseline Linux x64 is $0.006/minute, Windows x64 $0.010/minute and macOS $0.062/minute. Larger runners are always billed. These prices are timestamped assumptions—never hard-code them into long-lived business logic without re-reading billing documentation.
A public-repository lab can therefore be cost-free while still calculating a hypothetical private over-quota cost. This is useful because it teaches the formula without requiring a payment method: sum rounded job minutes by runner SKU, then multiply by the current rate. Keep included quota, budgets and actual invoices outside the benchmark unless you have authority to inspect them.
10. Cancellation and reruns can save work—or erase useful work
For fast-changing pull requests, concurrency with
cancel-in-progress: true can stop obsolete CI and
reduce waste. But cancellation is safe only when older results are
no longer required and jobs do not carry external side effects that
need completion/rollback. Deployment serialization uses different
semantics; do not apply “cancel old work” mechanically to production
rollout jobs.
Likewise, a rerun consumes more runner time. Preserve the first attempt and fix the causal layer before rerunning. Performance optimization that merely hides flaky failures behind automatic retries spends more and produces weaker evidence.
11. Read-only inspection first
# Exact run and attempt first.
gh run view RUN_ID --json databaseId,attempt,event,headSha,status,conclusion,createdAt,updatedAt,url
# Exact job/step timing and runner labels from the current REST API.
gh api \
-H 'Accept: application/vnd.github+json' \
-H 'X-GitHub-Api-Version: 2026-03-10' \
repos/{owner}/{repo}/actions/runs/RUN_ID/jobs \
--jq '.jobs[] | {id,name,status,conclusion,started_at,completed_at,labels,runner_name}'
# Cache metadata is read-only.
gh cache list --limit 100
For average queue time and trend data, use the repository or organization Actions Performance Metrics view when available. A single run is an observation; a distribution across equivalent runs is evidence.
12. The optimization acceptance rule
Accept an optimization only when it preserves the same source/input identity, required test/quality coverage, trust boundary and evidence obligations while improving a declared objective such as p50 wall-clock latency, p95 queue time, processing minutes or storage. If it changes the workload, test coverage, runner trust or release evidence, it is not a pure performance comparison and must be reviewed as a behavior change.
13. Lesson summary
GitHub Actions performance is a system property: demand, graph shape, runner capacity, setup, computation, I/O and billing interact. Optimize the measured bottleneck, not the fashionable feature. The next lesson builds a deliberately small benchmark where cache and bounded test parallelism can be evaluated separately.
Knowledge check
Why can four parallel jobs finish faster but cost more processing minutes?
They overlap in wall-clock time but each consumes its own runner setup and execution time; private hosted billing also rounds each job separately.
Is workflow created_at → dependent-job started_at an exact queue-time measurement?
No. Dependency waiting and scheduler/runner waiting are mixed together. Use performance metrics for queue time and treat raw timestamp differences as scoped signals.
What must remain true after adding a dependency cache?
A cold/missing cache must still produce the same correct build/test result; the cache cannot become a correctness prerequisite.
Why record repository visibility in benchmark evidence?
Current standard runner hardware and billing differ between public and private repositories, so visibility changes both performance and cost assumptions.
A 200-cell matrix is below the 256 limit. Is it therefore optimized?
No. Platform validity says nothing about startup overhead, queue pressure, cost, signal value or whether that many independent cells are necessary.
Official references and version notes
- GitHub Actions metrics — Current usage and performance metrics, including run time, queue time and failure-rate views.
- Viewing Actions metrics — Repository and organization Actions Usage/Performance Metrics and aggregation windows.
- Actions limits — Current matrix, concurrency, queue and job-duration limits; limits are explicitly subject to change.
- GitHub-hosted runners reference — Current public/private standard runner hardware, labels and isolation characteristics.
- Actions runner pricing — Current per-minute hosted runner prices and per-job minute rounding.
- GitHub Actions billing — Current plan allowances, free public standard-runner use, storage pricing and billing behavior.
- Concurrency — Concurrency groups, cancellation behavior and current queue:max semantics.
- Dependency caching reference — Cache identity, restore behavior, storage/eviction and rate-limit behavior.
- REST: workflow jobs — Job IDs, runner labels, started/completed timestamps and step timing evidence.
- REST: workflow runs — Run metadata, attempts, source SHA, status/conclusion and usage-related inspection.
- actions/checkout v7.0.1 — Full commit SHA used by executable lab examples.
- actions/setup-python v7.0.0 — Full commit SHA used for Python 3.13 setup in the lab.
- actions/cache v6.1.0 — Full commit SHA used for the explicit dependency-download cache.
- actions/upload-artifact v7.0.1 — Full commit SHA used only for tiny bounded benchmark evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.