Chapter 32Lesson 01~225 minutes

Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Core Concepts and Mental Model

Model GitHub Actions performance as event demand flowing through queues, runner capacity, setup, compute, cache/artifact I/O and billing so optimization starts from evidence rather than intuition.

PerformanceQueue timeCritical pathUsageCost model

Learning objectives

  • Explain queue time, execution time, critical path, processing time and billable usage as different measurements.
  • Separate event demand, job graph, runner capacity, setup/dependency time, compute time, cache/artifact I/O and external waits.
  • Inspect run/job/runner/cache evidence before proposing a performance change.
  • Explain why more parallelism, more caching and larger runners can each make cost or latency worse.
  • Build a cost model without confusing GitHub billing with self-hosted infrastructure total cost of ownership.

1. The practical problem: a faster workflow is not automatically a better workflow

Earlier chapters made GitHub Actions reproducible, secure, observable and controllable. Performance work is the next temptation: split every test into a matrix, cache everything, cancel old work and buy a bigger runner. Those changes can reduce latency, but they can also increase startup overhead, create stale-state risk, multiply rounded billable minutes, hide required tests or move a queue bottleneck from GitHub to your own runner fleet.

The operating rule for this chapter is therefore measure → explain → change one bounded variable → measure again → preserve correctness. You optimize the critical path and waste, not whichever number happens to look largest in one run.

2. Mental model: event demand becomes queue, work, I/O and cost

Every workflow begins with event volume: pushes, pull requests, schedules, dispatches and dependent workflows create runs. The workflow revision expands those runs into jobs and matrices. Jobs wait for dependencies and eligible runner capacity, then spend time on setup, dependency restoration, actual compute/test work and artifact/cache I/O. The longest dependency chain determines user-visible wall-clock latency, while the sum of runner processing time drives usage and, where applicable, billing.

Performance causality
flowchart TD
  A[Event volume] --> B[Workflow + matrix expansion]
  B --> C[Queued eligible jobs]
  C --> D[Runner availability]
  D --> E[Setup + dependency restore]
  E --> F[Build / test / analysis compute]
  F --> G[Artifact + cache I/O]
  G --> H[Job completion]
  H --> I[Workflow critical path]
  B --> J[Total runner processing time]
  H --> J
  J --> K[Plan allowance / billable usage]

The diagram has two endings because latency and usage are not the same objective. Four five-minute jobs that run perfectly in parallel may finish in roughly five minutes of compute wall time but consume roughly twenty processing minutes before per-job rounding and setup overhead. That can be a good trade when feedback latency matters; it is not a free optimization.

3. Record the performance state before changing it

State Evidence to capture Optimization question
Event/revision event, run ID/attempt, source SHA, workflow SHA Are we comparing the same workload and configuration?
Job graph job names, needs, matrix cells, skipped/cancelled state What is the actual critical path and parallelizable work?
Queue/capacity runner label/type, performance-metrics queue time, start timestamps Is waiting caused by eligible runner capacity rather than code?
Execution job/step started_at and completed_at, exit code Which setup/compute step dominates?
Cache primary key, hit state, restored path, bytes, retention Does reuse save more time than restore/save costs?
Artifacts artifact count, size, upload/download time, retention Is evidence I/O on the critical path or storage budget?
Waste reruns, cancellations, superseded commits, failed retries Can unnecessary work be prevented safely?
Billing assumptions visibility, plan, runner SKU, included quota, rounding Is a modeled dollar saving actually relevant to this repository?

4. Queue time, run time, critical path and processing time are different

Queue time is waiting for an eligible runner after a job becomes runnable. GitHub Actions Performance Metrics exposes average queue-time signals directly. The workflow-job REST payload exposes started_at and completed_at, but not a universal per-job queued_at field. Subtracting workflow created_at from a job start can be a useful dispatch-to-start signal for a root job, but it is not authoritative queue time for a dependent job because dependency waiting is mixed in.

Job duration is elapsed time from job start to completion. Critical path is the longest dependency chain that determines workflow completion. Total processing time is the sum of runner execution across jobs. Billable time depends on repository/runner type and current billing rules; on private repositories, GitHub-hosted job minutes are rounded up per job. Keep these quantities named correctly in every comparison.

5. Runner hardware is part of the benchmark input

As verified on September 10, 2026, standard ubuntu-24.04 differs by repository visibility: public repositories currently receive 4 vCPU and 16 GB RAM, while private repositories receive 2 vCPU and 8 GB RAM. Both have 14 GB SSD in the current reference. A timing comparison between a public lab and a private production repository is therefore not an apples-to-apples benchmark even when the workflow YAML is identical.

GitHub also offers ubuntu-slim, a 1-core containerized runner intended for lightweight automation with a 15-minute job limit and reduced privileges/tooling. Larger runners are Team/Enterprise features and are billed even for public repositories. Record the exact label and observed runner/image metadata rather than writing “Linux runner” in your evidence.

6. Parallelism has platform limits before it has design limits

Current GitHub.com limits include a maximum of 256 generated jobs for one matrix. Standard hosted concurrency is plan-dependent: the current reference lists 20 total concurrent jobs on Free, 40 on Pro, 60 on Team and 500 on Enterprise, with separate macOS caps. Larger-runner concurrency is a different capacity model. These are operational ceilings, not optimization targets.

A matrix of 200 tiny jobs can be technically valid and still be a poor design because runner startup, checkout, tool setup, cache traffic and per-job billing rounding dominate the work. Start with the workload graph: split only work that is independent and expensive enough to amortize job startup.

7. Cache economics: reuse only when correctness survives a miss

A dependency cache is a best-effort acceleration layer. The primary key must bind reuse to inputs that determine cache validity—typically OS/architecture, tool/runtime version and a lockfile digest. A cache hit can reduce download or compilation work; it must not become proof that dependencies are installed correctly or that tests can be skipped.

As of September 10, 2026, repositories include 10 GB of Actions cache storage by default. The default retention is 7 days; eligible settings can raise retention and storage limits, and storage above the included 10 GB can be billed. The performance question is therefore three-dimensional: time saved, correctness risk and storage/network cost.

8. Artifact I/O is evidence cost, not merely “free storage”

Artifacts are durable run outputs; caches are replaceable acceleration state. Uploading a 3 GB test tree from every matrix cell can dominate runtime and storage without improving diagnosis. Prefer small, meaningful evidence: failing logs, coverage reports, manifests and exact binaries that downstream jobs truly consume. Record artifact size and retention alongside upload time.

Current included artifact storage depends on plan and is shared with GitHub Packages. Extra shared artifact/package storage is billed separately from Actions cache storage. Optimization should shorten retention or reduce redundant payloads only after evidence obligations are understood.

9. Current cost model: model the repository you actually own

Standard GitHub-hosted runners are currently free and unlimited for public repositories. Private repositories receive plan-specific included minutes and then use per-minute pricing; as of September 10, 2026, baseline Linux x64 is $0.006/minute, Windows x64 $0.010/minute and macOS $0.062/minute. Larger runners are always billed. These prices are timestamped assumptions—never hard-code them into long-lived business logic without re-reading billing documentation.

A public-repository lab can therefore be cost-free while still calculating a hypothetical private over-quota cost. This is useful because it teaches the formula without requiring a payment method: sum rounded job minutes by runner SKU, then multiply by the current rate. Keep included quota, budgets and actual invoices outside the benchmark unless you have authority to inspect them.

10. Cancellation and reruns can save work—or erase useful work

For fast-changing pull requests, concurrency with cancel-in-progress: true can stop obsolete CI and reduce waste. But cancellation is safe only when older results are no longer required and jobs do not carry external side effects that need completion/rollback. Deployment serialization uses different semantics; do not apply “cancel old work” mechanically to production rollout jobs.

Likewise, a rerun consumes more runner time. Preserve the first attempt and fix the causal layer before rerunning. Performance optimization that merely hides flaky failures behind automatic retries spends more and produces weaker evidence.

11. Read-only inspection first

# Exact run and attempt first.
gh run view RUN_ID --json databaseId,attempt,event,headSha,status,conclusion,createdAt,updatedAt,url

# Exact job/step timing and runner labels from the current REST API.
gh api \
  -H 'Accept: application/vnd.github+json' \
  -H 'X-GitHub-Api-Version: 2026-03-10' \
  repos/{owner}/{repo}/actions/runs/RUN_ID/jobs \
  --jq '.jobs[] | {id,name,status,conclusion,started_at,completed_at,labels,runner_name}'

# Cache metadata is read-only.
gh cache list --limit 100

For average queue time and trend data, use the repository or organization Actions Performance Metrics view when available. A single run is an observation; a distribution across equivalent runs is evidence.

12. The optimization acceptance rule

Accept an optimization only when it preserves the same source/input identity, required test/quality coverage, trust boundary and evidence obligations while improving a declared objective such as p50 wall-clock latency, p95 queue time, processing minutes or storage. If it changes the workload, test coverage, runner trust or release evidence, it is not a pure performance comparison and must be reviewed as a behavior change.

13. Lesson summary

GitHub Actions performance is a system property: demand, graph shape, runner capacity, setup, computation, I/O and billing interact. Optimize the measured bottleneck, not the fashionable feature. The next lesson builds a deliberately small benchmark where cache and bounded test parallelism can be evaluated separately.

Next lesson

Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Guided Hands-On Workflow

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why can four parallel jobs finish faster but cost more processing minutes?

Is workflow created_at → dependent-job started_at an exact queue-time measurement?

What must remain true after adding a dependency cache?

Why record repository visibility in benchmark evidence?

A 200-cell matrix is below the 256 limit. Is it therefore optimized?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.