Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Configuration, Design Patterns, and Trade-Offs
Choose parallelism, runner size, caching, cancellation and hosted/self-hosted capacity using critical-path, trust, determinism, utilization and total-cost trade-offs.
Learning objectives
- Choose between more jobs, larger runners and sequential execution from workload shape rather than intuition.
- Treat cache hit rate as one input to a determinism/time/storage decision, not a success KPI by itself.
- Use cancellation only for superseded work whose external side effects and evidence obligations make cancellation safe.
- Compare GitHub-hosted price signals with self-hosted total cost of ownership and trust boundaries.
- Apply current plan, concurrency, storage and runner constraints without hard-coding volatile platform values into architecture.
1. Optimization is architecture under constraints
Lesson 2 changed two small variables in a controlled lab. Production optimization is harder because the constraints are different: a Java test suite may saturate CPU, a browser matrix may need memory, an integration suite may share one database, self-hosted runners may be scarce, macOS may be expensive, and deployment jobs may be intentionally serialized. The correct design minimizes the declared objective while preserving correctness, security, evidence and external-system safety.
2. More parallelism versus startup overhead
Parallelism helps when independent partitions are long enough that
overlap exceeds job-start overhead. It hurts when each partition is
tiny, duplicates expensive checkout/tool setup, competes for a
constrained external service, or consumes scarce concurrency that
blocks higher-value work. Matrix size and
max-parallel let you separate coverage from
simultaneous demand.
A useful design calculation is
parallelizable_work / cell_count versus measured
per-cell fixed overhead. If each test partition does 15 seconds of
work but runner/setup takes 45 seconds, splitting further increases
total processing and may not reduce the critical path.
3. Larger runner versus more standard runners
| Question | More standard runners | Larger runner |
|---|---|---|
| Workload shape | Many independent jobs/cells | One CPU/RAM-bound job that does not split cleanly |
| Feedback latency | Improves through overlap if concurrency exists | Improves one job if software scales with cores/RAM |
| Cost behavior | More processing jobs and per-job rounding | Higher per-minute SKU; always billed |
| Plan/access | Standard hosted available broadly | Currently Team/Enterprise Cloud |
| Evidence | job count, queue, per-cell duration | CPU/memory utilization + before/after duration |
| Risk | matrix explosion / contention | paying for idle cores or unchanged serial bottleneck |
Do not assume “8 cores is 4× faster than 2 cores.” Measure the exact workload. Compiler, test, I/O and lock contention determine scaling. A larger runner does not fix a dependency download bottleneck or an external API rate limit.
4. Cache versus determinism
A cache is justified when regeneration/download cost is meaningful and a precise key can express compatibility. The key should include the state that invalidates reuse. Restore prefixes may intentionally reuse older compatible content, but the consumer must reconcile against authoritative lock/config inputs afterward.
Reject caches for secrets, mutable deployment state, release evidence or opaque executable state that is hard to validate. If a cache can change the semantic outcome of the build, it has crossed from optimization into an undocumented dependency.
5. Cache capacity and storage cost are operational inputs
Current default cache storage is 10 GB per repository with a default 7-day retention. Users who opt into expanded cache storage can configure higher limits; current user-owned repository guidance allows up to 10 TB and paid cache storage above the included 10 GB is $0.07/GB-month. These settings are volatile and organization/enterprise caps may constrain them.
A higher limit can reduce eviction churn but also preserve low-value caches longer. Before buying more cache storage, inspect which keys are produced, size per key, reuse frequency and miss reasons. Fixing a cardinality explosion in keys is often cheaper than increasing the quota.
6. cancel-in-progress versus completed-work value
For pull-request CI, newer commits often supersede older ones. A branch/ref-scoped concurrency group can cancel obsolete runs and cut queue/processing waste. But cancellation is a semantic decision: an old security scan may still be required evidence, a release build may be immutable provenance, and a deployment may have already changed an external system. Use cancellation only where “newer run makes older unfinished run worthless” is actually true.
concurrency:
group: ci-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
Do not reuse this pattern blindly for deployments. Deployment serialization, approvals and rollback state require a different concurrency policy. A cancelled GitHub job does not automatically undo a cloud/Kubernetes side effect.
7. Hosted versus self-hosted: price is not total cost
GitHub currently does not charge Actions runner minutes for self-hosted jobs, but that does not make the runners free. Capacity planning must include VM/bare-metal/cloud cost, idle capacity, patching, image management, autoscaling, monitoring, security isolation, secrets exposure, incident response and staff time. A queue on self-hosted runners is commonly a fleet-capacity or label/routing problem under your control.
GitHub-hosted runners convert much of that operations burden into a per-minute managed service. The trade is not “$0 versus $0.006”; it is managed variable cost versus infrastructure and operational TCO under different trust and network requirements.
8. Current billing/allowance snapshot — timestamp it
| Plan / SKU | Current September 10, 2026 assumption | Design implication |
|---|---|---|
| Public + standard hosted | Free and unlimited standard-runner use | Good for disposable course lab; still avoid waste/abuse. |
| GitHub Free private | 2,000 included minutes/month; 500 MB artifact storage | Optimize high-volume private CI and evidence retention. |
| GitHub Pro private | 3,000 included minutes; 1 GB artifact storage | Model overage after allowance; repository owner is billed. |
| GitHub Team private | 3,000 included minutes; 2 GB artifact storage | Larger runners available but always billed. |
| GitHub Enterprise Cloud | 50,000 included minutes; 50 GB artifact storage | High concurrency does not remove need for cost governance. |
| Baseline Linux x64 | Current overage rate $0.006/min | Per-job rounding matters to matrix designs. |
| Cache | 10 GB/repo included by default; extra configured usage $0.07/GB-month | High-cardinality caches can become an explicit storage cost. |
Treat this table as a dated example, not an evergreen contract. Billing, hardware and plan allowances are exactly the kind of platform facts the master contract requires you to re-check before implementation.
9. Artifact retention versus evidence integrity
Shortening artifact retention can save storage, but only after you classify artifacts. Ephemeral benchmark traces can expire quickly. Release binaries, provenance, compliance evidence and incident material may need longer retention or a different durable system. “Delete artifacts to save cost” is not a valid recommendation until evidence requirements are known.
10. Human metrics versus machine evidence
Actions Usage Metrics and Performance Metrics are excellent trend views for workflows, jobs, repositories, runtime OS and runner types. Use them to find candidates: high average queue time, long jobs, frequent failures or high minute consumption. Then drill into exact run IDs, attempts, source SHAs and job timestamps before changing YAML. Aggregates identify where to look; exact run evidence explains why.
11. Worked decision table
| Observed evidence | Tempting change | Preferred first move | Why |
|---|---|---|---|
| 90% cache hit; install still 2% of critical path | Broaden cache restore keys | Do nothing to cache | No material bottleneck; broader reuse adds staleness risk. |
| Four 12-minute independent test groups in one job | Buy 16-core runner | Split into 4 cells, cap parallelism, measure | Work is naturally separable; standard runners may reduce wall time cheaply. |
| One memory-bound linker OOMs | Create 20 matrix cells | Evaluate larger runner or build redesign | Parallelism does not fix per-process memory requirement. |
| Self-hosted queue p95 18 min; runners 100% busy | Blame GitHub queue | Add/scale eligible capacity or schedule demand | Capacity is in the owned runner fleet. |
| PRs frequently superseded within 2 min | Run all old commits | Cancel safe obsolete PR CI | Avoids known worthless work if no side effects/evidence need. |
| macOS test matrix repeats equivalent cells | Cache more | Remove redundant coverage after test-owner review | The workload itself is duplicated; cache is not the primary issue. |
12. Performance changes need rollback criteria
Write a rollback condition before rollout: for example, “revert if p95 workflow latency worsens by >15%, test failure classification changes, cache-related correctness defects appear, or processing minutes rise >30% without an approved latency benefit.” Without a rollback threshold, teams tend to rationalize regressions after investing in an optimization.
13. Lesson summary
Parallelism, bigger runners, caches, cancellation and self-hosting are different levers. Each changes a different state: job graph, runner capacity, reuse state, superseded-work policy or infrastructure ownership. Measure the causal layer and keep billing assumptions timestamped. Lesson 4 now focuses on failures where teams optimize the wrong layer or destroy the evidence needed to prove it.
Knowledge check
When should a larger runner beat more matrix jobs conceptually?
When the bottleneck is one job that cannot split effectively and can use more CPU/RAM; prove scaling with workload/utilization evidence.
Why is cache hit rate not a sufficient KPI?
A high hit rate can save almost no critical-path time or restore stale/low-value state. Measure time saved, correctness and storage/network cost.
When is cancel-in-progress a good fit?
When newer work makes older unfinished work genuinely obsolete and cancellation cannot leave required evidence or external side effects in an unsafe state.
Why is a self-hosted runner not “free compute”?
GitHub runner minutes are unbilled, but infrastructure, idle capacity, patching, security, autoscaling, monitoring and operator time remain real TCO.
What makes the billing table safe to teach?
It is explicitly timestamped and paired with a requirement to re-check current GitHub documentation rather than treating the values as permanent.
Official references and version notes
- GitHub Actions metrics — Current usage and performance metrics, including run time, queue time and failure-rate views.
- Viewing Actions metrics — Repository and organization Actions Usage/Performance Metrics and aggregation windows.
- Actions limits — Current matrix, concurrency, queue and job-duration limits; limits are explicitly subject to change.
- GitHub-hosted runners reference — Current public/private standard runner hardware, labels and isolation characteristics.
- Actions runner pricing — Current per-minute hosted runner prices and per-job minute rounding.
- GitHub Actions billing — Current plan allowances, free public standard-runner use, storage pricing and billing behavior.
- Concurrency — Concurrency groups, cancellation behavior and current queue:max semantics.
- Dependency caching reference — Cache identity, restore behavior, storage/eviction and rate-limit behavior.
- REST: workflow jobs — Job IDs, runner labels, started/completed timestamps and step timing evidence.
- REST: workflow runs — Run metadata, attempts, source SHA, status/conclusion and usage-related inspection.
- actions/checkout v7.0.1 — Full commit SHA used by executable lab examples.
- actions/setup-python v7.0.0 — Full commit SHA used for Python 3.13 setup in the lab.
- actions/cache v6.1.0 — Full commit SHA used for the explicit dependency-download cache.
- actions/upload-artifact v7.0.1 — Full commit SHA used only for tiny bounded benchmark evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.