Chapter 32Lesson 03~225 minutes

Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Configuration, Design Patterns, and Trade-Offs

Choose parallelism, runner size, caching, cancellation and hosted/self-hosted capacity using critical-path, trust, determinism, utilization and total-cost trade-offs.

Trade-offsLarger runnersConcurrencyTCODeterminism

Learning objectives

  • Choose between more jobs, larger runners and sequential execution from workload shape rather than intuition.
  • Treat cache hit rate as one input to a determinism/time/storage decision, not a success KPI by itself.
  • Use cancellation only for superseded work whose external side effects and evidence obligations make cancellation safe.
  • Compare GitHub-hosted price signals with self-hosted total cost of ownership and trust boundaries.
  • Apply current plan, concurrency, storage and runner constraints without hard-coding volatile platform values into architecture.

1. Optimization is architecture under constraints

Lesson 2 changed two small variables in a controlled lab. Production optimization is harder because the constraints are different: a Java test suite may saturate CPU, a browser matrix may need memory, an integration suite may share one database, self-hosted runners may be scarce, macOS may be expensive, and deployment jobs may be intentionally serialized. The correct design minimizes the declared objective while preserving correctness, security, evidence and external-system safety.

2. More parallelism versus startup overhead

Parallelism helps when independent partitions are long enough that overlap exceeds job-start overhead. It hurts when each partition is tiny, duplicates expensive checkout/tool setup, competes for a constrained external service, or consumes scarce concurrency that blocks higher-value work. Matrix size and max-parallel let you separate coverage from simultaneous demand.

A useful design calculation is parallelizable_work / cell_count versus measured per-cell fixed overhead. If each test partition does 15 seconds of work but runner/setup takes 45 seconds, splitting further increases total processing and may not reduce the critical path.

3. Larger runner versus more standard runners

Question More standard runners Larger runner
Workload shape Many independent jobs/cells One CPU/RAM-bound job that does not split cleanly
Feedback latency Improves through overlap if concurrency exists Improves one job if software scales with cores/RAM
Cost behavior More processing jobs and per-job rounding Higher per-minute SKU; always billed
Plan/access Standard hosted available broadly Currently Team/Enterprise Cloud
Evidence job count, queue, per-cell duration CPU/memory utilization + before/after duration
Risk matrix explosion / contention paying for idle cores or unchanged serial bottleneck

Do not assume “8 cores is 4× faster than 2 cores.” Measure the exact workload. Compiler, test, I/O and lock contention determine scaling. A larger runner does not fix a dependency download bottleneck or an external API rate limit.

4. Cache versus determinism

A cache is justified when regeneration/download cost is meaningful and a precise key can express compatibility. The key should include the state that invalidates reuse. Restore prefixes may intentionally reuse older compatible content, but the consumer must reconcile against authoritative lock/config inputs afterward.

Reject caches for secrets, mutable deployment state, release evidence or opaque executable state that is hard to validate. If a cache can change the semantic outcome of the build, it has crossed from optimization into an undocumented dependency.

5. Cache capacity and storage cost are operational inputs

Current default cache storage is 10 GB per repository with a default 7-day retention. Users who opt into expanded cache storage can configure higher limits; current user-owned repository guidance allows up to 10 TB and paid cache storage above the included 10 GB is $0.07/GB-month. These settings are volatile and organization/enterprise caps may constrain them.

A higher limit can reduce eviction churn but also preserve low-value caches longer. Before buying more cache storage, inspect which keys are produced, size per key, reuse frequency and miss reasons. Fixing a cardinality explosion in keys is often cheaper than increasing the quota.

6. cancel-in-progress versus completed-work value

For pull-request CI, newer commits often supersede older ones. A branch/ref-scoped concurrency group can cancel obsolete runs and cut queue/processing waste. But cancellation is a semantic decision: an old security scan may still be required evidence, a release build may be immutable provenance, and a deployment may have already changed an external system. Use cancellation only where “newer run makes older unfinished run worthless” is actually true.

concurrency:
  group: ci-${{ github.event.pull_request.number || github.ref }}
  cancel-in-progress: true

Do not reuse this pattern blindly for deployments. Deployment serialization, approvals and rollback state require a different concurrency policy. A cancelled GitHub job does not automatically undo a cloud/Kubernetes side effect.

7. Hosted versus self-hosted: price is not total cost

GitHub currently does not charge Actions runner minutes for self-hosted jobs, but that does not make the runners free. Capacity planning must include VM/bare-metal/cloud cost, idle capacity, patching, image management, autoscaling, monitoring, security isolation, secrets exposure, incident response and staff time. A queue on self-hosted runners is commonly a fleet-capacity or label/routing problem under your control.

GitHub-hosted runners convert much of that operations burden into a per-minute managed service. The trade is not “$0 versus $0.006”; it is managed variable cost versus infrastructure and operational TCO under different trust and network requirements.

8. Current billing/allowance snapshot — timestamp it

Plan / SKU Current September 10, 2026 assumption Design implication
Public + standard hosted Free and unlimited standard-runner use Good for disposable course lab; still avoid waste/abuse.
GitHub Free private 2,000 included minutes/month; 500 MB artifact storage Optimize high-volume private CI and evidence retention.
GitHub Pro private 3,000 included minutes; 1 GB artifact storage Model overage after allowance; repository owner is billed.
GitHub Team private 3,000 included minutes; 2 GB artifact storage Larger runners available but always billed.
GitHub Enterprise Cloud 50,000 included minutes; 50 GB artifact storage High concurrency does not remove need for cost governance.
Baseline Linux x64 Current overage rate $0.006/min Per-job rounding matters to matrix designs.
Cache 10 GB/repo included by default; extra configured usage $0.07/GB-month High-cardinality caches can become an explicit storage cost.

Treat this table as a dated example, not an evergreen contract. Billing, hardware and plan allowances are exactly the kind of platform facts the master contract requires you to re-check before implementation.

9. Artifact retention versus evidence integrity

Shortening artifact retention can save storage, but only after you classify artifacts. Ephemeral benchmark traces can expire quickly. Release binaries, provenance, compliance evidence and incident material may need longer retention or a different durable system. “Delete artifacts to save cost” is not a valid recommendation until evidence requirements are known.

10. Human metrics versus machine evidence

Actions Usage Metrics and Performance Metrics are excellent trend views for workflows, jobs, repositories, runtime OS and runner types. Use them to find candidates: high average queue time, long jobs, frequent failures or high minute consumption. Then drill into exact run IDs, attempts, source SHAs and job timestamps before changing YAML. Aggregates identify where to look; exact run evidence explains why.

11. Worked decision table

Observed evidence Tempting change Preferred first move Why
90% cache hit; install still 2% of critical path Broaden cache restore keys Do nothing to cache No material bottleneck; broader reuse adds staleness risk.
Four 12-minute independent test groups in one job Buy 16-core runner Split into 4 cells, cap parallelism, measure Work is naturally separable; standard runners may reduce wall time cheaply.
One memory-bound linker OOMs Create 20 matrix cells Evaluate larger runner or build redesign Parallelism does not fix per-process memory requirement.
Self-hosted queue p95 18 min; runners 100% busy Blame GitHub queue Add/scale eligible capacity or schedule demand Capacity is in the owned runner fleet.
PRs frequently superseded within 2 min Run all old commits Cancel safe obsolete PR CI Avoids known worthless work if no side effects/evidence need.
macOS test matrix repeats equivalent cells Cache more Remove redundant coverage after test-owner review The workload itself is duplicated; cache is not the primary issue.

12. Performance changes need rollback criteria

Write a rollback condition before rollout: for example, “revert if p95 workflow latency worsens by >15%, test failure classification changes, cache-related correctness defects appear, or processing minutes rise >30% without an approved latency benefit.” Without a rollback threshold, teams tend to rationalize regressions after investing in an optimization.

13. Lesson summary

Parallelism, bigger runners, caches, cancellation and self-hosting are different levers. Each changes a different state: job graph, runner capacity, reuse state, superseded-work policy or infrastructure ownership. Measure the causal layer and keep billing assumptions timestamped. Lesson 4 now focuses on failures where teams optimize the wrong layer or destroy the evidence needed to prove it.

Next lesson

Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Diagnostics, Failure Modes, and Production Practices

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

When should a larger runner beat more matrix jobs conceptually?

Why is cache hit rate not a sufficient KPI?

When is cancel-in-progress a good fit?

Why is a self-hosted runner not “free compute”?

What makes the billing table safe to teach?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.