Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Guided Hands-On Workflow
Benchmark a disposable pipeline, add a correctness-preserving dependency cache, split independent test work with bounded matrix parallelism, and compare repeated measurements without paid runners.
Learning objectives
- Create a tiny public/disposable benchmark whose workload and source revision are stable across comparisons.
- Measure a sequential baseline before changing cache or job graph structure.
- Add a lockfile-bound dependency cache without making cache-hit state control correctness.
- Split four independent synthetic test partitions into a four-cell matrix capped at two concurrent cells.
- Compare repeated wall-clock, processing-time, queue/cache and modeled-cost evidence and explain noise.
1. Lab scenario and why it is intentionally small
Create a disposable public repository named
gha-performance-lab. Public visibility makes the
mandatory standard-runner path free, and all source/data is
synthetic. The workload has four independent “test shards” with
controlled durations plus one tiny locked Python dependency. The
benchmark is not intended to imitate a production test suite; it is
intended to make graph shape, startup overhead, cache state and
parallelism visible without cloud resources or proprietary code.
You will compare two workflow revisions against the same workload.
The baseline uses one job and no dependency cache. The optimized
workflow uses an explicit pip download cache and four matrix cells
with max-parallel: 2. Both execute all four required
shards. If an optimization makes fewer tests run, the comparison is
invalid.
2. Preflight: freeze the benchmark inputs
mkdir gha-performance-lab && cd gha-performance-lab
git init
git switch -c main
mkdir -p .github/workflows scripts
printf 'idna==3.10
' > requirements.lock
# Save the Python script from the next section as scripts/shard.py.
git add .
git status --short
python --version || true
Before the first run, record repository visibility, exact commit SHA, workflow revision, runner label, Python version, lockfile SHA-256 and the four shard names. Do not change the synthetic durations between baseline and optimized measurements.
3. The controlled synthetic workload
The script below represents four independent partitions. Sleeping is deliberate: it creates predictable elapsed work without consuming significant CPU or external services. The per-shard SHA-derived proof lets you verify that each named partition actually executed.
from __future__ import annotations
import argparse, hashlib, json, time
DURATIONS = {"unit-a": 2.8, "unit-b": 3.2, "integration-a": 4.4, "integration-b": 4.0}
p = argparse.ArgumentParser()
p.add_argument("shards", nargs="+")
args = p.parse_args()
unknown = sorted(set(args.shards) - set(DURATIONS))
if unknown:
raise SystemExit(f"unknown shard(s): {unknown}")
started = time.perf_counter()
for shard in args.shards:
t0 = time.perf_counter()
# Synthetic deterministic work: sleep represents an independent test partition.
time.sleep(DURATIONS[shard])
payload = hashlib.sha256((shard + "-synthetic-v1").encode()).hexdigest()[:12]
print(json.dumps({"shard": shard, "seconds": round(time.perf_counter()-t0, 3), "proof": payload}))
print(json.dumps({"total_seconds": round(time.perf_counter()-started, 3), "count": len(args.shards)}))
4. Baseline workflow: one job, all work, no cache
Save this as .github/workflows/perf-baseline.yml.
permissions: {} proves that the benchmark itself does
not need repository write authority. The artifact upload is a
GitHub-owned action pinned to a full SHA and stores only tiny
synthetic evidence for one day.
name: Performance Lab — Baseline
run-name: "perf-baseline / ${{ github.sha }}"
on:
workflow_dispatch:
permissions: {}
jobs:
baseline:
name: baseline-all-tests
runs-on: ubuntu-24.04
timeout-minutes: 10
steps:
- name: Checkout exact event revision
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
- name: Set up Python
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
with:
python-version: '3.13'
cache: ''
- name: Record runner and source identity
shell: bash
run: |
set -euo pipefail
python --version
printf 'run_id=%s\nattempt=%s\nsha=%s\nref=%s\nrunner=%s\n' \
"$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA" "$GITHUB_REF" "$RUNNER_NAME"
- name: Install locked dependency without cache
shell: bash
run: |
set -euo pipefail
python -m pip install --disable-pip-version-check -r requirements.lock
- name: Run the full synthetic suite sequentially
shell: bash
run: |
set -euo pipefail
mkdir -p evidence
/usr/bin/time -f 'elapsed=%e user=%U sys=%S maxrss_kb=%M' \
-o evidence/time.txt \
python scripts/shard.py unit-a unit-b integration-a integration-b \
| tee evidence/shards.jsonl
sha256sum requirements.lock scripts/shard.py > evidence/input-sha256.txt
du -sb evidence | tee evidence/bytes.txt
- name: Upload tiny benchmark evidence
if: ${{ always() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
with:
name: perf-baseline-${{ github.run_id }}-${{ github.run_attempt }}
path: evidence/
retention-days: 1
if-no-files-found: error
5. Run the baseline more than once
One run can be distorted by runner provisioning, network variance or transient package-index latency. Run the baseline at least three times from the same source SHA. Preserve every run ID and attempt; do not cherry-pick the fastest run.
# After publishing the disposable repository/workflow:
for i in 1 2 3; do
gh workflow run perf-baseline.yml --ref main
sleep 3
done
gh run list --workflow perf-baseline.yml --limit 10 --json databaseId,attempt,headSha,status,conclusion,createdAt,updatedAt,url
6. Identify the baseline critical path before optimizing
The baseline has only one job, so its critical path includes runner start, checkout, Python setup, dependency installation, all four shard durations and artifact upload. The synthetic shard section alone is about the sum of four durations. If package installation is only a small fraction of total runtime, a cache cannot produce a dramatic end-to-end win; if tests dominate, parallelism is the stronger hypothesis.
Use the job REST payload and log timestamps to identify which hypothesis is plausible. For repository-level trend data, compare the Actions Performance Metrics queue-time view rather than trying to infer exact queue time from workflow timestamps.
7. Optimization A + B: safe cache plus bounded matrix parallelism
The optimized workflow makes two changes: it caches only pip’s download/wheel cache using a key that includes OS, architecture, Python minor and lockfile digest; and it runs one shard per matrix job with at most two cells concurrently. The install command still runs after every restore so a hit never substitutes for dependency reconciliation.
name: Performance Lab — Optimized
run-name: "perf-optimized / ${{ github.sha }}"
on:
workflow_dispatch:
permissions: {}
jobs:
test:
name: test-${{ matrix.shard }}
runs-on: ubuntu-24.04
timeout-minutes: 10
strategy:
fail-fast: false
max-parallel: 2
matrix:
shard: [unit-a, unit-b, integration-a, integration-b]
steps:
- name: Checkout exact event revision
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
- name: Set up Python
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
with:
python-version: '3.13'
cache: ''
- name: Discover pip cache path
id: pip-cache
shell: bash
run: |
set -euo pipefail
echo "dir=$(python -m pip cache dir)" >> "$GITHUB_OUTPUT"
- name: Restore dependency-download cache
id: cache
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
with:
path: ${{ steps.pip-cache.outputs.dir }}
key: pip-${{ runner.os }}-${{ runner.arch }}-py313-${{ hashFiles('requirements.lock') }}
restore-keys: |
pip-${{ runner.os }}-${{ runner.arch }}-py313-
- name: Install locked dependency even on a cache hit
shell: bash
run: |
set -euo pipefail
python -m pip install --disable-pip-version-check -r requirements.lock
- name: Run one independent shard
shell: bash
env:
SHARD: ${{ matrix.shard }}
CACHE_HIT: ${{ steps.cache.outputs.cache-hit }}
run: |
set -euo pipefail
mkdir -p evidence
/usr/bin/time -f 'elapsed=%e user=%U sys=%S maxrss_kb=%M' \
-o "evidence/${SHARD}-time.txt" \
python scripts/shard.py "$SHARD" \
| tee "evidence/${SHARD}.jsonl"
printf 'cache_hit=%s\nrun_id=%s\nattempt=%s\nsha=%s\n' \
"$CACHE_HIT" "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA" \
> "evidence/${SHARD}-state.txt"
- name: Upload shard evidence
if: ${{ always() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
with:
name: perf-${{ matrix.shard }}-${{ github.run_id }}-${{ github.run_attempt }}
path: evidence/
retention-days: 1
if-no-files-found: error
required:
name: performance-lab-required
if: ${{ always() }}
needs: [test]
runs-on: ubuntu-24.04
steps:
- name: Preserve the same correctness gate
shell: bash
env:
RESULT: ${{ needs.test.result }}
run: |
set -euo pipefail
test "$RESULT" = success
echo "All four required shards completed successfully."
8. Why max-parallel is 2 instead of “as much as possible”
Four shards could request four runners at once, but the lab intentionally caps the matrix at two. This reveals that matrix cardinality and concurrency are separate controls. In production, a cap can reduce contention for databases, licenses, self-hosted capacity, cache services or an organization concurrency budget. The right value comes from measured throughput and resource constraints, not the matrix size itself.
9. Run cold and warm optimized measurements
Run the optimized workflow at least three times at the same source
revision. The first run may miss the cache; later runs may exact-hit
it. Preserve the cache-hit value from each cell rather
than labeling an entire workflow “warm” from memory.
for i in 1 2 3; do
gh workflow run perf-optimized.yml --ref main
sleep 3
done
gh run list --workflow perf-optimized.yml --limit 10 --json databaseId,attempt,headSha,status,conclusion,createdAt,updatedAt,url
gh cache list --limit 100
10. Measure job processing and a hypothetical private cost
Save the following local script as
scripts/summarize-jobs.py. It consumes the exact
run-jobs REST response. The dollar figure is explicitly a model for
private Linux usage beyond included quota at the September 10, 2026
baseline rate; the actual public lab remains free on standard hosted
runners.
#!/usr/bin/env python3
from __future__ import annotations
import json, math, sys
from datetime import datetime, timezone
def dt(s):
return datetime.fromisoformat(s.replace("Z", "+00:00"))
data=json.load(sys.stdin)
jobs=[j for j in data["jobs"] if j.get("started_at") and j.get("completed_at")]
processing=0.0
rounded_minutes=0
for j in jobs:
seconds=(dt(j["completed_at"])-dt(j["started_at"])).total_seconds()
processing += seconds
rounded_minutes += math.ceil(seconds/60)
print(f"{j['name']}\t{seconds:.1f}s\t{math.ceil(seconds/60)} modeled billable minute(s)")
print(f"processing_seconds={processing:.1f}")
print(f"rounded_job_minutes={rounded_minutes}")
print(f"hypothetical_private_linux_overage_usd={rounded_minutes*0.006:.4f}")
RUN_ID=123456789
GH_REPO=OWNER/gha-performance-lab
gh api \
-H 'Accept: application/vnd.github+json' \
-H 'X-GitHub-Api-Version: 2026-03-10' \
"repos/$GH_REPO/actions/runs/$RUN_ID/jobs?per_page=100" \
> run-$RUN_ID-jobs.json
python scripts/summarize-jobs.py < run-$RUN_ID-jobs.json
Do not compare the baseline’s one rounded billable job with optimized matrix jobs without including the matrix’s extra setup and the final aggregate job. Parallelism can lower wall-clock latency while raising rounded processing minutes. That may still be the desired trade.
11. Queue evidence: use the right source
For individual runs, preserve each root job’s
started_at, runner labels and workflow
created_at as scheduling context. Label any subtraction
as dispatch-to-start, not universal queue time. For average
queue time, use Actions Performance Metrics, which GitHub explicitly
defines for workflows/jobs/repositories/runner types. This avoids
blaming application code for an account-level concurrency bottleneck
or blaming GitHub for a starved self-hosted runner group.
12. Build a comparison table from repeated runs
| Metric | Baseline evidence | Optimized evidence | Interpretation |
|---|---|---|---|
| Source/workload | same SHA + four shard proofs | same SHA + four shard proofs | Must match before comparison is valid. |
| Workflow wall time | median of ≥3 equivalent runs | median of ≥3 equivalent runs | Primary feedback-latency signal. |
| Processing seconds | sum of job durations | sum across matrix + aggregate | Shows capacity consumed. |
| Cache state | none | per-cell miss/fallback/exact hit | Explains dependency timing, not correctness. |
| Matrix concurrency | 1 job | 4 cells, max-parallel 2 | Explains overlap and queue pressure. |
| Artifact bytes | tiny one-run evidence | four tiny evidence sets | Check I/O/storage overhead. |
| Modeled private overage | rounded jobs × $0.006/min | rounded jobs × $0.006/min | Model only; public lab has $0 standard-runner charge. |
13. Challenge: choose the layer, not the fashionable optimization
Suppose the optimized workflow has a 90% cache hit rate but median wall-clock time barely changes, while two matrix cells spend most of their time waiting for runners. Which layer should you investigate next? The answer is runner/concurrency capacity and graph shape—not a broader cache key. A broader key could make reuse less correct without addressing the measured bottleneck.
14. Cleanup
- Keep the six benchmark run IDs until the comparison and evidence packet are complete.
- Delete the one-day benchmark artifacts or let the one-day retention expire; do not delete first-failure evidence before analysis.
- Delete the disposable public repository when the exercise is complete.
- No cloud account, self-hosted runner, paid runner, secret or production registry should exist from this lab.
15. Lesson summary
You changed two bounded variables—dependency-download reuse and independent-test overlap—without changing required coverage. The next lesson generalizes the design choices: where more parallelism, larger runners, cancellation, caching or self-hosted capacity helps and where each shifts cost or risk elsewhere.
Knowledge check
Why run each profile at least three times?
To reduce the chance that one noisy runner/network/queue event determines the optimization decision; preserve the distribution and median rather than cherry-picking.
Why does the optimized job still run pip install on an exact cache hit?
The cache stores dependency download/wheel data, not proof of a correct installed environment. Installation/reconciliation remains part of correctness.
What does max-parallel: 2 control?
How many generated matrix jobs may run concurrently for that strategy; it does not reduce the number of required matrix cells.
Why can the optimized workflow’s modeled billable minutes increase?
Each matrix/aggregate job has setup time and private hosted billing rounds each job up separately even though jobs overlap in wall-clock time.
The cache hit rate is high but jobs wait for runners. What should change first?
Investigate eligible runner capacity/concurrency and critical-path scheduling; a broader cache key does not solve queueing and can weaken cache correctness.
Official references and version notes
- GitHub Actions metrics — Current usage and performance metrics, including run time, queue time and failure-rate views.
- Viewing Actions metrics — Repository and organization Actions Usage/Performance Metrics and aggregation windows.
- Actions limits — Current matrix, concurrency, queue and job-duration limits; limits are explicitly subject to change.
- GitHub-hosted runners reference — Current public/private standard runner hardware, labels and isolation characteristics.
- Actions runner pricing — Current per-minute hosted runner prices and per-job minute rounding.
- GitHub Actions billing — Current plan allowances, free public standard-runner use, storage pricing and billing behavior.
- Concurrency — Concurrency groups, cancellation behavior and current queue:max semantics.
- Dependency caching reference — Cache identity, restore behavior, storage/eviction and rate-limit behavior.
- REST: workflow jobs — Job IDs, runner labels, started/completed timestamps and step timing evidence.
- REST: workflow runs — Run metadata, attempts, source SHA, status/conclusion and usage-related inspection.
- actions/checkout v7.0.1 — Full commit SHA used by executable lab examples.
- actions/setup-python v7.0.0 — Full commit SHA used for Python 3.13 setup in the lab.
- actions/cache v6.1.0 — Full commit SHA used for the explicit dependency-download cache.
- actions/upload-artifact v7.0.1 — Full commit SHA used only for tiny bounded benchmark evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.