Chapter 32Lesson 02~270 minutes

Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Guided Hands-On Workflow

Benchmark a disposable pipeline, add a correctness-preserving dependency cache, split independent test work with bounded matrix parallelism, and compare repeated measurements without paid runners.

Hands-onBenchmarkingDependency cacheMatrixmax-parallel

Learning objectives

  • Create a tiny public/disposable benchmark whose workload and source revision are stable across comparisons.
  • Measure a sequential baseline before changing cache or job graph structure.
  • Add a lockfile-bound dependency cache without making cache-hit state control correctness.
  • Split four independent synthetic test partitions into a four-cell matrix capped at two concurrent cells.
  • Compare repeated wall-clock, processing-time, queue/cache and modeled-cost evidence and explain noise.

1. Lab scenario and why it is intentionally small

Create a disposable public repository named gha-performance-lab. Public visibility makes the mandatory standard-runner path free, and all source/data is synthetic. The workload has four independent “test shards” with controlled durations plus one tiny locked Python dependency. The benchmark is not intended to imitate a production test suite; it is intended to make graph shape, startup overhead, cache state and parallelism visible without cloud resources or proprietary code.

You will compare two workflow revisions against the same workload. The baseline uses one job and no dependency cache. The optimized workflow uses an explicit pip download cache and four matrix cells with max-parallel: 2. Both execute all four required shards. If an optimization makes fewer tests run, the comparison is invalid.

2. Preflight: freeze the benchmark inputs

mkdir gha-performance-lab && cd gha-performance-lab
git init
git switch -c main
mkdir -p .github/workflows scripts
printf 'idna==3.10
' > requirements.lock
# Save the Python script from the next section as scripts/shard.py.

git add .
git status --short
python --version || true

Before the first run, record repository visibility, exact commit SHA, workflow revision, runner label, Python version, lockfile SHA-256 and the four shard names. Do not change the synthetic durations between baseline and optimized measurements.

3. The controlled synthetic workload

The script below represents four independent partitions. Sleeping is deliberate: it creates predictable elapsed work without consuming significant CPU or external services. The per-shard SHA-derived proof lets you verify that each named partition actually executed.

from __future__ import annotations
import argparse, hashlib, json, time

DURATIONS = {"unit-a": 2.8, "unit-b": 3.2, "integration-a": 4.4, "integration-b": 4.0}

p = argparse.ArgumentParser()
p.add_argument("shards", nargs="+")
args = p.parse_args()
unknown = sorted(set(args.shards) - set(DURATIONS))
if unknown:
    raise SystemExit(f"unknown shard(s): {unknown}")

started = time.perf_counter()
for shard in args.shards:
    t0 = time.perf_counter()
    # Synthetic deterministic work: sleep represents an independent test partition.
    time.sleep(DURATIONS[shard])
    payload = hashlib.sha256((shard + "-synthetic-v1").encode()).hexdigest()[:12]
    print(json.dumps({"shard": shard, "seconds": round(time.perf_counter()-t0, 3), "proof": payload}))
print(json.dumps({"total_seconds": round(time.perf_counter()-started, 3), "count": len(args.shards)}))

4. Baseline workflow: one job, all work, no cache

Save this as .github/workflows/perf-baseline.yml. permissions: {} proves that the benchmark itself does not need repository write authority. The artifact upload is a GitHub-owned action pinned to a full SHA and stores only tiny synthetic evidence for one day.

name: Performance Lab — Baseline
run-name: "perf-baseline / ${{ github.sha }}"

on:
  workflow_dispatch:

permissions: {}

jobs:
  baseline:
    name: baseline-all-tests
    runs-on: ubuntu-24.04
    timeout-minutes: 10
    steps:
      - name: Checkout exact event revision
        uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1

      - name: Set up Python
        uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
        with:
          python-version: '3.13'
          cache: ''

      - name: Record runner and source identity
        shell: bash
        run: |
          set -euo pipefail
          python --version
          printf 'run_id=%s\nattempt=%s\nsha=%s\nref=%s\nrunner=%s\n' \
            "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA" "$GITHUB_REF" "$RUNNER_NAME"

      - name: Install locked dependency without cache
        shell: bash
        run: |
          set -euo pipefail
          python -m pip install --disable-pip-version-check -r requirements.lock

      - name: Run the full synthetic suite sequentially
        shell: bash
        run: |
          set -euo pipefail
          mkdir -p evidence
          /usr/bin/time -f 'elapsed=%e user=%U sys=%S maxrss_kb=%M' \
            -o evidence/time.txt \
            python scripts/shard.py unit-a unit-b integration-a integration-b \
            | tee evidence/shards.jsonl
          sha256sum requirements.lock scripts/shard.py > evidence/input-sha256.txt
          du -sb evidence | tee evidence/bytes.txt

      - name: Upload tiny benchmark evidence
        if: ${{ always() }}
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
        with:
          name: perf-baseline-${{ github.run_id }}-${{ github.run_attempt }}
          path: evidence/
          retention-days: 1
          if-no-files-found: error

5. Run the baseline more than once

One run can be distorted by runner provisioning, network variance or transient package-index latency. Run the baseline at least three times from the same source SHA. Preserve every run ID and attempt; do not cherry-pick the fastest run.

# After publishing the disposable repository/workflow:
for i in 1 2 3; do
  gh workflow run perf-baseline.yml --ref main
  sleep 3
done

gh run list --workflow perf-baseline.yml --limit 10   --json databaseId,attempt,headSha,status,conclusion,createdAt,updatedAt,url

6. Identify the baseline critical path before optimizing

The baseline has only one job, so its critical path includes runner start, checkout, Python setup, dependency installation, all four shard durations and artifact upload. The synthetic shard section alone is about the sum of four durations. If package installation is only a small fraction of total runtime, a cache cannot produce a dramatic end-to-end win; if tests dominate, parallelism is the stronger hypothesis.

Use the job REST payload and log timestamps to identify which hypothesis is plausible. For repository-level trend data, compare the Actions Performance Metrics queue-time view rather than trying to infer exact queue time from workflow timestamps.

7. Optimization A + B: safe cache plus bounded matrix parallelism

The optimized workflow makes two changes: it caches only pip’s download/wheel cache using a key that includes OS, architecture, Python minor and lockfile digest; and it runs one shard per matrix job with at most two cells concurrently. The install command still runs after every restore so a hit never substitutes for dependency reconciliation.

name: Performance Lab — Optimized
run-name: "perf-optimized / ${{ github.sha }}"

on:
  workflow_dispatch:

permissions: {}

jobs:
  test:
    name: test-${{ matrix.shard }}
    runs-on: ubuntu-24.04
    timeout-minutes: 10
    strategy:
      fail-fast: false
      max-parallel: 2
      matrix:
        shard: [unit-a, unit-b, integration-a, integration-b]
    steps:
      - name: Checkout exact event revision
        uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1

      - name: Set up Python
        uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
        with:
          python-version: '3.13'
          cache: ''

      - name: Discover pip cache path
        id: pip-cache
        shell: bash
        run: |
          set -euo pipefail
          echo "dir=$(python -m pip cache dir)" >> "$GITHUB_OUTPUT"

      - name: Restore dependency-download cache
        id: cache
        uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
        with:
          path: ${{ steps.pip-cache.outputs.dir }}
          key: pip-${{ runner.os }}-${{ runner.arch }}-py313-${{ hashFiles('requirements.lock') }}
          restore-keys: |
            pip-${{ runner.os }}-${{ runner.arch }}-py313-

      - name: Install locked dependency even on a cache hit
        shell: bash
        run: |
          set -euo pipefail
          python -m pip install --disable-pip-version-check -r requirements.lock

      - name: Run one independent shard
        shell: bash
        env:
          SHARD: ${{ matrix.shard }}
          CACHE_HIT: ${{ steps.cache.outputs.cache-hit }}
        run: |
          set -euo pipefail
          mkdir -p evidence
          /usr/bin/time -f 'elapsed=%e user=%U sys=%S maxrss_kb=%M' \
            -o "evidence/${SHARD}-time.txt" \
            python scripts/shard.py "$SHARD" \
            | tee "evidence/${SHARD}.jsonl"
          printf 'cache_hit=%s\nrun_id=%s\nattempt=%s\nsha=%s\n' \
            "$CACHE_HIT" "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA" \
            > "evidence/${SHARD}-state.txt"

      - name: Upload shard evidence
        if: ${{ always() }}
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
        with:
          name: perf-${{ matrix.shard }}-${{ github.run_id }}-${{ github.run_attempt }}
          path: evidence/
          retention-days: 1
          if-no-files-found: error

  required:
    name: performance-lab-required
    if: ${{ always() }}
    needs: [test]
    runs-on: ubuntu-24.04
    steps:
      - name: Preserve the same correctness gate
        shell: bash
        env:
          RESULT: ${{ needs.test.result }}
        run: |
          set -euo pipefail
          test "$RESULT" = success
          echo "All four required shards completed successfully."

8. Why max-parallel is 2 instead of “as much as possible”

Four shards could request four runners at once, but the lab intentionally caps the matrix at two. This reveals that matrix cardinality and concurrency are separate controls. In production, a cap can reduce contention for databases, licenses, self-hosted capacity, cache services or an organization concurrency budget. The right value comes from measured throughput and resource constraints, not the matrix size itself.

9. Run cold and warm optimized measurements

Run the optimized workflow at least three times at the same source revision. The first run may miss the cache; later runs may exact-hit it. Preserve the cache-hit value from each cell rather than labeling an entire workflow “warm” from memory.

for i in 1 2 3; do
  gh workflow run perf-optimized.yml --ref main
  sleep 3
done

gh run list --workflow perf-optimized.yml --limit 10   --json databaseId,attempt,headSha,status,conclusion,createdAt,updatedAt,url

gh cache list --limit 100

10. Measure job processing and a hypothetical private cost

Save the following local script as scripts/summarize-jobs.py. It consumes the exact run-jobs REST response. The dollar figure is explicitly a model for private Linux usage beyond included quota at the September 10, 2026 baseline rate; the actual public lab remains free on standard hosted runners.

#!/usr/bin/env python3
from __future__ import annotations
import json, math, sys
from datetime import datetime, timezone

def dt(s):
    return datetime.fromisoformat(s.replace("Z", "+00:00"))

data=json.load(sys.stdin)
jobs=[j for j in data["jobs"] if j.get("started_at") and j.get("completed_at")]
processing=0.0
rounded_minutes=0
for j in jobs:
    seconds=(dt(j["completed_at"])-dt(j["started_at"])).total_seconds()
    processing += seconds
    rounded_minutes += math.ceil(seconds/60)
    print(f"{j['name']}\t{seconds:.1f}s\t{math.ceil(seconds/60)} modeled billable minute(s)")
print(f"processing_seconds={processing:.1f}")
print(f"rounded_job_minutes={rounded_minutes}")
print(f"hypothetical_private_linux_overage_usd={rounded_minutes*0.006:.4f}")
RUN_ID=123456789
GH_REPO=OWNER/gha-performance-lab

gh api \
  -H 'Accept: application/vnd.github+json' \
  -H 'X-GitHub-Api-Version: 2026-03-10' \
  "repos/$GH_REPO/actions/runs/$RUN_ID/jobs?per_page=100" \
  > run-$RUN_ID-jobs.json

python scripts/summarize-jobs.py < run-$RUN_ID-jobs.json

Do not compare the baseline’s one rounded billable job with optimized matrix jobs without including the matrix’s extra setup and the final aggregate job. Parallelism can lower wall-clock latency while raising rounded processing minutes. That may still be the desired trade.

11. Queue evidence: use the right source

For individual runs, preserve each root job’s started_at, runner labels and workflow created_at as scheduling context. Label any subtraction as dispatch-to-start, not universal queue time. For average queue time, use Actions Performance Metrics, which GitHub explicitly defines for workflows/jobs/repositories/runner types. This avoids blaming application code for an account-level concurrency bottleneck or blaming GitHub for a starved self-hosted runner group.

12. Build a comparison table from repeated runs

Metric Baseline evidence Optimized evidence Interpretation
Source/workload same SHA + four shard proofs same SHA + four shard proofs Must match before comparison is valid.
Workflow wall time median of ≥3 equivalent runs median of ≥3 equivalent runs Primary feedback-latency signal.
Processing seconds sum of job durations sum across matrix + aggregate Shows capacity consumed.
Cache state none per-cell miss/fallback/exact hit Explains dependency timing, not correctness.
Matrix concurrency 1 job 4 cells, max-parallel 2 Explains overlap and queue pressure.
Artifact bytes tiny one-run evidence four tiny evidence sets Check I/O/storage overhead.
Modeled private overage rounded jobs × $0.006/min rounded jobs × $0.006/min Model only; public lab has $0 standard-runner charge.

13. Challenge: choose the layer, not the fashionable optimization

Suppose the optimized workflow has a 90% cache hit rate but median wall-clock time barely changes, while two matrix cells spend most of their time waiting for runners. Which layer should you investigate next? The answer is runner/concurrency capacity and graph shape—not a broader cache key. A broader key could make reuse less correct without addressing the measured bottleneck.

14. Cleanup

  • Keep the six benchmark run IDs until the comparison and evidence packet are complete.
  • Delete the one-day benchmark artifacts or let the one-day retention expire; do not delete first-failure evidence before analysis.
  • Delete the disposable public repository when the exercise is complete.
  • No cloud account, self-hosted runner, paid runner, secret or production registry should exist from this lab.

15. Lesson summary

You changed two bounded variables—dependency-download reuse and independent-test overlap—without changing required coverage. The next lesson generalizes the design choices: where more parallelism, larger runners, cancellation, caching or self-hosted capacity helps and where each shifts cost or risk elsewhere.

Next lesson

Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization: Configuration, Design Patterns, and Trade-Offs

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why run each profile at least three times?

Why does the optimized job still run pip install on an exact cache hit?

What does max-parallel: 2 control?

Why can the optimized workflow’s modeled billable minutes increase?

The cache hit rate is high but jobs wait for runners. What should change first?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.