Checkpoint Lab — Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization
Collect a repeatable baseline, make two bounded optimizations, compare wall-clock/queue/cache/usage evidence, reject one unsafe optimization, and preserve a defensible performance evidence packet.
Learning objectives
- Capture a three-run baseline and a repeated optimized profile against the same source/workload.
- Prove dependency caching and bounded matrix parallelism are optimizations rather than correctness dependencies.
- Compare wall-clock, job processing, queue signals, cache state, artifact size and hypothetical private cost.
- Reject a proposed optimization that cuts required test/security coverage and document why.
- Deliver a reproducible evidence packet with assumptions, limits, rollback criteria and the bridge to failure-recovery engineering.
1. Checkpoint mission
Use a disposable public repository named
gha-performance-checkpoint. Establish a sequential
no-cache baseline for the four synthetic shards, then apply exactly
two bounded optimizations: a precise pip download cache and a
four-cell test matrix capped at two concurrent cells. Execute all
required shards in both profiles. Collect enough evidence to say
whether wall-clock latency improved, whether processing minutes
changed, whether the cache was useful and whether any queue signal
dominates.
Finally, review and reject one “optimization” that would disable required integration coverage. The checkpoint passes only if the faster design is behaviorally equivalent at the required-check boundary.
2. Current assumptions and preflight
| Assumption | Checkpoint value / proof |
|---|---|
| Date / platform | GitHub.com behavior and pricing checked 2026-09-10. |
| Repository | Disposable public repository; standard hosted runner use is currently free. |
| Runner |
ubuntu-24.04; current public standard Linux is
4 vCPU / 16 GB / 14 GB SSD. Record actual runner/image
evidence.
|
| Python |
3.13 via setup-python v7.0.0 at
5fda3b95a4ea91299a34e894583c3862153e4b97.
|
| Checkout |
actions/checkout v7.0.1 at
3d3c42e5aac5ba805825da76410c181273ba90b1.
|
| Cache |
actions/cache v6.1.0 at
55cc8345863c7cc4c66a329aec7e433d2d1c52a9.
|
| Artifact evidence |
actions/upload-artifact v7.0.1 at
043fb46d1a93c77aae656e7c1c64a875d1fc6a0a,
retention 1 day.
|
| Permissions |
permissions: {}; no secret, OIDC, package,
environment or cloud permission.
|
| Cost model | Hypothetical private Linux overage uses current $0.006/min and per-job rounding; actual public lab standard-runner compute charge is $0. |
gh --version
gh auth status
gh repo view --json nameWithOwner,visibility,url
git rev-parse HEAD
git status --short
sha256sum requirements.lock scripts/shard.py
3. Predict at least two state changes before running
Write checkpoint-predictions.md before dispatch.
Include at least these four predictions and later mark each
verified/disproved:
- Prediction A: baseline and optimized runs execute the same four shard names at the same source SHA; only graph/cache behavior differs.
- Prediction B: the optimized warm profile can reduce wall-clock critical path through two-way overlap, but total processing/rounded-job minutes may increase because four test jobs plus one aggregate job each have startup overhead.
-
Prediction C: cache misses and exact hits both
pass because dependency installation always runs against
requirements.lock. - Prediction D: the proposed “skip integration-a/integration-b on ordinary PRs” change is rejected because it changes required correctness coverage rather than execution efficiency.
4. Exact workload fixture
Use the same requirements.lock and
scripts/shard.py from Lesson 2. The fixed shard names
and durations are part of the workload contract. Do not modify them
between the baseline and optimized samples.
requirements.lock
-----------------
idna==3.10
Required shards
---------------
unit-a
unit-b
integration-a
integration-b
from __future__ import annotations
import argparse, hashlib, json, time
DURATIONS = {"unit-a": 2.8, "unit-b": 3.2, "integration-a": 4.4, "integration-b": 4.0}
p = argparse.ArgumentParser()
p.add_argument("shards", nargs="+")
args = p.parse_args()
unknown = sorted(set(args.shards) - set(DURATIONS))
if unknown:
raise SystemExit(f"unknown shard(s): {unknown}")
started = time.perf_counter()
for shard in args.shards:
t0 = time.perf_counter()
# Synthetic deterministic work: sleep represents an independent test partition.
time.sleep(DURATIONS[shard])
payload = hashlib.sha256((shard + "-synthetic-v1").encode()).hexdigest()[:12]
print(json.dumps({"shard": shard, "seconds": round(time.perf_counter()-t0, 3), "proof": payload}))
print(json.dumps({"total_seconds": round(time.perf_counter()-started, 3), "count": len(args.shards)}))
5. Baseline workflow
Create .github/workflows/perf-baseline.yml exactly as
below. It runs all four shards sequentially and stores a tiny
one-day evidence artifact. Run it three times from the same commit.
name: Performance Lab — Baseline
run-name: "perf-baseline / ${{ github.sha }}"
on:
workflow_dispatch:
permissions: {}
jobs:
baseline:
name: baseline-all-tests
runs-on: ubuntu-24.04
timeout-minutes: 10
steps:
- name: Checkout exact event revision
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
- name: Set up Python
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
with:
python-version: '3.13'
cache: ''
- name: Record runner and source identity
shell: bash
run: |
set -euo pipefail
python --version
printf 'run_id=%s\nattempt=%s\nsha=%s\nref=%s\nrunner=%s\n' \
"$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA" "$GITHUB_REF" "$RUNNER_NAME"
- name: Install locked dependency without cache
shell: bash
run: |
set -euo pipefail
python -m pip install --disable-pip-version-check -r requirements.lock
- name: Run the full synthetic suite sequentially
shell: bash
run: |
set -euo pipefail
mkdir -p evidence
/usr/bin/time -f 'elapsed=%e user=%U sys=%S maxrss_kb=%M' \
-o evidence/time.txt \
python scripts/shard.py unit-a unit-b integration-a integration-b \
| tee evidence/shards.jsonl
sha256sum requirements.lock scripts/shard.py > evidence/input-sha256.txt
du -sb evidence | tee evidence/bytes.txt
- name: Upload tiny benchmark evidence
if: ${{ always() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
with:
name: perf-baseline-${{ github.run_id }}-${{ github.run_attempt }}
path: evidence/
retention-days: 1
if-no-files-found: error
6. Optimized workflow
Create .github/workflows/perf-checkpoint.yml. The job
graph changes, but the required test set does not.
max-parallel: 2 bounds demand and the final
performance-lab-required job fails unless every matrix
cell succeeds.
name: Performance Checkpoint — Optimized
run-name: "perf-checkpoint / ${{ github.sha }}"
on:
workflow_dispatch:
permissions: {}
jobs:
test:
name: test-${{ matrix.shard }}
runs-on: ubuntu-24.04
timeout-minutes: 10
strategy:
fail-fast: false
max-parallel: 2
matrix:
shard: [unit-a, unit-b, integration-a, integration-b]
steps:
- name: Checkout exact event revision
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
- name: Set up Python
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
with:
python-version: '3.13'
cache: ''
- name: Discover pip cache path
id: pip-cache
shell: bash
run: |
set -euo pipefail
echo "dir=$(python -m pip cache dir)" >> "$GITHUB_OUTPUT"
- name: Restore dependency-download cache
id: cache
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
with:
path: ${{ steps.pip-cache.outputs.dir }}
key: pip-${{ runner.os }}-${{ runner.arch }}-py313-${{ hashFiles('requirements.lock') }}
restore-keys: |
pip-${{ runner.os }}-${{ runner.arch }}-py313-
- name: Install locked dependency even on a cache hit
shell: bash
run: |
set -euo pipefail
python -m pip install --disable-pip-version-check -r requirements.lock
- name: Run one independent shard
shell: bash
env:
SHARD: ${{ matrix.shard }}
CACHE_HIT: ${{ steps.cache.outputs.cache-hit }}
run: |
set -euo pipefail
mkdir -p evidence
/usr/bin/time -f 'elapsed=%e user=%U sys=%S maxrss_kb=%M' \
-o "evidence/${SHARD}-time.txt" \
python scripts/shard.py "$SHARD" \
| tee "evidence/${SHARD}.jsonl"
printf 'cache_hit=%s\nrun_id=%s\nattempt=%s\nsha=%s\n' \
"$CACHE_HIT" "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA" \
> "evidence/${SHARD}-state.txt"
- name: Upload shard evidence
if: ${{ always() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
with:
name: perf-${{ matrix.shard }}-${{ github.run_id }}-${{ github.run_attempt }}
path: evidence/
retention-days: 1
if-no-files-found: error
required:
name: performance-lab-required
if: ${{ always() }}
needs: [test]
runs-on: ubuntu-24.04
steps:
- name: Preserve the same correctness gate
shell: bash
env:
RESULT: ${{ needs.test.result }}
run: |
set -euo pipefail
test "$RESULT" = success
echo "All four required shards completed successfully."
7. Execute a controlled measurement sequence
- Publish one commit containing the workload and both workflow files. Record its SHA.
- Dispatch the baseline three times; record every run ID/attempt and wait for terminal success.
- Dispatch the optimized workflow once with a cold/unknown cache state; record each matrix cell’s cache-hit output.
- Dispatch the optimized workflow at least two more times without changing the lockfile or Python/runtime inputs.
- Do not delete caches or artifacts until the comparison is complete.
- If queue conditions are visibly abnormal, keep the run as evidence and add more samples rather than silently discarding it.
for i in 1 2 3; do gh workflow run perf-baseline.yml --ref main; sleep 3; done
for i in 1 2 3; do gh workflow run perf-checkpoint.yml --ref main; sleep 3; done
gh run list --limit 20 --json databaseId,workflowName,attempt,headSha,status,conclusion,createdAt,updatedAt,url
8. Collect exact run/job/cache/artifact evidence
GH_REPO=OWNER/gha-performance-checkpoint
RUN_ID=123456789
gh run view "$RUN_ID" --json databaseId,attempt,event,headSha,status,conclusion,createdAt,updatedAt,url
gh api \
-H 'Accept: application/vnd.github+json' \
-H 'X-GitHub-Api-Version: 2026-03-10' \
"repos/$GH_REPO/actions/runs/$RUN_ID/jobs?per_page=100" \
> "checkpoint-evidence/run-$RUN_ID-jobs.json"
gh api \
-H 'X-GitHub-Api-Version: 2026-03-10' \
"repos/$GH_REPO/actions/runs/$RUN_ID/artifacts" \
> "checkpoint-evidence/run-$RUN_ID-artifacts.json"
gh cache list --limit 100 > checkpoint-evidence/cache-list.txt
For queue evidence, capture the repository Actions Performance
Metrics view or its exported CSV if available. For an individual
root job, you may also record workflow created_at and
job started_at as a dispatch-to-start signal, but label
it exactly that. Do not claim it is an authoritative queue duration
for dependent jobs.
9. Calculate processing and modeled cost
Use the Lesson 2 scripts/summarize-jobs.py script for
every sampled run. Then calculate median workflow wall time and
median processing seconds for each profile. If a public lab has zero
standard-runner charge, the hypothetical private Linux overage
column remains a model, not a bill.
#!/usr/bin/env python3
from __future__ import annotations
import json, math, sys
from datetime import datetime, timezone
def dt(s):
return datetime.fromisoformat(s.replace("Z", "+00:00"))
data=json.load(sys.stdin)
jobs=[j for j in data["jobs"] if j.get("started_at") and j.get("completed_at")]
processing=0.0
rounded_minutes=0
for j in jobs:
seconds=(dt(j["completed_at"])-dt(j["started_at"])).total_seconds()
processing += seconds
rounded_minutes += math.ceil(seconds/60)
print(f"{j['name']}\t{seconds:.1f}s\t{math.ceil(seconds/60)} modeled billable minute(s)")
print(f"processing_seconds={processing:.1f}")
print(f"rounded_job_minutes={rounded_minutes}")
print(f"hypothetical_private_linux_overage_usd={rounded_minutes*0.006:.4f}")
Record both latency and processing. If the optimized median wall time drops from 25 seconds to 15 seconds while processing rises from 25 to 35 seconds, the correct conclusion is “faster feedback at higher runner consumption,” not simply “40% cheaper/faster.”
10. Reject one unsafe optimization explicitly
A teammate proposes removing both integration shards on ordinary
pull requests. Do not run this as the accepted optimization. Record
it in rejected-optimization.md with the causal reason:
the job graph would be faster because the required workload is
smaller, so it violates the equivalence contract and creates a
quality risk.
# REJECTED PROPOSAL — changes correctness coverage.
matrix:
shard: [unit-a, unit-b] # integration-a and integration-b disappeared
A safe alternative is to keep all four required shards and optimize fixture startup, split independent integration work with bounded concurrency, or move additional non-required exhaustive coverage to scheduled workflows after explicit policy review.
11. Required before/after comparison
| Evidence | Baseline | Optimized cold | Optimized warm | Pass criterion |
|---|---|---|---|---|
| Source SHA | record | same | same | Exact match. |
| Required shards | 4 | 4 | 4 | No coverage loss. |
| Workflow wall time | median of 3 | record | median warm samples | Interpret with queue context. |
| Processing seconds | sum jobs | sum jobs | sum jobs | Explain increases/decreases. |
| Cache hit state | N/A | record each cell | record each cell | Miss and hit both correct. |
| Matrix / max-parallel | 1 sequential job | 4 / 2 | 4 / 2 | Bounded concurrency. |
| Artifact bytes | record | record | record | Tiny bounded evidence. |
| Modeled private cost | rounded × $0.006 | rounded × $0.006 | rounded × $0.006 | Clearly marked hypothetical. |
12. Required evidence packet
-
checkpoint-predictions.mdwith verified/disproved outcomes. - Exact repository visibility, source SHA and both workflow file SHAs/revisions.
- Run ID/attempt/event/source SHA/status/conclusion/URL for every sampled run.
- Job IDs/names/start/completion timestamps, runner labels/name and step timing for representative runs.
- Cache action SHA, primary key components, exact/fallback/miss state and cache metadata/size where available.
- Artifact IDs/names/sizes/digests or metadata for the tiny benchmark evidence and its 1-day retention assumption.
-
Matrix cardinality and
max-parallel; final required-check result proving all four shards executed. - Median wall-clock and processing-time comparison plus the hypothetical private billing calculation and timestamped $0.006/min assumption.
-
rejected-optimization.mdshowing why reduced test coverage was rejected. - Assumptions/limitations: public lab hardware differs from private standard Linux; queue is noisy; synthetic sleeps do not model CPU scaling; billing and limits can change.
13. Rollback and cleanup
- If optimized runs lose required coverage, restore the baseline graph immediately; performance evidence is invalid until equivalence returns.
- If the cache causes correctness anomalies, preserve cache metadata, bypass/re-key it, and rerun cold before deleting exact disposable cache entries.
- If matrix processing cost rises without an approved latency benefit, reduce granularity or restore sequential execution.
- Delete benchmark artifacts/caches/repository only after the evidence packet is reviewed; no production resource should exist.
- Do not interpret cleanup as rollback of an external system—this checkpoint intentionally has no deployment side effect.
14. What Chapter 32 adds — and the bridge to Chapter 33
Chapter 32 adds a measured optimization operating model: latency and usage are separate, queueing belongs to a capacity layer, caches are optional, matrix parallelism is bounded, pricing/limits are dated assumptions, and correctness/evidence define the acceptance boundary. Chapter 33 uses this foundation for failure recovery: reruns, idempotency, rollback and incident-safe automation must recover service without duplicating side effects or erasing the evidence that explained the failure.
Knowledge check
What proves the checkpoint optimized the pipeline rather than reduced the workload?
Baseline and optimized runs share the same source/workload identity and all four required shard proofs; the final aggregate check requires every matrix cell to succeed.
Why are three baseline and three optimized samples better than one pair?
They expose queue/network/runner noise and support median/distribution comparison instead of attributing a random fluctuation to the YAML change.
A warm cache hit makes install fail. What does that mean?
The cache has become incompatible or the workflow incorrectly treated restored data as authoritative. Preserve metadata, bypass/re-key, and keep installation/reconciliation as correctness.
Why is the private-cost column hypothetical?
The mandatory checkpoint is a public repository where standard hosted runners are currently free; the $0.006/min calculation models a private over-quota case only.
Which proposed optimization must be rejected even if it halves runtime?
Dropping required integration shards, because it changes quality/correctness coverage rather than making equivalent work execute more efficiently.
Official references and version notes
- GitHub Actions metrics — Current usage and performance metrics, including run time, queue time and failure-rate views.
- Viewing Actions metrics — Repository and organization Actions Usage/Performance Metrics and aggregation windows.
- Actions limits — Current matrix, concurrency, queue and job-duration limits; limits are explicitly subject to change.
- GitHub-hosted runners reference — Current public/private standard runner hardware, labels and isolation characteristics.
- Actions runner pricing — Current per-minute hosted runner prices and per-job minute rounding.
- GitHub Actions billing — Current plan allowances, free public standard-runner use, storage pricing and billing behavior.
- Concurrency — Concurrency groups, cancellation behavior and current queue:max semantics.
- Dependency caching reference — Cache identity, restore behavior, storage/eviction and rate-limit behavior.
- REST: workflow jobs — Job IDs, runner labels, started/completed timestamps and step timing evidence.
- REST: workflow runs — Run metadata, attempts, source SHA, status/conclusion and usage-related inspection.
- actions/checkout v7.0.1 — Full commit SHA used by executable lab examples.
- actions/setup-python v7.0.0 — Full commit SHA used for Python 3.13 setup in the lab.
- actions/cache v6.1.0 — Full commit SHA used for the explicit dependency-download cache.
- actions/upload-artifact v7.0.1 — Full commit SHA used only for tiny bounded benchmark evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.