Chapter 32Lesson 05~290 minutes

Checkpoint Lab — Performance, Queue Time, Parallelism, Caching, Usage, and Cost Optimization

Collect a repeatable baseline, make two bounded optimizations, compare wall-clock/queue/cache/usage evidence, reject one unsafe optimization, and preserve a defensible performance evidence packet.

CheckpointBefore/afterCost modelEvidence packetOptimization

Learning objectives

  • Capture a three-run baseline and a repeated optimized profile against the same source/workload.
  • Prove dependency caching and bounded matrix parallelism are optimizations rather than correctness dependencies.
  • Compare wall-clock, job processing, queue signals, cache state, artifact size and hypothetical private cost.
  • Reject a proposed optimization that cuts required test/security coverage and document why.
  • Deliver a reproducible evidence packet with assumptions, limits, rollback criteria and the bridge to failure-recovery engineering.

1. Checkpoint mission

Use a disposable public repository named gha-performance-checkpoint. Establish a sequential no-cache baseline for the four synthetic shards, then apply exactly two bounded optimizations: a precise pip download cache and a four-cell test matrix capped at two concurrent cells. Execute all required shards in both profiles. Collect enough evidence to say whether wall-clock latency improved, whether processing minutes changed, whether the cache was useful and whether any queue signal dominates.

Finally, review and reject one “optimization” that would disable required integration coverage. The checkpoint passes only if the faster design is behaviorally equivalent at the required-check boundary.

2. Current assumptions and preflight

Assumption Checkpoint value / proof
Date / platform GitHub.com behavior and pricing checked 2026-09-10.
Repository Disposable public repository; standard hosted runner use is currently free.
Runner ubuntu-24.04; current public standard Linux is 4 vCPU / 16 GB / 14 GB SSD. Record actual runner/image evidence.
Python 3.13 via setup-python v7.0.0 at 5fda3b95a4ea91299a34e894583c3862153e4b97.
Checkout actions/checkout v7.0.1 at 3d3c42e5aac5ba805825da76410c181273ba90b1.
Cache actions/cache v6.1.0 at 55cc8345863c7cc4c66a329aec7e433d2d1c52a9.
Artifact evidence actions/upload-artifact v7.0.1 at 043fb46d1a93c77aae656e7c1c64a875d1fc6a0a, retention 1 day.
Permissions permissions: {}; no secret, OIDC, package, environment or cloud permission.
Cost model Hypothetical private Linux overage uses current $0.006/min and per-job rounding; actual public lab standard-runner compute charge is $0.
gh --version
gh auth status
gh repo view --json nameWithOwner,visibility,url
git rev-parse HEAD
git status --short
sha256sum requirements.lock scripts/shard.py

3. Predict at least two state changes before running

Write checkpoint-predictions.md before dispatch. Include at least these four predictions and later mark each verified/disproved:

  • Prediction A: baseline and optimized runs execute the same four shard names at the same source SHA; only graph/cache behavior differs.
  • Prediction B: the optimized warm profile can reduce wall-clock critical path through two-way overlap, but total processing/rounded-job minutes may increase because four test jobs plus one aggregate job each have startup overhead.
  • Prediction C: cache misses and exact hits both pass because dependency installation always runs against requirements.lock.
  • Prediction D: the proposed “skip integration-a/integration-b on ordinary PRs” change is rejected because it changes required correctness coverage rather than execution efficiency.

4. Exact workload fixture

Use the same requirements.lock and scripts/shard.py from Lesson 2. The fixed shard names and durations are part of the workload contract. Do not modify them between the baseline and optimized samples.

requirements.lock
-----------------
idna==3.10

Required shards
---------------
unit-a
unit-b
integration-a
integration-b
from __future__ import annotations
import argparse, hashlib, json, time

DURATIONS = {"unit-a": 2.8, "unit-b": 3.2, "integration-a": 4.4, "integration-b": 4.0}

p = argparse.ArgumentParser()
p.add_argument("shards", nargs="+")
args = p.parse_args()
unknown = sorted(set(args.shards) - set(DURATIONS))
if unknown:
    raise SystemExit(f"unknown shard(s): {unknown}")

started = time.perf_counter()
for shard in args.shards:
    t0 = time.perf_counter()
    # Synthetic deterministic work: sleep represents an independent test partition.
    time.sleep(DURATIONS[shard])
    payload = hashlib.sha256((shard + "-synthetic-v1").encode()).hexdigest()[:12]
    print(json.dumps({"shard": shard, "seconds": round(time.perf_counter()-t0, 3), "proof": payload}))
print(json.dumps({"total_seconds": round(time.perf_counter()-started, 3), "count": len(args.shards)}))

5. Baseline workflow

Create .github/workflows/perf-baseline.yml exactly as below. It runs all four shards sequentially and stores a tiny one-day evidence artifact. Run it three times from the same commit.

name: Performance Lab — Baseline
run-name: "perf-baseline / ${{ github.sha }}"

on:
  workflow_dispatch:

permissions: {}

jobs:
  baseline:
    name: baseline-all-tests
    runs-on: ubuntu-24.04
    timeout-minutes: 10
    steps:
      - name: Checkout exact event revision
        uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1

      - name: Set up Python
        uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
        with:
          python-version: '3.13'
          cache: ''

      - name: Record runner and source identity
        shell: bash
        run: |
          set -euo pipefail
          python --version
          printf 'run_id=%s\nattempt=%s\nsha=%s\nref=%s\nrunner=%s\n' \
            "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA" "$GITHUB_REF" "$RUNNER_NAME"

      - name: Install locked dependency without cache
        shell: bash
        run: |
          set -euo pipefail
          python -m pip install --disable-pip-version-check -r requirements.lock

      - name: Run the full synthetic suite sequentially
        shell: bash
        run: |
          set -euo pipefail
          mkdir -p evidence
          /usr/bin/time -f 'elapsed=%e user=%U sys=%S maxrss_kb=%M' \
            -o evidence/time.txt \
            python scripts/shard.py unit-a unit-b integration-a integration-b \
            | tee evidence/shards.jsonl
          sha256sum requirements.lock scripts/shard.py > evidence/input-sha256.txt
          du -sb evidence | tee evidence/bytes.txt

      - name: Upload tiny benchmark evidence
        if: ${{ always() }}
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
        with:
          name: perf-baseline-${{ github.run_id }}-${{ github.run_attempt }}
          path: evidence/
          retention-days: 1
          if-no-files-found: error

6. Optimized workflow

Create .github/workflows/perf-checkpoint.yml. The job graph changes, but the required test set does not. max-parallel: 2 bounds demand and the final performance-lab-required job fails unless every matrix cell succeeds.

name: Performance Checkpoint — Optimized
run-name: "perf-checkpoint / ${{ github.sha }}"

on:
  workflow_dispatch:

permissions: {}

jobs:
  test:
    name: test-${{ matrix.shard }}
    runs-on: ubuntu-24.04
    timeout-minutes: 10
    strategy:
      fail-fast: false
      max-parallel: 2
      matrix:
        shard: [unit-a, unit-b, integration-a, integration-b]
    steps:
      - name: Checkout exact event revision
        uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1

      - name: Set up Python
        uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
        with:
          python-version: '3.13'
          cache: ''

      - name: Discover pip cache path
        id: pip-cache
        shell: bash
        run: |
          set -euo pipefail
          echo "dir=$(python -m pip cache dir)" >> "$GITHUB_OUTPUT"

      - name: Restore dependency-download cache
        id: cache
        uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
        with:
          path: ${{ steps.pip-cache.outputs.dir }}
          key: pip-${{ runner.os }}-${{ runner.arch }}-py313-${{ hashFiles('requirements.lock') }}
          restore-keys: |
            pip-${{ runner.os }}-${{ runner.arch }}-py313-

      - name: Install locked dependency even on a cache hit
        shell: bash
        run: |
          set -euo pipefail
          python -m pip install --disable-pip-version-check -r requirements.lock

      - name: Run one independent shard
        shell: bash
        env:
          SHARD: ${{ matrix.shard }}
          CACHE_HIT: ${{ steps.cache.outputs.cache-hit }}
        run: |
          set -euo pipefail
          mkdir -p evidence
          /usr/bin/time -f 'elapsed=%e user=%U sys=%S maxrss_kb=%M' \
            -o "evidence/${SHARD}-time.txt" \
            python scripts/shard.py "$SHARD" \
            | tee "evidence/${SHARD}.jsonl"
          printf 'cache_hit=%s\nrun_id=%s\nattempt=%s\nsha=%s\n' \
            "$CACHE_HIT" "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$GITHUB_SHA" \
            > "evidence/${SHARD}-state.txt"

      - name: Upload shard evidence
        if: ${{ always() }}
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
        with:
          name: perf-${{ matrix.shard }}-${{ github.run_id }}-${{ github.run_attempt }}
          path: evidence/
          retention-days: 1
          if-no-files-found: error

  required:
    name: performance-lab-required
    if: ${{ always() }}
    needs: [test]
    runs-on: ubuntu-24.04
    steps:
      - name: Preserve the same correctness gate
        shell: bash
        env:
          RESULT: ${{ needs.test.result }}
        run: |
          set -euo pipefail
          test "$RESULT" = success
          echo "All four required shards completed successfully."

7. Execute a controlled measurement sequence

  1. Publish one commit containing the workload and both workflow files. Record its SHA.
  2. Dispatch the baseline three times; record every run ID/attempt and wait for terminal success.
  3. Dispatch the optimized workflow once with a cold/unknown cache state; record each matrix cell’s cache-hit output.
  4. Dispatch the optimized workflow at least two more times without changing the lockfile or Python/runtime inputs.
  5. Do not delete caches or artifacts until the comparison is complete.
  6. If queue conditions are visibly abnormal, keep the run as evidence and add more samples rather than silently discarding it.
for i in 1 2 3; do gh workflow run perf-baseline.yml --ref main; sleep 3; done
for i in 1 2 3; do gh workflow run perf-checkpoint.yml --ref main; sleep 3; done

gh run list --limit 20   --json databaseId,workflowName,attempt,headSha,status,conclusion,createdAt,updatedAt,url

8. Collect exact run/job/cache/artifact evidence

GH_REPO=OWNER/gha-performance-checkpoint
RUN_ID=123456789

gh run view "$RUN_ID" --json databaseId,attempt,event,headSha,status,conclusion,createdAt,updatedAt,url

gh api \
  -H 'Accept: application/vnd.github+json' \
  -H 'X-GitHub-Api-Version: 2026-03-10' \
  "repos/$GH_REPO/actions/runs/$RUN_ID/jobs?per_page=100" \
  > "checkpoint-evidence/run-$RUN_ID-jobs.json"

gh api \
  -H 'X-GitHub-Api-Version: 2026-03-10' \
  "repos/$GH_REPO/actions/runs/$RUN_ID/artifacts" \
  > "checkpoint-evidence/run-$RUN_ID-artifacts.json"

gh cache list --limit 100 > checkpoint-evidence/cache-list.txt

For queue evidence, capture the repository Actions Performance Metrics view or its exported CSV if available. For an individual root job, you may also record workflow created_at and job started_at as a dispatch-to-start signal, but label it exactly that. Do not claim it is an authoritative queue duration for dependent jobs.

9. Calculate processing and modeled cost

Use the Lesson 2 scripts/summarize-jobs.py script for every sampled run. Then calculate median workflow wall time and median processing seconds for each profile. If a public lab has zero standard-runner charge, the hypothetical private Linux overage column remains a model, not a bill.

#!/usr/bin/env python3
from __future__ import annotations
import json, math, sys
from datetime import datetime, timezone

def dt(s):
    return datetime.fromisoformat(s.replace("Z", "+00:00"))

data=json.load(sys.stdin)
jobs=[j for j in data["jobs"] if j.get("started_at") and j.get("completed_at")]
processing=0.0
rounded_minutes=0
for j in jobs:
    seconds=(dt(j["completed_at"])-dt(j["started_at"])).total_seconds()
    processing += seconds
    rounded_minutes += math.ceil(seconds/60)
    print(f"{j['name']}\t{seconds:.1f}s\t{math.ceil(seconds/60)} modeled billable minute(s)")
print(f"processing_seconds={processing:.1f}")
print(f"rounded_job_minutes={rounded_minutes}")
print(f"hypothetical_private_linux_overage_usd={rounded_minutes*0.006:.4f}")

Record both latency and processing. If the optimized median wall time drops from 25 seconds to 15 seconds while processing rises from 25 to 35 seconds, the correct conclusion is “faster feedback at higher runner consumption,” not simply “40% cheaper/faster.”

10. Reject one unsafe optimization explicitly

A teammate proposes removing both integration shards on ordinary pull requests. Do not run this as the accepted optimization. Record it in rejected-optimization.md with the causal reason: the job graph would be faster because the required workload is smaller, so it violates the equivalence contract and creates a quality risk.

# REJECTED PROPOSAL — changes correctness coverage.
matrix:
  shard: [unit-a, unit-b]  # integration-a and integration-b disappeared

A safe alternative is to keep all four required shards and optimize fixture startup, split independent integration work with bounded concurrency, or move additional non-required exhaustive coverage to scheduled workflows after explicit policy review.

11. Required before/after comparison

Evidence Baseline Optimized cold Optimized warm Pass criterion
Source SHA record same same Exact match.
Required shards 4 4 4 No coverage loss.
Workflow wall time median of 3 record median warm samples Interpret with queue context.
Processing seconds sum jobs sum jobs sum jobs Explain increases/decreases.
Cache hit state N/A record each cell record each cell Miss and hit both correct.
Matrix / max-parallel 1 sequential job 4 / 2 4 / 2 Bounded concurrency.
Artifact bytes record record record Tiny bounded evidence.
Modeled private cost rounded × $0.006 rounded × $0.006 rounded × $0.006 Clearly marked hypothetical.

12. Required evidence packet

  • checkpoint-predictions.md with verified/disproved outcomes.
  • Exact repository visibility, source SHA and both workflow file SHAs/revisions.
  • Run ID/attempt/event/source SHA/status/conclusion/URL for every sampled run.
  • Job IDs/names/start/completion timestamps, runner labels/name and step timing for representative runs.
  • Cache action SHA, primary key components, exact/fallback/miss state and cache metadata/size where available.
  • Artifact IDs/names/sizes/digests or metadata for the tiny benchmark evidence and its 1-day retention assumption.
  • Matrix cardinality and max-parallel; final required-check result proving all four shards executed.
  • Median wall-clock and processing-time comparison plus the hypothetical private billing calculation and timestamped $0.006/min assumption.
  • rejected-optimization.md showing why reduced test coverage was rejected.
  • Assumptions/limitations: public lab hardware differs from private standard Linux; queue is noisy; synthetic sleeps do not model CPU scaling; billing and limits can change.

13. Rollback and cleanup

  • If optimized runs lose required coverage, restore the baseline graph immediately; performance evidence is invalid until equivalence returns.
  • If the cache causes correctness anomalies, preserve cache metadata, bypass/re-key it, and rerun cold before deleting exact disposable cache entries.
  • If matrix processing cost rises without an approved latency benefit, reduce granularity or restore sequential execution.
  • Delete benchmark artifacts/caches/repository only after the evidence packet is reviewed; no production resource should exist.
  • Do not interpret cleanup as rollback of an external system—this checkpoint intentionally has no deployment side effect.

14. What Chapter 32 adds — and the bridge to Chapter 33

Chapter 32 adds a measured optimization operating model: latency and usage are separate, queueing belongs to a capacity layer, caches are optional, matrix parallelism is bounded, pricing/limits are dated assumptions, and correctness/evidence define the acceptance boundary. Chapter 33 uses this foundation for failure recovery: reruns, idempotency, rollback and incident-safe automation must recover service without duplicating side effects or erasing the evidence that explained the failure.

Next lesson

Failure Recovery, Reruns, Idempotency, Rollbacks, and Incident-Safe Automation: Core Concepts and Mental Model

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

What proves the checkpoint optimized the pipeline rather than reduced the workload?

Why are three baseline and three optimized samples better than one pair?

A warm cache hit makes install fail. What does that mean?

Why is the private-cost column hypothetical?

Which proposed optimization must be rejected even if it halves runtime?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.