Checkpoint Lab — Large Repositories, Monorepos, Search, Actions Cost Controls, and Platform Performance
Prove path-aware optimization safely, inject one local routing failure, and produce a performance/cost evidence packet and operating policy for the production capstone.
Learning objectives
- Measure baseline clone/workflow/storage state and compare it with path-aware behavior.
- Prove a cross-cutting change still runs all required component checks.
- Inject and repair one routing defect locally without pushing unsafe configuration.
- Produce a performance/cost report and production operating policy that bridges into the capstone.
1. Checkpoint scenario and preflight
Reuse ch31-scale-lab from Lesson 2. If it no longer
exists, recreate the same disposable public repository. Mandatory
requirements are Git, GitHub CLI, Python 3, and permission to create
a public repository. Standard public GitHub-hosted runners are
sufficient; no payment method, organization, larger runner, LFS
object, PAT, or self-hosted runner is required.
libs/shared/version.txt, both API
and Web work should run; (3) breaking the router's shared dependency
locally should make its test fail before any unsafe
configuration reaches GitHub.
2. Capture baseline repository, clone, run, and storage evidence
OWNER="$(gh api -H 'X-GitHub-Api-Version: 2026-03-10' user --jq .login)"
FULL="$OWNER/ch31-scale-lab"
mkdir -p evidence
gh repo view "$FULL" --json nameWithOwner,diskUsage,defaultBranchRef > evidence/repository.json
gh workflow run ch31-baseline --repo "$FULL"
sleep 3
BASELINE_RUN="$(gh run list --repo "$FULL" --workflow ch31-baseline --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$BASELINE_RUN" --repo "$FULL" --exit-status
gh run view "$BASELINE_RUN" --repo "$FULL" --json databaseId,headSha,createdAt,updatedAt,jobs > evidence/baseline-run.json
gh cache list --repo "$FULL" --limit 100 --json id,key,sizeInBytes,createdAt,lastAccessedAt > evidence/caches.json
gh api -H 'X-GitHub-Api-Version: 2026-03-10' \
"repos/$FULL/actions/artifacts?per_page=100" > evidence/artifacts.json
gh api -H 'X-GitHub-Api-Version: 2026-03-10' rate_limit > evidence/rate-limit.json
Measure a clean clone separately from workflow time. Bash on Linux/macOS/Git Bash with GNU time can use:
rm -rf /tmp/ch31-clone
/usr/bin/time -p git clone "https://github.com/$FULL.git" /tmp/ch31-clone 2> evidence/clone-time.txt
cd /tmp/ch31-clone
git count-objects -vH | tee "$OLDPWD/evidence/git-objects.txt"
cd -
PowerShell alternative:
$sw = [System.Diagnostics.Stopwatch]::StartNew()
git clone "https://github.com/$env:FULL.git" "$env:TEMP\ch31-clone"
$sw.Stop()
$sw.Elapsed | Out-File evidence\clone-time.txt
Do not compare clone times across different networks/machines as if they were controlled experiments. The checkpoint establishes your own before/after record.
3. Prove the selective case
On the open api-only-change PR, add another API-only
edit if necessary and capture checks:
git switch api-only-change
printf '\n# checkpoint api-only\n' >> services/api/app.py
git add services/api/app.py
git commit -m 'Checkpoint API-only change'
git push
PR_NUM="$(gh pr view --repo "$FULL" --json number --jq .number)"
gh pr checks "$PR_NUM" --repo "$FULL" --watch
gh pr view "$PR_NUM" --repo "$FULL" --json number,headRefOid,baseRefOid,files,statusCheckRollup > evidence/api-only-pr.json
Verify prediction 1: the recorded check rollup must show the router/gate path while Web is skipped. If not, inspect router output and changed files before changing code.
4. Turn the same PR into a cross-cutting change
printf 'shared-v2\n' > libs/shared/version.txt
git add libs/shared/version.txt
git commit -m 'Change shared dependency'
git push
gh pr checks "$PR_NUM" --repo "$FULL" --watch
gh pr view "$PR_NUM" --repo "$FULL" --json number,headRefOid,baseRefOid,files,statusCheckRollup > evidence/shared-pr.json
Verify prediction 2: both component jobs should now run. This is the crucial counterexample that proves the optimization has not confused “path-local” with “dependency-local.”
5. Inject a routing defect locally—never push it
Copy the router outside the tracked tree, remove the shared-library prefix, and prove the test catches the defect:
cp tools/affected.py /tmp/affected-broken.py
python - <<'PY'
from pathlib import Path
p=Path('/tmp/affected-broken.py')
s=p.read_text()
s=s.replace('("libs/shared/", ".github/workflows/")', '(".github/workflows/",)')
p.write_text(s)
PY
python - <<'PY'
import importlib.util
spec=importlib.util.spec_from_file_location('broken','/tmp/affected-broken.py')
m=importlib.util.module_from_spec(spec); spec.loader.exec_module(m)
r=m.affected(['libs/shared/version.txt'])
print(r)
assert r['api'] and r['web'], 'EXPECTED FAILURE: shared change was not expanded'
PY
The assertion should fail. That is the intended failure. Now rerun the repository's real test:
python tests/test_affected.py
git status --short
Expected: repository test passes and no broken router is staged. This demonstrates a safe failure-injection discipline: test policy logic before pushing it.
6. Record search/API constraints as evidence, not assumptions
gh search code 'shared-v2' --repo "$FULL" --limit 20 --json path,repository,url > evidence/code-search.json || true
gh issue list --repo "$FULL" --search '"Performance evidence exercise"' --json number,title,state,url > evidence/issues.json
gh pr list --repo "$FULL" --state all --search '"API-only routing exercise"' --json number,title,state,url > evidence/pulls.json
gh api -H 'X-GitHub-Api-Version: 2026-03-10' rate_limit > evidence/rate-limit-final.json
Document whether code search had indexed the newest commit. A zero result is a limitation observation, not proof of absence.
7. Build a small platform performance/cost report
#!/usr/bin/env python3
import json
from pathlib import Path
from datetime import datetime
def load(name):
return json.loads(Path('evidence', name).read_text())
def seconds(a, b):
if not a or not b: return None
A=datetime.fromisoformat(a.replace('Z','+00:00'))
B=datetime.fromisoformat(b.replace('Z','+00:00'))
return round((B-A).total_seconds(), 2)
run=load('baseline-run.json')
caches=load('caches.json')
artifacts=load('artifacts.json').get('artifacts', [])
rates=load('rate-limit-final.json').get('resources', {})
jobs=[]
for j in run.get('jobs', []):
jobs.append({
'name': j.get('name'),
'seconds': seconds(j.get('startedAt'), j.get('completedAt')),
'conclusion': j.get('conclusion')
})
report={
'baseline_run_id': run.get('databaseId'),
'baseline_head_sha': run.get('headSha'),
'jobs': jobs,
'cache_bytes_visible': sum(x.get('sizeInBytes',0) or 0 for x in caches),
'artifact_bytes_visible': sum(x.get('size_in_bytes',0) or 0 for x in artifacts),
'rate_limit_snapshot': {k: rates.get(k) for k in ('core','search','code_search','graphql')},
'constraints': [
'Code Search is default-branch indexed discovery, not historical inventory.',
'Path-aware CI must model cross-cutting dependencies.',
'Cache/artifact retention is a separate storage lifecycle from runner time.'
],
'priorities': [
'Keep a stable always-running monorepo gate.',
'Measure clone/history before considering history surgery.',
'Track cache hit/cardinality and artifact consumers before expanding retention.',
'Cancel superseded PR work with scoped concurrency.'
]
}
Path('evidence/performance-cost-report.json').write_text(json.dumps(report, indent=2)+"\n")
print(json.dumps(report, indent=2))
Save as report.py outside the repository or under the
disposable lab, then run python report.py. The report
intentionally calls visible storage “visible bytes,” not “your
bill”: billing is time- and plan-dependent and may update
asynchronously.
8. Production performance and cost policy
| Control | Production rule | Evidence |
|---|---|---|
| Git history | No generated build outputs in normal Git; investigate large historical blobs quarterly. | Object inventory, clone/fetch timings. |
| Ownership | Directory/team CODEOWNERS; exceptions reviewed and kept small. | CODEOWNERS tests + representative PR review requests. |
| CI routing | Dependency-aware router; cross-cutting inputs documented/tested. | Router unit tests + changed-path outputs. |
| Required check | Require stable aggregate gate, not conditionally absent component workflows. | Ruleset/check name + PR rollup. |
| Runner | Standard runner by default; larger runner needs measured business case and budget owner. | Job duration, queue time, cost estimate. |
| Concurrency | Cancel superseded PR runs using PR-scoped key. | Run history/cancellations. |
| Cache | Cache only reproducible dependencies; monitor hit/cardinality/size. | Cache inventory and hit logs. |
| Artifacts | Upload only consumed evidence/output; explicit retention. | Artifact list, size, expiry, owner. |
| Search/API | Search for discovery; API/Git/audit for authoritative scope; respect rate limits. | Queries, refs, API headers/rate snapshots. |
| History rewrite | Exceptional coordinated operation after impact assessment; never a routine cleanup. | Approved runbook, backup/ref inventory, stakeholder sign-off. |
9. Cleanup and rollback
Close the disposable PR and archive the repository; do not delete it as part of the mandatory lab so the evidence remains reviewable:
gh pr close "$PR_NUM" --repo "$FULL" --comment 'Closing disposable Chapter 31 checkpoint.' || true
gh repo archive "$FULL" --yes
The tiny artifact expires after one day by workflow configuration. Cache entries are disposable and GitHub's normal eviction policy will remove unused entries; manual cache deletion is optional only after confirming no evidence consumer needs it. Archiving is reversible and safer than repository deletion.
10. Production operating model and bridge to Chapter 32
Chapter 31 adds a measurable scale plane to the operating model: Git history and binary lifecycle, dependency-aware CI routing, stable required checks, ownership at path scale, search/API scope, runner/concurrency behavior, and storage/cost evidence. Chapter 32 combines these controls with organization governance, supply-chain identity, security scanning, environments, APIs, and incident-ready evidence in the production capstone.
Knowledge check
What observation proves the API-only optimization is functioning?
The route/API/gate path succeeds while Web is intentionally skipped, and the PR check rollup records that state against the expected head SHA.
What observation proves the optimization still handles global dependencies?
After a shared-library change, both API and Web jobs run and the aggregate gate succeeds.
Why inject the broken router only in /tmp?
The goal is to prove policy tests detect a defect without pushing unsafe routing configuration into the hosted repository.
Why does the performance report call cache/artifact totals “visible bytes” rather than “cost”?
Billing depends on plan, retention over time, included allowances, and delayed usage reporting. A point-in-time inventory is an input to cost analysis, not the bill itself.
A required component workflow is permanently pending on docs-only PRs. What design error is likely?
The ruleset likely requires a workflow that can be skipped by event/path filtering. Require a stable always-running aggregate gate instead.
What does Chapter 31 add before the capstone?
A measured scale/cost operating layer: history and binary lifecycle, monorepo routing/ownership, stable gates, search/API scope, concurrency/runner choices, and cache/artifact retention evidence.
Further reading — current primary sources
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.