Checkpoint Lab — Caches, Cache Keys, Fallback Keys, Distributed Cache, Dependency Reuse, and Cache Correctness
The checkpoint lab designs a safe dependency cache, proves invalidation after a controlled input change, injects a stale/poisoned cache, and recovers by changing only the affected key/scope. The evidence packet records source, job, runner, key, hit/miss, timing, and independent dependency verification.
Learning objectives
- Design a lockfile- and toolchain-aware dependency cache with bounded fallback and safe update ownership.
- Predict and verify cache-key, hit/miss, timing, runner/backend, and dependency-integrity state across multiple runs.
- Inject and diagnose one stale/poisoned cache without clearing unrelated project or runner caches.
- Produce an evidence packet that proves dependency correctness independently from cache availability.
- Document cleanup/rollback and the controls required before using the same pattern in a production runner fleet.
1. Checkpoint scenario and success criteria
You will create one disposable branch, run a cold pipeline, run a warm pipeline, change the lockfile to force correct invalidation, then inject a stale synthetic cache and prove that validation repairs it without clearing unrelated keys. The lab never uses real credentials or production runners.
Success means: the cache improves the synthetic dependency step when compatible; an input change creates a new primary key; stale restored content is detected independently; unrelated cache directories/keys remain untouched; and the evidence packet can explain every state transition.
2. Preflight: record assumptions before mutating anything
- Disposable project:
glci-ch11-cache-lab. - Branch:
glci/ch11-cache-checkpoint. - GitLab tier: mandatory path uses Free-compatible cache syntax.
- Runner: any trusted disposable Shell or Docker runner capable of Python 3.11; do not register a new privileged runner solely for this lab.
-
Image assumption for Docker path:
python:3.11.9-slim-bookworm(pin a digest in stricter production environments). - No cloud/object-store credentials are required. Distributed-cache behavior may be observed on an existing authorized lab runner or simulated conceptually.
Before the first run, record git rev-parse HEAD,
sha256sum deps.lock, the selected
runner/executor/version, and whether protected/non-protected cache
separation is enabled. Do not record any runner authentication token
or backend secret.
3. Predict four state changes before running
-
Cold run: primary key misses; dependency step
reports
miss-or-stale; writer creates cache after successful verification. -
Warm run: same source-input digest selects the
same primary key; restore occurs; dependency step reports
verified-hit; elapsed synthetic install time drops. -
Lockfile change:
cache:key:filesproduces a different primary key; old data cannot silently satisfy the new validation contract. - Injected stale content: even if files are present under a candidate cache location, manifest-versus-lock verification rejects them and rebuilds before any authorized writer publishes the corrected key.
Write these predictions into
evidence/predictions.txt before the pipeline.
Prediction-first work prevents “I expected that” hindsight.
4. Checkpoint pipeline
stages: [prepare, verify]
default:
image: python:3.11.9-slim-bookworm
.cache-contract: &cache_contract
key:
files:
- deps.lock
prefix: deps-py311-linux-amd64
fallback_keys:
- deps-py311-main
- deps-py311-seed
paths:
- .deps-cache/
prepare_dependencies:
stage: prepare
cache:
<<: *cache_contract
policy: pull-push
script:
- mkdir -p evidence
- printf 'pipeline=%s\njob=%s\nsha=%s\nsource=%s\nrunner=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" "$CI_PIPELINE_SOURCE" "${CI_RUNNER_ID:-unknown}" > evidence/identity.txt
- sha256sum deps.lock | tee evidence/deps-lock.sha256
- python --version | tee evidence/python-version.txt
- python tools/simulate_deps.py | tee evidence/cache-observation.txt
artifacts:
when: always
expire_in: 7 days
paths:
- evidence/
verify_dependencies:
stage: verify
cache:
<<: *cache_contract
policy: pull
script:
- mkdir -p evidence
- python tools/simulate_deps.py | tee evidence/consumer-cache-observation.txt
- python - <<'PY'
import hashlib, json
lock = hashlib.sha256(open("deps.lock","rb").read()).hexdigest()
manifest = json.load(open(".deps-cache/manifest.json"))
assert manifest["lock_sha"] == lock
print("dependency_integrity=verified")
PY
artifacts:
when: always
expire_in: 7 days
paths:
- evidence/
The artifacts contain only non-secret evidence. The cache contains reconstructible synthetic dependency markers. The two mechanisms intentionally have different purposes.
5. Run 1 and Run 2: prove cold then warm behavior
Commit the lab and trigger the first pipeline. Preserve its
pipeline/job IDs and traces. Then trigger a second pipeline at the
same source state (manual/web is acceptable if your
workflow:rules permit it). Compare:
| Field | Run 1 | Run 2 | Interpretation |
|---|---|---|---|
| Source SHA | same or explicitly recorded | same or explicitly recorded | Avoids comparing different source states |
| deps.lock SHA-256 | same | same | Same correctness input |
| Primary key input | same | same | Expected same key |
| Runner/backend | recorded | recorded | Explains local/distributed availability |
| Trace result | likely miss | expected hit if available | Performance state only |
| simulate_deps state | miss-or-stale | verified-hit when compatible | Independent validation |
| Dependency-step time | ~2s synthetic | near-zero synthetic | Measured benefit |
6. Controlled invalidation: change deps.lock
Add gamma==3.0.0, commit, and trigger a third pipeline.
Predict a new primary key. Preserve the old pipeline evidence; do
not clear it. The first job under the new key may restore an ordered
fallback, but the manifest validation must detect the mismatch and
rebuild.
printf 'gamma==3.0.0
' >> deps.lock
git add deps.lock
git commit -m "lab: invalidate dependency cache correctly"
git push
Verify that the new pipeline’s lock digest differs from Run 1/2 and
that dependency integrity ends as verified. A slower
first run after a lock change is correct behavior.
7. Inject one stale/poisoned cache safely
Do not poison a real shared runner cache. Use the local simulation
directory or a dedicated disposable lab key. Copy an old
.deps-cache/manifest.json and package markers into the
location used for the current lockfile, then execute the verifier.
It should report miss-or-stale, rebuild, and rewrite
only the current lab directory.
mkdir -p .lab-poison/current
cp -a evidence-old-cache/. .lab-poison/current/ 2>/dev/null || true
# Faithful local check: place stale data into the working cache only for this lab.
rm -rf .deps-cache
cp -a .lab-poison/current .deps-cache
python tools/simulate_deps.py
python - <<'PY'
import hashlib, json
lock = hashlib.sha256(open("deps.lock","rb").read()).hexdigest()
manifest = json.load(open(".deps-cache/manifest.json"))
assert manifest["lock_sha"] == lock
print("recovery=verified")
PY
~/.cache,
runner cache roots, object-storage prefixes, or project-wide cache
namespaces as part of this exercise. Recovery targets only the
disposable working directory/key created by the lab.
8. Required evidence packet
Create a small evidence packet that allows another engineer to distinguish correctness from performance:
- pipeline ID/source/ref/SHA for each run;
- job IDs/names and runner ID/version/executor;
-
deps.lockSHA-256 and Python/toolchain version; - compiled primary key inputs and configured fallback keys;
- protected/non-protected cache scope assumption;
- trace classification: primary hit, fallback hit, miss, upload, or backend failure;
- cache archive size/transfer timing if available without secrets;
-
simulate_deps.pystate and independentdependency_integrity=verifiedresult; - an assumptions/limitations note stating whether backend behavior was real Runner cache or local simulation.
9. Cleanup and rollback
Delete only the disposable branch/project if you created it solely
for the lab, and remove the local synthetic directories
.deps-cache/, .local-cache/, and
.lab-poison/. If you intentionally created a dedicated
cache key on a shared lab runner, leave unrelated keys untouched. If
you use GitLab’s “Clear runner caches” UI in a disposable project,
document that the operation increments the internal cache index
rather than proving old backend objects were physically deleted.
rm -rf .deps-cache .local-cache .lab-poison
git switch main
git branch -D glci/ch11-cache-checkpoint 2>/dev/null || true
git push origin --delete glci/ch11-cache-checkpoint 2>/dev/null || true
10. What this adds to a production operating model
Chapter 11 adds a disciplined rule: cache is mutable, optional, runner-managed performance state. Production pipelines should derive keys from compatibility inputs, minimize writers, preserve trust boundaries, measure transfer economics, and verify dependencies independently. Build/release evidence stays in artifacts/packages/registries with source/digest identity.
Chapter 12 will build on this by moving from cache/dataflow concerns into reusable CI configuration patterns, where the same principle applies: reuse should be explicit, versioned, reviewable, and bounded by a clear contract.
Knowledge check
After the lockfile changes, why is a cache miss a success rather than a failure?
Because the old dependency state is no longer compatible. Correct invalidation sacrifices one warm run to preserve correctness.
A fallback cache restores quickly but its manifest digest differs from deps.lock. What should the job do?
Reject the stale content for correctness purposes, reconstruct/refresh dependencies from the current lockfile, verify them, and only then let an authorized writer publish the new primary cache.
Which evidence proves the checkpoint succeeded independently of cache transport?
Exact source SHA, lockfile digest, toolchain identity, validation output, tests, and retained artifacts/reports. A cache hit/miss is only performance evidence.
Why should unrelated caches remain untouched during recovery?
Broad clearing destroys warm state and evidence outside the fault boundary. The repair should target the exact key/scope shown to be stale or poisoned.
What design change belongs in Chapter 12 rather than this chapter?
Versioned reusable CI configuration/includes/components. Chapter 11 defines cache correctness; Chapter 12 will define how pipeline configuration itself is reused safely.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. Cache behavior is version-sensitive at both GitLab and GitLab Runner layers. Re-check the deployed GitLab/Runner versions for Self-Managed or Dedicated installations, especially when runner topology, object storage, protected-cache behavior, or cache archive implementation differs from GitLab.com.
- Caching in GitLab CI/CD — cache-versus-artifact boundary, fallback keys, protected-cache separation, availability, storage, clearing, and troubleshooting.
-
CI/CD YAML syntax reference
— authoritative
cache,cache:key,cache:key:files,cache:key:files_commits,cache:key:prefix,cache:fallback_keys,cache:policy,cache:when, andcache:unprotectsemantics. - CI/CD caching examples — dependency-manager patterns and lockfile-aware examples.
- GitLab Runner advanced configuration — distributed cache backend configuration, sharing, paths, and archive limits.
- Speed up job execution — distributed cache backends and transfer-performance considerations.
- Job artifacts — the retained-output mechanism that must not be confused with cache.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.