Chapter 11Lesson 05~190 minutes

Checkpoint Lab — Caches, Cache Keys, Fallback Keys, Distributed Cache, Dependency Reuse, and Cache Correctness

The checkpoint lab designs a safe dependency cache, proves invalidation after a controlled input change, injects a stale/poisoned cache, and recovers by changing only the affected key/scope. The evidence packet records source, job, runner, key, hit/miss, timing, and independent dependency verification.

CacheCache keysCorrectnessGitLab RunnerDependency reuse

Learning objectives

  • Design a lockfile- and toolchain-aware dependency cache with bounded fallback and safe update ownership.
  • Predict and verify cache-key, hit/miss, timing, runner/backend, and dependency-integrity state across multiple runs.
  • Inject and diagnose one stale/poisoned cache without clearing unrelated project or runner caches.
  • Produce an evidence packet that proves dependency correctness independently from cache availability.
  • Document cleanup/rollback and the controls required before using the same pattern in a production runner fleet.

1. Checkpoint scenario and success criteria

You will create one disposable branch, run a cold pipeline, run a warm pipeline, change the lockfile to force correct invalidation, then inject a stale synthetic cache and prove that validation repairs it without clearing unrelated keys. The lab never uses real credentials or production runners.

Success means: the cache improves the synthetic dependency step when compatible; an input change creates a new primary key; stale restored content is detected independently; unrelated cache directories/keys remain untouched; and the evidence packet can explain every state transition.

2. Preflight: record assumptions before mutating anything

  • Disposable project: glci-ch11-cache-lab.
  • Branch: glci/ch11-cache-checkpoint.
  • GitLab tier: mandatory path uses Free-compatible cache syntax.
  • Runner: any trusted disposable Shell or Docker runner capable of Python 3.11; do not register a new privileged runner solely for this lab.
  • Image assumption for Docker path: python:3.11.9-slim-bookworm (pin a digest in stricter production environments).
  • No cloud/object-store credentials are required. Distributed-cache behavior may be observed on an existing authorized lab runner or simulated conceptually.

Before the first run, record git rev-parse HEAD, sha256sum deps.lock, the selected runner/executor/version, and whether protected/non-protected cache separation is enabled. Do not record any runner authentication token or backend secret.

3. Predict four state changes before running

  1. Cold run: primary key misses; dependency step reports miss-or-stale; writer creates cache after successful verification.
  2. Warm run: same source-input digest selects the same primary key; restore occurs; dependency step reports verified-hit; elapsed synthetic install time drops.
  3. Lockfile change: cache:key:files produces a different primary key; old data cannot silently satisfy the new validation contract.
  4. Injected stale content: even if files are present under a candidate cache location, manifest-versus-lock verification rejects them and rebuilds before any authorized writer publishes the corrected key.

Write these predictions into evidence/predictions.txt before the pipeline. Prediction-first work prevents “I expected that” hindsight.

4. Checkpoint pipeline

stages: [prepare, verify]

default:
  image: python:3.11.9-slim-bookworm

.cache-contract: &cache_contract
  key:
    files:
      - deps.lock
    prefix: deps-py311-linux-amd64
  fallback_keys:
    - deps-py311-main
    - deps-py311-seed
  paths:
    - .deps-cache/

prepare_dependencies:
  stage: prepare
  cache:
    <<: *cache_contract
    policy: pull-push
  script:
    - mkdir -p evidence
    - printf 'pipeline=%s\njob=%s\nsha=%s\nsource=%s\nrunner=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" "$CI_PIPELINE_SOURCE" "${CI_RUNNER_ID:-unknown}" > evidence/identity.txt
    - sha256sum deps.lock | tee evidence/deps-lock.sha256
    - python --version | tee evidence/python-version.txt
    - python tools/simulate_deps.py | tee evidence/cache-observation.txt
  artifacts:
    when: always
    expire_in: 7 days
    paths:
      - evidence/

verify_dependencies:
  stage: verify
  cache:
    <<: *cache_contract
    policy: pull
  script:
    - mkdir -p evidence
    - python tools/simulate_deps.py | tee evidence/consumer-cache-observation.txt
    - python - <<'PY'
      import hashlib, json
      lock = hashlib.sha256(open("deps.lock","rb").read()).hexdigest()
      manifest = json.load(open(".deps-cache/manifest.json"))
      assert manifest["lock_sha"] == lock
      print("dependency_integrity=verified")
      PY
  artifacts:
    when: always
    expire_in: 7 days
    paths:
      - evidence/

The artifacts contain only non-secret evidence. The cache contains reconstructible synthetic dependency markers. The two mechanisms intentionally have different purposes.

5. Run 1 and Run 2: prove cold then warm behavior

Commit the lab and trigger the first pipeline. Preserve its pipeline/job IDs and traces. Then trigger a second pipeline at the same source state (manual/web is acceptable if your workflow:rules permit it). Compare:

Field Run 1 Run 2 Interpretation
Source SHA same or explicitly recorded same or explicitly recorded Avoids comparing different source states
deps.lock SHA-256 same same Same correctness input
Primary key input same same Expected same key
Runner/backend recorded recorded Explains local/distributed availability
Trace result likely miss expected hit if available Performance state only
simulate_deps state miss-or-stale verified-hit when compatible Independent validation
Dependency-step time ~2s synthetic near-zero synthetic Measured benefit

6. Controlled invalidation: change deps.lock

Add gamma==3.0.0, commit, and trigger a third pipeline. Predict a new primary key. Preserve the old pipeline evidence; do not clear it. The first job under the new key may restore an ordered fallback, but the manifest validation must detect the mismatch and rebuild.

printf 'gamma==3.0.0
' >> deps.lock
git add deps.lock
git commit -m "lab: invalidate dependency cache correctly"
git push

Verify that the new pipeline’s lock digest differs from Run 1/2 and that dependency integrity ends as verified. A slower first run after a lock change is correct behavior.

7. Inject one stale/poisoned cache safely

Do not poison a real shared runner cache. Use the local simulation directory or a dedicated disposable lab key. Copy an old .deps-cache/manifest.json and package markers into the location used for the current lockfile, then execute the verifier. It should report miss-or-stale, rebuild, and rewrite only the current lab directory.

mkdir -p .lab-poison/current
cp -a evidence-old-cache/. .lab-poison/current/ 2>/dev/null || true
# Faithful local check: place stale data into the working cache only for this lab.
rm -rf .deps-cache
cp -a .lab-poison/current .deps-cache
python tools/simulate_deps.py
python - <<'PY'
import hashlib, json
lock = hashlib.sha256(open("deps.lock","rb").read()).hexdigest()
manifest = json.load(open(".deps-cache/manifest.json"))
assert manifest["lock_sha"] == lock
print("recovery=verified")
PY
Guardrail: never delete ~/.cache, runner cache roots, object-storage prefixes, or project-wide cache namespaces as part of this exercise. Recovery targets only the disposable working directory/key created by the lab.

8. Required evidence packet

Create a small evidence packet that allows another engineer to distinguish correctness from performance:

  • pipeline ID/source/ref/SHA for each run;
  • job IDs/names and runner ID/version/executor;
  • deps.lock SHA-256 and Python/toolchain version;
  • compiled primary key inputs and configured fallback keys;
  • protected/non-protected cache scope assumption;
  • trace classification: primary hit, fallback hit, miss, upload, or backend failure;
  • cache archive size/transfer timing if available without secrets;
  • simulate_deps.py state and independent dependency_integrity=verified result;
  • an assumptions/limitations note stating whether backend behavior was real Runner cache or local simulation.

9. Cleanup and rollback

Delete only the disposable branch/project if you created it solely for the lab, and remove the local synthetic directories .deps-cache/, .local-cache/, and .lab-poison/. If you intentionally created a dedicated cache key on a shared lab runner, leave unrelated keys untouched. If you use GitLab’s “Clear runner caches” UI in a disposable project, document that the operation increments the internal cache index rather than proving old backend objects were physically deleted.

rm -rf .deps-cache .local-cache .lab-poison
git switch main
git branch -D glci/ch11-cache-checkpoint 2>/dev/null || true
git push origin --delete glci/ch11-cache-checkpoint 2>/dev/null || true

10. What this adds to a production operating model

Chapter 11 adds a disciplined rule: cache is mutable, optional, runner-managed performance state. Production pipelines should derive keys from compatibility inputs, minimize writers, preserve trust boundaries, measure transfer economics, and verify dependencies independently. Build/release evidence stays in artifacts/packages/registries with source/digest identity.

Chapter 12 will build on this by moving from cache/dataflow concerns into reusable CI configuration patterns, where the same principle applies: reuse should be explicit, versioned, reviewable, and bounded by a clear contract.

Knowledge check

After the lockfile changes, why is a cache miss a success rather than a failure?

A fallback cache restores quickly but its manifest digest differs from deps.lock. What should the job do?

Which evidence proves the checkpoint succeeded independently of cache transport?

Why should unrelated caches remain untouched during recovery?

What design change belongs in Chapter 12 rather than this chapter?

Next lesson

Chapter 12 — reusable CI configuration

Carry the same explicit-contract mindset into YAML reuse, includes/templates/components, and versioned configuration interfaces.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-11. Cache behavior is version-sensitive at both GitLab and GitLab Runner layers. Re-check the deployed GitLab/Runner versions for Self-Managed or Dedicated installations, especially when runner topology, object storage, protected-cache behavior, or cache archive implementation differs from GitLab.com.

  • Caching in GitLab CI/CD — cache-versus-artifact boundary, fallback keys, protected-cache separation, availability, storage, clearing, and troubleshooting.
  • CI/CD YAML syntax reference — authoritative cache, cache:key, cache:key:files, cache:key:files_commits, cache:key:prefix, cache:fallback_keys, cache:policy, cache:when, and cache:unprotect semantics.
  • CI/CD caching examples — dependency-manager patterns and lockfile-aware examples.
  • GitLab Runner advanced configuration — distributed cache backend configuration, sharing, paths, and archive limits.
  • Speed up job execution — distributed cache backends and transfer-performance considerations.
  • Job artifacts — the retained-output mechanism that must not be confused with cache.
Current behavior used by this chapter: GitLab documents a maximum of four caches per job and up to five per-cache fallback keys. Caches are restored before artifacts. Cache availability is not guaranteed. Protected and non-protected cache namespaces are separated by default, and current documentation records role/ref-sensitive protected-key suffix behavior introduced during the GitLab 18.x series. Runner 18.1 also changed cache archiving so symlinks are no longer followed in relevant edge cases.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.