Chapter 11Lesson 02~175 minutes

Caches, Cache Keys, Fallback Keys, Distributed Cache, Dependency Reuse, and Cache Correctness: Guided Hands-On Workflow and Core Operations

This guided lab uses synthetic dependency data so you can measure a cold run, create a lockfile-aware cache, observe a warm hit, exercise ordered fallback keys, force a miss by changing the lockfile, and prove that correctness checks still work even when cache contents are stale.

CacheCache keysCorrectnessGitLab RunnerDependency reuse

Learning objectives

  • Measure a synthetic dependency install with caching disabled and record a cold-run baseline.
  • Add a lockfile-aware key using cache:key:files and a toolchain prefix, then observe a warm restore.
  • Exercise ordered fallback keys and prove that restored fallback content is validated before use.
  • Force a correct miss by changing the lockfile and compare job timing/evidence before and after.
  • Complete a challenge that selects the correct cache/configuration/runner layer instead of copying a cleanup command.

1. Lab scenario: a dependency cache you can reason about

Create a disposable project named glci-ch11-cache-lab with branch glci/ch11-cache. The lab uses a synthetic dependency installer rather than real credentials or paid package services. The installer reads deps.lock, computes its SHA-256 digest, and stores only synthetic package markers under .deps-cache/. It intentionally sleeps on a cold install so timing differences are visible.

Assumptions: current GitLab Free-compatible pipeline syntax; GitLab Runner with a Shell or Docker executor; Python 3.11+ in the job image or host. If no runner is available, the local simulation later in this lesson reproduces the cache-key and validation logic without pretending to reproduce GitLab’s backend transport.

2. Create deterministic synthetic inputs and the verifier

The lockfile is the correctness input. The cache directory is derived, disposable state. The installer never accepts a cache solely because files exist; it compares the stored lock digest with the current lock digest before reusing it.

# deps.lock
alpha==1.0.0
beta==2.0.0
# tools/simulate_deps.py
from pathlib import Path
import hashlib, json, sys, time

lock = Path("deps.lock").read_bytes()
lock_sha = hashlib.sha256(lock).hexdigest()
cache = Path(".deps-cache")
manifest = cache / "manifest.json"

valid = False
if manifest.exists():
    try:
        data = json.loads(manifest.read_text())
        valid = data.get("lock_sha") == lock_sha
    except Exception:
        valid = False

start = time.monotonic()
if valid:
    state = "verified-hit"
else:
    state = "miss-or-stale"
    cache.mkdir(parents=True, exist_ok=True)
    time.sleep(2.0)  # deterministic stand-in for dependency download
    (cache / "alpha.pkg").write_text("alpha synthetic payload
")
    (cache / "beta.pkg").write_text("beta synthetic payload
")
    manifest.write_text(json.dumps({"lock_sha": lock_sha}, sort_keys=True))

elapsed = time.monotonic() - start
print(f"cache_state={state}")
print(f"lock_sha={lock_sha}")
print(f"elapsed_seconds={elapsed:.3f}")
if json.loads(manifest.read_text())["lock_sha"] != lock_sha:
    raise SystemExit("dependency cache verification failed")

3. Measure the no-cache baseline first

The first job disables any inherited cache with cache: []. That creates a performance baseline and proves the workload works without cache. This is an important production property: if cache storage is unavailable, the job becomes slower, not wrong.

stages: [baseline, warm, verify]

baseline_no_cache:
  stage: baseline
  image: python:3.11.9-slim-bookworm
  cache: []
  script:
    - rm -rf .deps-cache
    - python tools/simulate_deps.py
    - test -f .deps-cache/manifest.json

Record the job ID, source SHA, runner ID/version, deps.lock digest, cache_state=miss-or-stale, and elapsed time. Do not compare only the total pipeline duration because runner queue time is a different state.

4. Add a lockfile-aware primary cache and make one writer

The writer uses cache:key:files so the key changes when deps.lock changes. A fixed prefix records the runtime family used by the lab. It uses pull-push because this job is responsible for warming/updating the cache.

warm_dependencies:
  stage: warm
  image: python:3.11.9-slim-bookworm
  cache:
    key:
      files:
        - deps.lock
      prefix: deps-py311-linux-amd64
    fallback_keys:
      - deps-py311-main
      - deps-py311-seed
    paths:
      - .deps-cache/
    policy: pull-push
  script:
    - python tools/simulate_deps.py
    - sha256sum deps.lock
    - cat .deps-cache/manifest.json

On the first run, expect a primary miss and likely fallback misses. The script recreates verified synthetic dependencies, then Runner archives the cache after the job according to policy. On a later pipeline with the same lockfile/key, Runner can restore it before the script starts.

5. Add a read-only cache consumer

Most consumers should not race to mutate shared cache state. The verification job uses the same key but policy: pull. It restores cache if available, runs the independent manifest/lock verification, and never uploads changes.

verify_cached_dependencies:
  stage: verify
  image: python:3.11.9-slim-bookworm
  cache:
    key:
      files:
        - deps.lock
      prefix: deps-py311-linux-amd64
    fallback_keys:
      - deps-py311-main
      - deps-py311-seed
    paths:
      - .deps-cache/
    policy: pull
  script:
    - python tools/simulate_deps.py
    - python - <<'PY'
      import hashlib, json
      lock = hashlib.sha256(open("deps.lock","rb").read()).hexdigest()
      data = json.load(open(".deps-cache/manifest.json"))
      assert data["lock_sha"] == lock
      print("dependency_integrity=verified")
      PY

Run the pipeline a second time without changing deps.lock. A warm restore should reduce the script’s synthetic install time. The job remains correct because the manifest check would reject incompatible content even if a fallback archive were restored.

6. Exercise a fallback without weakening validation

To observe fallback behavior safely, create a branch whose exact primary key has not yet been populated but whose fallback key exists. The trace should show the ordered lookup attempts. The restored fallback may contain useful files, but the script decides whether they match the current lockfile. If not, it rebuilds before the writer uploads the branch’s primary key.

Evidence rule: capture the trace lines that identify the attempted cache keys, but do not upload the entire cache archive as “proof.” Correctness evidence is the lock digest and validation result, not the cache contents.

7. Force a correct miss by changing a correctness input

Edit deps.lock to add gamma==3.0.0, commit, and run a new pipeline. Because the content hash changes, the primary cache:key:files value changes. That is a correct invalidation. A fallback may still be restored, but simulate_deps.py detects the old manifest and performs the slower rebuild.

printf 'gamma==3.0.0
' >> deps.lock
git add deps.lock
git commit -m "lab: change dependency input"
git push

Compare the two pipeline/job IDs, exact SHAs, lock digests, trace lookup results, script cache_state, elapsed time, and final validation result. This is a stronger learning artifact than simply seeing “Restoring cache… Successfully extracted cache.”

8. Free/local simulation path when no GitLab Runner is available

A local simulation cannot reproduce GitLab’s cache transport or protected-key suffixes, but it can faithfully teach key derivation and stale-cache rejection. Use the lock digest plus toolchain prefix as a directory name, copy a previously created .deps-cache into that directory, and run the verifier. Change the lockfile and confirm the script rejects the old manifest.

LOCK_SHA=$(sha256sum deps.lock | awk '{print $1}')
KEY="deps-py311-linux-amd64-$LOCK_SHA"
mkdir -p ".local-cache/$KEY"
cp -a .deps-cache/. ".local-cache/$KEY/"
printf 'local_key=%s
' "$KEY"
python tools/simulate_deps.py

Document the limitation: this demonstrates correctness logic, not GitLab Runner cache upload/download, object-storage credentials, or fleet sharing.

9. Inspect runner/distributed-cache behavior without mutating it

If you administer the lab runner, record whether cache is local or configured under Runner’s [runners.cache] section. Do not publish credentials. A distributed backend may be S3/S3-compatible, GCS, or Azure Blob. Record only non-secret fields such as backend type, path prefix, shared flag, runner version, executor, and whether the job trace shows a remote download/upload.

Evidence Example safe value Why it matters
Runner ID 42 / Docker / version X.Y.Z Correlates cache behavior with execution infrastructure
Backend type s3 / gcs / azure / local Explains cross-runner availability
Key input digest SHA-256 of deps.lock Proves invalidation input
Primary/fallback result primary hit / fallback hit / miss Explains restored state
Elapsed dependency step 0.03s vs 2.00s synthetic Quantifies benefit independently from queue time
Integrity result verified Proves cache did not replace correctness checks

10. Challenge: choose the layer, not the command

A feature branch is slow after the organization adds a second standalone runner. The pipeline YAML and lockfile key are unchanged, and jobs alternate between runners. Which layer should you investigate first: application code, lockfile, artifact transfer, runner/cache topology, or protected variable scope? Explain which evidence would distinguish “correct cache miss because another runner has no archive” from “wrong key” and from “distributed backend outage.”

A good answer starts with runner ID/executor, cache trace, backend configuration state, and key input digest. It does not start by clearing all caches.

Knowledge check

Why does baseline_no_cache matter even after the cache works?

Why is policy: pull useful for test jobs?

What should happen after deps.lock changes?

Why can two standalone runners produce inconsistent hit rates?

What is the safe response to a fallback hit with a mismatched manifest?

Next lesson

Configuration, design choices, and tradeoffs

Choose key granularity, writer/reader policy, protected scope, and local versus distributed cache from trust and measured cost.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-11. Cache behavior is version-sensitive at both GitLab and GitLab Runner layers. Re-check the deployed GitLab/Runner versions for Self-Managed or Dedicated installations, especially when runner topology, object storage, protected-cache behavior, or cache archive implementation differs from GitLab.com.

  • Caching in GitLab CI/CD — cache-versus-artifact boundary, fallback keys, protected-cache separation, availability, storage, clearing, and troubleshooting.
  • CI/CD YAML syntax reference — authoritative cache, cache:key, cache:key:files, cache:key:files_commits, cache:key:prefix, cache:fallback_keys, cache:policy, cache:when, and cache:unprotect semantics.
  • CI/CD caching examples — dependency-manager patterns and lockfile-aware examples.
  • GitLab Runner advanced configuration — distributed cache backend configuration, sharing, paths, and archive limits.
  • Speed up job execution — distributed cache backends and transfer-performance considerations.
  • Job artifacts — the retained-output mechanism that must not be confused with cache.
Current behavior used by this chapter: GitLab documents a maximum of four caches per job and up to five per-cache fallback keys. Caches are restored before artifacts. Cache availability is not guaranteed. Protected and non-protected cache namespaces are separated by default, and current documentation records role/ref-sensitive protected-key suffix behavior introduced during the GitLab 18.x series. Runner 18.1 also changed cache archiving so symlinks are no longer followed in relevant edge cases.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.