Caches, Cache Keys, Fallback Keys, Distributed Cache, Dependency Reuse, and Cache Correctness: Guided Hands-On Workflow and Core Operations
This guided lab uses synthetic dependency data so you can measure a cold run, create a lockfile-aware cache, observe a warm hit, exercise ordered fallback keys, force a miss by changing the lockfile, and prove that correctness checks still work even when cache contents are stale.
Learning objectives
- Measure a synthetic dependency install with caching disabled and record a cold-run baseline.
- Add a lockfile-aware key using cache:key:files and a toolchain prefix, then observe a warm restore.
- Exercise ordered fallback keys and prove that restored fallback content is validated before use.
- Force a correct miss by changing the lockfile and compare job timing/evidence before and after.
- Complete a challenge that selects the correct cache/configuration/runner layer instead of copying a cleanup command.
1. Lab scenario: a dependency cache you can reason about
Create a disposable project named
glci-ch11-cache-lab with branch
glci/ch11-cache. The lab uses a synthetic dependency
installer rather than real credentials or paid package services. The
installer reads deps.lock, computes its SHA-256 digest,
and stores only synthetic package markers under
.deps-cache/. It intentionally sleeps on a cold install
so timing differences are visible.
2. Create deterministic synthetic inputs and the verifier
The lockfile is the correctness input. The cache directory is derived, disposable state. The installer never accepts a cache solely because files exist; it compares the stored lock digest with the current lock digest before reusing it.
# deps.lock
alpha==1.0.0
beta==2.0.0
# tools/simulate_deps.py
from pathlib import Path
import hashlib, json, sys, time
lock = Path("deps.lock").read_bytes()
lock_sha = hashlib.sha256(lock).hexdigest()
cache = Path(".deps-cache")
manifest = cache / "manifest.json"
valid = False
if manifest.exists():
try:
data = json.loads(manifest.read_text())
valid = data.get("lock_sha") == lock_sha
except Exception:
valid = False
start = time.monotonic()
if valid:
state = "verified-hit"
else:
state = "miss-or-stale"
cache.mkdir(parents=True, exist_ok=True)
time.sleep(2.0) # deterministic stand-in for dependency download
(cache / "alpha.pkg").write_text("alpha synthetic payload
")
(cache / "beta.pkg").write_text("beta synthetic payload
")
manifest.write_text(json.dumps({"lock_sha": lock_sha}, sort_keys=True))
elapsed = time.monotonic() - start
print(f"cache_state={state}")
print(f"lock_sha={lock_sha}")
print(f"elapsed_seconds={elapsed:.3f}")
if json.loads(manifest.read_text())["lock_sha"] != lock_sha:
raise SystemExit("dependency cache verification failed")
3. Measure the no-cache baseline first
The first job disables any inherited cache with
cache: []. That creates a performance baseline and
proves the workload works without cache. This is an important
production property: if cache storage is unavailable, the job
becomes slower, not wrong.
stages: [baseline, warm, verify]
baseline_no_cache:
stage: baseline
image: python:3.11.9-slim-bookworm
cache: []
script:
- rm -rf .deps-cache
- python tools/simulate_deps.py
- test -f .deps-cache/manifest.json
Record the job ID, source SHA, runner ID/version,
deps.lock digest,
cache_state=miss-or-stale, and elapsed time. Do not
compare only the total pipeline duration because runner queue time
is a different state.
4. Add a lockfile-aware primary cache and make one writer
The writer uses cache:key:files so the key changes when
deps.lock changes. A fixed prefix records the runtime
family used by the lab. It uses pull-push because this
job is responsible for warming/updating the cache.
warm_dependencies:
stage: warm
image: python:3.11.9-slim-bookworm
cache:
key:
files:
- deps.lock
prefix: deps-py311-linux-amd64
fallback_keys:
- deps-py311-main
- deps-py311-seed
paths:
- .deps-cache/
policy: pull-push
script:
- python tools/simulate_deps.py
- sha256sum deps.lock
- cat .deps-cache/manifest.json
On the first run, expect a primary miss and likely fallback misses. The script recreates verified synthetic dependencies, then Runner archives the cache after the job according to policy. On a later pipeline with the same lockfile/key, Runner can restore it before the script starts.
5. Add a read-only cache consumer
Most consumers should not race to mutate shared cache state. The
verification job uses the same key but policy: pull. It
restores cache if available, runs the independent manifest/lock
verification, and never uploads changes.
verify_cached_dependencies:
stage: verify
image: python:3.11.9-slim-bookworm
cache:
key:
files:
- deps.lock
prefix: deps-py311-linux-amd64
fallback_keys:
- deps-py311-main
- deps-py311-seed
paths:
- .deps-cache/
policy: pull
script:
- python tools/simulate_deps.py
- python - <<'PY'
import hashlib, json
lock = hashlib.sha256(open("deps.lock","rb").read()).hexdigest()
data = json.load(open(".deps-cache/manifest.json"))
assert data["lock_sha"] == lock
print("dependency_integrity=verified")
PY
Run the pipeline a second time without changing
deps.lock. A warm restore should reduce the script’s
synthetic install time. The job remains correct because the manifest
check would reject incompatible content even if a fallback archive
were restored.
6. Exercise a fallback without weakening validation
To observe fallback behavior safely, create a branch whose exact primary key has not yet been populated but whose fallback key exists. The trace should show the ordered lookup attempts. The restored fallback may contain useful files, but the script decides whether they match the current lockfile. If not, it rebuilds before the writer uploads the branch’s primary key.
7. Force a correct miss by changing a correctness input
Edit deps.lock to add gamma==3.0.0,
commit, and run a new pipeline. Because the content hash changes,
the primary cache:key:files value changes. That is a
correct invalidation. A fallback may still be restored, but
simulate_deps.py detects the old manifest and performs
the slower rebuild.
printf 'gamma==3.0.0
' >> deps.lock
git add deps.lock
git commit -m "lab: change dependency input"
git push
Compare the two pipeline/job IDs, exact SHAs, lock digests, trace
lookup results, script cache_state, elapsed time, and
final validation result. This is a stronger learning artifact than
simply seeing “Restoring cache… Successfully extracted cache.”
8. Free/local simulation path when no GitLab Runner is available
A local simulation cannot reproduce GitLab’s cache transport or
protected-key suffixes, but it can faithfully teach key derivation
and stale-cache rejection. Use the lock digest plus toolchain prefix
as a directory name, copy a previously created
.deps-cache into that directory, and run the verifier.
Change the lockfile and confirm the script rejects the old manifest.
LOCK_SHA=$(sha256sum deps.lock | awk '{print $1}')
KEY="deps-py311-linux-amd64-$LOCK_SHA"
mkdir -p ".local-cache/$KEY"
cp -a .deps-cache/. ".local-cache/$KEY/"
printf 'local_key=%s
' "$KEY"
python tools/simulate_deps.py
Document the limitation: this demonstrates correctness logic, not GitLab Runner cache upload/download, object-storage credentials, or fleet sharing.
9. Inspect runner/distributed-cache behavior without mutating it
If you administer the lab runner, record whether cache is local or
configured under Runner’s [runners.cache] section. Do
not publish credentials. A distributed backend may be
S3/S3-compatible, GCS, or Azure Blob. Record only non-secret fields
such as backend type, path prefix, shared flag, runner version,
executor, and whether the job trace shows a remote download/upload.
| Evidence | Example safe value | Why it matters |
|---|---|---|
| Runner | ID 42 / Docker / version X.Y.Z | Correlates cache behavior with execution infrastructure |
| Backend type | s3 / gcs / azure / local | Explains cross-runner availability |
| Key input digest | SHA-256 of deps.lock | Proves invalidation input |
| Primary/fallback result | primary hit / fallback hit / miss | Explains restored state |
| Elapsed dependency step | 0.03s vs 2.00s synthetic | Quantifies benefit independently from queue time |
| Integrity result | verified | Proves cache did not replace correctness checks |
10. Challenge: choose the layer, not the command
A feature branch is slow after the organization adds a second standalone runner. The pipeline YAML and lockfile key are unchanged, and jobs alternate between runners. Which layer should you investigate first: application code, lockfile, artifact transfer, runner/cache topology, or protected variable scope? Explain which evidence would distinguish “correct cache miss because another runner has no archive” from “wrong key” and from “distributed backend outage.”
A good answer starts with runner ID/executor, cache trace, backend configuration state, and key input digest. It does not start by clearing all caches.
Knowledge check
Why does baseline_no_cache matter even after the cache works?
It proves the job has a correct cache-independent path and provides a timing baseline. Cache loss should degrade performance, not correctness.
Why is policy: pull useful for test jobs?
It lets them benefit from shared cache state without racing to overwrite it after every parallel consumer run.
What should happen after deps.lock changes?
The lockfile-derived key should change, causing a primary miss. Any fallback content must be revalidated against the new lock before reuse.
Why can two standalone runners produce inconsistent hit rates?
Without distributed/shared cache, each runner may have different local archives even when the pipeline key is identical.
What is the safe response to a fallback hit with a mismatched manifest?
Treat the restored data as stale, rebuild/refresh according to the current lockfile, and write only the intended new key if the job is an authorized cache writer.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. Cache behavior is version-sensitive at both GitLab and GitLab Runner layers. Re-check the deployed GitLab/Runner versions for Self-Managed or Dedicated installations, especially when runner topology, object storage, protected-cache behavior, or cache archive implementation differs from GitLab.com.
- Caching in GitLab CI/CD — cache-versus-artifact boundary, fallback keys, protected-cache separation, availability, storage, clearing, and troubleshooting.
-
CI/CD YAML syntax reference
— authoritative
cache,cache:key,cache:key:files,cache:key:files_commits,cache:key:prefix,cache:fallback_keys,cache:policy,cache:when, andcache:unprotectsemantics. - CI/CD caching examples — dependency-manager patterns and lockfile-aware examples.
- GitLab Runner advanced configuration — distributed cache backend configuration, sharing, paths, and archive limits.
- Speed up job execution — distributed cache backends and transfer-performance considerations.
- Job artifacts — the retained-output mechanism that must not be confused with cache.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.