Chapter 11Lesson 04~155 minutes

Caches, Cache Keys, Fallback Keys, Distributed Cache, Dependency Reuse, and Cache Correctness: Diagnostics, Failure Modes, Security, and Performance

Cache failures often look like application or dependency failures even though the causal state lives in the cache key, runner/backend, or trust boundary. This lesson preserves first-failure evidence and diagnoses stale, poisoned, oversized, misplaced, and incorrectly trusted caches without deleting unrelated data.

CacheCache keysCorrectnessGitLab RunnerDependency reuse

Learning objectives

  • Diagnose a cache hit that hides stale dependencies without assuming the cache is correct.
  • Recognize key designs that omit lockfile/toolchain inputs and shared-key patterns that permit cache poisoning.
  • Measure when archive/transfer cost makes a cache slower than a cold dependency restore.
  • Separate cache failures from artifact, runner, network, and application failures using an evidence-first sequence.
  • Repair only the affected cache key/scope while preserving pipeline/job IDs and first-failure evidence.

1. Evidence-first diagnostic sequence

Preserve the original pipeline ID, job ID, source SHA, pipeline source, compiled cache configuration, runner identity/version/executor, lockfile/toolchain fingerprints, first cache-related trace lines, and the first dependency/test failure. Then classify the failure before changing anything:

  1. Did the pipeline/job exist with the expected rules and exact SHA?
  2. Which cache key and fallbacks were compiled?
  3. Which runner/backend handled the job?
  4. Was there a hit, fallback hit, miss, upload, or transfer error?
  5. Did restored dependency data pass independent lock/toolchain verification?
  6. Did the job fail in dependency restore, application tests, artifact handling, or a later external system?

Only after this chain should you change a key or clear a narrowly identified lab cache.

2. Failure mode: “cache hit” treated as correctness proof

A trace says the cache extracted successfully, but tests fail with an API that exists only in an older dependency version. The extraction is evidence of transport success, not dependency correctness. Compare the current lockfile digest with a manifest inside the restored cache or let the dependency manager perform its lock/checksum validation.

Least-destructive fix: correct the key/validation contract and populate a new key. Do not suppress the test and do not call the old cache “known good” merely because earlier pipelines were green.

3. Failure mode: the key ignores a correctness input

This intentionally broken configuration reuses one static key across lockfile changes and runtime upgrades:

default:
  image: python:3.12-slim
  cache:
    key: dependencies
    paths:
      - .deps-cache/
    policy: pull-push

test:
  script:
    - python tools/simulate_deps.py
    - pytest -q

The YAML is valid. The problem is semantic: the key omits both deps.lock content and runtime family. Repair it by deriving from the lockfile and adding a toolchain prefix. Preserve the failed job’s ID and trace before rerunning.

cache:
  key:
    files:
      - deps.lock
    prefix: deps-py312-linux-amd64
  paths:
    - .deps-cache/
  policy: pull

4. Failure mode: an untrusted branch poisons a shared cache

An unprotected branch can write a key later consumed by a protected release pipeline because cache separation was disabled for convenience. The issue is not “bad package manager luck”; it is a trust-boundary violation. Restore protected/non-protected separation, restrict cache writers, and validate dependencies before use.

Do not test this against an organizational shared runner or production dependency cache. Use the disposable Chapter 11 project with synthetic package markers. Evidence should record which branch/ref and actor wrote the lab key, never real credentials.

5. Failure mode: the cache costs more than it saves

A large archive can dominate job time through compression, upload, object-storage download, and extraction. Measure:

Measurement Why it matters
Cache archive size Predicts network/storage/compression cost
Restore duration Direct critical-path cost before script
Upload duration Writer job tail latency and backend cost
Dependency cold-install duration Upper bound on work cache can save
Hit rate by runner/key Shows whether topology/key design provides reuse
Queue time Must be excluded from cache transfer conclusions

Then remove low-value paths or split independent caches. Do not cache the entire workspace or build outputs simply because selecting a narrow dependency directory takes more thought.

6. Failure mode: cache mistaken for artifact

A package job places dist/app.tar.gz into cache and a deploy job restores it in a later pipeline. This may appear to work until the key is overwritten, evicted, restored from a fallback, or produced by the wrong source. The fix is architectural: publish the intended output as a job artifact or immutable registry/package object, record its digest, and let cache remain dependency acceleration.

7. Failure mode: same key, different paths or runners

Current GitLab troubleshooting documentation highlights two common mismatches: jobs using the same key with different cached paths can overwrite one another’s archive, and multiple standalone runners without shared/distributed cache can produce uneven hit behavior. Inspect runner IDs and key/path definitions before touching application code.

8. Clearing is recovery, not diagnosis

GitLab’s “Clear runner caches” operation changes the internal cache name/index for subsequent jobs; old data is not necessarily physically deleted from runner/object storage. A broad clear can remove warm state unrelated to the fault and erase useful reproduction conditions. Prefer changing the affected key or correcting its inputs. Use the UI clear only in the disposable lab when you can name exactly what you are invalidating and why.

9. Security-sensitive shortcuts to reject

  • Do not put secrets, tokens, private keys, or credential files in cache.
  • Do not set cache:unprotect: true merely to improve hit rate across trust boundaries.
  • Do not expose distributed-cache backend credentials to ordinary job logs.
  • Do not mount broad host directories into privileged runners to “share cache faster.”
  • Do not disable TLS or object-storage certificate validation for cache troubleshooting.
  • Do not let untrusted jobs overwrite caches consumed by privileged deployment jobs.

10. Compact recovery playbook

  1. Freeze the failed pipeline/job IDs and collect first-failure evidence.
  2. Compute the exact lockfile/toolchain fingerprints expected by the job.
  3. Confirm the compiled primary/fallback keys and protected scope.
  4. Correlate runner ID/executor with local/distributed backend state.
  5. Verify restored dependency contents independently.
  6. If the key contract is wrong, create a corrected new key rather than mutating old evidence.
  7. If the backend is unavailable, run the cache-independent path and repair infrastructure separately.
  8. Retry only the smallest safe job/pipeline scope after the causal fix.

Knowledge check

A job trace says “Successfully extracted cache” but tests use an old dependency. Which layer failed?

Why is “clear all runner caches” a poor first step?

How do you diagnose uneven hit rates across two runners?

What is the right fix when a deployable binary is being passed through cache?

Why can a fallback hit be dangerous across trust scopes?

Next lesson

Checkpoint lab

Design, measure, invalidate, poison safely, recover narrowly, and produce an evidence packet that proves correctness independent of cache.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-11. Cache behavior is version-sensitive at both GitLab and GitLab Runner layers. Re-check the deployed GitLab/Runner versions for Self-Managed or Dedicated installations, especially when runner topology, object storage, protected-cache behavior, or cache archive implementation differs from GitLab.com.

  • Caching in GitLab CI/CD — cache-versus-artifact boundary, fallback keys, protected-cache separation, availability, storage, clearing, and troubleshooting.
  • CI/CD YAML syntax reference — authoritative cache, cache:key, cache:key:files, cache:key:files_commits, cache:key:prefix, cache:fallback_keys, cache:policy, cache:when, and cache:unprotect semantics.
  • CI/CD caching examples — dependency-manager patterns and lockfile-aware examples.
  • GitLab Runner advanced configuration — distributed cache backend configuration, sharing, paths, and archive limits.
  • Speed up job execution — distributed cache backends and transfer-performance considerations.
  • Job artifacts — the retained-output mechanism that must not be confused with cache.
Current behavior used by this chapter: GitLab documents a maximum of four caches per job and up to five per-cache fallback keys. Caches are restored before artifacts. Cache availability is not guaranteed. Protected and non-protected cache namespaces are separated by default, and current documentation records role/ref-sensitive protected-key suffix behavior introduced during the GitLab 18.x series. Runner 18.1 also changed cache archiving so symlinks are no longer followed in relevant edge cases.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.