Chapter 11Lesson 03~145 minutes

Caches, Cache Keys, Fallback Keys, Distributed Cache, Dependency Reuse, and Cache Correctness: Configuration, Design Choices, and Tradeoffs

Cache design is a tradeoff among reuse, isolation, transfer cost, poisoning risk, runner topology, and developer feedback time. This lesson turns those tradeoffs into explicit choices about key granularity, multiple caches, fallback scope, pull/push policy, and local versus distributed storage.

CacheCache keysCorrectnessGitLab RunnerDependency reuse

Learning objectives

  • Choose between one cache and multiple independent caches based on dependency ownership and invalidation boundaries.
  • Compare exact keys, lockfile-derived keys, prefixes, and fallback keys for speed, auditability, and poisoning risk.
  • Use pull, push, and pull-push policies deliberately so most consumers do not mutate shared cache state.
  • Decide when local cache is sufficient and when a runner fleet needs distributed cache.
  • Apply a worked decision matrix that includes trust, tier/offering, cost, rollback, and observable evidence.

1. Design from invalidation and trust boundaries, not from cache syntax

A cache design is good when a reviewer can answer four questions: what makes the content reusable, who can write it, where it is stored, and what independent evidence proves the job is correct if the cache is stale or missing. These questions scale from one runner to a fleet.

2. One cache versus multiple caches

GitLab currently supports up to four cache entries per job. Multiple caches are useful when dependency families have different invalidation inputs or sizes—for example Python package downloads and frontend package downloads. They are harmful when split mechanically into many archives whose compression/transfer overhead exceeds their saved work.

Choice Benefits Costs / risks Evidence to collect
Single cache Simple config, one lookup/archive Unrelated data invalidates or overwrites together; large transfers Archive size, restore/upload time, hit rate
Multiple caches Independent invalidation and policies More lookups, fallback traffic, config complexity Per-key size, per-key hit rate, critical-path effect

3. Exact keys, lockfile keys, prefixes, and fallback keys

A static exact key is appropriate only when some external process changes the key deliberately. A branch key isolates branches but can waste reuse when lockfiles are identical. cache:key:files often gives a better dependency contract. A prefix can add toolchain/OS/architecture identity. Fallback keys provide warm starts but broaden the set of possible restored inputs and therefore increase the importance of validation.

Pattern Good fit Failure mode to watch
Static key Explicitly versioned shared seed cache Never invalidates after dependency/toolchain change
$CI_COMMIT_REF_SLUG Branch-local mutable cache Low reuse; branch names do not encode dependency compatibility
key:files lockfile Dependency cache Toolchain/ABI omitted unless included via prefix
key:files + prefix Dependency + toolchain/platform contract Prefix becomes stale when image/runtime changes
fallback_keys Feature branch warm start Fallback from broader trust/compatibility scope

4. Pull, push, and pull-push: separate readers from writers

Cache poisoning risk grows with the number of writers. A common production pattern is to let a dedicated dependency-preparation job use pull-push while parallel test jobs use pull. A special cache-warmer can use push after reconstructing dependencies from trusted sources. This reduces upload contention and makes mutation ownership reviewable.

Policy choice is independent from pipeline authorization. A low-trust branch with pull-push can still be dangerous if it reaches a shared key. Protected-cache separation, runner trust, key scope, and dependency verification must align.

5. Local versus distributed cache

Local cache is operationally simple and can be extremely fast when the same runner executes related jobs. It is unreliable as a fleet-wide optimization because another runner may not have the archive. Distributed cache moves the archive to object storage reachable by multiple runners. Current GitLab Runner docs support S3/S3-compatible, GCS, and Azure backends.

Topology When it fits Operational state you own
Single persistent runner + local cache Small trusted project, stable host Runner disk capacity, cleanup, host lifecycle
Multiple persistent runners + no sharing Only when misses are acceptable Uneven cache state and hit rate across runners
Runner fleet + distributed object cache Autoscaling or multi-runner reuse Backend credentials, lifecycle rules, transfer cost, availability
Shared network directory Specialized homogeneous fleets Filesystem availability, locking/performance, mount security

6. Protected separation versus broad reuse

Default protected/non-protected separation should be viewed as a security feature. Disabling it or using cache:unprotect may improve reuse but increases the writer population that can influence a cache later read by trusted jobs. If you cross that boundary, compensate with content-addressed packages, strong package-manager verification, read-only consumers, and a narrow set of trusted writers.

7. Speed versus archive/transfer cost

A cache is a net win only when saved dependency work exceeds compression, upload, download, extraction, and backend latency. A 3 GB cache that saves 20 seconds of dependency work but adds 45 seconds of transfer has negative value. Measure the critical path, not only cache hit percentage.

Also consider backend lifecycle cost: distributed caches accumulate mutable archives. Apply object-storage lifecycle rules and runner-side limits instead of assuming GitLab project artifact retention settings manage runner cache objects.

8. Cached downloads versus deployable outputs

Use cache for dependency reuse. Use artifacts/packages/container registry objects for outputs that need producer identity, retention, promotion, or deployment. Do not take a binary from a mutable dependency cache and call it a release simply because its key contains a source SHA. If the output matters to deployment, Chapter 10’s build-once/digest-verification model applies.

9. Worked scenario: choose a cache for a polyglot monorepo

A repository has Python and Node services. Python jobs use CPython 3.11 on Linux; Node jobs use Node 24. Both share an autoscaled runner fleet. Feature branches are untrusted, while main produces release candidates. Choose an approach and justify it.

Decision Recommended pattern Reason / evidence
Cache split Two caches: Python + Node Independent lockfiles/invalidation; within four-cache limit
Key Lockfile content + toolchain prefix Reuse tracks dependency compatibility, not branch name
Fallback Default-branch fallback only within same trust scope Warm start without crossing protected boundary
Policy Writer pull-push; test jobs pull Reduces race/poisoning surface
Backend Distributed object cache Autoscaled workers are ephemeral
Correctness Package-manager/manifest verification + tests Cache availability/hit is not proof
Release output Artifact/registry digest, not cache Preserves producer identity and promotion contract

10. Decision checklist

  1. List correctness-relevant dependency and toolchain inputs.
  2. Decide which inputs become key content versus key prefix.
  3. Define writer and reader jobs separately.
  4. Keep protected/unprotected scopes separated unless there is a reviewed reason not to.
  5. Choose local/distributed storage from actual runner topology.
  6. Measure archive size, hit rate, transfer time, and end-to-end critical-path impact.
  7. Verify that every consumer remains correct after a miss and rejects incompatible restored data.
  8. Keep deployable outputs in artifacts/registry/package systems with immutable identity.

Knowledge check

Why might multiple caches improve correctness as well as speed?

When is policy: pull preferable to pull-push?

What problem does distributed cache solve?

Why is a source-SHA cache key still not a release identity?

Which metric matters more than raw cache hit rate?

Next lesson

Diagnostics, failure modes, security, and performance

Diagnose stale, poisoned, oversized, unavailable, and incorrectly trusted caches while preserving first-failure evidence.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-11. Cache behavior is version-sensitive at both GitLab and GitLab Runner layers. Re-check the deployed GitLab/Runner versions for Self-Managed or Dedicated installations, especially when runner topology, object storage, protected-cache behavior, or cache archive implementation differs from GitLab.com.

  • Caching in GitLab CI/CD — cache-versus-artifact boundary, fallback keys, protected-cache separation, availability, storage, clearing, and troubleshooting.
  • CI/CD YAML syntax reference — authoritative cache, cache:key, cache:key:files, cache:key:files_commits, cache:key:prefix, cache:fallback_keys, cache:policy, cache:when, and cache:unprotect semantics.
  • CI/CD caching examples — dependency-manager patterns and lockfile-aware examples.
  • GitLab Runner advanced configuration — distributed cache backend configuration, sharing, paths, and archive limits.
  • Speed up job execution — distributed cache backends and transfer-performance considerations.
  • Job artifacts — the retained-output mechanism that must not be confused with cache.
Current behavior used by this chapter: GitLab documents a maximum of four caches per job and up to five per-cache fallback keys. Caches are restored before artifacts. Cache availability is not guaranteed. Protected and non-protected cache namespaces are separated by default, and current documentation records role/ref-sensitive protected-key suffix behavior introduced during the GitLab 18.x series. Runner 18.1 also changed cache archiving so symlinks are no longer followed in relevant edge cases.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.