Caches, Cache Keys, Fallback Keys, Distributed Cache, Dependency Reuse, and Cache Correctness: Configuration, Design Choices, and Tradeoffs
Cache design is a tradeoff among reuse, isolation, transfer cost, poisoning risk, runner topology, and developer feedback time. This lesson turns those tradeoffs into explicit choices about key granularity, multiple caches, fallback scope, pull/push policy, and local versus distributed storage.
Learning objectives
- Choose between one cache and multiple independent caches based on dependency ownership and invalidation boundaries.
- Compare exact keys, lockfile-derived keys, prefixes, and fallback keys for speed, auditability, and poisoning risk.
- Use pull, push, and pull-push policies deliberately so most consumers do not mutate shared cache state.
- Decide when local cache is sufficient and when a runner fleet needs distributed cache.
- Apply a worked decision matrix that includes trust, tier/offering, cost, rollback, and observable evidence.
1. Design from invalidation and trust boundaries, not from cache syntax
A cache design is good when a reviewer can answer four questions: what makes the content reusable, who can write it, where it is stored, and what independent evidence proves the job is correct if the cache is stale or missing. These questions scale from one runner to a fleet.
2. One cache versus multiple caches
GitLab currently supports up to four cache entries per job. Multiple caches are useful when dependency families have different invalidation inputs or sizes—for example Python package downloads and frontend package downloads. They are harmful when split mechanically into many archives whose compression/transfer overhead exceeds their saved work.
| Choice | Benefits | Costs / risks | Evidence to collect |
|---|---|---|---|
| Single cache | Simple config, one lookup/archive | Unrelated data invalidates or overwrites together; large transfers | Archive size, restore/upload time, hit rate |
| Multiple caches | Independent invalidation and policies | More lookups, fallback traffic, config complexity | Per-key size, per-key hit rate, critical-path effect |
3. Exact keys, lockfile keys, prefixes, and fallback keys
A static exact key is appropriate only when some external process
changes the key deliberately. A branch key isolates branches but can
waste reuse when lockfiles are identical.
cache:key:files often gives a better dependency
contract. A prefix can add toolchain/OS/architecture identity.
Fallback keys provide warm starts but broaden the set of possible
restored inputs and therefore increase the importance of validation.
| Pattern | Good fit | Failure mode to watch |
|---|---|---|
| Static key | Explicitly versioned shared seed cache | Never invalidates after dependency/toolchain change |
| $CI_COMMIT_REF_SLUG | Branch-local mutable cache | Low reuse; branch names do not encode dependency compatibility |
| key:files lockfile | Dependency cache | Toolchain/ABI omitted unless included via prefix |
| key:files + prefix | Dependency + toolchain/platform contract | Prefix becomes stale when image/runtime changes |
| fallback_keys | Feature branch warm start | Fallback from broader trust/compatibility scope |
4. Pull, push, and pull-push: separate readers from writers
Cache poisoning risk grows with the number of writers. A common
production pattern is to let a dedicated dependency-preparation job
use pull-push while parallel test jobs use
pull. A special cache-warmer can use
push after reconstructing dependencies from trusted
sources. This reduces upload contention and makes mutation ownership
reviewable.
Policy choice is independent from pipeline authorization. A
low-trust branch with pull-push can still be dangerous
if it reaches a shared key. Protected-cache separation, runner
trust, key scope, and dependency verification must align.
5. Local versus distributed cache
Local cache is operationally simple and can be extremely fast when the same runner executes related jobs. It is unreliable as a fleet-wide optimization because another runner may not have the archive. Distributed cache moves the archive to object storage reachable by multiple runners. Current GitLab Runner docs support S3/S3-compatible, GCS, and Azure backends.
| Topology | When it fits | Operational state you own |
|---|---|---|
| Single persistent runner + local cache | Small trusted project, stable host | Runner disk capacity, cleanup, host lifecycle |
| Multiple persistent runners + no sharing | Only when misses are acceptable | Uneven cache state and hit rate across runners |
| Runner fleet + distributed object cache | Autoscaling or multi-runner reuse | Backend credentials, lifecycle rules, transfer cost, availability |
| Shared network directory | Specialized homogeneous fleets | Filesystem availability, locking/performance, mount security |
6. Protected separation versus broad reuse
Default protected/non-protected separation should be viewed as a
security feature. Disabling it or using
cache:unprotect may improve reuse but increases the
writer population that can influence a cache later read by trusted
jobs. If you cross that boundary, compensate with content-addressed
packages, strong package-manager verification, read-only consumers,
and a narrow set of trusted writers.
7. Speed versus archive/transfer cost
A cache is a net win only when saved dependency work exceeds compression, upload, download, extraction, and backend latency. A 3 GB cache that saves 20 seconds of dependency work but adds 45 seconds of transfer has negative value. Measure the critical path, not only cache hit percentage.
Also consider backend lifecycle cost: distributed caches accumulate mutable archives. Apply object-storage lifecycle rules and runner-side limits instead of assuming GitLab project artifact retention settings manage runner cache objects.
8. Cached downloads versus deployable outputs
Use cache for dependency reuse. Use artifacts/packages/container registry objects for outputs that need producer identity, retention, promotion, or deployment. Do not take a binary from a mutable dependency cache and call it a release simply because its key contains a source SHA. If the output matters to deployment, Chapter 10’s build-once/digest-verification model applies.
9. Worked scenario: choose a cache for a polyglot monorepo
A repository has Python and Node services. Python jobs use CPython
3.11 on Linux; Node jobs use Node 24. Both share an autoscaled
runner fleet. Feature branches are untrusted, while
main produces release candidates. Choose an approach
and justify it.
| Decision | Recommended pattern | Reason / evidence |
|---|---|---|
| Cache split | Two caches: Python + Node | Independent lockfiles/invalidation; within four-cache limit |
| Key | Lockfile content + toolchain prefix | Reuse tracks dependency compatibility, not branch name |
| Fallback | Default-branch fallback only within same trust scope | Warm start without crossing protected boundary |
| Policy | Writer pull-push; test jobs pull | Reduces race/poisoning surface |
| Backend | Distributed object cache | Autoscaled workers are ephemeral |
| Correctness | Package-manager/manifest verification + tests | Cache availability/hit is not proof |
| Release output | Artifact/registry digest, not cache | Preserves producer identity and promotion contract |
10. Decision checklist
- List correctness-relevant dependency and toolchain inputs.
- Decide which inputs become key content versus key prefix.
- Define writer and reader jobs separately.
- Keep protected/unprotected scopes separated unless there is a reviewed reason not to.
- Choose local/distributed storage from actual runner topology.
- Measure archive size, hit rate, transfer time, and end-to-end critical-path impact.
- Verify that every consumer remains correct after a miss and rejects incompatible restored data.
- Keep deployable outputs in artifacts/registry/package systems with immutable identity.
Knowledge check
Why might multiple caches improve correctness as well as speed?
They can give independently versioned dependency families separate keys and update policies, reducing accidental overwrite and over-broad invalidation.
When is policy: pull preferable to pull-push?
For read-only consumers such as parallel tests that should reuse cache state but should not all mutate the shared key.
What problem does distributed cache solve?
Availability/reuse across runners or ephemeral workers. It does not prove dependency correctness and adds backend trust/cost/availability state.
Why is a source-SHA cache key still not a release identity?
Cache is mutable performance state without the producer/retention/promotion semantics of an artifact or registry/package object.
Which metric matters more than raw cache hit rate?
End-to-end/critical-path improvement while preserving correctness; transfer and queue overhead can make a high hit rate still unhelpful.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. Cache behavior is version-sensitive at both GitLab and GitLab Runner layers. Re-check the deployed GitLab/Runner versions for Self-Managed or Dedicated installations, especially when runner topology, object storage, protected-cache behavior, or cache archive implementation differs from GitLab.com.
- Caching in GitLab CI/CD — cache-versus-artifact boundary, fallback keys, protected-cache separation, availability, storage, clearing, and troubleshooting.
-
CI/CD YAML syntax reference
— authoritative
cache,cache:key,cache:key:files,cache:key:files_commits,cache:key:prefix,cache:fallback_keys,cache:policy,cache:when, andcache:unprotectsemantics. - CI/CD caching examples — dependency-manager patterns and lockfile-aware examples.
- GitLab Runner advanced configuration — distributed cache backend configuration, sharing, paths, and archive limits.
- Speed up job execution — distributed cache backends and transfer-performance considerations.
- Job artifacts — the retained-output mechanism that must not be confused with cache.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.