Dependency Caching, Cache Keys, Restore Strategies, and Performance: Core Concepts and Mental Model
Chapter 13 treated artifacts as explicit run evidence. Chapter 14 deliberately changes the persistence contract: a dependency cache is disposable acceleration state that may be absent, stale, partially matched or evicted, so build correctness must remain anchored in the lockfile, toolchain and source revision rather than the cache itself.
Learning objectives
- Explain why a cache is optional acceleration state rather than authoritative build evidence.
- Build precise cache keys from correctness-relevant inputs such as OS, architecture, toolchain and dependency lock hash.
- Interpret exact hits, restore-key matches, misses, cache versions and branch/ref scope correctly.
- Distinguish caches from artifacts, job outputs and installed environments.
- Identify cache poisoning, secret leakage, eviction and stale-state risks before optimizing latency.
1. The practical problem: repeated dependency work is expensive, but correctness cannot depend on reuse
A fresh GitHub-hosted runner starts without the dependency downloads from the previous run. Re-downloading the same wheels, npm tarballs or Maven dependencies wastes time and network bandwidth. Caching solves that performance problem by saving selected reusable files and restoring them into a later run.
The critical boundary is that a cache is best effort. It can miss because the key changed, the cache version changed, branch scope prevents access, the entry expired, storage pressure evicted it, or a low-trust run has read-only access. A correct workflow must still be able to reconstruct the required dependencies from source-controlled inputs when no cache is available.
2. Causal model: dependency contract → key/version/scope → restore → install/build → optional save
First define the dependency contract: lockfile or equivalent manifest, exact toolchain, operating system/architecture when binaries differ, and any configuration that changes what is downloaded or compiled. The workflow derives a primary key from those inputs. The cache service searches that key and cache version in the permitted scope; if there is no exact match, ordered restore prefixes may provide older acceleration state. The package manager then reconciles restored files with the current dependency contract. Only after a successful job may the current cache action save a new primary-key entry.
flowchart TD
A[Lockfile + toolchain + OS/arch] --> B[Primary cache key]
B --> C[Cache key + version + ref scope]
C --> D{Exact match?}
D -->|Yes| E[Restore exact cache]
D -->|No| F{Restore-key prefix match?}
F -->|Yes| G[Restore older/partial cache]
F -->|No| H[Start without cache]
E --> I[Package manager validates current inputs]
G --> I
H --> I
I --> J[Build/test]
J --> K{Job succeeds and key absent?}
K -->|Yes| L[Save new cache entry]
K -->|No| M[No new cache save]
The package manager or build system remains the correctness authority. Restored bytes are an optimization input, not proof that the current lockfile has been satisfied.
3. Record the state that explains a cache decision
| State | Question | Useful evidence |
|---|---|---|
| Action/runtime | Which cache action and runtime executed? | full action SHA; runner version for self-hosted infrastructure |
| Primary key | Which exact key did this run request? | expanded key in workflow/logs |
| Restore keys | Which prefixes were eligible, in what order? | workflow manifest + cache action log |
| Cache version | Was path/compression metadata compatible? | cache list/API version field |
| Scope/ref | Which branch/tag/PR scopes could be searched? | event/ref/base branch + cache metadata ref |
| Dependency input | What lockfile/toolchain/OS/arch should determine correctness? | lock hash; Python/Node/Java version; runner OS/arch |
| Hit result | Exact hit, fallback or complete miss? | cache-hit plus restore log |
| Restored path | What reusable directory entered the runner? | configured path and post-restore listing |
| Save eligibility | Was a new primary-key cache allowed and successful? | job conclusion + post-action warning/log |
| Performance | Did restore reduce the measured critical operation? | cold/warm durations, cache size and network context |
4. A key is part of the correctness boundary
A useful dependency key changes when a correctness-relevant input changes. For a Python wheel/download cache, a reasonable key includes runner OS/architecture, Python major/minor and a lockfile hash. The exact syntax is less important than the invariant: two runs that can require incompatible cached bytes must not be forced to share the same exact key.
- id: cache
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
with:
path: ${{ steps.pip.outputs.cache_dir }}
key: >-
pip-${{ runner.os }}-${{ runner.arch }}-py${{ steps.py.outputs.minor }}-${{ steps.lock.outputs.sha256 }}
restore-keys: |
pip-${{ runner.os }}-${{ runner.arch }}-py${{ steps.py.outputs.minor }}-
A key such as pip-cache can return an exact hit after
the lockfile or Python version changes. That does not mean the new
dependency set is installed or valid; it only means that bytes
exist under that stale key.
5. Exact hit, fallback and miss are three different observations
Current actions/cache first looks for the primary key
and cache version, then prefix matches, then ordered
restore-keys. If the current branch has no permitted
match, GitHub can repeat the lookup against the default branch
subject to cache-scope restrictions. An exact primary-key match
reports cache-hit as 'true'. A restored
prefix/restore-key match reports 'false'. No restored
cache yields an empty output.
That distinction matters. A fallback cache is intentionally stale relative to the current primary key, so the dependency tool must fill or replace missing data. Even an exact hit normally does not justify skipping dependency installation when the cached path is merely a package-manager download cache.
6. Same text key can still miss: cache version and scope
GitHub also stamps cache entries with a cache version derived from the cached paths and compression/archive behavior. This prevents a run from restoring an entry it cannot correctly unpack into the requested path. Branch/tag/PR scope adds another boundary: caches are not a repository-wide anonymous bucket.
When a key appears “identical” but restore still misses, inspect the
cache version and ref before deleting
anything. A missing cache can be correct behavior rather than
corruption.
7. Cache, artifact and output solve different persistence problems
| Mechanism | Primary purpose | Correctness expectation | Typical lifetime/identity |
|---|---|---|---|
| Job output | small control-plane values between jobs | exact value from this workflow run | run-scoped expression data |
| Artifact | explicit files/evidence from a run | consumer should bind to run/SHA/digest | artifact ID/digest + retention |
| Cache | reusable acceleration state across runs | workflow must work if absent/stale | key/version/ref scope + eviction |
8. Restored cache contents are untrusted input
Caches can be readable from broader workflow contexts than the job
that created them, including pull-request scenarios that can restore
base/default-branch caches. Never place access tokens, private keys,
.env files or other secrets in cached paths. A poisoned
cache can also become an execution path if a trusted workflow
executes binaries or scripts restored from a cache without
validation.
Prefer caching package-manager download stores or deterministic intermediate state whose consumers revalidate it. Be especially cautious with compiled build-output caches that embed environment variables, compiler flags, absolute paths or trust-sensitive generated code.
9. Read-only cache inspection before tuning or deletion
Cache metadata can explain a surprising restore without mutating anything. In an authorized disposable repository, inspect IDs, keys, refs, versions, sizes and access times before changing keys or deleting entries.
GH_REPO='OWNER/gha-cache-lab'
gh api -H 'X-GitHub-Api-Version: 2026-03-10' "repos/$GH_REPO/actions/caches?per_page=100" --jq '.actions_caches[] | {id,ref,key,version,size_in_bytes,created_at,last_accessed_at}'
10. Performance state is also storage state
GitHub currently removes cache entries after the configured inactivity period and evicts older entries when repository cache storage exceeds its limit. The default remains 7 days of inactivity and 10 GB per repository, while eligible plans can configure larger limits. Therefore “the warm run will always hit next week” is not a valid invariant.
Measure cache hit rate, restore duration, install duration and cache size before increasing storage. A large cache that is slow to transfer can make a workflow slower than rebuilding or re-downloading.
Knowledge check
Why is a cache miss not a workflow correctness failure?
Because cache availability is only a performance optimization. The workflow should reconstruct dependencies from the lockfile/toolchain when no cache is available.
What does cache-hit == false mean for
actions/cache?
A cache was restored through a prefix/restore-key match rather than an exact primary-key match. The restored state is intentionally older or less specific.
Why can the same key text still fail to restore?
Cache version and ref/branch scope are also part of cache matching and compatibility.
Should an exact hit on the pip download cache allow skipping
pip install?
Normally no. The pip cache contains reusable downloads/wheels, not proof that the fresh runner has the dependencies installed in the required environment.
Why are secrets forbidden in cache paths?
Authorized or low-trust workflow contexts may be able to restore caches, and masking does not protect secret bytes stored inside an archive.
Official references and version notes
- GitHub Docs — Dependency caching — cache purpose, artifact distinction and cache-security boundary.
- GitHub Docs — Dependency caching reference — key matching, restore keys, branch scope, versioning, limits and eviction.
- GitHub Docs — Managing caches — read-only inspection and exact cache deletion.
-
actions/cache v6.1.0
— explicit cache action pinned to
55cc8345863c7cc4c66a329aec7e433d2d1c52a9. -
actions/setup-python v7.0.0
— Python setup action pinned to
5fda3b95a4ea91299a34e894583c3862153e4b97. -
setup-python dependency caching
— built-in pip/pipenv/Poetry caching and
cache-dependency-path.
Version-sensitive behavior was rechecked on
2026-09-09. Mandatory examples target GitHub.com
and ubuntu-24.04. actions/cache v6.1.0
is pinned by full commit SHA and runs on Node 24; keep self-hosted
runners current enough to execute Node 24 actions (GitHub
documents runner 2.327.1 as the Node 24 minimum in current
official action guidance). Cache entries are scoped by key, cache
version and ref/branch visibility. Exact primary-key matches
report cache-hit == 'true'; prefix/restore-key
matches report 'false'; a total miss yields an empty
value. The cache action's post-save runs on successful jobs; do
not rely on deprecated save-always. GitHub currently
defaults cache inactivity retention to 7 days and repository cache
storage to 10 GB; eligible repositories/organizations can
configure higher limits, so performance/storage assumptions must
be recorded rather than hard-coded.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.