Chapter 14Lesson 01~155 minutes

Dependency Caching, Cache Keys, Restore Strategies, and Performance: Core Concepts and Mental Model

Chapter 13 treated artifacts as explicit run evidence. Chapter 14 deliberately changes the persistence contract: a dependency cache is disposable acceleration state that may be absent, stale, partially matched or evicted, so build correctness must remain anchored in the lockfile, toolchain and source revision rather than the cache itself.

Dependency cacheCache keyCache versionRestore keysCorrectness boundary

Learning objectives

  • Explain why a cache is optional acceleration state rather than authoritative build evidence.
  • Build precise cache keys from correctness-relevant inputs such as OS, architecture, toolchain and dependency lock hash.
  • Interpret exact hits, restore-key matches, misses, cache versions and branch/ref scope correctly.
  • Distinguish caches from artifacts, job outputs and installed environments.
  • Identify cache poisoning, secret leakage, eviction and stale-state risks before optimizing latency.

1. The practical problem: repeated dependency work is expensive, but correctness cannot depend on reuse

A fresh GitHub-hosted runner starts without the dependency downloads from the previous run. Re-downloading the same wheels, npm tarballs or Maven dependencies wastes time and network bandwidth. Caching solves that performance problem by saving selected reusable files and restoring them into a later run.

The critical boundary is that a cache is best effort. It can miss because the key changed, the cache version changed, branch scope prevents access, the entry expired, storage pressure evicted it, or a low-trust run has read-only access. A correct workflow must still be able to reconstruct the required dependencies from source-controlled inputs when no cache is available.

2. Causal model: dependency contract → key/version/scope → restore → install/build → optional save

First define the dependency contract: lockfile or equivalent manifest, exact toolchain, operating system/architecture when binaries differ, and any configuration that changes what is downloaded or compiled. The workflow derives a primary key from those inputs. The cache service searches that key and cache version in the permitted scope; if there is no exact match, ordered restore prefixes may provide older acceleration state. The package manager then reconciles restored files with the current dependency contract. Only after a successful job may the current cache action save a new primary-key entry.

Cache lookup and correctness boundary
flowchart TD
  A[Lockfile + toolchain + OS/arch] --> B[Primary cache key]
  B --> C[Cache key + version + ref scope]
  C --> D{Exact match?}
  D -->|Yes| E[Restore exact cache]
  D -->|No| F{Restore-key prefix match?}
  F -->|Yes| G[Restore older/partial cache]
  F -->|No| H[Start without cache]
  E --> I[Package manager validates current inputs]
  G --> I
  H --> I
  I --> J[Build/test]
  J --> K{Job succeeds and key absent?}
  K -->|Yes| L[Save new cache entry]
  K -->|No| M[No new cache save]

The package manager or build system remains the correctness authority. Restored bytes are an optimization input, not proof that the current lockfile has been satisfied.

3. Record the state that explains a cache decision

State Question Useful evidence
Action/runtime Which cache action and runtime executed? full action SHA; runner version for self-hosted infrastructure
Primary key Which exact key did this run request? expanded key in workflow/logs
Restore keys Which prefixes were eligible, in what order? workflow manifest + cache action log
Cache version Was path/compression metadata compatible? cache list/API version field
Scope/ref Which branch/tag/PR scopes could be searched? event/ref/base branch + cache metadata ref
Dependency input What lockfile/toolchain/OS/arch should determine correctness? lock hash; Python/Node/Java version; runner OS/arch
Hit result Exact hit, fallback or complete miss? cache-hit plus restore log
Restored path What reusable directory entered the runner? configured path and post-restore listing
Save eligibility Was a new primary-key cache allowed and successful? job conclusion + post-action warning/log
Performance Did restore reduce the measured critical operation? cold/warm durations, cache size and network context

4. A key is part of the correctness boundary

A useful dependency key changes when a correctness-relevant input changes. For a Python wheel/download cache, a reasonable key includes runner OS/architecture, Python major/minor and a lockfile hash. The exact syntax is less important than the invariant: two runs that can require incompatible cached bytes must not be forced to share the same exact key.

- id: cache
  uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
  with:
    path: ${{ steps.pip.outputs.cache_dir }}
    key: >-
      pip-${{ runner.os }}-${{ runner.arch }}-py${{ steps.py.outputs.minor }}-${{ steps.lock.outputs.sha256 }}
    restore-keys: |
      pip-${{ runner.os }}-${{ runner.arch }}-py${{ steps.py.outputs.minor }}-
Do not omit the dependency contract

A key such as pip-cache can return an exact hit after the lockfile or Python version changes. That does not mean the new dependency set is installed or valid; it only means that bytes exist under that stale key.

5. Exact hit, fallback and miss are three different observations

Current actions/cache first looks for the primary key and cache version, then prefix matches, then ordered restore-keys. If the current branch has no permitted match, GitHub can repeat the lookup against the default branch subject to cache-scope restrictions. An exact primary-key match reports cache-hit as 'true'. A restored prefix/restore-key match reports 'false'. No restored cache yields an empty output.

That distinction matters. A fallback cache is intentionally stale relative to the current primary key, so the dependency tool must fill or replace missing data. Even an exact hit normally does not justify skipping dependency installation when the cached path is merely a package-manager download cache.

6. Same text key can still miss: cache version and scope

GitHub also stamps cache entries with a cache version derived from the cached paths and compression/archive behavior. This prevents a run from restoring an entry it cannot correctly unpack into the requested path. Branch/tag/PR scope adds another boundary: caches are not a repository-wide anonymous bucket.

When a key appears “identical” but restore still misses, inspect the cache version and ref before deleting anything. A missing cache can be correct behavior rather than corruption.

7. Cache, artifact and output solve different persistence problems

Mechanism Primary purpose Correctness expectation Typical lifetime/identity
Job output small control-plane values between jobs exact value from this workflow run run-scoped expression data
Artifact explicit files/evidence from a run consumer should bind to run/SHA/digest artifact ID/digest + retention
Cache reusable acceleration state across runs workflow must work if absent/stale key/version/ref scope + eviction

8. Restored cache contents are untrusted input

Caches can be readable from broader workflow contexts than the job that created them, including pull-request scenarios that can restore base/default-branch caches. Never place access tokens, private keys, .env files or other secrets in cached paths. A poisoned cache can also become an execution path if a trusted workflow executes binaries or scripts restored from a cache without validation.

Prefer caching package-manager download stores or deterministic intermediate state whose consumers revalidate it. Be especially cautious with compiled build-output caches that embed environment variables, compiler flags, absolute paths or trust-sensitive generated code.

9. Read-only cache inspection before tuning or deletion

Cache metadata can explain a surprising restore without mutating anything. In an authorized disposable repository, inspect IDs, keys, refs, versions, sizes and access times before changing keys or deleting entries.

GH_REPO='OWNER/gha-cache-lab'

gh api   -H 'X-GitHub-Api-Version: 2026-03-10'   "repos/$GH_REPO/actions/caches?per_page=100"   --jq '.actions_caches[] | {id,ref,key,version,size_in_bytes,created_at,last_accessed_at}'

10. Performance state is also storage state

GitHub currently removes cache entries after the configured inactivity period and evicts older entries when repository cache storage exceeds its limit. The default remains 7 days of inactivity and 10 GB per repository, while eligible plans can configure larger limits. Therefore “the warm run will always hit next week” is not a valid invariant.

Measure cache hit rate, restore duration, install duration and cache size before increasing storage. A large cache that is slow to transfer can make a workflow slower than rebuilding or re-downloading.

Knowledge check

Why is a cache miss not a workflow correctness failure?

What does cache-hit == false mean for actions/cache?

Why can the same key text still fail to restore?

Should an exact hit on the pip download cache allow skipping pip install?

Why are secrets forbidden in cache paths?

Next lesson

Measure a real cold/warm dependency cache

Lesson 2 creates a tiny disposable Python scenario and makes every key, fallback and timing observation visible.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked on 2026-09-09. Mandatory examples target GitHub.com and ubuntu-24.04. actions/cache v6.1.0 is pinned by full commit SHA and runs on Node 24; keep self-hosted runners current enough to execute Node 24 actions (GitHub documents runner 2.327.1 as the Node 24 minimum in current official action guidance). Cache entries are scoped by key, cache version and ref/branch visibility. Exact primary-key matches report cache-hit == 'true'; prefix/restore-key matches report 'false'; a total miss yields an empty value. The cache action's post-save runs on successful jobs; do not rely on deprecated save-always. GitHub currently defaults cache inactivity retention to 7 days and repository cache storage to 10 GB; eligible repositories/organizations can configure higher limits, so performance/storage assumptions must be recorded rather than hard-coded.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.