Dependency Caching, Cache Keys, Restore Strategies, and Performance: Diagnostics, Failure Modes, and Production Practices
Cache incidents are often misdiagnosed as random runner or network failures. This lesson preserves first-failure evidence and separates key design, scope/version mismatches, poisoned or secret-bearing contents, stale generated outputs, incorrect cache-hit assumptions, and obsolete runner/action runtimes before any cache deletion.
Learning objectives
- Use an evidence-first sequence to separate cache matching problems from dependency, runner, network and build failures.
- Diagnose keys that omit lockfile, toolchain, OS or architecture inputs.
- Recognize secret-bearing and poisoned caches as security incidents rather than performance glitches.
- Explain why a cache hit does not prove installed dependencies or build correctness.
- Repair stale or incompatible caches without erasing the original run/cache evidence.
1. Evidence-first diagnostic sequence
Before rerunning or deleting a cache, preserve the run ID/attempt,
event/ref/SHA, workflow revision, expanded cache key, restore-key
list, cache-hit, cache ID/version/ref/size,
runner/toolchain and the first failing dependency/build log. Then
determine whether the failure belongs to key selection,
scope/version matching, cache contents, package-manager
reconciliation, runner/toolchain, network or the application itself.
This ordering matters because a rerun can create or access a different cache and make the original cause harder to prove.
2. Failure A — key omits the lockfile/toolchain/OS
A key such as dependencies is stable precisely when the
dependency contract changes. An exact hit under that key is
misleading: it proves only that the stale key exists.
# Broken: different lockfiles/toolchains collapse into one exact key.
- id: cache
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
with:
path: ${{ steps.pip.outputs.cache_dir }}
key: dependencies
Repair by deriving the exact key from the dimensions that can change cached bytes. Preserve the old key and failing run in the incident note; do not rewrite history by only showing the repaired run.
3. Failure B — cache contains secrets or trust-sensitive executable state
If a cache path includes credentials, private configuration or executable hooks from a low-trust workflow, treat it as a potential credential/data exposure or poisoning incident. Stop relying on the entry, rotate/revoke any exposed credentials, preserve exact cache identity and audit which refs could restore it. Masking has no effect on bytes inside the cache.
Do not “fix” this by encrypting a secret with a key stored in the same workflow. Sensitive data belongs in the appropriate secret system and should not be cached.
4. Failure C — using cache-hit as proof that installation can be skipped
This intentionally broken example caches pip's download store, then skips dependency installation on an exact hit. The first run may succeed because it installs packages; the next fresh runner gets an exact cache hit but no installed environment, so application import fails.
- id: cache
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9
with:
path: ${{ steps.pip.outputs.cache_dir }}
key: pip-${{ runner.os }}-${{ steps.lock.outputs.sha256 }}
# Broken for a download cache: a hit does not populate .lab-site.
- name: Install dependencies
if: steps.cache.outputs.cache-hit != 'true'
run: python -m pip install --target .lab-site -r requirements.lock
- name: Verify fresh-runner environment
run: PYTHONPATH="$PWD/.lab-site" python -S -c 'import idna'
The cache service worked. The workflow's meaning was wrong: it treated cached downloads as an installed environment. Repair the dataflow by always reconciling the current lockfile, not by increasing cache permissions or retrying the restore.
5. Failure D — overly broad restore key
A restore prefix such as pip- can select entries from
an incompatible Python version or architecture if the scope/version
permits it. The action log may show a successful restore and
cache-hit=false, yet the restored state can be useless
or risky.
Repair by narrowing restore prefixes to compatibility boundaries such as OS/architecture/toolchain, or remove fallback entirely when consumers cannot safely reconcile stale state.
6. Failure E — generated build output hides environment dependence
A cached generated directory can depend on flags or SDKs that never appear in the key. Symptoms include intermittent test failures, binaries that differ from clean builds or machine-specific absolute paths. Compare a clean-cache build with the cached build and record toolchain/configuration digests before deciding the cache service is corrupt.
7. Failure F — stale self-hosted runner or cache action runtime
Current actions/cache v6 runs on Node 24. A self-hosted
runner that cannot execute current Node 24 actions fails before
cache semantics are even relevant. Preserve runner version and
action SHA, update the disposable/staged runner, and rerun the
smallest equivalent scope. Do not pin an obsolete insecure cache
action indefinitely as the “fix.”
8. Reconcile workflow logs with cache-service metadata
GH_REPO='OWNER/gha-cache-lab'
# Read-only evidence. Filter the output locally by exact key/ref from the failed run.
gh api -H 'X-GitHub-Api-Version: 2026-03-10' "repos/$GH_REPO/actions/caches?per_page=100" --jq '.actions_caches[] | {id,ref,key,version,size_in_bytes,created_at,last_accessed_at}'
Use the metadata to answer: Did an entry exist? Was it in the expected ref scope? Did its version match the requested path/archive format? Was the entry old or unexpectedly large? Only then decide whether to bypass, re-key or delete it.
9. Least-destructive repair order
- Preserve first-failure logs and cache metadata.
- Prove a clean-cache run can reconstruct dependencies.
- Fix the key or restore strategy if it omitted correctness inputs.
- Fix consumer behavior so restored state is validated/reconciled.
- Rotate credentials if sensitive content entered a cache.
- Delete only exact poisoned/stale lab entries after evidence is retained.
- Rerun the smallest equivalent workflow and compare source SHA, key, hit state and output.
Knowledge check
An exact cache hit is followed by
ModuleNotFoundError on a fresh runner. What should
you inspect first?
Whether the cached path actually contains an installed environment or only package-manager downloads, and whether the workflow wrongly skipped installation based on cache-hit.
Why preserve the cache ID/version/ref before deletion?
Those fields can prove which entry was restored and whether key/version/scope—not random network behavior—caused the failure.
What is the correct response if a secret was cached?
Treat it as exposure: stop using the cache, preserve exact evidence, rotate/revoke the credential and remove only the affected cache after the exposure path is understood.
Why is retrying a restore not a strong fix for a stale key?
The same key contract remains wrong. A retry may simply restore the same incompatible entry and erase the diagnostic value of the first run.
What does a Node/runtime error on self-hosted infrastructure suggest?
The failure may be runner/action-runtime compatibility rather than cache matching; record runner version and update the runner instead of downgrading security indefinitely.
Official references and version notes
- GitHub Docs — Dependency caching — cache purpose, artifact distinction and cache-security boundary.
- GitHub Docs — Dependency caching reference — key matching, restore keys, branch scope, versioning, limits and eviction.
- GitHub Docs — Managing caches — read-only inspection and exact cache deletion.
-
actions/cache v6.1.0
— explicit cache action pinned to
55cc8345863c7cc4c66a329aec7e433d2d1c52a9. -
actions/setup-python v7.0.0
— Python setup action pinned to
5fda3b95a4ea91299a34e894583c3862153e4b97. -
setup-python dependency caching
— built-in pip/pipenv/Poetry caching and
cache-dependency-path.
Version-sensitive behavior was rechecked on
2026-09-09. Mandatory examples target GitHub.com
and ubuntu-24.04. actions/cache v6.1.0
is pinned by full commit SHA and runs on Node 24; keep self-hosted
runners current enough to execute Node 24 actions (GitHub
documents runner 2.327.1 as the Node 24 minimum in current
official action guidance). Cache entries are scoped by key, cache
version and ref/branch visibility. Exact primary-key matches
report cache-hit == 'true'; prefix/restore-key
matches report 'false'; a total miss yields an empty
value. The cache action's post-save runs on successful jobs; do
not rely on deprecated save-always. GitHub currently
defaults cache inactivity retention to 7 days and repository cache
storage to 10 GB; eligible repositories/organizations can
configure higher limits, so performance/storage assumptions must
be recorded rather than hard-coded.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.