Chapter 14Lesson 02~205 minutes

Dependency Caching, Cache Keys, Restore Strategies, and Performance: Guided Hands-On Workflow

This lesson builds a disposable Python dependency-cache workflow. You will record the lockfile hash and toolchain, restore the pip download cache with a precise key, compare cold and warm runs, force a lockfile-driven miss, observe a restore-key fallback, and prove that installation still validates the dependency contract.

pip cacheCold/warm runsLock hashcache-hitMeasured duration

Learning objectives

  • Create a deterministic Python dependency input using exact versions and a recorded lockfile SHA-256.
  • Cache the pip download directory with OS/architecture/toolchain/lockfile dimensions using a pinned cache action.
  • Measure cold and warm install durations without asserting that a warm run must be faster.
  • Force a new precise key and observe safe restore-key fallback behavior.
  • Verify that dependency installation still runs and satisfies the current lockfile after exact or fallback cache restoration.

1. Disposable scenario and assumptions

Create a throwaway GitHub.com repository named gha-cache-lab. The mandatory path uses only workflow_dispatch, GitHub-hosted ubuntu-24.04, Python 3.13, PyPI public packages and the GitHub cache service. No repository write permission, secret, cloud account, container registry or self-hosted runner is required.

The workflow intentionally generates a tiny lockfile from a typed input so you can repeat the experiment without committing several throwaway dependency revisions. In production, the equivalent lockfile should normally be source-controlled and reviewed.

2. Progressive workflow: precise key, safe fallback and measured install

The workflow creates one of two exact-version dependency sets, records the lockfile hash, pins Python, discovers the actual pip cache directory, constructs a precise cache key and restores that directory. It then runs pip install regardless of cache-hit state because the cached path is only pip's download/wheel cache.

name: dependency-cache-lab
on:
  workflow_dispatch:
    inputs:
      lock_variant:
        description: Dependency input to model
        required: true
        type: choice
        options: [v1, v2]
        default: v1
permissions: {}

jobs:
  cache-lab:
    runs-on: ubuntu-24.04
    steps:
      - name: Pin Python toolchain
        uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
        with:
          python-version: '3.13'

      - id: lock
        name: Create and hash synthetic lockfile
        shell: bash
        env:
          LOCK_VARIANT: ${{ inputs.lock_variant }}
        run: |
          if [[ "$LOCK_VARIANT" == 'v1' ]]; then
            printf 'idna==3.10
' > requirements.lock
          else
            printf 'idna==3.10
packaging==25.0
' > requirements.lock
          fi
          sha="$(sha256sum requirements.lock | awk '{print $1}')"
          echo "sha256=$sha" >> "$GITHUB_OUTPUT"
          cat requirements.lock

      - id: pip
        name: Inspect Python and pip cache inputs
        shell: bash
        run: |
          cache_dir="$(python -m pip cache dir)"
          minor="$(python -c 'import sys; print(f"{sys.version_info.major}.{sys.version_info.minor}")')"
          echo "cache_dir=$cache_dir" >> "$GITHUB_OUTPUT"
          echo "minor=$minor" >> "$GITHUB_OUTPUT"
          python --version
          python -m pip --version
          printf 'pip_cache_dir=%s
' "$cache_dir"

      - id: cache
        name: Restore pip download cache
        uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
        with:
          path: ${{ steps.pip.outputs.cache_dir }}
          key: >-
            pip-${{ runner.os }}-${{ runner.arch }}-py${{ steps.pip.outputs.minor }}-${{ steps.lock.outputs.sha256 }}
          restore-keys: |
            pip-${{ runner.os }}-${{ runner.arch }}-py${{ steps.pip.outputs.minor }}-

      - id: install
        name: Install current dependency contract and measure it
        shell: bash
        run: |
          rm -rf .lab-site
          start="$(python -c 'import time; print(time.time_ns())')"
          python -m pip install             --disable-pip-version-check             --no-input             --target .lab-site             -r requirements.lock
          end="$(python -c 'import time; print(time.time_ns())')"
          ms="$(( (end - start) / 1000000 ))"
          echo "install_ms=$ms" >> "$GITHUB_OUTPUT"
          printf 'install_ms=%s
' "$ms"

      - name: Verify the lockfile was satisfied
        shell: bash
        run: |
          PYTHONPATH="$PWD/.lab-site" python -S - <<'PY'
          import idna
          print('idna', idna.__version__)
          try:
              import packaging
              print('packaging', packaging.__version__)
          except ModuleNotFoundError:
              print('packaging not requested in v1')
          PY

      - name: Record cache evidence
        shell: bash
        env:
          CACHE_HIT: ${{ steps.cache.outputs.cache-hit }}
          LOCK_SHA: ${{ steps.lock.outputs.sha256 }}
          INSTALL_MS: ${{ steps.install.outputs.install_ms }}
        run: |
          {
            echo '### Dependency cache evidence'
            echo "- run: $GITHUB_RUN_ID attempt $GITHUB_RUN_ATTEMPT"
            echo "- source SHA: $GITHUB_SHA"
            echo "- lock SHA-256: $LOCK_SHA"
            echo "- cache-hit: ${CACHE_HIT:-<empty>}"
            echo "- install ms: $INSTALL_MS"
            echo "- runner: $RUNNER_OS / $RUNNER_ARCH"
          } >> "$GITHUB_STEP_SUMMARY"

3. Run sequence: interpret state before chasing speed

Run Input Expected cache observation Why
A lock_variant=v1 usually miss on a fresh lab; cache-hit empty no precise key exists yet; successful job can save it
B lock_variant=v1 again exact hit; cache-hit=true same OS/arch/Python/lock hash
C lock_variant=v2 new primary key; may restore v1 by prefix and report false lock hash changed but compatible pip download state can accelerate reconciliation
D lock_variant=v2 again exact v2 hit after successful C new precise key should now exist

Do not assert that B or D must be faster. Network conditions, cache transfer size, CDN locality, pip behavior and runner load can dominate a tiny example. The evidence is the measured duration plus hit state, not a predetermined performance claim.

4. Why restore-key fallback is safe here

The fallback is constrained to the same OS, architecture and Python major/minor. It may contain wheels downloaded for the v1 lockfile. When v2 adds packaging==25.0, pip still reads the current lockfile and downloads anything missing. The restored directory can accelerate work, but cannot replace dependency reconciliation.

If instead the cached path were a fully assembled virtual environment and the workflow skipped installation on fallback, the stale state could become a correctness bug. Lesson 3 compares that trade-off directly.

5. Inspect service metadata read-only

After two or more runs, inspect cache records to connect the workflow's expanded key to GitHub's cache state.

GH_REPO='OWNER/gha-cache-lab'

gh api   -H 'X-GitHub-Api-Version: 2026-03-10'   "repos/$GH_REPO/actions/caches?per_page=100"   --jq '.actions_caches[] | {id,ref,key,version,size_in_bytes,created_at,last_accessed_at}'

Record the cache ID/key/version/ref, but do not delete the first entry until you have compared the cold/warm and fallback runs.

6. Trust and permission boundary

The mandatory workflow is manual and contains no untrusted fork code. That keeps the lab focused on cache semantics. In pull-request workflows, restored cache contents must be treated as untrusted input, and low-trust runs may be restore-only. Do not grant write permissions merely to “make cache save work.” The cache service uses its own runtime token; repository GITHUB_TOKEN permissions are not a reason to use write-all.

7. Mini challenge: choose the missing key dimension

Suppose a project compiles native wheels differently on x64 and arm64 but uses the same lockfile and Python version. Which layer should change? Add runner.arch to the primary key rather than adding an arbitrary cache epoch or a broad restore prefix. Then predict how many distinct exact keys two architectures will create.

Knowledge check

Why does Run C report cache-hit=false when a cache was restored?

Why does the workflow always run pip install?

What should be compared between cold and warm runs?

Why include OS, architecture and Python version in the key?

When should you delete the lab cache?

Next lesson

Choose the cache architecture intentionally

Lesson 3 compares explicit cache actions with setup-action caching, exact keys with prefixes, and dependency caches with riskier build-output caches.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked on 2026-09-09. Mandatory examples target GitHub.com and ubuntu-24.04. actions/cache v6.1.0 is pinned by full commit SHA and runs on Node 24; keep self-hosted runners current enough to execute Node 24 actions (GitHub documents runner 2.327.1 as the Node 24 minimum in current official action guidance). Cache entries are scoped by key, cache version and ref/branch visibility. Exact primary-key matches report cache-hit == 'true'; prefix/restore-key matches report 'false'; a total miss yields an empty value. The cache action's post-save runs on successful jobs; do not rely on deprecated save-always. GitHub currently defaults cache inactivity retention to 7 days and repository cache storage to 10 GB; eligible repositories/organizations can configure higher limits, so performance/storage assumptions must be recorded rather than hard-coded. The tiny PyPI dependency set is synthetic and public; PyPI availability is an external network dependency, so the checkpoint also includes a local simulation path.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.