Chapter 17Lesson 03~160 minutes

Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data: Configuration, Design Choices, and Tradeoffs

Storage choices encode operational policy. Caches optimize regeneration; artifacts preserve run-scoped outputs; packages and release assets publish versioned distribution units; logs and summaries explain execution; machine-readable reports feed automation. This lesson turns those different lifecycles into an explicit data-governance decision.

Data lifecycleInvalidationNaming & retentionEvidence design

Learning objectives

  • Choose cache, workflow artifact, package, release asset, log, summary, or structured report from the data lifecycle rather than convenience.
  • Design cache keys and restore-key fallback that balance reuse, correctness, trust, and churn.
  • Design artifact naming/retention conventions that preserve provenance while controlling storage.
  • Separate human-readable test evidence from machine-consumable evidence and explain when both are required.
  • Evaluate plan/deployment differences, especially GitHub.com artifact backends versus GitHub Enterprise Server.

Availability: The required design work is plan-neutral and the mandatory examples use GitHub Free/public repositories. GitHub.com upload-artifact@v4+ uses the newer artifact backend; the current official upload-artifact documentation states that v4+ is not supported on GitHub Enterprise Server, which uses its server-supported artifact action line. GHES operators must follow the action/backend version documented for their server release.

1. Start with the lifecycle question, not the YAML keyword

Ask whether the data is required for correctness, regenerable, run-scoped evidence, or published distribution. That immediately eliminates common mistakes: a cache cannot be the only copy of a build; a run artifact should not become a permanent dependency registry; logs should not carry large binary output.

2. Cache versus artifact versus package/release asset

Decision factor Cache Workflow artifact Package / Release asset
Primary purpose Speed up regeneration Preserve/pass run output Publish a version to consumers
Expected on miss/expiry Build regenerates Evidence/output may be unavailable after retention Consumers expect governed availability
Identity Key + version + ref scope Artifact ID/name/digest + run Version/tag/package digest/release governance
Trust posture Restored bytes untrusted; validate Treat as build output; verify provenance/digest Promotion/signing/provenance policy may apply
Typical retention Eviction-driven/short Days to months by policy Release/package lifecycle
Good example npm/pip/compiler cache coverage report, built binary, debug bundle v2.4.0 package, release installer

3. Cache key granularity: correctness first, reuse second

A useful key normally includes the dimensions that make restored data compatible: operating system or architecture when relevant, toolchain/runtime major version, and a content hash of the lockfile or dependency manifest. Include too little and stale/incompatible content can contaminate builds. Include too much—such as the commit SHA—and every run misses, turning the cache into expensive storage with no acceleration.

# Balanced: OS + runtime generation + dependency content
key: deps-${{ runner.os }}-py3-${{ hashFiles('requirements.lock') }}
restore-keys: |
  deps-${{ runner.os }}-py3-

# Usually too broad for correctness:
# key: deps

# Usually too specific for reuse:
# key: deps-${{ github.sha }}

Restore keys are an explicit compatibility claim. If an older prefix-matched cache can safely seed downloads, use it. If restored bytes must match exactly, omit broad fallback and accept the miss.

4. Cache trust policy must follow trigger trust

Current GitHub cache protections reduce writes from low-trust trigger types into the default-branch scope, and pull-request caches are scoped to the PR merge ref. Still, a workflow should not execute cached scripts blindly. Prefer caching package-manager download stores or compile intermediates whose correctness is verified by manifests/build steps, not authorization tokens or opaque executable bootstrap code.

For privileged deployment workflows, a strong pattern is to restore only caches populated by trusted build triggers—or avoid cache consumption for security-sensitive material entirely.

5. Artifact naming and retention should answer “what run produced this?”

Names should be deterministic enough to understand and unique enough to avoid ambiguity: linux-amd64-test-report-${run_id} is easier to operate than artifact. Include matrix dimensions when several jobs upload similar outputs. Do not put secrets, customer identifiers, or private branch descriptions into names because names themselves become metadata.

Retention should be derived from the operational need: short-lived debugging output, longer-lived release-candidate evidence, or audit-required test records. If an audit policy requires evidence longer than Actions retention permits, export that evidence to an approved system of record rather than pretending a 90-day workflow artifact is permanent.

6. Compression and file semantics are performance choices

Text compresses well but costs CPU at higher compression levels. Already-compressed binaries often waste CPU when recompressed. Current upload-artifact v7 also supports direct single-file upload with archive: false. If Unix permissions/case semantics matter, bundle them deliberately before artifact upload because normal zipped artifact extraction does not preserve executable bits as a source-control/archive format would.

7. Human evidence and machine evidence should reinforce each other

Audience Surface What it should contain
Developer scanning run Job summary Counts, key failures, cache/artifact links, source SHA
Developer fixing code Annotation Small actionable warning/error with file/line when meaningful
Investigator Logs Chronological detailed trace with selected safe diagnostics
Automation / analytics JSON/JUnit/XML/other schema artifact Structured tests, tool/schema version, source identity

Do not force humans to parse megabytes of logs for a test count, and do not make machines scrape Markdown summaries. Publish both when the workflow serves both audiences.

8. Worked scenarios

Scenario Choice Why / tradeoff
Open-source library CI downloads dependencies repeatedly Content-keyed cache Fast, free public-runner path; miss remains safe.
Integration tests generate a 2 MB JUnit report Artifact + concise summary Machine evidence retained; humans see the outcome immediately.
Nightly build creates 5 GB debug dump Short retention + low/no compression + explicit sensitivity review Controls storage and upload CPU; avoid keeping huge evidence by default.
Release pipeline produces customer installer Promote to governed Release asset/package; keep CI artifact only as intermediate evidence Distribution lifecycle and rollback should not depend on run retention.
Privileged workflow wants cache written by issue-comment-triggered job Reject or use trusted restore-only strategy Prevent low-trust data from becoming privileged executable input.

9. Storage and cost policy: optimize the cause, not only the quota

Cache thrashing often means the key space is too fragmented or the repository is caching expensive but low-value directories. Artifact growth often means every run stores redundant binaries or debug data for too long. Before buying more storage, inspect cache count/size/last access and artifact retention/usefulness. Increase limits only when the workload justifies it.

Current cache defaults are a 7-day inactivity retention and 10 GB repository limit; customizable cache settings are opt-in and can raise retention/size on eligible accounts. Because these product limits can evolve, production runbooks should point operators to current settings/API instead of hard-coding a permanent capacity assumption.

10. Supply-chain rule: stored CI data is input when consumed later

The moment another job or later workflow restores a cache or downloads an artifact, stored output becomes input. Apply the same questions you would to a dependency: who produced it, from which SHA/event/actor, under what permissions, and can you verify its expected digest/schema? Artifact storage does not convert untrusted output into trusted code.

11. Lesson summary

Choose the surface from the data lifecycle. Cache keys express compatibility and trust. Artifacts express run-scoped evidence and transfer. Packages/releases express publication. Summaries and annotations serve humans; structured reports serve machines. Retention, compression, naming, cost, and sensitivity must follow the operational purpose.

Knowledge check

A cache key includes the commit SHA. What likely problem does that create?

When is a broad restore key dangerous?

Why should a JUnit/JSON report often accompany a job summary?

A release binary currently lives only as a 5-day workflow artifact. What governance improvement is needed?

Does GitHub.com upload-artifact v7 behavior automatically apply to every GHES deployment?

Next lesson

Next: Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data: Diagnostics, Failure Modes, Security, and Performance

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.