Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data: Concepts, Architecture, and Mental Model
Chapter 16 treated runners as the compute trust plane. Chapter 17 focuses on the data those runners create, reuse, and leave behind. A cache is not an artifact, a workflow artifact is not a release asset, a job summary is not a test database, and none of these should be assumed safe merely because GitHub stores them.
Learning objectives
- Separate workflow artifacts, dependency caches, release assets, packages, logs, annotations, summaries, and machine-readable test evidence by purpose and lifecycle.
- Explain cache keys, restore-key fallback, branch/ref scope, cache immutability, invalidation, eviction, and cache-poisoning risk.
- Explain artifact naming, upload immutability, SHA-256 artifact digests, compression, retention, access, and the difference between an artifact archive and the files inside it.
- Choose what belongs in logs, annotations, job summaries, and structured reports without dumping secrets or excessive context.
- Reason about evidence retention and storage cost as production controls rather than cleanup chores.
Availability: The mandatory chapter path uses GitHub.com, GitHub Free, a disposable public personal repository, and standard GitHub-hosted runners. Public-repository artifact/log retention can currently be configured from 1 to 90 days. Cache-limit customization is a separate opt-in capability that currently requires a payment method on file or an eligible paid plan; the lab does not require it.
1. The problem: CI creates several kinds of “stored data,” but they are not interchangeable
A workflow can become unreliable even when every command is syntactically correct. One job expects a file that only existed on another runner. A cache restores dependencies for the wrong lockfile. A test report disappears before an incident review. A log contains a transformed credential that secret masking did not recognize. Or a release team downloads a workflow artifact and treats it as if it were a governed package.
The fix is a lifecycle model. For each piece of data ask: why does it exist, who produces it, who may consume it, how is it identified, how long should it exist, can it be regenerated, and what happens if an attacker controls it?
2. Mental model: regeneration data, run evidence, and distribution data follow different lifecycles
flowchart TD
SRC["Source SHA + lockfile"] --> C["Dependency cache performance hint"]
SRC --> JOB["Workflow job"]
C --> JOB
JOB --> A["Workflow artifact run-scoped file evidence"]
JOB --> L["Logs + annotations + summary execution evidence"]
JOB --> T["Machine-readable test report"]
A --> NEXT["Later job / operator"]
T --> TOOL["Reporter / downstream automation"]
A --> PUB["Promotion decision"]
PUB --> P["Package or release asset versioned distribution"]
A cache helps recreate a build and must be safely disposable. Workflow artifacts and test reports preserve evidence from one run. Logs/summaries explain the run. Packages and release assets are longer-lived distribution objects and should be promoted deliberately rather than confused with temporary CI storage.
The arrows are data-flow relationships, not permission inheritance. A file uploaded as an Actions artifact does not become a Release asset or package automatically. Likewise, restoring a cache does not prove the cached bytes are correct; the workflow must be able to regenerate or validate them.
3. Choose the storage object from the operational question
| Object | Primary purpose | Identity/lifecycle | Production rule |
|---|---|---|---|
| Dependency cache | Avoid re-downloading/recomputing reusable inputs | Key + cache version + ref scope; evictable | A miss must be survivable. Treat restored bytes as untrusted input. |
| Workflow artifact | Pass files between jobs or preserve run output | Artifact ID/name/digest; tied to run/repository; expires | Use deterministic names and retention; never assume it is a release channel. |
| Release asset | Attach downloadable files to a GitHub Release | Release/tag/version governance | Use when the version is intentionally published to consumers. |
| Package | Distribute package/container versions through a registry | Package/version/digest + registry permissions | Use for dependency consumption and promotion policy. |
| Workflow log | Detailed execution trace | Run/job/step; retained with Actions data | Keep diagnostically useful, not secret-rich or endlessly verbose. |
| Annotation | Point a warning/error/notice to run or file context | Run-scoped UI evidence | Use for actionable findings, not bulk reports. |
| Job summary | Human-readable Markdown overview | Per-step summary files grouped per job | Surface decisions/results; link to deeper evidence. |
| Test report | Machine-readable result such as JSON/JUnit/SARIF-like data | Schema + source SHA + test/tool version | Preserve enough structure for automation and audit. |
4. Cache keys are a reproducibility contract, not a folder backup
actions/cache first searches for an exact key match. If
there is no exact match, it can use prefix matching and ordered
restore-keys; when no suitable cache exists and the job
succeeds, it can save a new cache under the requested key. Existing
cache content is not edited in place—changing dependency inputs
should produce a new key.
- name: Restore synthetic dependency cache
uses: actions/cache@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
id: dependency-cache
with:
path: .chapter17-cache
key: c17-${{ runner.os }}-${{ hashFiles('dependency.lock') }}
restore-keys: |
c17-${{ runner.os }}-
The content-derived hash is the important part: a lockfile change changes the exact key. Broad restore keys can be useful for performance, but a partial match may intentionally return older compatible data. The consumer therefore must still validate or regenerate what correctness requires.
Security boundary: GitHub explicitly warns that caches can be read by workflows with access to their scope and that restored cache contents are not signed or verified. Do not cache credentials. Treat restored files as untrusted. GitHub now restricts default-branch cache writes for low-trust trigger types, but that platform protection does not make cache contents inherently trustworthy.
5. Cache scope and eviction explain “why did I miss?”
Cache lookup is scoped by branch/tag/ref rules. A workflow can restore caches from its current branch and the default branch when allowed; pull-request caches are created against the pull-request merge ref and are intentionally limited. Different workflows in the same repository can share a cache when they fall inside the same scope, which is another reason keys need an application/runtime/dependency namespace.
Current default cache management removes entries not accessed for more than seven days and uses a default repository cache-size limit of 10 GB, with oldest-accessed entries evicted as space is needed. These are operational limits, not durability guarantees. A build that cannot tolerate a cache eviction is incorrectly designed.
6. Workflow artifacts preserve run output and are immutable creation events
Current actions/upload-artifact versions create an
artifact with an ID, URL, and SHA-256 artifact digest. Since the v4
artifact architecture, an artifact is not incrementally appended by
later jobs: names must be unique for separate creations, and an
overwrite operation means delete-and-create semantics rather than
mutation of the same evidence object.
- name: Upload deterministic report
id: upload
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: c17-report-${{ github.run_id }}
path: dist/report.txt
if-no-files-found: error
retention-days: 5
compression-level: 6
The upload action currently exposes artifact-id,
artifact-url, and artifact-digest. The
digest identifies the uploaded artifact representation. If you also
care about a build file itself, compute and record that file’s
digest before upload, then verify it after download.
Current upload-artifact v7 can also upload one file unzipped with
archive: false. That is useful when direct browser
viewing or preserving a pre-compressed file is the goal, but the
mandatory lab uses the default archive behavior so learners see
normal artifact extraction.
7. Naming, compression, retention, and limits are part of the evidence contract
Artifact names should describe source/run/purpose without embedding secrets. Compression trades CPU time for storage/transfer size. The current upload action supports compression levels 0–9 and currently limits a single job to 500 artifact creations. Zipped uploads do not preserve Unix executable permission bits; when permissions matter, package the files deliberately (for example into a tar archive) before upload.
GitHub currently retains workflow artifacts and logs for 90 days by
default. Public repositories can configure 1–90 days; private
repositories can configure 1–400 days, subject to
organization/enterprise maximums. Changing a repository retention
setting applies to new artifacts/logs, not retroactively to existing
ones. A per-artifact retention-days value cannot exceed
the enclosing policy.
8. Logs, annotations, summaries, and test reports answer different questions
Logs are a chronological diagnostic trace. An annotation is a small
actionable marker such as a warning tied to a file/line. A job
summary is Markdown written through
GITHUB_STEP_SUMMARY and shown prominently on the run
page. A machine-readable report is structured data another tool can
reliably parse.
printf '%s
' '::notice file=dependency.lock,line=1::Chapter 17 dependency input was validated'
printf '%s
' '### Chapter 17 evidence' >> "$GITHUB_STEP_SUMMARY"
printf '%s
' '- Cache state: recorded separately' >> "$GITHUB_STEP_SUMMARY"
printf '%s
' '- Test report: test-results/result.json' >> "$GITHUB_STEP_SUMMARY"
Do not dump whole github, env, runner, or
secret contexts to make “debugging easier.” Good observability
selects stable fields—source SHA, run ID, cache hit/miss, artifact
ID/digest, tool version, and test counts—and omits unrelated
environment data.
9. Data sensitivity: storage is not a sanitization boundary
Artifacts and caches can accidentally collect
.env files, cloud configuration, signing material,
browser profiles, debug dumps, or transformed credentials. The
current upload action does not include hidden files by default,
which reduces one class of mistake, but it is not a substitute for
an allowlist.
Secret masking is also not a general data-loss prevention system. A secret may be transformed, encoded, partially printed, split across lines, or embedded in a structured dump in a form that exact-value masking does not catch. Prevent sensitive output at the source; do not rely on redaction after logging.
10. Read-only inspection: prove what exists before creating more data
REPO="OWNER/REPOSITORY"
RUN_ID="123456789"
gh run view "$RUN_ID" -R "$REPO" --json headSha,status,conclusion,createdAt,startedAt,updatedAt,url
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/runs/$RUN_ID/artifacts?per_page=100" --jq '.artifacts[] | {id,name,size_in_bytes,expired,created_at,expires_at,digest}'
gh cache list -R "$REPO" --limit 30 --json id,key,ref,sizeInBytes,createdAt,lastAccessedAt
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/permissions/artifact-and-log-retention"
These are different inventories: the run owns artifact evidence,
while caches are repository/ref-scoped acceleration data. The
retention endpoint shows policy, not the expiration date of a
particular artifact; use the artifact API’s
expires_at for the object itself.
11. DevOps connection: CI data is evidence and billable storage
Release decisions depend on trustworthy evidence. If the only test output is a noisy log that expires before an incident, the organization cannot reconstruct what was validated. If caches churn due over-granular keys, latency and storage cost rise. If retention is excessive, sensitive and obsolete build output accumulates. Data lifecycle therefore belongs in the same operational review as workflow permissions and runner selection.
12. Lesson summary
Caches are disposable performance hints keyed to reproducible inputs. Workflow artifacts preserve files from a run and can bridge job isolation. Logs, annotations, summaries, and structured reports expose different views of execution evidence. Packages and release assets are separate publication objects. Retention, naming, digest identity, and sensitivity are part of the control plane—not post-processing details.
Knowledge check
A build fails when the dependency cache is evicted. What is the design defect?
The cache has become a required source of truth. A correct build must be able to regenerate or redownload dependencies when the cache is absent.
Why is an artifact digest not automatically the same as the digest of the file you built?
The artifact digest identifies the uploaded artifact representation/archive. Record and verify a file-level digest separately when the file identity matters.
A pull-request workflow can read a cache from the base/default branch. Should cached executables be trusted automatically?
No. Cache contents are not signed/verified simply because GitHub stored them. Treat restored files as untrusted and design writes/keys/triggers accordingly.
What is the difference between a job summary and a test report?
A summary is human-oriented Markdown shown on the run page. A test report is structured evidence intended for reliable machine consumption or later analysis.
You reduce repository retention from 90 to 14 days. Do existing artifacts instantly adopt 14 days?
No. GitHub documents that customized retention applies to new artifacts/logs, not retroactively to existing objects.
Further reading — current official GitHub sources
- GitHub Docs — Workflow artifacts concepts
- GitHub Docs — Store and share data with workflow artifacts
- GitHub Docs — Dependency caching reference
- GitHub Docs — Managing caches
- GitHub Docs — Removing workflow artifacts and retention
- GitHub Docs — Workflow commands, annotations, and job summaries
- GitHub Docs — Repository Actions settings and retention
- GitHub REST — Actions artifacts
- GitHub REST — Actions cache
- GitHub REST — Actions permissions / retention
- GitHub CLI — gh cache list
- GitHub CLI — gh run download
- GitHub CLI — gh run view
- Official action — actions/upload-artifact
- Official action — actions/download-artifact
- Official action — actions/cache
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.