Chapter 17Lesson 01~180 minutes

Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data: Concepts, Architecture, and Mental Model

Chapter 16 treated runners as the compute trust plane. Chapter 17 focuses on the data those runners create, reuse, and leave behind. A cache is not an artifact, a workflow artifact is not a release asset, a job summary is not a test database, and none of these should be assumed safe merely because GitHub stores them.

Artifacts vs cachesLogs & summariesRetentionData sensitivity

Learning objectives

  • Separate workflow artifacts, dependency caches, release assets, packages, logs, annotations, summaries, and machine-readable test evidence by purpose and lifecycle.
  • Explain cache keys, restore-key fallback, branch/ref scope, cache immutability, invalidation, eviction, and cache-poisoning risk.
  • Explain artifact naming, upload immutability, SHA-256 artifact digests, compression, retention, access, and the difference between an artifact archive and the files inside it.
  • Choose what belongs in logs, annotations, job summaries, and structured reports without dumping secrets or excessive context.
  • Reason about evidence retention and storage cost as production controls rather than cleanup chores.

Availability: The mandatory chapter path uses GitHub.com, GitHub Free, a disposable public personal repository, and standard GitHub-hosted runners. Public-repository artifact/log retention can currently be configured from 1 to 90 days. Cache-limit customization is a separate opt-in capability that currently requires a payment method on file or an eligible paid plan; the lab does not require it.

1. The problem: CI creates several kinds of “stored data,” but they are not interchangeable

A workflow can become unreliable even when every command is syntactically correct. One job expects a file that only existed on another runner. A cache restores dependencies for the wrong lockfile. A test report disappears before an incident review. A log contains a transformed credential that secret masking did not recognize. Or a release team downloads a workflow artifact and treats it as if it were a governed package.

The fix is a lifecycle model. For each piece of data ask: why does it exist, who produces it, who may consume it, how is it identified, how long should it exist, can it be regenerated, and what happens if an attacker controls it?

2. Mental model: regeneration data, run evidence, and distribution data follow different lifecycles

Concept / workflow diagram
              flowchart TD
                SRC["Source SHA + lockfile"] --> C["Dependency cache performance hint"]
                SRC --> JOB["Workflow job"]
                C --> JOB
                JOB --> A["Workflow artifact run-scoped file evidence"]
                JOB --> L["Logs + annotations + summary execution evidence"]
                JOB --> T["Machine-readable test report"]
                A --> NEXT["Later job / operator"]
                T --> TOOL["Reporter / downstream automation"]
                A --> PUB["Promotion decision"]
                PUB --> P["Package or release asset versioned distribution"]
            

A cache helps recreate a build and must be safely disposable. Workflow artifacts and test reports preserve evidence from one run. Logs/summaries explain the run. Packages and release assets are longer-lived distribution objects and should be promoted deliberately rather than confused with temporary CI storage.

The arrows are data-flow relationships, not permission inheritance. A file uploaded as an Actions artifact does not become a Release asset or package automatically. Likewise, restoring a cache does not prove the cached bytes are correct; the workflow must be able to regenerate or validate them.

3. Choose the storage object from the operational question

Object Primary purpose Identity/lifecycle Production rule
Dependency cache Avoid re-downloading/recomputing reusable inputs Key + cache version + ref scope; evictable A miss must be survivable. Treat restored bytes as untrusted input.
Workflow artifact Pass files between jobs or preserve run output Artifact ID/name/digest; tied to run/repository; expires Use deterministic names and retention; never assume it is a release channel.
Release asset Attach downloadable files to a GitHub Release Release/tag/version governance Use when the version is intentionally published to consumers.
Package Distribute package/container versions through a registry Package/version/digest + registry permissions Use for dependency consumption and promotion policy.
Workflow log Detailed execution trace Run/job/step; retained with Actions data Keep diagnostically useful, not secret-rich or endlessly verbose.
Annotation Point a warning/error/notice to run or file context Run-scoped UI evidence Use for actionable findings, not bulk reports.
Job summary Human-readable Markdown overview Per-step summary files grouped per job Surface decisions/results; link to deeper evidence.
Test report Machine-readable result such as JSON/JUnit/SARIF-like data Schema + source SHA + test/tool version Preserve enough structure for automation and audit.

4. Cache keys are a reproducibility contract, not a folder backup

actions/cache first searches for an exact key match. If there is no exact match, it can use prefix matching and ordered restore-keys; when no suitable cache exists and the job succeeds, it can save a new cache under the requested key. Existing cache content is not edited in place—changing dependency inputs should produce a new key.

- name: Restore synthetic dependency cache
  uses: actions/cache@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
  id: dependency-cache
  with:
    path: .chapter17-cache
    key: c17-${{ runner.os }}-${{ hashFiles('dependency.lock') }}
    restore-keys: |
      c17-${{ runner.os }}-

The content-derived hash is the important part: a lockfile change changes the exact key. Broad restore keys can be useful for performance, but a partial match may intentionally return older compatible data. The consumer therefore must still validate or regenerate what correctness requires.

Security boundary: GitHub explicitly warns that caches can be read by workflows with access to their scope and that restored cache contents are not signed or verified. Do not cache credentials. Treat restored files as untrusted. GitHub now restricts default-branch cache writes for low-trust trigger types, but that platform protection does not make cache contents inherently trustworthy.

5. Cache scope and eviction explain “why did I miss?”

Cache lookup is scoped by branch/tag/ref rules. A workflow can restore caches from its current branch and the default branch when allowed; pull-request caches are created against the pull-request merge ref and are intentionally limited. Different workflows in the same repository can share a cache when they fall inside the same scope, which is another reason keys need an application/runtime/dependency namespace.

Current default cache management removes entries not accessed for more than seven days and uses a default repository cache-size limit of 10 GB, with oldest-accessed entries evicted as space is needed. These are operational limits, not durability guarantees. A build that cannot tolerate a cache eviction is incorrectly designed.

6. Workflow artifacts preserve run output and are immutable creation events

Current actions/upload-artifact versions create an artifact with an ID, URL, and SHA-256 artifact digest. Since the v4 artifact architecture, an artifact is not incrementally appended by later jobs: names must be unique for separate creations, and an overwrite operation means delete-and-create semantics rather than mutation of the same evidence object.

- name: Upload deterministic report
  id: upload
  uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
  with:
    name: c17-report-${{ github.run_id }}
    path: dist/report.txt
    if-no-files-found: error
    retention-days: 5
    compression-level: 6

The upload action currently exposes artifact-id, artifact-url, and artifact-digest. The digest identifies the uploaded artifact representation. If you also care about a build file itself, compute and record that file’s digest before upload, then verify it after download.

Current upload-artifact v7 can also upload one file unzipped with archive: false. That is useful when direct browser viewing or preserving a pre-compressed file is the goal, but the mandatory lab uses the default archive behavior so learners see normal artifact extraction.

7. Naming, compression, retention, and limits are part of the evidence contract

Artifact names should describe source/run/purpose without embedding secrets. Compression trades CPU time for storage/transfer size. The current upload action supports compression levels 0–9 and currently limits a single job to 500 artifact creations. Zipped uploads do not preserve Unix executable permission bits; when permissions matter, package the files deliberately (for example into a tar archive) before upload.

GitHub currently retains workflow artifacts and logs for 90 days by default. Public repositories can configure 1–90 days; private repositories can configure 1–400 days, subject to organization/enterprise maximums. Changing a repository retention setting applies to new artifacts/logs, not retroactively to existing ones. A per-artifact retention-days value cannot exceed the enclosing policy.

8. Logs, annotations, summaries, and test reports answer different questions

Logs are a chronological diagnostic trace. An annotation is a small actionable marker such as a warning tied to a file/line. A job summary is Markdown written through GITHUB_STEP_SUMMARY and shown prominently on the run page. A machine-readable report is structured data another tool can reliably parse.

printf '%s
' '::notice file=dependency.lock,line=1::Chapter 17 dependency input was validated'
printf '%s
' '### Chapter 17 evidence' >> "$GITHUB_STEP_SUMMARY"
printf '%s
' '- Cache state: recorded separately' >> "$GITHUB_STEP_SUMMARY"
printf '%s
' '- Test report: test-results/result.json' >> "$GITHUB_STEP_SUMMARY"

Do not dump whole github, env, runner, or secret contexts to make “debugging easier.” Good observability selects stable fields—source SHA, run ID, cache hit/miss, artifact ID/digest, tool version, and test counts—and omits unrelated environment data.

9. Data sensitivity: storage is not a sanitization boundary

Artifacts and caches can accidentally collect .env files, cloud configuration, signing material, browser profiles, debug dumps, or transformed credentials. The current upload action does not include hidden files by default, which reduces one class of mistake, but it is not a substitute for an allowlist.

Secret masking is also not a general data-loss prevention system. A secret may be transformed, encoded, partially printed, split across lines, or embedded in a structured dump in a form that exact-value masking does not catch. Prevent sensitive output at the source; do not rely on redaction after logging.

10. Read-only inspection: prove what exists before creating more data

REPO="OWNER/REPOSITORY"
RUN_ID="123456789"

gh run view "$RUN_ID" -R "$REPO"   --json headSha,status,conclusion,createdAt,startedAt,updatedAt,url

gh api -H "X-GitHub-Api-Version: 2026-03-10"   "repos/$REPO/actions/runs/$RUN_ID/artifacts?per_page=100"   --jq '.artifacts[] | {id,name,size_in_bytes,expired,created_at,expires_at,digest}'

gh cache list -R "$REPO" --limit 30   --json id,key,ref,sizeInBytes,createdAt,lastAccessedAt

gh api -H "X-GitHub-Api-Version: 2026-03-10"   "repos/$REPO/actions/permissions/artifact-and-log-retention"

These are different inventories: the run owns artifact evidence, while caches are repository/ref-scoped acceleration data. The retention endpoint shows policy, not the expiration date of a particular artifact; use the artifact API’s expires_at for the object itself.

11. DevOps connection: CI data is evidence and billable storage

Release decisions depend on trustworthy evidence. If the only test output is a noisy log that expires before an incident, the organization cannot reconstruct what was validated. If caches churn due over-granular keys, latency and storage cost rise. If retention is excessive, sensitive and obsolete build output accumulates. Data lifecycle therefore belongs in the same operational review as workflow permissions and runner selection.

12. Lesson summary

Caches are disposable performance hints keyed to reproducible inputs. Workflow artifacts preserve files from a run and can bridge job isolation. Logs, annotations, summaries, and structured reports expose different views of execution evidence. Packages and release assets are separate publication objects. Retention, naming, digest identity, and sensitivity are part of the control plane—not post-processing details.

Knowledge check

A build fails when the dependency cache is evicted. What is the design defect?

Why is an artifact digest not automatically the same as the digest of the file you built?

A pull-request workflow can read a cache from the base/default branch. Should cached executables be trusted automatically?

What is the difference between a job summary and a test report?

You reduce repository retention from 90 to 14 days. Do existing artifacts instantly adopt 14 days?

Next lesson

Next: Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data: Guided Hands-On Workflow and Core Operations

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.