Chapter 13Lesson 01~155 minutes

Artifacts, Retention, Cross-Job Data, and Workflow Result Management: Core Concepts and Mental Model

Chapter 12 used per-cell artifacts as compatibility evidence. Chapter 13 makes that transport explicit: what exactly leaves a runner, what identity GitHub assigns it, how provenance and retention are recorded, and why an artifact is neither a cache nor a job output.

Workflow artifactsProvenanceArtifact IDDigestRetention

Learning objectives

  • Explain the artifact lifecycle from runner filesystem selection through GitHub storage, consumer download and expiry/deletion.
  • Distinguish artifact identity, artifact digest and file/content digest from job outputs and caches.
  • Record producing run/job/SHA, selected paths, sensitivity classification and retention as provenance evidence.
  • Explain hidden-file, compression and overwrite behavior without treating masking or file selection as a security boundary.
  • Choose artifacts only when durable cross-job/run evidence is the intended data transport.

1. The practical problem: the runner disappears but evidence must remain

GitHub-hosted jobs execute on fresh, temporary machines. A build may produce a binary, test report, screenshot, crash dump or manifest that must outlive that runner. Chapter 5 taught job outputs for small structured values; Chapter 12 used artifacts to retain matrix-cell evidence. This chapter defines the full artifact lifecycle so the stored object can be traced back to the exact run and commit that produced it.

An artifact is not “whatever is in the workspace.” It is an explicit selection of paths uploaded to GitHub's artifact service. That selection, its sensitivity, its producing SHA and run, and its retention policy are all part of the data contract.

2. Causal model: filesystem → selection → artifact record → consumer

First, a job creates files on its runner. The upload action expands the configured paths, applies hidden-file policy, packages/compresses the selected content when archival mode is used, and creates a run-scoped artifact record. GitHub assigns an artifact ID and digest. A later job, user or API consumer selects that record, downloads it, verifies integrity, and interprets it in the context of the producing run/SHA. Eventually repository policy expires it or an authorized actor deletes it.

Artifact lifecycle and provenance
flowchart TD
  A[Runner job filesystem] --> B[Explicit selected paths]
  B --> C[Hidden-file + compression policy]
  C --> D[Upload action / artifact service]
  D --> E[Artifact ID + name + digest + size]
  E --> F[Producing run / job / SHA]
  E --> G[Later job / user / REST API]
  G --> H[Download + digest validation]
  H --> I[Consumer verification / use]
  E --> J[Retention expiry or explicit deletion]

The arrows are trust boundaries. The runner decides what files exist. The workflow decides which paths are eligible. The artifact service gives the upload a durable identity. The consumer must still verify that the identity belongs to the expected run and SHA; a valid digest of the wrong run is still the wrong input.

3. Name the artifact state before moving data

State Question Evidence
Artifact name What human-readable label was assigned? workflow manifest + run UI
Artifact ID Which immutable service record is this? upload output / REST metadata
Producer Which run, attempt, job and SHA created it? run ID, attempt, head SHA, job logs
Selected paths Exactly which files were eligible for upload? workflow path patterns + pre-upload listing
Hidden-file policy Were dot-prefixed files excluded or explicitly included? action input + selected-file evidence
Compression/archive How was the payload packaged? action inputs + artifact metadata
Digest and size What object did the artifact service record? upload output + REST digest/size_in_bytes
Retention When is it expected to expire? retention-days + REST expires_at
Sensitivity May this content be retained or downloaded by repo readers? data classification / review
Consumer Which job/user/API downloaded it and why? download log / API audit evidence

4. Name, ID and digest solve different problems

The name is convenient routing metadata. The artifact ID identifies one artifact record in the repository. The artifact digest records the SHA-256 digest of the artifact payload. For a production evidence packet, keep all three plus the producing run ID and head SHA.

- name: Upload evidence
  id: upload
  uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
  with:
    name: test-evidence-${{ github.run_id }}
    path: evidence/
    if-no-files-found: error
    retention-days: 7
    include-hidden-files: false

- name: Record artifact identity
  shell: bash
  run: |
    printf 'artifact_id=%s
' '${{ steps.upload.outputs.artifact-id }}'
    printf 'artifact_digest=%s
' '${{ steps.upload.outputs.artifact-digest }}'
Digest is not provenance

A matching digest proves the downloaded artifact matches the service-recorded payload. It does not by itself prove who built the content, which source SHA was intended, or whether the workflow was trusted. Keep run/SHA evidence; for release-grade signed provenance, artifact attestations are a separate mechanism.

5. Hidden-file behavior reduces accidents; it does not classify data for you

Current upload-artifact excludes dot-prefixed hidden files and files under dot-prefixed directories by default. This helps prevent accidental upload of files such as .env, but a sensitive file named credentials.txt is not hidden. Conversely, a harmless hidden report may be intentionally excluded unless you opt in.

When include-hidden-files: true is genuinely required, validate the selection and exclude sensitive patterns explicitly. Never rely on log masking to make a secret safe inside an artifact: masking only affects log rendering.

6. Retention is a policy ceiling, not a promise of forever

retention-days requests an expiry for a newly uploaded artifact, bounded by repository, organization and enterprise policy. GitHub's current public-repository range is 1–90 days; private repositories can be configured up to 400 days where higher-level policy allows it. Deleting the workflow run also deletes its artifacts.

Release records that must survive a compliance period should not silently depend on a default 90-day artifact. Choose the retention intentionally, copy durable evidence to an approved records system when required, and document the lifecycle.

7. Artifact, job output and cache are different transports

Mechanism Use it for Identity/lifetime Do not use it for
Job output small structured control data such as an ID, version or JSON selector workflow-run dataflow; size constrained large binaries, logs or durable evidence
Artifact build/test files and explicit run evidence that must cross jobs or outlive the job artifact ID, digest, retention/expiry dependency acceleration or secret storage
Cache reusable dependency/tool data intended to improve performance key/restore-key based reusable cache state release evidence or authoritative build output

8. Read-only inspection before changing retention or names

Before editing a workflow, capture the current run and artifact records. The REST artifact object includes the artifact ID, name, size, digest, expiry state, timestamps and producing workflow-run identity including head SHA.

# Read-only: authenticated repository inspection
GH_REPO='OWNER/gha-artifact-lab'
ARTIFACT_ID='123456789'

gh api   -H 'X-GitHub-Api-Version: 2026-03-10'   "repos/$GH_REPO/actions/artifacts/$ARTIFACT_ID"   --jq '{id,name,size_in_bytes,digest,expired,created_at,expires_at,workflow_run}'

Knowledge check

Why is an artifact name insufficient as provenance?

Does the default hidden-file exclusion make artifact upload safe for secrets?

What does an artifact digest establish?

When should a cache replace an artifact?

Why can a 90-day default be dangerous for release evidence?

Next lesson

Move one artifact safely between jobs

Lesson 2 builds a disposable pipeline that creates, uploads, identifies, downloads, verifies, inspects and optionally deletes one lab artifact.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked on 2026-09-09. Mandatory examples target GitHub.com and ubuntu-24.04. actions/upload-artifact v7.0.1 and actions/download-artifact v8.0.1 are pinned by full commit SHA. Upload v7 excludes dot-prefixed hidden files by default, supports compression levels 0–9, returns artifact ID/URL/SHA-256 digest, and treats overwrite as delete-and-create rather than in-place mutation. Download v8 can select by artifact ID and defaults digest mismatch handling to error. Artifact/log retention is policy bounded: GitHub documents 1–90 days for public repositories and up to 400 days for private repositories when repository/organization/enterprise policy permits. Upload-artifact v4+ is not supported on GitHub Enterprise Server; GHES users must use an appliance-compatible artifact action/backend and must not copy GitHub.com majors blindly.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.