Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data: Guided Hands-On Workflow and Core Operations
You will make workflow data observable in one disposable repository. The workflow uses a content-derived cache key, creates a deterministic small artifact, passes it to a separate job, verifies its file digest, writes a human-readable job summary, emits one safe annotation, and leaves enough metadata to explain exactly what GitHub stored and for how long.
Learning objectives
- Create a disposable workflow that demonstrates a cache miss, cache hit, artifact upload, cross-job download, digest verification, annotation, and job summary.
-
Inspect run artifacts, cache entries, logs, and retention metadata
through UI,
gh, and versioned REST output. - Change repository artifact/log retention safely in a disposable repository, prove the setting, and restore the original value.
- Distinguish an artifact file digest from the artifact object digest and prove causality from one exact source SHA.
- Choose the correct data surface in a short challenge instead of copying commands blindly.
Mandatory path: GitHub.com, GitHub Free, repository owner/write access, GitHub CLI authenticated to GitHub.com, Git, and a disposable public repository. The workflow contains no secrets and only reads repository contents. Official actions are pinned to full verified commit SHAs.
1. Preflight: create one repository with one deterministic dependency input
Commands are Bash/Git Bash. PowerShell users can create the same two
files with Set-Content; the GitHub CLI commands are
unchanged.
gh --version
gh auth status --active --hostname github.com
OWNER="$(gh api -H "X-GitHub-Api-Version: 2026-03-10" user --jq .login)"
REPO="$OWNER/atlas-c17-workflow-data-lab"
gh repo create "$REPO" --public --clone --add-readme
cd atlas-c17-workflow-data-lab
DEFAULT_BRANCH="$(gh repo view "$REPO" --json defaultBranchRef --jq .defaultBranchRef.name)"
printf 'demo-dependency=1
' > dependency.lock
git add dependency.lock
git commit -m "build: add deterministic dependency input"
git push origin "$DEFAULT_BRANCH"
The lockfile is deliberately synthetic. It avoids a flaky external registry while preserving the important cache property: cache identity must derive from a declared dependency input.
2. Inspect storage state before the first run
gh cache list -R "$REPO" --limit 30 --json id,key,ref,sizeInBytes,createdAt,lastAccessedAt
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/permissions/artifact-and-log-retention"
A fresh repository should have no caches. Record the current retention days because you will restore them later. This is policy evidence before mutation.
3. Create the workflow: cache → deterministic build → artifact → second job
The build job publishes two identities: a file SHA-256 and a run-unique artifact name. The artifact is then downloaded by another job and the file digest is independently verified. The cache is only an acceleration layer; the deterministic dependency file remains the source of truth.
name: Chapter 17 workflow data lab
on:
workflow_dispatch:
inputs:
break_artifact_path:
description: Intentionally select a missing artifact path
required: true
type: boolean
default: false
permissions:
contents: read
env:
ARTIFACT_RETENTION_DAYS: 5
jobs:
build:
runs-on: ubuntu-24.04
outputs:
file_sha256: ${{ steps.build.outputs.file_sha256 }}
artifact_name: ${{ steps.names.outputs.artifact_name }}
cache_hit: ${{ steps.cache.outputs.cache-hit }}
steps:
- name: Checkout exact source
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Restore deterministic dependency cache
id: cache
uses: actions/cache@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
with:
path: .chapter17-cache
key: c17-${{ runner.os }}-${{ hashFiles('dependency.lock') }}
restore-keys: |
c17-${{ runner.os }}-
- name: Materialize synthetic dependency when cache misses
if: steps.cache.outputs.cache-hit != 'true'
run: |
mkdir -p .chapter17-cache
cp dependency.lock .chapter17-cache/dependency.lock
printf 'materialized_from=%s\n' "$(sha256sum dependency.lock | awk '{print $1}')" > .chapter17-cache/metadata.txt
- name: Build deterministic report and safe test result
id: build
run: |
mkdir -p dist test-results artifact-staging
DEP_SHA="$(sha256sum dependency.lock | awk '{print $1}')"
printf 'source_sha=%s\ndependency_sha=%s\n' "$GITHUB_SHA" "$DEP_SHA" > dist/report.txt
FILE_SHA="$(sha256sum dist/report.txt | awk '{print $1}')"
printf 'file_sha256=%s\n' "$FILE_SHA" >> "$GITHUB_OUTPUT"
printf '{"tests":1,"passed":1,"failed":0,"source_sha":"%s"}\n' "$GITHUB_SHA" > test-results/result.json
printf '%s\n' '::notice file=dependency.lock,line=1::Dependency input validated'
- name: Select run-unique artifact name
id: names
run: |
printf 'artifact_name=c17-report-%s\n' "$GITHUB_RUN_ID" >> "$GITHUB_OUTPUT"
- name: Upload expected artifact
id: upload
if: ${{ !inputs.break_artifact_path }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: ${{ steps.names.outputs.artifact_name }}
path: artifact-staging/
if-no-files-found: error
retention-days: ${{ env.ARTIFACT_RETENTION_DAYS }}
compression-level: 6
- name: Intentionally broken upload path
if: ${{ inputs.break_artifact_path }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: c17-broken-${{ github.run_id }}
path: artifact-staging/does-not-exist.txt
if-no-files-found: error
retention-days: ${{ env.ARTIFACT_RETENTION_DAYS }}
- name: Write job summary
if: ${{ always() }}
env:
CACHE_HIT: ${{ steps.cache.outputs.cache-hit }}
FILE_SHA: ${{ steps.build.outputs.file_sha256 }}
ARTIFACT_DIGEST: ${{ steps.upload.outputs.artifact-digest }}
run: |
printf '%s\n' '### Chapter 17 build evidence' >> "$GITHUB_STEP_SUMMARY"
printf '%s\n' "- Source SHA: $GITHUB_SHA" >> "$GITHUB_STEP_SUMMARY"
printf '%s\n' "- Exact cache hit: $CACHE_HIT" >> "$GITHUB_STEP_SUMMARY"
printf '%s\n' "- File SHA-256: $FILE_SHA" >> "$GITHUB_STEP_SUMMARY"
if [ -n "$ARTIFACT_DIGEST" ]; then
printf '%s\n' "- Artifact digest: $ARTIFACT_DIGEST" >> "$GITHUB_STEP_SUMMARY"
else
printf '%s\n' '- Artifact digest: unavailable because upload failed or was skipped' >> "$GITHUB_STEP_SUMMARY"
fi
consume:
needs: build
runs-on: ubuntu-24.04
steps:
- name: Download exact artifact from this run
uses: actions/download-artifact@70fc10c6e5e1ce46ad2ea6f2b72d43f7d47b13c3 # v8.0.0
with:
name: ${{ needs.build.outputs.artifact_name }}
path: received
- name: Verify file identity
env:
EXPECTED_SHA: ${{ needs.build.outputs.file_sha256 }}
run: |
test -f received/report.txt
ACTUAL_SHA="$(sha256sum received/report.txt | awk '{print $1}')"
printf 'expected_sha=%s\nactual_sha=%s\n' "$EXPECTED_SHA" "$ACTUAL_SHA"
test "$ACTUAL_SHA" = "$EXPECTED_SHA"
- name: Consumer summary
run: |
printf '%s\n' '### Chapter 17 consume evidence' >> "$GITHUB_STEP_SUMMARY"
printf '%s\n' '- Artifact crossed a job boundary and file digest matched.' >> "$GITHUB_STEP_SUMMARY"
Notice three guardrails: if-no-files-found: error turns
a missing artifact into a visible failure; artifact names include
the run ID; and the only logged fields are safe identity/evidence
values. No environment or context dump is used.
4. First run: predict a cache miss and a new artifact object
Prediction A: because the repository has no cache, the exact cache key will miss and the workflow will materialize the synthetic dependency. Prediction B: the successful build will create one artifact, and the consume job will receive the same report bytes through GitHub artifact storage—not through a shared runner filesystem.
mkdir -p .github/workflows
# Save the YAML as .github/workflows/ch17-data.yml
git add .github/workflows/ch17-data.yml
git commit -m "ci: add Chapter 17 workflow data lab"
git push origin "$DEFAULT_BRANCH"
WORKFLOW_SHA="$(git rev-parse HEAD)"
gh workflow run ch17-data.yml -R "$REPO" --ref "$DEFAULT_BRANCH" -f break_artifact_path=false
sleep 3
RUN1="$(gh run list -R "$REPO" --workflow ch17-data.yml --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$RUN1" -R "$REPO" --exit-status
Open the run summary in the UI as well. You should see two jobs, a notice annotation, two job summaries, and one artifact. The consumer job is proof that job isolation was crossed explicitly through artifact storage.
5. Inspect artifacts, logs, cache metadata, and exact source revision
gh run view "$RUN1" -R "$REPO" --json headSha,status,conclusion,createdAt,startedAt,updatedAt,jobs,url --jq '{headSha,status,conclusion,jobs:[.jobs[]|{name,conclusion,startedAt,completedAt}]}'
gh run view "$RUN1" -R "$REPO" --log
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/runs/$RUN1/artifacts?per_page=100" --jq '.artifacts[] | {id,name,size_in_bytes,expired,created_at,expires_at,digest}'
gh cache list -R "$REPO" --limit 30 --json id,key,ref,sizeInBytes,createdAt,lastAccessedAt
printf 'expected_workflow_sha=%s
' "$WORKFLOW_SHA"
The run headSha should equal the commit you dispatched.
The artifact API exposes object metadata including expiration and
digest. The cache list exposes a separate key/ref/size inventory.
6. Independently download the artifact outside the workflow
rm -rf downloaded-c17
mkdir downloaded-c17
gh run download "$RUN1" -R "$REPO" -D downloaded-c17
find downloaded-c17 -maxdepth 3 -type f -print
sha256sum downloaded-c17/*/report.txt 2>/dev/null || sha256sum downloaded-c17/report.txt
gh run download extracts artifacts by artifact name
when multiple artifacts are present. Compare the file SHA printed
here with the build and consume logs. This is a third independent
verification path.
7. Second run: prove an exact cache hit without changing dependency input
Dispatch the same workflow revision again. The cache key derives
from OS and dependency.lock, so the second run should
report an exact cache hit. A new run still creates a new artifact
because artifact identity and cache identity have different
purposes.
gh workflow run ch17-data.yml -R "$REPO" --ref "$DEFAULT_BRANCH" -f break_artifact_path=false
sleep 3
RUN2="$(gh run list -R "$REPO" --workflow ch17-data.yml --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$RUN2" -R "$REPO" --exit-status
gh run view "$RUN2" -R "$REPO" --log | grep -E 'Cache restored|Cache hit|Exact match|expected_sha|actual_sha' || true
Do not fail the lesson merely because a log phrase changes—use the job summary and cache action output as the authoritative state. The grep is only a convenience for locating evidence.
8. Temporarily change repository artifact/log retention, then restore it
This setting is an administrative policy mutation, so use only the disposable repository. It affects new artifacts/logs; it does not rewrite the expiry of the artifact you already created. Store the original value before changing anything.
ORIGINAL_DAYS="$(gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/permissions/artifact-and-log-retention" --jq .days)"
printf 'original_days=%s
' "$ORIGINAL_DAYS"
gh api --method PUT -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/permissions/artifact-and-log-retention" -F days=14
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/permissions/artifact-and-log-retention"
# Restore immediately after inspection.
gh api --method PUT -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/permissions/artifact-and-log-retention" -F days="$ORIGINAL_DAYS"
Why restore immediately? Retention is evidence policy. A short lab should not leave a repository-level setting changed accidentally. The workflow already demonstrates per-artifact retention with five days.
9. Challenge: choose cache, artifact, summary, or package
For each requirement, choose the primary surface and justify the choice before revealing the model answer:
| Requirement | Best primary surface | Reason |
|---|---|---|
| Reuse downloaded compiler dependencies next run | Cache | Regenerable acceleration data; build must survive a miss. |
| Pass a generated binary from build job to test job | Workflow artifact | Cross-job file transfer plus run evidence. |
| Show reviewers 12 passed / 1 failed immediately | Job summary + targeted annotation | Human-visible result without opening large logs. |
| Publish version 2.4.0 for downstream dependency consumption | Package or governed Release asset | Distribution lifecycle, not temporary CI evidence. |
10. Verification and cleanup
- RUN1 used the expected source SHA and created one cache plus one artifact.
- RUN2 reused the exact cache key while creating a separate run artifact.
- Consumer file SHA matched producer file SHA.
- Artifact API showed
expires_atand digest. - Logs/summaries contained selected evidence only; no secret/context dump.
- Repository retention was restored to the recorded original value.
gh workflow disable ch17-data.yml -R "$REPO"
gh workflow list -R "$REPO" --all --json name,path,state
gh repo archive "$REPO" --yes
gh repo view "$REPO" --json nameWithOwner,isArchived,url
11. Lesson summary
You observed cache and artifact state as different GitHub resources, proved a cache miss/hit cycle, transferred deterministic bytes between isolated jobs, verified file identity, downloaded run evidence with the CLI, and rehearsed retention policy safely. The next lesson turns these mechanics into design decisions for real pipelines.
Knowledge check
Why does the second run create a new artifact even though it hits the same cache?
The cache is reusable acceleration state keyed to dependencies. The artifact is run-specific output/evidence and should have a separate identity.
Why use if-no-files-found: error for a required
build artifact?
A warning-only upload could leave the workflow green while preserving no required evidence. Failing makes the missing path causal and visible.
What proves a file really crossed the job boundary?
The downstream job downloads the named artifact from GitHub storage and independently verifies the file SHA against the producer output.
Why restore the retention setting after the lab?
It is repository policy that affects future evidence. The lab should not leave a persistent governance change behind.
Would a cache be appropriate for the only copy of an audit report?
No. Caches are evictable and optimized for regenerable data, not durable audit evidence.
Further reading — current official GitHub sources
- GitHub Docs — Workflow artifacts concepts
- GitHub Docs — Store and share data with workflow artifacts
- GitHub Docs — Dependency caching reference
- GitHub Docs — Managing caches
- GitHub Docs — Removing workflow artifacts and retention
- GitHub Docs — Workflow commands, annotations, and job summaries
- GitHub Docs — Repository Actions settings and retention
- GitHub REST — Actions artifacts
- GitHub REST — Actions cache
- GitHub REST — Actions permissions / retention
- GitHub CLI — gh cache list
- GitHub CLI — gh run download
- GitHub CLI — gh run view
- Official action — actions/upload-artifact
- Official action — actions/download-artifact
- Official action — actions/cache
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.