Checkpoint Lab — Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data
This checkpoint combines Chapters 13–17 into one observable CI data lifecycle. You will predict cache and artifact state, run a deterministic build/consume graph, force a cache miss by changing the dependency input, deliberately break an artifact path, repair it from logs, and finish with a written retention and sensitivity policy suitable for production review.
Learning objectives
- Build a two-job workflow that uses a content-derived dependency cache and a deterministic run artifact with independent file-digest verification.
- Predict and verify cache miss → hit → miss behavior as dependency input changes.
- Inject one artifact-path failure, preserve the failed run, interpret it, and repair the actual data contract.
- Produce a trigger/run/cache/artifact evidence table plus retention and sensitivity policy.
- Close Chapter 17 with production controls that prepare for reusable workflows and actions in Chapter 18.
Checkpoint assumptions: GitHub.com, GitHub Free, one disposable public personal repository, standard GitHub-hosted Ubuntu runner, GitHub CLI, Git, and repository owner/write access. No secrets, packages, release assets, paid storage, private runner, or external service are required.
1. Scenario and predictions
Create a small validation pipeline for a fictional build. The
dependency input is dependency.lock. The cache contains
only a synthetic materialized dependency directory. The staged run
artifact contains report.txt plus the machine-readable
result.json test record. The downstream job must prove
the same file bytes crossed the job boundary.
| Prediction | Before action | Expected after action |
|---|---|---|
| P1 | No cache exists for dependency.lock v1 | First successful run creates a cache under content-derived key. |
| P2 | No run artifact exists | Successful normal run creates one run-unique artifact and consumer verifies file SHA. |
| P3 | Dependency input unchanged | Second normal run reports exact cache hit. |
| P4 | dependency.lock changes to v2 | Next run uses a different exact key and misses before saving new cache. |
| P5 | break_artifact_path=true | Upload fails; no expected artifact object; consumer is skipped. |
2. Setup and preflight
gh auth status --active --hostname github.com
OWNER="$(gh api -H "X-GitHub-Api-Version: 2026-03-10" user --jq .login)"
REPO="$OWNER/atlas-c17-checkpoint"
gh repo create "$REPO" --public --clone --add-readme
cd atlas-c17-checkpoint
DEFAULT_BRANCH="$(gh repo view "$REPO" --json defaultBranchRef --jq .defaultBranchRef.name)"
printf 'demo-dependency=1
' > dependency.lock
mkdir -p .github/workflows
Before committing, confirm the repository has no secrets you intend to exercise and no existing cache/artifact history that would confuse the experiment.
gh secret list -R "$REPO" --json name --jq 'length'
gh cache list -R "$REPO" --json id,key,ref,sizeInBytes
3. Install the checkpoint workflow
name: Chapter 17 workflow data lab
on:
workflow_dispatch:
inputs:
break_artifact_path:
description: Intentionally select a missing artifact path
required: true
type: boolean
default: false
permissions:
contents: read
env:
ARTIFACT_RETENTION_DAYS: 5
jobs:
build:
runs-on: ubuntu-24.04
outputs:
file_sha256: ${{ steps.build.outputs.file_sha256 }}
artifact_name: ${{ steps.names.outputs.artifact_name }}
cache_hit: ${{ steps.cache.outputs.cache-hit }}
steps:
- name: Checkout exact source
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Restore deterministic dependency cache
id: cache
uses: actions/cache@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
with:
path: .chapter17-cache
key: c17-${{ runner.os }}-${{ hashFiles('dependency.lock') }}
restore-keys: |
c17-${{ runner.os }}-
- name: Materialize synthetic dependency when cache misses
if: steps.cache.outputs.cache-hit != 'true'
run: |
mkdir -p .chapter17-cache
cp dependency.lock .chapter17-cache/dependency.lock
printf 'materialized_from=%s\n' "$(sha256sum dependency.lock | awk '{print $1}')" > .chapter17-cache/metadata.txt
- name: Build deterministic report and safe test result
id: build
run: |
mkdir -p dist test-results artifact-staging
DEP_SHA="$(sha256sum dependency.lock | awk '{print $1}')"
printf 'source_sha=%s\ndependency_sha=%s\n' "$GITHUB_SHA" "$DEP_SHA" > dist/report.txt
FILE_SHA="$(sha256sum dist/report.txt | awk '{print $1}')"
printf 'file_sha256=%s\n' "$FILE_SHA" >> "$GITHUB_OUTPUT"
printf '{"tests":1,"passed":1,"failed":0,"source_sha":"%s"}\n' "$GITHUB_SHA" > test-results/result.json
printf '%s\n' '::notice file=dependency.lock,line=1::Dependency input validated'
- name: Select run-unique artifact name
id: names
run: |
printf 'artifact_name=c17-report-%s\n' "$GITHUB_RUN_ID" >> "$GITHUB_OUTPUT"
- name: Upload expected artifact
id: upload
if: ${{ !inputs.break_artifact_path }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: ${{ steps.names.outputs.artifact_name }}
path: artifact-staging/
if-no-files-found: error
retention-days: ${{ env.ARTIFACT_RETENTION_DAYS }}
compression-level: 6
- name: Intentionally broken upload path
if: ${{ inputs.break_artifact_path }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: c17-broken-${{ github.run_id }}
path: artifact-staging/does-not-exist.txt
if-no-files-found: error
retention-days: ${{ env.ARTIFACT_RETENTION_DAYS }}
- name: Write job summary
if: ${{ always() }}
env:
CACHE_HIT: ${{ steps.cache.outputs.cache-hit }}
FILE_SHA: ${{ steps.build.outputs.file_sha256 }}
ARTIFACT_DIGEST: ${{ steps.upload.outputs.artifact-digest }}
run: |
printf '%s\n' '### Chapter 17 build evidence' >> "$GITHUB_STEP_SUMMARY"
printf '%s\n' "- Source SHA: $GITHUB_SHA" >> "$GITHUB_STEP_SUMMARY"
printf '%s\n' "- Exact cache hit: $CACHE_HIT" >> "$GITHUB_STEP_SUMMARY"
printf '%s\n' "- File SHA-256: $FILE_SHA" >> "$GITHUB_STEP_SUMMARY"
if [ -n "$ARTIFACT_DIGEST" ]; then
printf '%s\n' "- Artifact digest: $ARTIFACT_DIGEST" >> "$GITHUB_STEP_SUMMARY"
else
printf '%s\n' '- Artifact digest: unavailable because upload failed or was skipped' >> "$GITHUB_STEP_SUMMARY"
fi
consume:
needs: build
runs-on: ubuntu-24.04
steps:
- name: Download exact artifact from this run
uses: actions/download-artifact@70fc10c6e5e1ce46ad2ea6f2b72d43f7d47b13c3 # v8.0.0
with:
name: ${{ needs.build.outputs.artifact_name }}
path: received
- name: Verify file identity
env:
EXPECTED_SHA: ${{ needs.build.outputs.file_sha256 }}
run: |
test -f received/report.txt
ACTUAL_SHA="$(sha256sum received/report.txt | awk '{print $1}')"
printf 'expected_sha=%s\nactual_sha=%s\n' "$EXPECTED_SHA" "$ACTUAL_SHA"
test "$ACTUAL_SHA" = "$EXPECTED_SHA"
- name: Consumer summary
run: |
printf '%s\n' '### Chapter 17 consume evidence' >> "$GITHUB_STEP_SUMMARY"
printf '%s\n' '- Artifact crossed a job boundary and file digest matched.' >> "$GITHUB_STEP_SUMMARY"
The build job keeps all repository permissions read-only. Cache and artifact actions use full pinned commit SHAs. The intentional failure is controlled by a typed Boolean input; it cannot select an arbitrary filesystem path.
# Save the YAML as .github/workflows/ch17-checkpoint.yml
git add dependency.lock .github/workflows/ch17-checkpoint.yml
git commit -m "ci: add Chapter 17 workflow data checkpoint"
git push origin "$DEFAULT_BRANCH"
BASE_SHA="$(git rev-parse HEAD)"
printf 'base_sha=%s
' "$BASE_SHA"
4. Run 1 — verify P1 and P2: cache miss plus artifact transfer
gh workflow run ch17-checkpoint.yml -R "$REPO" --ref "$DEFAULT_BRANCH" -f break_artifact_path=false
sleep 3
RUN1="$(gh run list -R "$REPO" --workflow ch17-checkpoint.yml --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$RUN1" -R "$REPO" --exit-status
gh run view "$RUN1" -R "$REPO" --log
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/runs/$RUN1/artifacts?per_page=100" --jq '.artifacts[] | {id,name,size_in_bytes,expires_at,digest}'
gh cache list -R "$REPO" --json id,key,ref,sizeInBytes,createdAt,lastAccessedAt
Record the exact cache key, artifact ID/name/digest, file SHA from logs, and run source SHA. The consumer should succeed only after the artifact download and file-digest comparison.
5. Run 2 — verify P3: exact cache hit
gh workflow run ch17-checkpoint.yml -R "$REPO" --ref "$DEFAULT_BRANCH" -f break_artifact_path=false
sleep 3
RUN2="$(gh run list -R "$REPO" --workflow ch17-checkpoint.yml --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$RUN2" -R "$REPO" --exit-status
gh run view "$RUN2" -R "$REPO" --log
The artifact is new because it is run evidence. The cache should be the same exact key and report a hit because neither OS nor lockfile changed.
6. Change dependency input — verify P4: a new exact cache key misses
printf 'demo-dependency=2
' > dependency.lock
git add dependency.lock
git commit -m "build: change Chapter 17 dependency input"
git push origin "$DEFAULT_BRANCH"
V2_SHA="$(git rev-parse HEAD)"
gh workflow run ch17-checkpoint.yml -R "$REPO" --ref "$DEFAULT_BRANCH" -f break_artifact_path=false
sleep 3
RUN3="$(gh run list -R "$REPO" --workflow ch17-checkpoint.yml --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$RUN3" -R "$REPO" --exit-status
gh cache list -R "$REPO" --sort created_at --order desc --limit 10 --json id,key,ref,sizeInBytes,createdAt,lastAccessedAt
Because hashFiles(dependency.lock) changed, the exact
key changes. A broad restore prefix may seed older data, but
cache-hit is true only for an exact match. The workflow
rewrites the synthetic dependency directory after a non-exact
hit/miss as needed by the lab logic.
7. Inject the artifact-path failure — verify P5 without hiding the cause
Dispatch the same V2 revision with the failure input. Expect the upload step to fail because its required file path is absent. Do not re-run immediately; preserve and inspect the failed attempt first.
gh workflow run ch17-checkpoint.yml -R "$REPO" --ref "$DEFAULT_BRANCH" -f break_artifact_path=true
sleep 3
BROKEN="$(gh run list -R "$REPO" --workflow ch17-checkpoint.yml --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
# Expected non-zero result; preserve it rather than suppressing the evidence.
gh run watch "$BROKEN" -R "$REPO" || true
gh run view "$BROKEN" -R "$REPO" --json headSha,status,conclusion,jobs,url
gh run view "$BROKEN" -R "$REPO" --log
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/runs/$BROKEN/artifacts?per_page=100" --jq '{total_count,artifacts:[.artifacts[]|{id,name,digest}]}'
Interpretation: the producer created the approved
artifact-staging/ directory, but the selected upload
contract asks for artifact-staging/does-not-exist.txt.
The expected artifact is therefore absent and the dependent consumer
cannot run. The repair is to return to
break_artifact_path=false (or correct the path in
production), not to change if-no-files-found to a
warning.
8. Repair and independently verify the final healthy state
gh workflow run ch17-checkpoint.yml -R "$REPO" --ref "$DEFAULT_BRANCH" -f break_artifact_path=false
sleep 3
RECOVERED="$(gh run list -R "$REPO" --workflow ch17-checkpoint.yml --event workflow_dispatch --limit 1 --json databaseId --jq '.[0].databaseId')"
gh run watch "$RECOVERED" -R "$REPO" --exit-status
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/actions/runs/$RECOVERED/artifacts?per_page=100" --jq '.artifacts[] | {id,name,size_in_bytes,expires_at,digest}'
rm -rf verify-download
mkdir verify-download
gh run download "$RECOVERED" -R "$REPO" -D verify-download
find verify-download -maxdepth 3 -type f -print -exec sha256sum {} \;
The recovered run must use V2 SHA, create the expected artifact, and pass the downstream digest check. The broken run remains in history as evidence of the failure and diagnosis.
9. Produce the trigger/data lifecycle test matrix
| Run | Source SHA / dependency input | Expected cache | Expected artifact | Consumer | Evidence |
|---|---|---|---|---|---|
| RUN1 | v1 / demo-dependency=1 | Miss → save | Created | Success | Logs + summary + API + cache list |
| RUN2 | v1 unchanged | Exact hit | New artifact | Success | Cache-hit output + new artifact ID |
| RUN3 | v2 / demo-dependency=2 | New exact-key miss | Created | Success | Two distinct content-derived cache keys |
| BROKEN | v2 | Likely hit | Required artifact absent | Skipped | Upload error + zero expected artifact |
| RECOVERED | v2 | Hit | Created | Success | Download + file SHA verification |
10. Write the production retention and sensitivity policy
Add WORKFLOW_DATA_POLICY.md with explicit answers:
- Cache: allowed paths, key dimensions, trusted writers, restore-key policy, regeneration guarantee, no secrets.
-
Artifacts: naming convention, required paths use
if-no-files-found: error, file/artifact digest requirements, retention class. - Logs: selected identifiers only; no context/env dumps; incident response for credential leakage.
- Test evidence: concise summary for humans plus machine-readable schema when downstream automation/audit requires it.
- Retention: debugging, release-candidate, and audit classes with owner/review cadence; export to a system of record when Actions maximum retention is insufficient.
- Deletion: who may delete run artifacts/caches and what evidence must be preserved first.
- Cost: periodic review of cache size/last access and artifact retention/size; fix churn before buying capacity.
11. Final verification checklist
- All healthy runs used
contents: readonly. - Official actions were pinned to full verified commit SHAs.
- The cache key changed only when declared dependency input changed.
- Cache loss never became a correctness failure.
- Every healthy run produced a unique artifact and downstream file digest matched.
- The broken upload failed loudly and preserved the original failed run.
- No artifact/cache/log contained secrets, hidden config, or whole contexts.
- Artifact expiration metadata was captured from the API.
- The written policy distinguishes caches, run artifacts, release/package distribution, logs, summaries, and structured reports.
12. Cleanup / rollback
Do not delete the broken run during the exercise; it is the diagnostic evidence. At the end, disable the workflow and archive the disposable repository. Archive preserves evidence while preventing casual mutation.
gh workflow disable ch17-checkpoint.yml -R "$REPO"
gh workflow list -R "$REPO" --all --json name,path,state
gh repo archive "$REPO" --yes
gh repo view "$REPO" --json nameWithOwner,isArchived,url
Destructive option not required: Deleting artifacts, caches, workflow runs, or the repository can reclaim storage but destroys evidence. The mandatory checkpoint does not require those operations.
13. What Chapter 17 adds to the production GitHub operating model
Chapter 17 adds the workflow data plane: every cache, artifact, log, summary, and test report now has an identity, producer, consumer, trust classification, retention, and cleanup rule. Combined with Chapters 13–16, you can explain not only why a workflow ran and where it executed, but what evidence it produced and whether later jobs may safely consume it.
Chapter 18 builds on this by reducing duplicated automation through reusable workflows and custom actions. Reuse increases supply-chain leverage, so the data contracts and trust rules established here become even more important.
14. Checkpoint summary
You verified cache miss/hit/invalidation behavior, transferred and hashed a deterministic artifact across jobs, produced safe human/machine evidence, preserved an intentional artifact-path failure, and wrote a retention/sensitivity policy. CI data is now a governed lifecycle rather than incidental files left behind by runners.
Knowledge check
RUN3 changes only dependency.lock. Why should the exact cache key change?
The key contains hashFiles(dependency.lock);
dependency content is part of the cache compatibility contract.
BROKEN has a failed upload and skipped consumer. Why not change
if-no-files-found to warn?
The artifact is required for the downstream job. Warning would hide a broken producer/transfer contract instead of fixing it.
Why keep the failed run until the checkpoint ends?
It preserves original logs, step state, source SHA, and artifact absence needed to prove the diagnosis rather than reconstructing it from memory.
A cache hit makes tests pass, but deleting the cache makes them fail. What conclusion follows?
The build is incorrectly dependent on cache state. The cache must remain optional and reproducible from declared inputs.
What should happen if workflow evidence must survive longer than the repository’s supported Actions retention?
Export approved machine-readable evidence to an appropriate durable system of record with its own access/retention controls.
What does Chapter 18 add next?
Reusable workflows and actions—shared automation units whose inputs, outputs, permissions, versions, and supply-chain trust must be governed.
Further reading — current official GitHub sources
- GitHub Docs — Workflow artifacts concepts
- GitHub Docs — Store and share data with workflow artifacts
- GitHub Docs — Dependency caching reference
- GitHub Docs — Managing caches
- GitHub Docs — Removing workflow artifacts and retention
- GitHub Docs — Workflow commands, annotations, and job summaries
- GitHub Docs — Repository Actions settings and retention
- GitHub REST — Actions artifacts
- GitHub REST — Actions cache
- GitHub REST — Actions permissions / retention
- GitHub CLI — gh cache list
- GitHub CLI — gh run download
- GitHub CLI — gh run view
- Official action — actions/upload-artifact
- Official action — actions/download-artifact
- Official action — actions/cache
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.