Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data: Diagnostics, Failure Modes, Security, and Performance
Workflow-data failures are deceptive because the workflow may still be green: a stale cache can make tests run against the wrong dependency state, a warning-only artifact upload can silently preserve nothing, and exact-string secret masking can miss transformed data. Diagnosis must reconstruct key, ref, path, artifact ID/digest, retention, and trust source before deleting or retrying anything.
Learning objectives
- Diagnose stale and poisoned caches from key/ref/trigger evidence instead of deleting all caches blindly.
- Diagnose missing/wrong artifact paths and accidental sensitive-file inclusion while preserving the original run evidence.
- Explain why retention can erase incident evidence and why exact-value masking is not complete secret protection.
- Interpret one intentionally broken upload failure and repair only the actual path contract.
- Connect data volume and cache churn to performance/storage cost without generic optimization advice.
Diagnostic rule: Preserve the failed run URL/ID, source SHA, workflow revision, cache key/ref, artifact name/path, action versions, and retention metadata before deleting caches/artifacts or re-running. Re-running first can replace the easiest evidence of the original cause.
1. Diagnostic sequence: evidence → scope → object state → least-destructive correction → verification
- Preserve: run ID/attempt, logs, job summary, source SHA, workflow YAML revision, artifact/cache API metadata.
- Scope: repository, event/ref, job, cache key/version, artifact path/name, consumer.
- Inspect: permissions, trigger trust, cache ref, artifact existence/expiry, retention, exact action output.
- Correct minimally: fix key/path/trust/retention policy without force pushes, broad cache deletion, or hiding failures.
- Verify independently: compare next run plus API/CLI object state.
2. Failure: the cache key is too broad and stale dependencies contaminate builds
Symptom: a lockfile changes but logs still show an exact cache hit
under a key such as deps-linux. Preserve the run, then
compare the lockfile SHA to the cache key. The cache has no
dependency identity, so GitHub is doing exactly what the key
requested.
gh cache list -R "$REPO" --key deps-linux --json id,key,ref,sizeInBytes,createdAt,lastAccessedAt
git show "$FAILED_SHA:dependency.lock"
git show "$KNOWN_GOOD_SHA:dependency.lock"
Repair: version the key with the dependency manifest hash. Do not “solve” the issue by deleting every repository cache; that destroys evidence and other workflows’ acceleration without fixing the key design.
3. Failure: low-trust workflow data reaches a privileged cache consumer
Cache poisoning means attacker-controlled workflow execution creates or influences cache content that a more privileged workflow later restores and executes/trusts. GitHub now restricts writes into default-branch cache scope for several low-trust trigger types, and PR caches use the PR merge ref, but privileged workflows should still treat restored bytes as untrusted.
Inspect event/ref and cache origin. Prefer trusted cache writers, restore-only behavior for low-trust jobs, content-derived keys, and caches that contain package-manager downloads/intermediates rather than executable scripts with authority. Never store credentials in cache paths.
4. Intentionally broken example: required artifact path does not exist
In the checkpoint workflow, dispatch with
break_artifact_path=true. The build creates
dist/report.txt, but the upload action is intentionally
pointed at dist/does-not-exist.txt with
if-no-files-found: error. Expected result: upload
fails, build job fails, and the dependent consumer job is skipped.
Representative failure interpretation:
Error: No files were found with the provided path: dist/does-not-exist.txt.
No artifacts will be uploaded.
Meaning:
- The runner executed far enough to reach artifact upload.
- The artifact object was never created.
- The dependent job has no valid artifact contract to consume.
- Fix the path/producer contract; do not weaken the upload to "warn" merely to make CI green.
Preserve the failed run and confirm the actual filesystem producer
from earlier log steps. Repair by restoring the correct
path: dist/report.txt or fixing the producer if that
file truly should have existed elsewhere.
5. Failure: artifact glob captures credentials or configuration
A broad path such as path: . is dangerous. It can
collect test fixtures, local auth files, generated configuration, or
hidden content when hidden-file upload is enabled. Before upload,
construct an explicit staging directory containing only approved
output and inspect it.
mkdir -p artifact-staging
cp dist/report.txt artifact-staging/
find artifact-staging -maxdepth 2 -type f -print
# Optional content checks belong here before upload.
If a real credential was uploaded, treat it as exposed: revoke/rotate first, then delete the artifact/run and investigate. Deletion alone does not invalidate a credential already seen or downloaded.
6. Failure: evidence expires before audit or incident review
Retention assumptions must be written into the pipeline policy. If the repository keeps logs/artifacts for 30 days but an audit or incident-review requirement is 180 days, the correct fix is an approved evidence export/archive strategy—not wishful thinking that Actions storage is permanent.
Use the artifact API expires_at to prove object expiry.
Repository retention setting changes are not retroactive, so
changing policy after an incident does not resurrect evidence that
already expired.
7. Failure: logs are noisy or secret redaction gives false confidence
Dumping complete environment/context objects makes logs harder to use and expands the leakage surface. Secret masking generally protects values GitHub knows as secrets/masked strings, but derived representations—base64, URL-encoded values, substrings, transformed JSON, hashes used improperly, or concatenated fragments—can escape exact matching.
Repair logging at the producer: print selected non-sensitive identifiers, group noisy diagnostics, keep machine data in a structured report, and avoid printing secret-derived data at all. If a secret leaks, revoke/rotate first.
8. Performance and storage failures should be traced to lifecycle design
| Symptom | Likely causal question | Targeted correction |
|---|---|---|
| Cache thrashing | Are keys too unique or total cache size too small for active key set? | Reduce unnecessary key dimensions/cache paths; inspect last-access/size before changing quota. |
| Large slow uploads | Is data already compressed or low-value? | Lower compression for incompressible files; reduce/stage only required evidence. |
| Storage growth | Are artifacts retained longer than operational need? | Set per-artifact/repository retention from policy; delete only disposable evidence. |
| Logs hard to diagnose | Is structured test data being emitted as text? | Move machine data to a report; summarize counts and annotate actionable failures. |
9. Security-sensitive/destructive operations in this chapter
Artifact deletion, workflow-run deletion, cache force-deletion, and retention shortening can destroy evidence. Use them only on disposable resources or under an approved evidence-retention process. Repository deletion, force-push/history rewrite, token creation, runner registration, policy bypass, package deletion, and privilege changes are not required anywhere in this chapter.
10. Lesson summary
Diagnose workflow data from identity and scope: cache key/ref/trigger, artifact path/ID/digest, source SHA, retention, and producer/consumer jobs. The safe correction fixes the contract that was wrong rather than deleting evidence, broadening trust, or downgrading failures to warnings.
Knowledge check
A required artifact upload warns “no files found” but the run is green. What should you change?
For required evidence, use
if-no-files-found: error and fix the producer/path
contract rather than accepting the warning.
Why is deleting all caches a poor first response to stale dependencies?
It destroys evidence and other useful caches while leaving the overly broad key design unchanged.
A credential appears base64-encoded in logs even though the raw secret was masked. What is the first incident step?
Revoke/rotate the credential. Then remove exposed evidence where appropriate and fix logging so secret-derived values are never emitted.
Can increasing retention today restore an artifact that expired yesterday?
No. Retention changes do not resurrect deleted/expired evidence and are not retroactive to existing objects.
Why can a cache become a supply-chain input?
A later workflow restores bytes produced earlier. If those bytes are executed or trusted, their producer/trust path matters like any other dependency.
Further reading — current official GitHub sources
- GitHub Docs — Workflow artifacts concepts
- GitHub Docs — Store and share data with workflow artifacts
- GitHub Docs — Dependency caching reference
- GitHub Docs — Managing caches
- GitHub Docs — Removing workflow artifacts and retention
- GitHub Docs — Workflow commands, annotations, and job summaries
- GitHub Docs — Repository Actions settings and retention
- GitHub REST — Actions artifacts
- GitHub REST — Actions cache
- GitHub REST — Actions permissions / retention
- GitHub CLI — gh cache list
- GitHub CLI — gh run download
- GitHub CLI — gh run view
- Official action — actions/upload-artifact
- Official action — actions/download-artifact
- Official action — actions/cache
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.