Chapter 17Lesson 04~185 minutes

Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data: Diagnostics, Failure Modes, Security, and Performance

Workflow-data failures are deceptive because the workflow may still be green: a stale cache can make tests run against the wrong dependency state, a warning-only artifact upload can silently preserve nothing, and exact-string secret masking can miss transformed data. Diagnosis must reconstruct key, ref, path, artifact ID/digest, retention, and trust source before deleting or retrying anything.

DiagnosticsCache poisoningSensitive outputEvidence expiry

Learning objectives

  • Diagnose stale and poisoned caches from key/ref/trigger evidence instead of deleting all caches blindly.
  • Diagnose missing/wrong artifact paths and accidental sensitive-file inclusion while preserving the original run evidence.
  • Explain why retention can erase incident evidence and why exact-value masking is not complete secret protection.
  • Interpret one intentionally broken upload failure and repair only the actual path contract.
  • Connect data volume and cache churn to performance/storage cost without generic optimization advice.

Diagnostic rule: Preserve the failed run URL/ID, source SHA, workflow revision, cache key/ref, artifact name/path, action versions, and retention metadata before deleting caches/artifacts or re-running. Re-running first can replace the easiest evidence of the original cause.

1. Diagnostic sequence: evidence → scope → object state → least-destructive correction → verification

  1. Preserve: run ID/attempt, logs, job summary, source SHA, workflow YAML revision, artifact/cache API metadata.
  2. Scope: repository, event/ref, job, cache key/version, artifact path/name, consumer.
  3. Inspect: permissions, trigger trust, cache ref, artifact existence/expiry, retention, exact action output.
  4. Correct minimally: fix key/path/trust/retention policy without force pushes, broad cache deletion, or hiding failures.
  5. Verify independently: compare next run plus API/CLI object state.

2. Failure: the cache key is too broad and stale dependencies contaminate builds

Symptom: a lockfile changes but logs still show an exact cache hit under a key such as deps-linux. Preserve the run, then compare the lockfile SHA to the cache key. The cache has no dependency identity, so GitHub is doing exactly what the key requested.

gh cache list -R "$REPO" --key deps-linux   --json id,key,ref,sizeInBytes,createdAt,lastAccessedAt

git show "$FAILED_SHA:dependency.lock"
git show "$KNOWN_GOOD_SHA:dependency.lock"

Repair: version the key with the dependency manifest hash. Do not “solve” the issue by deleting every repository cache; that destroys evidence and other workflows’ acceleration without fixing the key design.

3. Failure: low-trust workflow data reaches a privileged cache consumer

Cache poisoning means attacker-controlled workflow execution creates or influences cache content that a more privileged workflow later restores and executes/trusts. GitHub now restricts writes into default-branch cache scope for several low-trust trigger types, and PR caches use the PR merge ref, but privileged workflows should still treat restored bytes as untrusted.

Inspect event/ref and cache origin. Prefer trusted cache writers, restore-only behavior for low-trust jobs, content-derived keys, and caches that contain package-manager downloads/intermediates rather than executable scripts with authority. Never store credentials in cache paths.

4. Intentionally broken example: required artifact path does not exist

In the checkpoint workflow, dispatch with break_artifact_path=true. The build creates dist/report.txt, but the upload action is intentionally pointed at dist/does-not-exist.txt with if-no-files-found: error. Expected result: upload fails, build job fails, and the dependent consumer job is skipped.

Representative failure interpretation:
Error: No files were found with the provided path: dist/does-not-exist.txt.
No artifacts will be uploaded.

Meaning:
- The runner executed far enough to reach artifact upload.
- The artifact object was never created.
- The dependent job has no valid artifact contract to consume.
- Fix the path/producer contract; do not weaken the upload to "warn" merely to make CI green.

Preserve the failed run and confirm the actual filesystem producer from earlier log steps. Repair by restoring the correct path: dist/report.txt or fixing the producer if that file truly should have existed elsewhere.

5. Failure: artifact glob captures credentials or configuration

A broad path such as path: . is dangerous. It can collect test fixtures, local auth files, generated configuration, or hidden content when hidden-file upload is enabled. Before upload, construct an explicit staging directory containing only approved output and inspect it.

mkdir -p artifact-staging
cp dist/report.txt artifact-staging/
find artifact-staging -maxdepth 2 -type f -print
# Optional content checks belong here before upload.

If a real credential was uploaded, treat it as exposed: revoke/rotate first, then delete the artifact/run and investigate. Deletion alone does not invalidate a credential already seen or downloaded.

6. Failure: evidence expires before audit or incident review

Retention assumptions must be written into the pipeline policy. If the repository keeps logs/artifacts for 30 days but an audit or incident-review requirement is 180 days, the correct fix is an approved evidence export/archive strategy—not wishful thinking that Actions storage is permanent.

Use the artifact API expires_at to prove object expiry. Repository retention setting changes are not retroactive, so changing policy after an incident does not resurrect evidence that already expired.

7. Failure: logs are noisy or secret redaction gives false confidence

Dumping complete environment/context objects makes logs harder to use and expands the leakage surface. Secret masking generally protects values GitHub knows as secrets/masked strings, but derived representations—base64, URL-encoded values, substrings, transformed JSON, hashes used improperly, or concatenated fragments—can escape exact matching.

Repair logging at the producer: print selected non-sensitive identifiers, group noisy diagnostics, keep machine data in a structured report, and avoid printing secret-derived data at all. If a secret leaks, revoke/rotate first.

8. Performance and storage failures should be traced to lifecycle design

Symptom Likely causal question Targeted correction
Cache thrashing Are keys too unique or total cache size too small for active key set? Reduce unnecessary key dimensions/cache paths; inspect last-access/size before changing quota.
Large slow uploads Is data already compressed or low-value? Lower compression for incompressible files; reduce/stage only required evidence.
Storage growth Are artifacts retained longer than operational need? Set per-artifact/repository retention from policy; delete only disposable evidence.
Logs hard to diagnose Is structured test data being emitted as text? Move machine data to a report; summarize counts and annotate actionable failures.

9. Security-sensitive/destructive operations in this chapter

Artifact deletion, workflow-run deletion, cache force-deletion, and retention shortening can destroy evidence. Use them only on disposable resources or under an approved evidence-retention process. Repository deletion, force-push/history rewrite, token creation, runner registration, policy bypass, package deletion, and privilege changes are not required anywhere in this chapter.

10. Lesson summary

Diagnose workflow data from identity and scope: cache key/ref/trigger, artifact path/ID/digest, source SHA, retention, and producer/consumer jobs. The safe correction fixes the contract that was wrong rather than deleting evidence, broadening trust, or downgrading failures to warnings.

Knowledge check

A required artifact upload warns “no files found” but the run is green. What should you change?

Why is deleting all caches a poor first response to stale dependencies?

A credential appears base64-encoded in logs even though the raw secret was masked. What is the first incident step?

Can increasing retention today restore an artifact that expired yesterday?

Why can a cache become a supply-chain input?

Next lesson

Next: Checkpoint Lab — Artifacts, Dependency Caching, Logs, Test Reports, Retention, and Workflow Data

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.