Chapter 16Lesson 04~255 minutes

Artifacts, Reports, Cache, Dependencies, Retention, and Data Flow Between Jobs: Diagnostics, Failure Modes, Security, and Performance

Preserve producer/consumer evidence and diagnose missing, expired, unsafe, poisoned, or misrouted data before changing configuration.

DiagnosticsExpired artifactsCache poisoningDAG data flowSecurityRecovery

Learning objectives

  • Use an evidence-first diagnostic sequence for artifact/cache/report failures.
  • Distinguish missing, expired, inaccessible, and not-requested artifacts.
  • Recognize cache poisoning and unsafe artifact content as security incidents.
  • Repair DAG artifact-flow regressions without reverting to broad implicit downloads.
  • Connect retention/storage failures to recovery and audit requirements.
Availability baseline (verified 2026-08-21 against current GitLab documentation). Ordinary job artifacts, report artifacts such as JUnit, job caches, dependencies, and needs:artifacts are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. If artifacts:expire_in is omitted, the instance default controls expiry; GitLab also keeps artifacts from the most recent successful pipeline on each ref by default unless that behavior is disabled. Cache is a performance optimization, not an authoritative release/evidence store. Protected and non-protected refs use separate caches by default; disabling that boundary or using cache:unprotect broadens who can read/write the same cache and must be a deliberate trust decision. Hosted storage/compute quotas and billing are volatile, so labs use tiny files and a no-runner fixture path.

1. Diagnostic sequence: preserve identity before changing YAML

When a downstream job says a file is missing, do not immediately add cache, retry, or broaden artifacts. Preserve: GitLab offering/version, project, ref/SHA, pipeline ID/source, producer/consumer job IDs, relevant YAML expansion, artifact metadata/expiry, runner identity, and the exact error. Then classify the failure as production, storage, access, scheduling/data-flow, or cache.

2. Intentionally broken example: the artifact exists, but the consumer does not request it

Start from a safe fixture:

build:
  stage: build
  script:
    - mkdir -p out
    - echo synthetic > out/result.txt
  artifacts:
    paths: [out/result.txt]

consumer:
  stage: test
  dependencies: []
  script:
    - test -f out/result.txt

The producer succeeds and uploads its artifact. The consumer fails with a non-zero test exit because dependencies: [] explicitly disables downloads. Preserve the failing log and producer artifact metadata. The repair is not a retry; change the declared data contract:

consumer:
  stage: test
  dependencies:
    - build
  script:
    - test -f out/result.txt

Validate and rerun. The same producer artifact now appears because the consumer asked for it.

3. Failure: artifact expired before a consumer or incident review

GitLab can report that a job could not retrieve needed artifacts when expected dependencies are gone, expired, or inaccessible. First verify producer job status and expiry. If expiry is the cause, extending retention fixes future pipelines; it does not resurrect deleted bytes. Rebuild from the original commit only if that process is deterministic and authorized.

For audit evidence, the durable fix may be promotion to a release/package/registry/evidence store rather than arbitrarily increasing every job artifact’s lifetime.

4. Failure: DAG conversion downloads the wrong or no artifact

A consumer used to receive all earlier-stage artifacts. After adding needs: [compile], it loses schema.json from schema. That is expected: needs narrowed artifact flow. Add the true producer edge with artifacts: true, or refactor to one canonical producer. Do not “fix” this with cache because that removes pipeline identity.

5. Failure: parallel producers overwrite same artifact path

Matrix jobs can each publish dist/output.bin. A consumer that downloads multiple producer artifacts may see one overwrite another in the workspace. Give outputs unique paths/names by matrix coordinates or select one producer with needs:parallel:matrix. Verify hashes after extraction.

6. Failure: report uploads but GitLab does not interpret it

A JUnit file can exist as an artifact and still fail parsing because of invalid XML or unsupported structure. Inspect the job log and pipeline Tests view. Validate the report format locally/with the test framework. Do not change artifact access or runner tags to solve a schema problem.

7. Cache miss: usually performance, not correctness

A cache miss after runner replacement is normal if no distributed cache exists. The job should redownload/rebuild and continue. If the job fails because the cache is empty, the job has accidentally elevated cache to authoritative state. Repair the bootstrap path first, then tune cache topology.

8. Security failure: untrusted branch overwrites trusted cache

Suppose a shared cache key stores generated executables and an untrusted branch can push to it. A protected/default-branch job later executes that cache. Preserve cache key, pipeline/ref/user/runner evidence, then stop consuming the shared cache. Rotate/rebuild any affected generated tools from trusted source, restore protected/non-protected separation, and use immutable/verifiable dependencies where possible.

Do not “inspect” a suspected poisoned executable by running it. Treat it as untrusted evidence. Hash/quarantine it and rebuild the cache from a trusted source.

9. Security incident: artifact contains credentials or private build context

Artifact access control does not undo secret disclosure. If a real credential was uploaded, incident response begins by revoking/rotating the credential, then restricting/removing the artifact where permitted, reviewing access/audit evidence, and correcting the pipeline so the secret never enters artifact paths. History/log cleanup is secondary to credential invalidation.

For private build context/customer data, follow data-handling policy and minimize retention/download exposure. Do not copy the artifact into chat/tickets just to debug it.

10. Destructive correction: artifact deletion

The Job Artifacts API can delete one job’s artifacts; current docs require Maintainer/Owner. Deletion includes archived files, metadata, and report artifacts, and it cannot be recovered. Deleting all eligible project artifacts is asynchronous and does not delete job logs. Use these operations only after exact project/job identity and retention obligations are checked.

Lab policy. The chapter does not require destructive artifact deletion. Short expiry + disposable project/branch is the default cleanup. If practicing deletion, use only synthetic data in a project created for the exercise.

11. Reliability failure: retention policy is shorter than recovery time

A release incident is discovered 45 days later but the candidate artifact expired after seven. This is not a CI syntax failure; the organization’s evidence/recovery model is incomplete. Define which outputs need long-term retention and promote them to a durable release/package/registry/evidence surface. Keep short-lived transient artifacts short.

12. Self-Managed storage failure

On Self-Managed, an object-storage outage, permission change, lifecycle rule, or capacity problem can make artifacts unavailable even when GitLab metadata still references them. Job authors should capture API/UI symptoms and escalate with IDs/timestamps; administrators inspect object storage, GitLab logs, configuration and backup/recovery. Do not ask application developers to “fix S3” from CI YAML.

13. Performance failure: every job downloads every artifact

Large broad downloads consume network, runner startup time, and storage egress. Measure artifact sizes and which consumers actually need them. Explicit dependencies/needs:artifacts can reduce transfer. Do not trade correctness for speed by moving authoritative outputs to cache.

14. Verification after repair

  • Producer job/pipeline/SHA is unchanged or intentionally rebuilt from the same source.
  • Consumer downloads exactly the required artifact set.
  • Checksums/content identity match expectation.
  • Report appears in the intended GitLab view and script status remains meaningful.
  • Cache miss does not change correctness.
  • Sensitive data is absent from artifact/cache paths and logs.
  • Retention/access match documented policy.

Knowledge check

A consumer fails after dependencies: []. Should you retry the pipeline?

Can increasing expire_in recover an artifact already deleted?

A cache miss causes a build to fail. What design smell does that reveal?

What is the first response when a real credential is found in an artifact?

Why might a report artifact exist but not appear in GitLab’s report UI?

What should a Self-Managed job author do if artifact metadata exists but object storage is unavailable?

Summary

Artifact/cache failures become tractable when you preserve producer/consumer identity and classify the data plane first. Missing downloads are usually contract problems, cache misses are performance events, expired artifacts are retention/recovery problems, and sensitive uploads are security incidents. Repair the narrowest layer and re-prove data identity.

Official references

Next lesson

Checkpoint the complete data-flow model

Lesson 5 predicts files per job, produces an artifact/report/cache, triggers a safe cache or data-flow miss, verifies authority and retention, captures evidence, and cleans up without touching valuable project data.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.