Chapter 16Lesson 03~235 minutes

Artifacts, Reports, Cache, Dependencies, Retention, and Data Flow Between Jobs: Configuration, Design Choices, and Tradeoffs

Turn artifact/cache mechanics into maintainable policy for retention, dependency clarity, storage cost, and cache trust boundaries.

DesignAccessRetentionObject storageCache isolationTradeoffs

Learning objectives

  • Choose artifact versus cache from authority and rebuildability rather than convenience.
  • Set retention from recovery/audit needs instead of defaulting to permanent storage.
  • Choose automatic versus explicit artifact downloads based on dependency clarity.
  • Design cache keys/policies around trust and compatibility boundaries.
  • Separate job-author decisions from Self-Managed storage administration.
Availability baseline (verified 2026-08-21 against current GitLab documentation). Ordinary job artifacts, report artifacts such as JUnit, job caches, dependencies, and needs:artifacts are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. If artifacts:expire_in is omitted, the instance default controls expiry; GitLab also keeps artifacts from the most recent successful pipeline on each ref by default unless that behavior is disabled. Cache is a performance optimization, not an authoritative release/evidence store. Protected and non-protected refs use separate caches by default; disabling that boundary or using cache:unprotect broadens who can read/write the same cache and must be a deliberate trust decision. Hosted storage/compute quotas and billing are volatile, so labs use tiny files and a no-runner fixture path.

1. Artifact versus cache: ask what failure you can tolerate

An artifact failure means an expected pipeline output is missing. A cache failure should mean only “the job may be slower.” That difference is the most useful design rule in the chapter.

Question Artifact Cache
Can loss change release/evidence correctness? Usually yes. Should be no.
Must data bind to producer job/SHA? Usually yes. Not necessarily.
Should later pipelines reuse it? Only deliberately through APIs/other pipeline patterns. Often yes.
Can it be recomputed cheaply? Maybe, but retention/identity still matter. Expected.
Primary optimization? Traceable data transfer. Speed/cost.

2. Short retention versus audit/reproducibility

Retention is not “longer is safer.” Keeping every transient artifact forever expands storage cost, sensitive-data exposure, and discovery scope. Expiring release evidence too quickly destroys incident reconstruction. Define artifact classes, then map each class to retention and promotion.

Class Example Suggested policy reasoning
Ephemeral diagnostics Temporary logs/screenshots from a test. Short retention; enough for routine debugging.
Review evidence JUnit/coverage/report used during an MR. Retain through review plus an agreed buffer.
Release candidate Build output tied to a commit/pipeline. Retain until promotion/rejection; then promote durable identity or expire.
Release payload Published package/image/release asset. Use the appropriate durable distribution surface and organizational retention policy.
Compliance evidence Signed/verified evidence required by policy. Retention derives from policy/legal requirement; do not rely on accidental keep-latest behavior.

3. Broad automatic downloads versus explicit dependencies

Automatic earlier-stage artifact downloads are convenient in small pipelines. As pipelines grow, they create hidden coupling and unnecessary network/storage transfer. Explicit dependencies or needs makes the consumer contract visible and reduces accidental downloads.

The tradeoff is maintenance: every real producer/consumer edge must remain accurate. Use explicit flow when artifacts are large, sensitive, expensive, or correctness-critical. Keep stage-default behavior only when the stage truly represents a shared data boundary and the set is small/stable.

4. dependencies versus needs

Choice Use when Scheduling effect Artifact effect
dependencies You keep stage barriers but want a narrower artifact set. None; stage order remains. Downloads only listed earlier-stage jobs.
needs You need an explicit DAG dependency and possibly earlier start. Creates DAG edge. Downloads artifacts only from needed jobs unless disabled.
dependencies: [] Job should receive no earlier artifacts. None. Disables automatic artifact download.
needs: artifacts:false Job depends on producer completion but not its files. Keeps DAG edge. No artifact download from that edge.

Do not use both keywords in one job as a clever merge of behaviors. A single clear scheduling/data model is easier to validate.

5. Cache key design: compatibility + trust

A good cache key changes when cached data is no longer compatible. Lockfile-based keys are stronger than “one cache forever.” Branch/ref keys reduce cross-branch influence. Protected/non-protected separation reduces trust crossing.

cache:
  key:
    files:
      - requirements.lock
  paths:
    - .cache/pip/
  policy: pull

For a default-branch cache builder, a deliberate pattern is pull-push on the trusted/default branch and pull on feature branches. That limits which contexts may update shared acceleration state.

6. Shared cache performance versus poisoning risk

Disabling protected/non-protected separation can increase reuse, but it also allows a less-trusted branch to influence data consumed by a trusted branch. cache:unprotect: true is therefore a trust-policy change, not a performance toggle. If cache content can execute code—compiled tools, hooks, generated scripts—the poisoning impact is especially high.

7. Cache topology versus portability

A cache hit depends on runner/cache topology. A single persistent runner can keep local cache. An autoscaled fleet usually needs distributed cache. Two unrelated self-managed runners without a shared backend may both use the same YAML key and still miss. Design correctness so it does not depend on a particular runner retaining local state.

8. Reports versus ordinary downloadable files

Report configuration should exist because GitLab needs structured information, not because the team wants a zip file. Some report features are Free while security-product reports have separate tier/product semantics covered later. For Chapter 16, JUnit is the mandatory Free example because its semantics are stable and safe.

If humans also need to browse the XML, add it to artifacts:paths deliberately. Remember that storing a report for platform ingestion and storing it as a browsable artifact can increase storage/access surface.

9. Artifact access is not secret management

artifacts:access controls UI/API download access. It does not transform a file into a secret. Pipeline users with sufficient CI access, downstream mechanics, administrators, backups, or storage operators can introduce additional exposure paths. Secrets belong in secret/identity systems; sensitive generated evidence should be minimized/redacted before upload.

10. Job-author versus platform-admin decisions

Layer Job author owns Platform administrator owns
Artifact Paths, naming, expiry, access, downstream flow. Maximum sizes, storage backend, cleanup workers, object storage lifecycle/backup.
Cache Keys, paths, policy, trust boundary. Runner fleet, distributed cache backend, bucket lifecycle/network/encryption.
Reports Correct schema/path and data minimization. Platform availability/limits/upgrades.

On GitLab.com/Dedicated, users consume managed storage behavior; on Self-Managed, administrators must additionally verify object storage and backup/recovery configuration. Do not require those admin steps for this course chapter.

11. Worked scenario: a 12-person service team

The team builds one service, runs unit tests on every MR, and releases weekly. Choose data surfaces:

Need Choice Rationale across quality attributes
Dependency downloads Cache keyed by lockfile; feature branches pull, trusted builder pushes. Performance without making cache authoritative; narrows poisoning risk.
Compiled candidate Artifact named by SHA, retained through release decision. Traceability/reproducibility.
JUnit results JUnit report, short/medium retention. Maintainable review UX and evidence.
Published release Promote to package/container/release surface in later chapters. Durable distribution, identity, access, retention.
Diagnostic dump containing customer data Do not upload by default; create sanitized/minimized evidence. Security/privacy over convenience.

12. Decision checklist

  • What is the authoritative source if this object disappears?
  • Which exact producer job/SHA must the consumer trust?
  • How long is the object operationally/legal useful?
  • Who may download or overwrite it?
  • Can an untrusted ref write data later consumed by a trusted ref?
  • How much network/storage/runner time does the design consume?
  • What happens if the cache/storage backend is empty?

Knowledge check

A generated binary must be used for a release candidate. Artifact or cache?

Why might explicit dependencies improve security?

When is a shared cache key dangerous?

What is the platform administrator’s role in Self-Managed artifact storage?

Why is artifacts:access not a secret manager?

Summary

Data-flow design is a policy problem: authority, identity, trust, retention, and cost determine the correct GitLab object. Artifacts carry pipeline outputs, reports feed platform interpretation, caches accelerate reproducible work, and explicit edges keep consumers from depending on accidental data.

Official references

Next lesson

Diagnose when the data contract breaks

Lesson 4 preserves missing/expired artifact evidence, separates cache misses from correctness failures, diagnoses cache poisoning and sensitive uploads, and repairs DAG data-flow mistakes without hiding their original causes.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.