Artifacts, Reports, Cache, Dependencies, Retention, and Data Flow Between Jobs: Configuration, Design Choices, and Tradeoffs
Turn artifact/cache mechanics into maintainable policy for retention, dependency clarity, storage cost, and cache trust boundaries.
Learning objectives
- Choose artifact versus cache from authority and rebuildability rather than convenience.
- Set retention from recovery/audit needs instead of defaulting to permanent storage.
- Choose automatic versus explicit artifact downloads based on dependency clarity.
- Design cache keys/policies around trust and compatibility boundaries.
- Separate job-author decisions from Self-Managed storage administration.
dependencies, and needs:artifacts are
available in GitLab Free/Premium/Ultimate across GitLab.com,
Self-Managed, and Dedicated. If artifacts:expire_in is
omitted, the instance default controls expiry; GitLab also keeps
artifacts from the most recent successful pipeline on each ref by
default unless that behavior is disabled. Cache is a performance
optimization, not an authoritative release/evidence store. Protected
and non-protected refs use separate caches by default; disabling that
boundary or using cache:unprotect broadens who can
read/write the same cache and must be a deliberate trust decision.
Hosted storage/compute quotas and billing are volatile, so labs use
tiny files and a no-runner fixture path.
1. Artifact versus cache: ask what failure you can tolerate
An artifact failure means an expected pipeline output is missing. A cache failure should mean only “the job may be slower.” That difference is the most useful design rule in the chapter.
| Question | Artifact | Cache |
|---|---|---|
| Can loss change release/evidence correctness? | Usually yes. | Should be no. |
| Must data bind to producer job/SHA? | Usually yes. | Not necessarily. |
| Should later pipelines reuse it? | Only deliberately through APIs/other pipeline patterns. | Often yes. |
| Can it be recomputed cheaply? | Maybe, but retention/identity still matter. | Expected. |
| Primary optimization? | Traceable data transfer. | Speed/cost. |
2. Short retention versus audit/reproducibility
Retention is not “longer is safer.” Keeping every transient artifact forever expands storage cost, sensitive-data exposure, and discovery scope. Expiring release evidence too quickly destroys incident reconstruction. Define artifact classes, then map each class to retention and promotion.
| Class | Example | Suggested policy reasoning |
|---|---|---|
| Ephemeral diagnostics | Temporary logs/screenshots from a test. | Short retention; enough for routine debugging. |
| Review evidence | JUnit/coverage/report used during an MR. | Retain through review plus an agreed buffer. |
| Release candidate | Build output tied to a commit/pipeline. | Retain until promotion/rejection; then promote durable identity or expire. |
| Release payload | Published package/image/release asset. | Use the appropriate durable distribution surface and organizational retention policy. |
| Compliance evidence | Signed/verified evidence required by policy. | Retention derives from policy/legal requirement; do not rely on accidental keep-latest behavior. |
3. Broad automatic downloads versus explicit dependencies
Automatic earlier-stage artifact downloads are convenient in small
pipelines. As pipelines grow, they create hidden coupling and
unnecessary network/storage transfer. Explicit
dependencies or needs makes the consumer
contract visible and reduces accidental downloads.
The tradeoff is maintenance: every real producer/consumer edge must remain accurate. Use explicit flow when artifacts are large, sensitive, expensive, or correctness-critical. Keep stage-default behavior only when the stage truly represents a shared data boundary and the set is small/stable.
4. dependencies versus needs
| Choice | Use when | Scheduling effect | Artifact effect |
|---|---|---|---|
dependencies |
You keep stage barriers but want a narrower artifact set. | None; stage order remains. | Downloads only listed earlier-stage jobs. |
needs |
You need an explicit DAG dependency and possibly earlier start. | Creates DAG edge. | Downloads artifacts only from needed jobs unless disabled. |
dependencies: [] |
Job should receive no earlier artifacts. | None. | Disables automatic artifact download. |
needs: artifacts:false |
Job depends on producer completion but not its files. | Keeps DAG edge. | No artifact download from that edge. |
Do not use both keywords in one job as a clever merge of behaviors. A single clear scheduling/data model is easier to validate.
5. Cache key design: compatibility + trust
A good cache key changes when cached data is no longer compatible. Lockfile-based keys are stronger than “one cache forever.” Branch/ref keys reduce cross-branch influence. Protected/non-protected separation reduces trust crossing.
cache:
key:
files:
- requirements.lock
paths:
- .cache/pip/
policy: pull
For a default-branch cache builder, a deliberate pattern is
pull-push on the trusted/default branch and
pull on feature branches. That limits which contexts
may update shared acceleration state.
6. Shared cache performance versus poisoning risk
Disabling protected/non-protected separation can increase reuse, but
it also allows a less-trusted branch to influence data consumed by a
trusted branch. cache:unprotect: true is therefore a
trust-policy change, not a performance toggle. If cache content can
execute code—compiled tools, hooks, generated scripts—the poisoning
impact is especially high.
8. Reports versus ordinary downloadable files
Report configuration should exist because GitLab needs structured information, not because the team wants a zip file. Some report features are Free while security-product reports have separate tier/product semantics covered later. For Chapter 16, JUnit is the mandatory Free example because its semantics are stable and safe.
If humans also need to browse the XML, add it to
artifacts:paths deliberately. Remember that storing a
report for platform ingestion and storing it as a browsable artifact
can increase storage/access surface.
9. Artifact access is not secret management
artifacts:access controls UI/API download access. It
does not transform a file into a secret. Pipeline users with
sufficient CI access, downstream mechanics, administrators, backups,
or storage operators can introduce additional exposure paths.
Secrets belong in secret/identity systems; sensitive generated
evidence should be minimized/redacted before upload.
10. Job-author versus platform-admin decisions
| Layer | Job author owns | Platform administrator owns |
|---|---|---|
| Artifact | Paths, naming, expiry, access, downstream flow. | Maximum sizes, storage backend, cleanup workers, object storage lifecycle/backup. |
| Cache | Keys, paths, policy, trust boundary. | Runner fleet, distributed cache backend, bucket lifecycle/network/encryption. |
| Reports | Correct schema/path and data minimization. | Platform availability/limits/upgrades. |
On GitLab.com/Dedicated, users consume managed storage behavior; on Self-Managed, administrators must additionally verify object storage and backup/recovery configuration. Do not require those admin steps for this course chapter.
11. Worked scenario: a 12-person service team
The team builds one service, runs unit tests on every MR, and releases weekly. Choose data surfaces:
| Need | Choice | Rationale across quality attributes |
|---|---|---|
| Dependency downloads | Cache keyed by lockfile; feature branches pull, trusted builder pushes. | Performance without making cache authoritative; narrows poisoning risk. |
| Compiled candidate | Artifact named by SHA, retained through release decision. | Traceability/reproducibility. |
| JUnit results | JUnit report, short/medium retention. | Maintainable review UX and evidence. |
| Published release | Promote to package/container/release surface in later chapters. | Durable distribution, identity, access, retention. |
| Diagnostic dump containing customer data | Do not upload by default; create sanitized/minimized evidence. | Security/privacy over convenience. |
12. Decision checklist
- What is the authoritative source if this object disappears?
- Which exact producer job/SHA must the consumer trust?
- How long is the object operationally/legal useful?
- Who may download or overwrite it?
- Can an untrusted ref write data later consumed by a trusted ref?
- How much network/storage/runner time does the design consume?
- What happens if the cache/storage backend is empty?
Knowledge check
A generated binary must be used for a release candidate. Artifact or cache?
Artifact, because identity and reliable retrieval matter; promote it later to a durable release/package/registry surface as appropriate.
Why might explicit dependencies improve security?
They reduce accidental download of unrelated or sensitive artifacts and make producer/consumer relationships auditable.
When is a shared cache key dangerous?
When less-trusted jobs can write content that more-trusted jobs later consume, especially executable/generated content.
What is the platform administrator’s role in Self-Managed artifact storage?
Configure/operate storage limits, object storage, lifecycle, backup/recovery and cleanup; job authors still own artifact content/retention intent.
Why is artifacts:access not a secret
manager?
It restricts UI/API artifact downloads but does not eliminate all CI/admin/storage exposure paths or provide secret lifecycle semantics.
Summary
Data-flow design is a policy problem: authority, identity, trust, retention, and cost determine the correct GitLab object. Artifacts carry pipeline outputs, reports feed platform interpretation, caches accelerate reproducible work, and explicit edges keep consumers from depending on accidental data.
Official references
- GitLab Docs — Job artifacts
- GitLab Docs — Job artifacts troubleshooting
- GitLab Docs — Job Artifacts API
- GitLab Docs — CI/CD artifacts reports types
- GitLab Docs — Unit test reports
- GitLab Docs — Unit test report examples
- GitLab Docs — CI/CD caching
- GitLab Docs — CI/CD caching examples
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — Pass dotenv variables to specific jobs
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — Job artifacts administration
- GitLab Docs — Object storage
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.