Artifacts, Reports, Cache, Dependencies, Retention, and Data Flow Between Jobs: Concepts, Architecture, and Mental Model
Treat build outputs, reports, and caches as distinct data planes with explicit authority, lifetime, access, integrity, and downstream flow.
Learning objectives
- Distinguish ordinary artifacts, report artifacts, cache, and repository source by authority and lifecycle.
-
Explain default artifact download behavior and how
dependenciesorneeds:artifactschanges it. - Reason about retention, access, object storage, sensitivity, integrity, and storage cost.
- Explain why report files are both stored artifacts and structured input to GitLab features.
- Inspect pipeline/job artifact metadata before changing data-flow configuration.
dependencies, and needs:artifacts are
available in GitLab Free/Premium/Ultimate across GitLab.com,
Self-Managed, and Dedicated. If artifacts:expire_in is
omitted, the instance default controls expiry; GitLab also keeps
artifacts from the most recent successful pipeline on each ref by
default unless that behavior is disabled. Cache is a performance
optimization, not an authoritative release/evidence store. Protected
and non-protected refs use separate caches by default; disabling that
boundary or using cache:unprotect broadens who can
read/write the same cache and must be a deliberate trust decision.
Hosted storage/compute quotas and billing are volatile, so labs use
tiny files and a no-runner fixture path.
1. The problem: “a file exists somewhere in CI” is not a data contract
Chapter 15 made job ordering explicit. Chapter 16 adds the second half of that model: what data crosses each job boundary. A build output, a JUnit XML report, a dependency cache, and a cloned repository can all appear as files inside a runner workspace, but they do not mean the same thing and they do not have the same lifetime.
If a team treats cache as a release store, a cache miss becomes a release outage. If it treats every artifact as permanent evidence, storage grows without a retention policy. If it uploads credentials in a report, GitLab faithfully stores the sensitive file. Correct CI data flow therefore starts by classifying the data before choosing YAML.
2. Mental model: four data planes around one job
flowchart LR G[(Git repository/ref)] -- checkout --> J[Job workspace] C[(Cache store)] -- restore before artifacts --> J A[(Earlier job artifacts)] -- download --> J J -- archive authoritative output --> O[(Job artifacts)] J -- structured report --> R[(GitLab report ingestion)] J -- reusable acceleration state --> C O -- explicit downstream flow --> D[Later job]
The repository supplies versioned source. Cache is reusable
acceleration state and can span pipelines. Job artifacts are
pipeline outputs stored by GitLab. Report artifacts are files with
a schema GitLab parses to enrich pipeline/MR features. Downstream
jobs may receive artifacts automatically by stage or explicitly
through dependencies/needs.
3. Classify data by authority, consumer, and lifetime
| Data type | Primary purpose | Typical lifetime | Authoritative? | Example |
|---|---|---|---|---|
| Repository file | Versioned source/configuration. | Git history. | Yes for source/config. | Lockfile, source code, CI YAML. |
| Ordinary artifact | Carry a build/test output through pipeline or expose it for download. | Pipeline/release/audit policy. | Often yes for that pipeline output. | Binary, generated docs, manifest. |
| Report artifact | Provide structured data GitLab interprets. | Enough for review/audit policy. | Authoritative for the report producer, not necessarily for release payload. | JUnit XML, coverage report, dotenv report. |
| Cache | Avoid repeating downloads/build preparation. | Disposable and rebuildable. | No. | Package-manager download directory. |
A simple test helps: if this object disappears, can I deterministically rebuild or redownload it without changing the release identity? If yes, cache may be appropriate. If no, it belongs in an artifact/package/release/registry system with explicit identity and retention.
4. Job artifacts: archived outputs with lifecycle and access rules
A job artifact is created after the job script finishes. Paths are relative to the project working directory. GitLab receives the archive and associates it with the job, commit, pipeline, and ref. That identity is valuable: it lets you answer “which pipeline produced this file?”
build:
stage: build
script:
- mkdir -p out
- printf 'commit=%s\n' "$CI_COMMIT_SHA" > out/build.txt
artifacts:
name: "build-$CI_COMMIT_SHORT_SHA"
paths:
- out/build.txt
expire_in: 7 days
access: developer
artifacts:access currently supports all,
developer, maintainer, and
none for UI/API downloads. It also applies to report
artifacts. This is a download-access control; it is not a universal
secrecy boundary for all downstream CI paths.
5. Report artifacts: stored files plus platform interpretation
artifacts:reports tells GitLab that a file has a known
schema. For JUnit, GitLab parses test cases and renders them in
pipeline/MR interfaces. The report does not decide whether the job
succeeds: the script exit code still does that. This separation is
useful when you want a failing test command to fail the job
and preserve the report for diagnosis.
test:
stage: test
script:
- ./run-tests-and-write-junit.sh
artifacts:
when: always
reports:
junit: reports/junit.xml
paths:
- reports/junit.xml
expire_in: 14 days
GitLab documents report artifacts as uploaded regardless of job
result. Adding the same report file under
artifacts:paths makes it browsable/downloadable as an
ordinary artifact too; treat that as a storage/access decision, not
a requirement for report ingestion.
6. Cache: acceleration state, not evidence
A cache may be reused by later jobs and later pipelines. Runner topology determines where it lives: local runner storage, a distributed cache such as object storage, or another runner-supported cache backend. It is intentionally weaker than artifact identity.
test:
cache:
key:
files:
- package-lock.json
paths:
- .npm/
policy: pull-push
script:
- npm ci --cache .npm --prefer-offline
- npm test
The safe production assumption is “cache can be absent, stale, or evicted.” The job must be able to reconstruct it. A cache key should capture compatibility inputs such as a lockfile, platform, toolchain, or branch trust boundary.
7. Restore order matters
GitLab Runner restores caches before artifacts. If cache and artifact paths overlap, the later artifact extraction can overwrite files restored from cache. Avoid overlapping authoritative and disposable paths; otherwise debugging becomes a question of extraction order rather than application behavior.
8. Default artifact download versus explicit flow
Without needs or dependencies, later-stage
jobs normally fetch artifacts from jobs in earlier stages. That
convenience can hide dependencies. dependencies narrows
which earlier-stage artifacts a job downloads.
dependencies: [] downloads none.
With a needs DAG, the model changes: the job may start
before an entire previous stage is complete, so it can only download
artifacts from jobs named in needs. Use
artifacts: true (default for a normal needs edge) or
false deliberately.
package:
needs:
- job: build
artifacts: true
- job: lint
artifacts: false
script:
- test -f out/build.txt
- ./package.sh
9. Retention is part of reproducibility
artifacts:expire_in starts its clock when GitLab stores
the artifact. If it is omitted, the instance default applies. GitLab
also keeps artifacts from the latest successful pipeline on each ref
by default unless the project/instance setting is disabled.
Therefore “expire in seven days” does not necessarily mean “every
copy disappears exactly seven days later.”
Retention should map to a business reason: fast feedback artifacts
can be short-lived; release evidence may need longer retention or
promotion to a release/package/registry system designed for durable
distribution. Do not use expire_in: never as a
substitute for an evidence-retention policy.
10. Sensitive content and cache poisoning are different risks
Artifacts are stored and downloadable according to project/access
rules, so a file containing a credential is already an incident.
Cache poisoning is different: an untrusted pipeline writes content
under a cache key that a more trusted job later consumes. GitLab
separates protected and non-protected ref caches by default, and
current cache keys receive protection-related suffixes. Disabling
that boundary or using cache:unprotect: true should
require a documented trust argument.
11. GitLab control plane versus storage backend
On GitLab.com, the service owns the storage implementation. On Self-Managed, administrators may configure local storage or object storage for artifacts and runner-side distributed cache separately. A job author should reason about logical lifecycle/access; a platform administrator additionally owns bucket lifecycle, encryption, capacity, replication, and recovery. Do not blend those responsibilities into one “CI storage” concept.
12. Read-only inspection first
Before changing YAML, open Build → Artifacts and the relevant job/pipeline. Record job ID, pipeline ID, SHA, artifact filenames/types, expiry where visible, and whether the report appears in GitLab’s structured UI. Machine-readable inspection can stay metadata-only:
PROJECT_ID="12345678"
PIPELINE_ID="123456789"
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
--jq '.[] | {id,name,status,stage,ref,artifacts_file,artifacts}'
Do not download or print sensitive artifacts merely to prove that they exist. Presence, size, type, checksum, and access behavior are usually enough.
13. DevOps connection: make every edge a data contract
A dependable pipeline states both the scheduling dependency and the data dependency. The producer identifies an immutable output; the consumer names exactly what it needs; retention matches recovery/audit needs; cache accelerates reconstruction but is never the only copy. That model survives runner churn, cache eviction, DAG refactors, and incident review.
Knowledge check
Why is a cache a poor place for the only copy of a release binary?
Cache is disposable acceleration state and can miss or be evicted; a release payload needs durable identity and retention as an artifact/package/release/registry object.
What changes when a job starts using
needs?
Its scheduling becomes DAG-based and artifact downloads are limited to needed jobs rather than every previous-stage job.
Does a JUnit report make a failing test job fail?
No. The script exit status controls job success; the JUnit report supplies structured test data for GitLab.
Why is
cache:unprotect: true security-sensitive?
It allows protected and unprotected contexts to share cache content, expanding who can influence data trusted jobs may consume.
If expire_in is absent, what determines artifact
expiry?
The GitLab instance default, subject also to the keep-latest-successful-artifacts behavior unless disabled.
What is the safest first inspection for a suspicious artifact?
Inspect metadata, producer job/pipeline/SHA, access, type and retention; avoid dumping potentially sensitive file contents.
Summary
Repository source, artifacts, reports, and caches are distinct data
planes. Artifacts carry identified pipeline outputs, reports feed
structured GitLab features, and cache trades integrity guarantees
for speed. Explicit dependencies/needs,
bounded retention, download access, and cache trust boundaries turn
those files into a production-grade data-flow model.
Official references
- GitLab Docs — Job artifacts
- GitLab Docs — Job artifacts troubleshooting
- GitLab Docs — Job Artifacts API
- GitLab Docs — CI/CD artifacts reports types
- GitLab Docs — Unit test reports
- GitLab Docs — Unit test report examples
- GitLab Docs — CI/CD caching
- GitLab Docs — CI/CD caching examples
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — Pass dotenv variables to specific jobs
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — Job artifacts administration
- GitLab Docs — Object storage
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.