Artifacts, Reports, Retention, Dependencies, needs:artifacts, and Cross-Job Data Flow: Concepts, Architecture, and Mental Model
A job workspace is temporary; evidence that later jobs or humans must inspect needs an explicit transfer and retention contract. This lesson builds that contract around producer identity, selected paths and typed reports, upload, retention/access, downstream download, digest verification, and the boundary between an artifact and a release/package/registry object.
Learning objectives
- Explain why a runner workspace is not durable cross-job state and why artifacts are an explicit transfer/evidence contract.
- Identify producer pipeline/job/SHA, artifact name/path/type/digest, expiry/access policy, consumer job, and transfer mechanism.
- Distinguish generic artifact archives from typed artifacts:reports and from caches, registries, packages, and releases.
- Explain default previous-stage downloading, dependencies, and needs:artifacts without conflating ordering with file transfer.
- Preserve artifact identity and sensitivity evidence without placing secrets in artifacts or logs.
1. The practical problem: a successful job can still lose its output
Chapter 09 made job dependencies explicit. That graph answers when a consumer may run, but not automatically which files the consumer should receive or how long those files remain available. A runner workspace can disappear after a job, a Docker container can be destroyed, and a Kubernetes pod can be removed. If a file matters after the producing job, the pipeline needs an explicit retained-output contract.
GitLab job artifacts provide that contract for pipeline outputs and evidence. A producer selects paths or typed reports from its project workspace, GitLab stores them with the producing job, and later jobs or authorized users can retrieve them according to retention and access rules.
2. Mental model: workspace → selected outputs → retained evidence → verified consumer
Read the model from left to right. The producer begins with a
checked-out source revision and temporary workspace. Its script
creates files. The artifacts declaration selects only
intended paths/reports. Runner uploads those outputs to GitLab,
where the artifact record inherits producer identity and
retention/access policy. A consumer then downloads through the
pipeline transfer mechanism, verifies identity/digest, and uses the
output.
flowchart TD
A[Source ref + exact SHA] --> B[Producer job workspace]
B --> C[Selected artifacts: paths / reports]
C --> D[GitLab artifact record
producer job + pipeline + expiry/access]
D --> E[Consumer download
default / dependencies / needs:artifacts]
E --> F[Digest + identity verification]
F --> G[Review / test / package / deploy input]
The arrows are causal. If the producer did not create the path, upload fails or collects nothing. If the retention window expires, a later consumer may fail even though the original build was green. If the consumer rebuilds instead of downloading, it has created a different object whose equivalence must not be assumed.
3. Keep artifact state separate from every neighboring state
| State | Evidence to preserve | What it does not prove |
|---|---|---|
| Repository/source | CI_COMMIT_SHA, ref, repository tree |
That a build output was uploaded or retained |
| Compiled pipeline |
Artifact declarations, dependencies/needs
edges
|
That files actually existed at upload time |
| Producer job | Pipeline/job ID, status, runner/executor, logs | That a consumer downloaded the intended bytes |
| Artifact/archive | Name, selected paths, size/digest, expiry/access, producer identity | That a typed report was ingested or a deployment used it |
| Typed report | Report type/path and GitLab report UI/API result | That the report is a deployable build artifact |
| Consumer job | Downloaded path, producer ID, verified digest | That external target state changed |
| Deployment/external | Environment/deployment/target identity and health | That the bytes came from the intended producer unless linked by digest |
4. Generic artifact versus typed report
A generic artifact is primarily a retained file/archive. A typed
report tells GitLab how to interpret a file for a feature such as
test results, coverage, code quality, or security. The same
underlying file can sometimes be listed in both
artifacts:reports and artifacts:paths when
you want GitLab-native ingestion and convenient
browsing/download.
| Need | Use | Example |
|---|---|---|
| Pass a build output | artifacts:paths |
dist/app.txt or a compiled binary |
| Show unit-test results | artifacts:reports:junit |
JUnit XML |
| Browse the JUnit file too | Report + path | Declare JUnit under reports and paths |
| Reuse dependencies for speed | Cache, not artifact | Package-manager cache keyed by correctness inputs |
| Distribute a long-lived version | Package/container registry or release asset | Version/digest promotion rather than ordinary job retention |
artifacts:reports are uploaded for report processing
even when the job fails. Do not interpret “report was uploaded” as
“job passed.”
5. Three same-pipeline fetching models
Without an explicit restriction, jobs in later stages download
artifacts from all earlier-stage jobs. This is convenient for small
pipelines but can become ambiguous and wasteful.
dependencies narrows which earlier-stage jobs provide
artifacts while retaining normal stage scheduling. A job that uses
needs no longer relies on that default previous-stage
download model; its needs entries can specify whether
each needed producer’s artifacts are fetched.
| Mechanism | Ordering model | Artifact source | Use when |
|---|---|---|---|
| Default | Stages | All earlier-stage artifact producers | Tiny pipeline where broad implicit transfer is acceptable |
dependencies |
Stages | Named earlier-stage jobs | Keep stage barriers but narrow downloads |
needs:artifacts |
DAG via needs |
Named needed jobs with artifact flag | Start as soon as producer completes and receive only intended DAG inputs |
needs and
dependencies in the same job. Encode one coherent
scheduling/transfer model.
6. Give every artifact an evidence identity
A filename such as app.zip is not enough. Two pipelines
can produce files with the same name. The minimum evidence tuple is
producer project + pipeline ID + job ID/name + exact source SHA +
artifact logical name/path + cryptographic digest + retention/access
assumptions.
build_app:
stage: build
script:
- mkdir -p dist evidence
- printf 'release-candidate:%s\n' "$CI_COMMIT_SHA" > dist/app.txt
- sha256sum dist/app.txt > evidence/app.sha256
- printf 'pipeline=%s\njob=%s\nsha=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" > evidence/producer.txt
artifacts:
name: "app-$CI_PIPELINE_ID-$CI_JOB_ID"
paths:
- dist/app.txt
- evidence/app.sha256
- evidence/producer.txt
expire_in: 7 days
The example prints identifiers, not secrets. It records enough provenance for a later job to prove which bytes it received.
7. Retention is part of correctness
artifacts:expire_in starts its retention clock when
GitLab stores the artifact. If omitted, GitLab uses the
instance-wide default. Current GitLab also normally keeps artifacts
from the most recent successful pipeline for each ref even when an
expire_in value exists, unless that keep-latest
behavior is disabled. Therefore “expires in seven days” is not a
universal statement about actual deletion time.
Retention should outlive every intended consumer: later pipeline stages, human review, deployment windows, incident investigation, compliance evidence, or rollback. Conversely, long retention of large or sensitive outputs has storage and exposure cost.
8. Access controls reduce exposure; they do not make artifacts a secret vault
artifacts:access controls who can download a job’s
artifacts through the GitLab UI/API. Current values include
all, developer, maintainer,
and none. This control also applies to report
artifacts, but it is not a substitute for preventing secrets from
entering the artifact in the first place.
Job-token access has its own authorization model. A UI/API artifact access setting and CI job-token cross-project permissions are separate security states.
9. Read-only inspection before changing artifact flow
- Record pipeline source, ref, and exact SHA.
- Open the compiled configuration and identify every artifact producer and declared consumer.
- Record producer job ID/status and the exact artifact paths/reports.
- Inspect the job’s artifact metadata/UI and retention/access policy without deleting or keeping anything yet.
-
For a consumer, identify whether transfer is default,
dependencies, orneeds:artifacts. - Compare producer and consumer digests when the file is correctness-relevant.
Pipeline source/ref/SHA
Pipeline + job ID/name
Path/name/digest/type
Expiry + keep-latest assumption
UI/API access + job-token context
Transfer mechanism + verified digest
10. Misconceptions to remove now
| Misconception | Correction |
|---|---|
| “The next job can read the previous runner filesystem.” | Runner workspaces are not a portable data contract; upload/download artifacts explicitly. |
| “A report is the release binary.” | A typed report is evidence for GitLab features; deploy the intended build artifact/digest. |
| “A cache is a trusted build artifact.” | Caches are performance state and can be missing/stale; artifacts carry intended pipeline outputs. |
| “Same filename means same build.” | Verify producer job/SHA and digest. |
| “expire_in is the exact deletion moment.” | Keep-latest and cleanup scheduling affect actual retention. |
| “Private artifacts may contain secrets.” | Do not place secrets in artifacts at all. |
11. Micro-lab: predict the transfer before running
Given one build job that produces dist/app.txt and
test/report.xml, predict the answer for each consumer:
-
A later-stage job with no
dependencies: which earlier artifacts does it fetch by default? -
A later-stage job with
dependencies: [build_app]: what does it fetch? -
A DAG job with
needs: [{job: build_app, artifacts: true}]: when can it start and what does it receive? -
A DAG job with the same need but
artifacts: false: what control dependency remains and what data dependency disappears?
Do not execute yet. The goal is to separate ordering from transfer before syntax becomes habit.
Knowledge check
Why is a successful producer job not enough to prove downstream reproducibility?
Because you still need proof that the intended paths were uploaded, retained, downloaded by the consumer, and match the producer’s source/digest identity.
What is the difference between an artifact path and a typed report?
A path retains a file/archive; a typed report also tells GitLab how to interpret the file for report-specific UI or processing.
When should dependencies be preferred over default fetching?
When stage scheduling is correct but the consumer should download artifacts only from explicitly named earlier-stage producers.
What changes when a consumer uses needs?
It can run according to explicit DAG dependencies instead of stage barriers, and artifact transfer should be declared on the needed producers rather than relying on broad previous-stage fetching.
Why should a deployment record a digest?
A digest proves the deployed bytes match the intended producer artifact instead of a same-named or rebuilt substitute.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. Artifact/report retention, access, API behavior, and cross-job transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.
- Job artifacts — creation, paths, expiry, download, artifact browsing, access, latest-success retention, and default previous-stage fetching.
-
CI/CD YAML syntax reference
— authoritative
artifacts,artifacts:access,artifacts:expire_in,dependencies, andneeds:artifactssemantics. - CI/CD artifacts report types — typed report ingestion and report-specific GitLab UI behavior.
- Unit test reports — JUnit report configuration and display.
- Job Artifacts API — artifact archive/file/report download, keep, and delete operations with exact job identity.
- Troubleshooting job artifacts — artifact expiry and upload/report problems.
- Caching in GitLab CI/CD — artifact-versus-cache boundary.
-
Make jobs start earlier with
needs— DAG dependency edges and artifact-transfer interaction.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.