Artifacts, Reports, Retention, Dependencies, needs:artifacts, and Cross-Job Data Flow: Configuration, Design Choices, and Tradeoffs
Artifact design is a data-contract decision. This lesson compares generic artifacts with typed reports, broad versus allowlisted paths, short versus long retention, UI/API access controls, same-pipeline transfer versus durable registry/package promotion, and the audit evidence required for each choice.
Learning objectives
- Choose generic artifacts or typed reports according to whether the output is for machines, humans, GitLab UI ingestion, or deployment.
- Choose retention and access settings from evidence lifetime, sensitivity, storage cost, and recovery requirements.
- Keep artifact paths narrow and distinguish artifact transfer from package/registry promotion for durable distribution.
- Explain current access semantics and why artifact download authorization is separate from CI_JOB_TOKEN authorization.
- Document source SHA, producer job, digest, retention, access, and consumer assumptions as part of a production artifact contract.
1. Artifact design is interface design
When one job produces files for another job, the producer has an
interface. Its contract includes file names/paths, format, source
identity, digest, sensitivity, retention window, access assumptions,
and who is allowed to consume it. A broad
artifacts: paths: [.] is not a convenient interface—it
is an undocumented export of the workspace.
Design the artifact contract so a reviewer can answer: what was produced, by which source/job, for whom, for how long, and how do we know it is the same object?
2. Generic artifacts versus typed reports
| Question | Generic artifact | Typed report |
|---|---|---|
| Primary purpose | Retain/download files | Feed a GitLab feature with structured data |
| Examples | Build output, logs, evidence manifest | JUnit, coverage, code quality, security, dotenv |
| Human browsing | Yes when retained as paths |
Add artifacts:paths if direct browsing is
needed
|
| Success meaning | Upload exists | Report upload/ingestion is distinct from job pass/fail |
| Deployable output? | Possibly, if contract says so | Usually evidence/metadata, not the release binary |
Do not deploy a JUnit XML simply because it is an artifact. Conversely, a deployable binary can have a separate test report that proves validation without being the deployable object itself.
3. Short versus long retention
Retention should match the consumer horizon. A two-hour debug trace may not need months of storage. A release candidate needed for a week-long approval window must not expire after one hour. Compliance evidence may need a dedicated retention strategy beyond ordinary CI artifact storage.
| Output | Typical concern | Design direction |
|---|---|---|
| Temporary debug file | Exposure + storage | Short retention, restricted access, no secrets |
| MR test report | Review window | Retention long enough for review/incident triage |
| Release candidate | Promotion/rollback identity | Prefer durable package/container registry or release asset after verification |
| Generated documentation preview | Reviewer access | Bound retention and explicit access; treat rendered HTML as untrusted content |
4. Broad paths versus allowlisted outputs
Narrow paths are both a security control and a maintainability control. A workspace may contain source, temporary files, downloaded dependencies, debug logs, generated credentials, and tool caches. Export only the files the next consumer genuinely needs.
artifacts:
paths:
- dist/app.bin
- evidence/app.sha256
- evidence/provenance.txt
This is preferable to archiving the entire project directory. If a new required file is introduced, update the interface deliberately and review its sensitivity.
5. Access is a separate authorization decision
artifacts:access controls UI/API download eligibility.
Current documented values are all,
developer, maintainer, and
none. The maintainer option is a newer
value, so older GitLab installations require compatibility review.
Do not assume this setting blocks every CI-to-CI retrieval route. Job-token authorization and project CI/CD visibility are separate controls. For sensitive-but-nonsecret evidence, combine narrow paths, suitable access, short retention, and project membership policy.
6. Default fetching, dependencies, or needs:artifacts?
| Situation | Choice | Reason |
|---|---|---|
| Small linear pipeline, every later job really needs everything | Default fetching | Lowest configuration overhead |
| Stage barriers are correct; consumer needs only named producers | dependencies |
Narrow data transfer without changing scheduling |
| Consumer can run as soon as named producer finishes | needs:artifacts |
Encode both DAG dependency and intentional artifact transfer |
| Consumer needs no artifacts |
dependencies: [] or appropriate
needs with artifacts disabled
|
Avoid waste and accidental coupling |
needs and dependencies in one job. A
reader should not have to reconcile two different transfer models.
7. Cross-job transfer is not the same as durable promotion
Job artifacts are excellent same-pipeline outputs and evidence, but release distribution often needs a longer-lived, versioned object with repository semantics. For a production package or container, promote the exact verified artifact/digest into the appropriate package/container registry or release process rather than relying indefinitely on a branch’s latest job artifact.
That promotion is a side effect and belongs to later course chapters. The Chapter 10 rule is simpler: build once, preserve identity, and never rebuild silently merely because a later consumer cannot find the original output.
8. Download by exact job versus “latest by ref”
The Artifacts API can target a specific job ID or resolve an
artifact by job name + ref from the latest successful pipeline.
Those are different identity guarantees. For incident response or
release evidence, prefer exact job/pipeline identity when you
already know it. A ref such as main moves over time.
Current API behavior also supports direct report downloads by job ID/file type. That capability should not encourage scripts to select ambiguous “latest” artifacts for deployment.
9. Worked scenario: four outputs, four different contracts
A service pipeline emits: a Linux binary, a JUnit report, a short debug trace, and dependency cache data. Choose separately:
| Output | Correct state | Retention/access | Consumer |
|---|---|---|---|
| Linux binary | Artifact now; registry/package at promotion | Enough for approval; restricted according to project policy | Verification/package/deploy jobs |
| JUnit XML | Typed report (+ path if browsing) | Review/triage window | GitLab test UI and reviewers |
| Debug trace | Generic artifact | Short, restricted, sanitized | Troubleshooter |
| Package-manager downloads | Cache | Performance lifetime, keyed by lockfile/toolchain | Future matching jobs/pipelines |
The decision is not “everything is an artifact.” It is “each data type gets the state system that matches its correctness and lifecycle role.”
10. Decision checklist
Can I name producer project/pipeline/job/SHA?
Can I verify a digest or structured report identity?
Are paths narrowly allowlisted?
Will it exist for every intended consumer?
Who can download it, and through which mechanism?
Should this become a durable package/registry/release object instead?
Knowledge check
Why might a report also appear under artifacts:paths?
Typed reports feed GitLab features; adding the file under paths can also make the raw report directly browsable/downloadable.
When is long artifact retention a bad default?
When it increases storage cost or exposure without a real audit/recovery need; retention should match the consumer horizon.
Why is broad workspace archiving risky?
It can unintentionally include source, caches, debug data, generated credentials, or other files outside the intended contract.
Why is “latest artifact on main” weaker than an exact job ID?
The branch moves and later pipelines can replace which artifact “latest” resolves to; exact job/SHA/digest gives stable provenance.
When should an output move to a registry/package/release process?
When it becomes a durable, versioned distribution/promotion object rather than ordinary same-pipeline evidence.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. Artifact/report retention, access, API behavior, and cross-job transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.
- Job artifacts — creation, paths, expiry, download, artifact browsing, access, latest-success retention, and default previous-stage fetching.
-
CI/CD YAML syntax reference
— authoritative
artifacts,artifacts:access,artifacts:expire_in,dependencies, andneeds:artifactssemantics. - CI/CD artifacts report types — typed report ingestion and report-specific GitLab UI behavior.
- Unit test reports — JUnit report configuration and display.
- Job Artifacts API — artifact archive/file/report download, keep, and delete operations with exact job identity.
- Troubleshooting job artifacts — artifact expiry and upload/report problems.
- Caching in GitLab CI/CD — artifact-versus-cache boundary.
-
Make jobs start earlier with
needs— DAG dependency edges and artifact-transfer interaction.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.