Chapter 10Lesson 01~155 minutes

Artifacts, Reports, Retention, Dependencies, needs:artifacts, and Cross-Job Data Flow: Concepts, Architecture, and Mental Model

A job workspace is temporary; evidence that later jobs or humans must inspect needs an explicit transfer and retention contract. This lesson builds that contract around producer identity, selected paths and typed reports, upload, retention/access, downstream download, digest verification, and the boundary between an artifact and a release/package/registry object.

ArtifactsReportsRetentionData flowEvidence

Learning objectives

  • Explain why a runner workspace is not durable cross-job state and why artifacts are an explicit transfer/evidence contract.
  • Identify producer pipeline/job/SHA, artifact name/path/type/digest, expiry/access policy, consumer job, and transfer mechanism.
  • Distinguish generic artifact archives from typed artifacts:reports and from caches, registries, packages, and releases.
  • Explain default previous-stage downloading, dependencies, and needs:artifacts without conflating ordering with file transfer.
  • Preserve artifact identity and sensitivity evidence without placing secrets in artifacts or logs.

1. The practical problem: a successful job can still lose its output

Chapter 09 made job dependencies explicit. That graph answers when a consumer may run, but not automatically which files the consumer should receive or how long those files remain available. A runner workspace can disappear after a job, a Docker container can be destroyed, and a Kubernetes pod can be removed. If a file matters after the producing job, the pipeline needs an explicit retained-output contract.

GitLab job artifacts provide that contract for pipeline outputs and evidence. A producer selects paths or typed reports from its project workspace, GitLab stores them with the producing job, and later jobs or authorized users can retrieve them according to retention and access rules.

Core invariant: downstream correctness should point to a specific producer pipeline/job/SHA and a verified artifact digest, not to “whatever file happens to exist” in a runner workspace or to a silently rebuilt substitute.

2. Mental model: workspace → selected outputs → retained evidence → verified consumer

Read the model from left to right. The producer begins with a checked-out source revision and temporary workspace. Its script creates files. The artifacts declaration selects only intended paths/reports. Runner uploads those outputs to GitLab, where the artifact record inherits producer identity and retention/access policy. A consumer then downloads through the pipeline transfer mechanism, verifies identity/digest, and uses the output.

Artifact lifecycle and data contract
            flowchart TD
              A[Source ref + exact SHA] --> B[Producer job workspace]
              B --> C[Selected artifacts: paths / reports]
              C --> D[GitLab artifact record
            producer job + pipeline + expiry/access]
              D --> E[Consumer download
            default / dependencies / needs:artifacts]
              E --> F[Digest + identity verification]
              F --> G[Review / test / package / deploy input]
          

The arrows are causal. If the producer did not create the path, upload fails or collects nothing. If the retention window expires, a later consumer may fail even though the original build was green. If the consumer rebuilds instead of downloading, it has created a different object whose equivalence must not be assumed.

3. Keep artifact state separate from every neighboring state

State Evidence to preserve What it does not prove
Repository/source CI_COMMIT_SHA, ref, repository tree That a build output was uploaded or retained
Compiled pipeline Artifact declarations, dependencies/needs edges That files actually existed at upload time
Producer job Pipeline/job ID, status, runner/executor, logs That a consumer downloaded the intended bytes
Artifact/archive Name, selected paths, size/digest, expiry/access, producer identity That a typed report was ingested or a deployment used it
Typed report Report type/path and GitLab report UI/API result That the report is a deployable build artifact
Consumer job Downloaded path, producer ID, verified digest That external target state changed
Deployment/external Environment/deployment/target identity and health That the bytes came from the intended producer unless linked by digest
Terminology: this course calls artifacts “immutable-ish” because a given uploaded artifact is tied to a job record, but retention, deletion, authorization, download-by-ref selection, and rebuilding are all mutable operational concerns. For release identity, preserve digest + producer job + source SHA.

4. Generic artifact versus typed report

A generic artifact is primarily a retained file/archive. A typed report tells GitLab how to interpret a file for a feature such as test results, coverage, code quality, or security. The same underlying file can sometimes be listed in both artifacts:reports and artifacts:paths when you want GitLab-native ingestion and convenient browsing/download.

Need Use Example
Pass a build output artifacts:paths dist/app.txt or a compiled binary
Show unit-test results artifacts:reports:junit JUnit XML
Browse the JUnit file too Report + path Declare JUnit under reports and paths
Reuse dependencies for speed Cache, not artifact Package-manager cache keyed by correctness inputs
Distribute a long-lived version Package/container registry or release asset Version/digest promotion rather than ordinary job retention
Current behavior: artifacts created for artifacts:reports are uploaded for report processing even when the job fails. Do not interpret “report was uploaded” as “job passed.”

5. Three same-pipeline fetching models

Without an explicit restriction, jobs in later stages download artifacts from all earlier-stage jobs. This is convenient for small pipelines but can become ambiguous and wasteful. dependencies narrows which earlier-stage jobs provide artifacts while retaining normal stage scheduling. A job that uses needs no longer relies on that default previous-stage download model; its needs entries can specify whether each needed producer’s artifacts are fetched.

Mechanism Ordering model Artifact source Use when
Default Stages All earlier-stage artifact producers Tiny pipeline where broad implicit transfer is acceptable
dependencies Stages Named earlier-stage jobs Keep stage barriers but narrow downloads
needs:artifacts DAG via needs Named needed jobs with artifact flag Start as soon as producer completes and receive only intended DAG inputs
Design rule: do not combine needs and dependencies in the same job. Encode one coherent scheduling/transfer model.

6. Give every artifact an evidence identity

A filename such as app.zip is not enough. Two pipelines can produce files with the same name. The minimum evidence tuple is producer project + pipeline ID + job ID/name + exact source SHA + artifact logical name/path + cryptographic digest + retention/access assumptions.

build_app:
  stage: build
  script:
    - mkdir -p dist evidence
    - printf 'release-candidate:%s\n' "$CI_COMMIT_SHA" > dist/app.txt
    - sha256sum dist/app.txt > evidence/app.sha256
    - printf 'pipeline=%s\njob=%s\nsha=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" > evidence/producer.txt
  artifacts:
    name: "app-$CI_PIPELINE_ID-$CI_JOB_ID"
    paths:
      - dist/app.txt
      - evidence/app.sha256
      - evidence/producer.txt
    expire_in: 7 days

The example prints identifiers, not secrets. It records enough provenance for a later job to prove which bytes it received.

7. Retention is part of correctness

artifacts:expire_in starts its retention clock when GitLab stores the artifact. If omitted, GitLab uses the instance-wide default. Current GitLab also normally keeps artifacts from the most recent successful pipeline for each ref even when an expire_in value exists, unless that keep-latest behavior is disabled. Therefore “expires in seven days” is not a universal statement about actual deletion time.

Retention should outlive every intended consumer: later pipeline stages, human review, deployment windows, incident investigation, compliance evidence, or rollback. Conversely, long retention of large or sensitive outputs has storage and exposure cost.

Failure mode: a long-running pipeline can reach a later job after an early artifact has expired. Treat retention as a data dependency, not housekeeping.

8. Access controls reduce exposure; they do not make artifacts a secret vault

artifacts:access controls who can download a job’s artifacts through the GitLab UI/API. Current values include all, developer, maintainer, and none. This control also applies to report artifacts, but it is not a substitute for preventing secrets from entering the artifact in the first place.

Job-token access has its own authorization model. A UI/API artifact access setting and CI job-token cross-project permissions are separate security states.

Never artifact secrets: masked variables, tokens, private keys, cloud credentials, runner auth tokens, and OIDC tokens do not become safe merely because the artifact is private or short-lived.

9. Read-only inspection before changing artifact flow

  1. Record pipeline source, ref, and exact SHA.
  2. Open the compiled configuration and identify every artifact producer and declared consumer.
  3. Record producer job ID/status and the exact artifact paths/reports.
  4. Inspect the job’s artifact metadata/UI and retention/access policy without deleting or keeping anything yet.
  5. For a consumer, identify whether transfer is default, dependencies, or needs:artifacts.
  6. Compare producer and consumer digests when the file is correctness-relevant.
Source identity

Pipeline source/ref/SHA

Producer identity

Pipeline + job ID/name

Artifact identity

Path/name/digest/type

Lifecycle

Expiry + keep-latest assumption

Authorization

UI/API access + job-token context

Consumer proof

Transfer mechanism + verified digest

10. Misconceptions to remove now

Misconception Correction
“The next job can read the previous runner filesystem.” Runner workspaces are not a portable data contract; upload/download artifacts explicitly.
“A report is the release binary.” A typed report is evidence for GitLab features; deploy the intended build artifact/digest.
“A cache is a trusted build artifact.” Caches are performance state and can be missing/stale; artifacts carry intended pipeline outputs.
“Same filename means same build.” Verify producer job/SHA and digest.
“expire_in is the exact deletion moment.” Keep-latest and cleanup scheduling affect actual retention.
“Private artifacts may contain secrets.” Do not place secrets in artifacts at all.

11. Micro-lab: predict the transfer before running

Given one build job that produces dist/app.txt and test/report.xml, predict the answer for each consumer:

  • A later-stage job with no dependencies: which earlier artifacts does it fetch by default?
  • A later-stage job with dependencies: [build_app]: what does it fetch?
  • A DAG job with needs: [{job: build_app, artifacts: true}]: when can it start and what does it receive?
  • A DAG job with the same need but artifacts: false: what control dependency remains and what data dependency disappears?

Do not execute yet. The goal is to separate ordering from transfer before syntax becomes habit.

Knowledge check

Why is a successful producer job not enough to prove downstream reproducibility?

What is the difference between an artifact path and a typed report?

When should dependencies be preferred over default fetching?

What changes when a consumer uses needs?

Why should a deployment record a digest?

Next lesson

Guided hands-on workflow and core operations

Produce a deterministic artifact and JUnit-like report, then compare default, dependencies, and needs:artifacts transfer paths with concrete evidence.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-11. Artifact/report retention, access, API behavior, and cross-job transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.