Chapter 16Lesson 01~250 minutes

Artifacts, Reports, Cache, Dependencies, Retention, and Data Flow Between Jobs: Concepts, Architecture, and Mental Model

Treat build outputs, reports, and caches as distinct data planes with explicit authority, lifetime, access, integrity, and downstream flow.

ArtifactsReportsCacheData flowRetentionIntegrity

Learning objectives

  • Distinguish ordinary artifacts, report artifacts, cache, and repository source by authority and lifecycle.
  • Explain default artifact download behavior and how dependencies or needs:artifacts changes it.
  • Reason about retention, access, object storage, sensitivity, integrity, and storage cost.
  • Explain why report files are both stored artifacts and structured input to GitLab features.
  • Inspect pipeline/job artifact metadata before changing data-flow configuration.
Availability baseline (verified 2026-08-21 against current GitLab documentation). Ordinary job artifacts, report artifacts such as JUnit, job caches, dependencies, and needs:artifacts are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. If artifacts:expire_in is omitted, the instance default controls expiry; GitLab also keeps artifacts from the most recent successful pipeline on each ref by default unless that behavior is disabled. Cache is a performance optimization, not an authoritative release/evidence store. Protected and non-protected refs use separate caches by default; disabling that boundary or using cache:unprotect broadens who can read/write the same cache and must be a deliberate trust decision. Hosted storage/compute quotas and billing are volatile, so labs use tiny files and a no-runner fixture path.

1. The problem: “a file exists somewhere in CI” is not a data contract

Chapter 15 made job ordering explicit. Chapter 16 adds the second half of that model: what data crosses each job boundary. A build output, a JUnit XML report, a dependency cache, and a cloned repository can all appear as files inside a runner workspace, but they do not mean the same thing and they do not have the same lifetime.

If a team treats cache as a release store, a cache miss becomes a release outage. If it treats every artifact as permanent evidence, storage grows without a retention policy. If it uploads credentials in a report, GitLab faithfully stores the sensitive file. Correct CI data flow therefore starts by classifying the data before choosing YAML.

2. Mental model: four data planes around one job

Job workspace and data flow
flowchart LR
  G[(Git repository/ref)] -- checkout --> J[Job workspace]
  C[(Cache store)] -- restore before artifacts --> J
  A[(Earlier job artifacts)] -- download --> J
  J -- archive authoritative output --> O[(Job artifacts)]
  J -- structured report --> R[(GitLab report ingestion)]
  J -- reusable acceleration state --> C
  O -- explicit downstream flow --> D[Later job]

The repository supplies versioned source. Cache is reusable acceleration state and can span pipelines. Job artifacts are pipeline outputs stored by GitLab. Report artifacts are files with a schema GitLab parses to enrich pipeline/MR features. Downstream jobs may receive artifacts automatically by stage or explicitly through dependencies/needs.

3. Classify data by authority, consumer, and lifetime

Data type Primary purpose Typical lifetime Authoritative? Example
Repository file Versioned source/configuration. Git history. Yes for source/config. Lockfile, source code, CI YAML.
Ordinary artifact Carry a build/test output through pipeline or expose it for download. Pipeline/release/audit policy. Often yes for that pipeline output. Binary, generated docs, manifest.
Report artifact Provide structured data GitLab interprets. Enough for review/audit policy. Authoritative for the report producer, not necessarily for release payload. JUnit XML, coverage report, dotenv report.
Cache Avoid repeating downloads/build preparation. Disposable and rebuildable. No. Package-manager download directory.

A simple test helps: if this object disappears, can I deterministically rebuild or redownload it without changing the release identity? If yes, cache may be appropriate. If no, it belongs in an artifact/package/release/registry system with explicit identity and retention.

4. Job artifacts: archived outputs with lifecycle and access rules

A job artifact is created after the job script finishes. Paths are relative to the project working directory. GitLab receives the archive and associates it with the job, commit, pipeline, and ref. That identity is valuable: it lets you answer “which pipeline produced this file?”

build:
  stage: build
  script:
    - mkdir -p out
    - printf 'commit=%s\n' "$CI_COMMIT_SHA" > out/build.txt
  artifacts:
    name: "build-$CI_COMMIT_SHORT_SHA"
    paths:
      - out/build.txt
    expire_in: 7 days
    access: developer

artifacts:access currently supports all, developer, maintainer, and none for UI/API downloads. It also applies to report artifacts. This is a download-access control; it is not a universal secrecy boundary for all downstream CI paths.

5. Report artifacts: stored files plus platform interpretation

artifacts:reports tells GitLab that a file has a known schema. For JUnit, GitLab parses test cases and renders them in pipeline/MR interfaces. The report does not decide whether the job succeeds: the script exit code still does that. This separation is useful when you want a failing test command to fail the job and preserve the report for diagnosis.

test:
  stage: test
  script:
    - ./run-tests-and-write-junit.sh
  artifacts:
    when: always
    reports:
      junit: reports/junit.xml
    paths:
      - reports/junit.xml
    expire_in: 14 days

GitLab documents report artifacts as uploaded regardless of job result. Adding the same report file under artifacts:paths makes it browsable/downloadable as an ordinary artifact too; treat that as a storage/access decision, not a requirement for report ingestion.

6. Cache: acceleration state, not evidence

A cache may be reused by later jobs and later pipelines. Runner topology determines where it lives: local runner storage, a distributed cache such as object storage, or another runner-supported cache backend. It is intentionally weaker than artifact identity.

test:
  cache:
    key:
      files:
        - package-lock.json
    paths:
      - .npm/
    policy: pull-push
  script:
    - npm ci --cache .npm --prefer-offline
    - npm test

The safe production assumption is “cache can be absent, stale, or evicted.” The job must be able to reconstruct it. A cache key should capture compatibility inputs such as a lockfile, platform, toolchain, or branch trust boundary.

7. Restore order matters

GitLab Runner restores caches before artifacts. If cache and artifact paths overlap, the later artifact extraction can overwrite files restored from cache. Avoid overlapping authoritative and disposable paths; otherwise debugging becomes a question of extraction order rather than application behavior.

8. Default artifact download versus explicit flow

Without needs or dependencies, later-stage jobs normally fetch artifacts from jobs in earlier stages. That convenience can hide dependencies. dependencies narrows which earlier-stage artifacts a job downloads. dependencies: [] downloads none.

With a needs DAG, the model changes: the job may start before an entire previous stage is complete, so it can only download artifacts from jobs named in needs. Use artifacts: true (default for a normal needs edge) or false deliberately.

package:
  needs:
    - job: build
      artifacts: true
    - job: lint
      artifacts: false
  script:
    - test -f out/build.txt
    - ./package.sh

9. Retention is part of reproducibility

artifacts:expire_in starts its clock when GitLab stores the artifact. If it is omitted, the instance default applies. GitLab also keeps artifacts from the latest successful pipeline on each ref by default unless the project/instance setting is disabled. Therefore “expire in seven days” does not necessarily mean “every copy disappears exactly seven days later.”

Retention should map to a business reason: fast feedback artifacts can be short-lived; release evidence may need longer retention or promotion to a release/package/registry system designed for durable distribution. Do not use expire_in: never as a substitute for an evidence-retention policy.

10. Sensitive content and cache poisoning are different risks

Artifacts are stored and downloadable according to project/access rules, so a file containing a credential is already an incident. Cache poisoning is different: an untrusted pipeline writes content under a cache key that a more trusted job later consumes. GitLab separates protected and non-protected ref caches by default, and current cache keys receive protection-related suffixes. Disabling that boundary or using cache:unprotect: true should require a documented trust argument.

11. GitLab control plane versus storage backend

On GitLab.com, the service owns the storage implementation. On Self-Managed, administrators may configure local storage or object storage for artifacts and runner-side distributed cache separately. A job author should reason about logical lifecycle/access; a platform administrator additionally owns bucket lifecycle, encryption, capacity, replication, and recovery. Do not blend those responsibilities into one “CI storage” concept.

12. Read-only inspection first

Before changing YAML, open Build → Artifacts and the relevant job/pipeline. Record job ID, pipeline ID, SHA, artifact filenames/types, expiry where visible, and whether the report appears in GitLab’s structured UI. Machine-readable inspection can stay metadata-only:

PROJECT_ID="12345678"
PIPELINE_ID="123456789"

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
  --jq '.[] | {id,name,status,stage,ref,artifacts_file,artifacts}'

Do not download or print sensitive artifacts merely to prove that they exist. Presence, size, type, checksum, and access behavior are usually enough.

13. DevOps connection: make every edge a data contract

A dependable pipeline states both the scheduling dependency and the data dependency. The producer identifies an immutable output; the consumer names exactly what it needs; retention matches recovery/audit needs; cache accelerates reconstruction but is never the only copy. That model survives runner churn, cache eviction, DAG refactors, and incident review.

Knowledge check

Why is a cache a poor place for the only copy of a release binary?

What changes when a job starts using needs?

Does a JUnit report make a failing test job fail?

Why is cache:unprotect: true security-sensitive?

If expire_in is absent, what determines artifact expiry?

What is the safest first inspection for a suspicious artifact?

Summary

Repository source, artifacts, reports, and caches are distinct data planes. Artifacts carry identified pipeline outputs, reports feed structured GitLab features, and cache trades integrity guarantees for speed. Explicit dependencies/needs, bounded retention, download access, and cache trust boundaries turn those files into a production-grade data-flow model.

Official references

Next lesson

Build and inspect artifact, report, and cache flows

Lesson 2 creates a tiny deterministic pipeline, verifies report interpretation and cache hit/miss behavior, then narrows downstream artifact downloads without introducing secrets or external services.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.