Chapter 16Lesson 02~275 minutes

Artifacts, Reports, Cache, Dependencies, Retention, and Data Flow Between Jobs: Guided Hands-On Workflow and Core Operations

Create and verify a tiny artifact, parsed JUnit report, disposable cache, and selective artifact downloads without secrets or external services.

Hands-onJUnitdependenciesneeds:artifactsCache hit/missCleanup

Learning objectives

  • Create a tiny deterministic artifact and bind it to the producing job/pipeline/SHA.
  • Create a harmless JUnit report and distinguish report ingestion from job success.
  • Create a rebuildable cache and verify miss/hit behavior from runner logs.
  • Narrow artifact downloads with dependencies and needs:artifacts.
  • Inspect retention/access and clean up only disposable data.
Availability baseline (verified 2026-08-21 against current GitLab documentation). Ordinary job artifacts, report artifacts such as JUnit, job caches, dependencies, and needs:artifacts are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. If artifacts:expire_in is omitted, the instance default controls expiry; GitLab also keeps artifacts from the most recent successful pipeline on each ref by default unless that behavior is disabled. Cache is a performance optimization, not an authoritative release/evidence store. Protected and non-protected refs use separate caches by default; disabling that boundary or using cache:unprotect broadens who can read/write the same cache and must be a deliberate trust decision. Hosted storage/compute quotas and billing are volatile, so labs use tiny files and a no-runner fixture path.

1. Disposable scenario and preflight

Use a disposable GitLab Free project or branch such as ch16/data-flow-lab. The lab writes only tiny text/XML files and a synthetic cache marker. It requires no package registry, cloud account, external database, secret variable, or production runner.

Preflight item Required state Why
Project/ref Disposable project or branch ch16/data-flow-lab. Artifact/cache cleanup must not affect valuable evidence.
Role Developer for commits/pipelines; Maintainer only if you choose artifact-deletion API cleanup. Most learning does not require elevated access.
Runner Any eligible runner with a basic shell, or CI Lint + supplied evidence fixture. Hosted compute quota is not guaranteed.
Secrets None. Do not add variables or credentials. Artifacts/cache are intentionally inspectable in this lab.
Storage Files are a few bytes; use short expire_in. Keeps storage/cost negligible.

2. Create the three-job data-flow pipeline

The first job produces the authoritative build artifact and a harmless cache directory. The second creates a JUnit report while consuming only the build artifact. The third packages evidence and demonstrates that cache state is optional.

stages: [build, test, verify]

default:
  cache:
    key: "ch16-$CI_COMMIT_REF_SLUG"
    paths:
      - .chapter16-cache/

build_output:
  stage: build
  script:
    - mkdir -p out .chapter16-cache
    - printf 'commit=%s\n' "$CI_COMMIT_SHA" > out/build.txt
    - printf 'cache-created-by=%s\n' "$CI_PIPELINE_ID" > .chapter16-cache/marker.txt
    - sha256sum out/build.txt > out/build.txt.sha256
  artifacts:
    name: "ch16-build-$CI_COMMIT_SHORT_SHA"
    paths:
      - out/build.txt
      - out/build.txt.sha256
    expire_in: 2 days
    access: developer

test_report:
  stage: test
  dependencies:
    - build_output
  script:
    - test -f out/build.txt
    - mkdir -p reports
    - |
      cat > reports/junit.xml <<'EOF'
      <?xml version="1.0" encoding="UTF-8"?>
      <testsuite name="chapter16" tests="1" failures="0">
        <testcase classname="dataflow" name="artifact_received"/>
      </testsuite>
      EOF
  artifacts:
    when: always
    reports:
      junit: reports/junit.xml
    paths:
      - reports/junit.xml
    expire_in: 2 days

verify_flow:
  stage: verify
  needs:
    - job: build_output
      artifacts: true
    - job: test_report
      artifacts: false
  script:
    - test -f out/build.txt
    - sha256sum -c out/build.txt.sha256
    - |
      if test -f .chapter16-cache/marker.txt; then
        echo "cache-state=present"
      else
        echo "cache-state=miss-rebuildable"
      fi

3. Predict state before the first run

Question Prediction
What must test_report receive? Only artifacts from build_output, because dependencies narrows the automatic previous-stage set.
Will verify_flow download JUnit XML? No; its needs edge to test_report sets artifacts: false.
Is cache marker authoritative? No. A miss is acceptable and should not invalidate the build artifact.
What proves build identity? Artifact content/checksum + producer job/pipeline + CI_COMMIT_SHA.
What determines test job success? Its script exit status, not the mere existence/ingestion of JUnit XML.

4. Validate before committing

Use Pipeline Editor/CI Lint to validate the configuration. Then inspect the staged diff so artifact/cache paths are exactly the synthetic directories.

git switch -c ch16/data-flow-lab
git add .gitlab-ci.yml
git diff --cached --check
git diff --cached
git commit -m "ch16: add artifact report cache lab"
git push -u origin ch16/data-flow-lab

5. First run: capture artifact identity

Record pipeline ID and commit SHA. In Build → Jobs → build_output, verify the artifact archive exists and the job log contains the upload step. Do not infer success from the file being visible in a later job; bind it to the producer first.

PROJECT_ID="12345678"
PIPELINE_ID="123456789"

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
  --jq '.[] | {id,name,status,stage,ref,artifacts_file,artifacts}'

6. Verify report interpretation independently

Open the pipeline details and its Tests tab/summary. The single synthetic test case should be displayed as passed. Also inspect test_report job metadata. This proves two different facts: the XML was uploaded, and GitLab parsed it as JUnit.

If the Tests view is empty, inspect the report path and XML format before changing runner or permissions. A valid artifact archive is not proof of a valid report schema.

7. Observe cache miss and hit without depending on either

On the first run, runner logs may show a cache miss followed by cache creation. Push a trivial change that does not alter the cache key and run again. On compatible runner/cache topology, the second run may show cache restoration. If it still misses, the pipeline remains correct because no authoritative output depends on the cache.

No-runner/unstable-cache path. Use the deterministic fixture: first log contains “Checking cache … cache file does not exist” and later upload; second log may contain “Successfully extracted cache.” The learning objective is to classify the cache as optional, not to guarantee a particular hosted backend.

8. Prove downstream file availability

In test_report, out/build.txt exists because dependencies names the producer. In verify_flow, build files exist because needs:artifacts is enabled for build_output. JUnit XML should not be present through the test_report edge because that edge disables artifact download.

Do not use ls -laR or an environment dump as evidence in real pipelines; such broad output can reveal unrelated files. Check only expected paths.

9. Inspect expiry and access

On the Build → Artifacts page or job details, verify the artifacts are associated with the expected jobs and have the configured short expiry. access: developer restricts UI/API downloads to Developer-or-higher roles, but downstream CI semantics are separate. Never describe artifact access as encryption or a secret manager.

10. Convert the test job to a DAG edge

Replace dependencies on test_report with an explicit needs edge and validate again:

test_report:
  stage: test
  needs:
    - job: build_output
      artifacts: true
  script:
    - test -f out/build.txt
    - ./write-safe-junit-fixture.sh

This change does two things: it permits earlier scheduling when the producer completes, and it makes the artifact dependency explicit. Do not combine needs and dependencies in the same job as a normal design pattern; choose the model that matches scheduling/data intent.

11. Challenge: choose the correct surface

Classify each requirement before writing YAML:

Requirement Correct surface Reason
Keep a release binary for traceability. Artifact now; release/package/registry for durable distribution. Identity and retention matter.
Avoid downloading dependencies repeatedly. Cache. Rebuildable acceleration state.
Show test cases in GitLab. JUnit report artifact. GitLab parses a structured schema.
Prevent a downstream job from downloading unrelated artifacts. dependencies or needs:artifacts. Defines explicit data flow.
Limit who can download a non-secret diagnostic bundle. artifacts:access. UI/API download access control.

12. Cleanup

The safest cleanup is to revert/remove the synthetic CI file on the disposable branch and then delete only that exact branch after verifying its identity. Individual artifact deletion is irreversible and normally requires Maintainer/Owner; it is optional in this lab because short expiry already bounds storage.

git switch main
git ls-remote --heads origin refs/heads/ch16/data-flow-lab
# Delete only after confirming this exact disposable branch:
git push origin --delete ch16/data-flow-lab
git branch -D ch16/data-flow-lab
Artifact deletion is destructive. If practicing the Job Artifacts API, use only a disposable project/job whose ID you verified twice. Deleting artifacts cannot be undone, and deleting all project artifacts is asynchronous and may preserve the latest successful artifacts according to current retention rules.

Knowledge check

Why can the second pipeline be correct even if its cache still misses?

What two facts must be verified for JUnit?

What does artifacts: false on a needs edge do?

Why avoid broad recursive directory listings in real CI evidence?

When is deleting artifacts during the lab unnecessary?

Summary

You created three different data behaviors and verified them independently: an identified artifact, a parsed report, and a disposable cache. You then constrained downloads with explicit dependency edges and inspected retention/access without exposing sensitive values.

Official references

Next lesson

Turn mechanics into policy choices

Lesson 3 decides when to prefer artifacts or cache, how long evidence should live, when explicit downloads improve governance, and when shared caches create unacceptable trust coupling.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.