Artifacts, Reports, Retention, Dependencies, needs:artifacts, and Cross-Job Data Flow: Guided Hands-On Workflow and Core Operations
Build a tiny synthetic binary once, save its digest and a JUnit-like report, then consume the same bytes downstream. You will compare GitLab’s default previous-stage artifact fetching with dependencies and needs:artifacts, inspect retained metadata, and prove exactly which producer supplied each consumer.
Learning objectives
- Create a disposable producer that emits a small deterministic file, SHA-256 manifest, and JUnit-like XML report.
- Observe GitLab’s default artifact download behavior in later stages and narrow it with dependencies.
- Use needs:artifacts to combine an explicit DAG edge with intentional same-pipeline artifact transfer.
- Inspect producer/consumer IDs, SHA, artifact metadata, report ingestion, expiry, and digest evidence.
- Complete a challenge that selects the correct transfer mechanism rather than rebuilding or reaching into another workspace.
1. Disposable scenario and assumptions
Create or reuse only a throwaway project such as
glci-artifacts-lab. Use a branch
glci/ch10-artifacts. No real credentials, packages,
releases, registries, or production systems are required. The
pipeline writes synthetic text files and JUnit XML only.
Assumptions verified 2026-09-11: ordinary job
artifacts and JUnit reports are available on Free/Premium/Ultimate;
jobs in later stages fetch all previous-stage artifacts by default;
dependencies narrows that stage-based fetching;
needs:artifacts controls artifact fetching from
explicit DAG needs.
.gitlab-ci.yml before the lab.
2. Add one synthetic source file
The “application” is deliberately trivial so the lesson can focus on GitLab data flow rather than a build tool.
mkdir -p src
printf 'hello-artifact-lab\n' > src/message.txt
git add src/message.txt
git commit -m 'ch10: add synthetic source'
Record the resulting commit SHA. Every artifact in this lab should ultimately be traceable to that exact revision.
3. Start with one producer, one report producer, and broad default fetching
The build job creates a deterministic artifact plus a
digest/provenance record. The test job writes a tiny JUnit XML file
and declares it as both a typed report and a browsable path. The
inspect_default_fetch job has no dependencies
declaration, so in a later stage it can receive artifacts from all
earlier-stage artifact producers.
stages: [build, test, inspect]
default:
image: alpine:3.22.1
build_app:
stage: build
script:
- mkdir -p dist evidence
- cp src/message.txt dist/app.txt
- sha256sum dist/app.txt > evidence/app.sha256
- printf 'pipeline=%s\njob=%s\nsha=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" > evidence/producer.txt
artifacts:
name: "app-$CI_PIPELINE_ID-$CI_JOB_ID"
paths:
- dist/app.txt
- evidence/app.sha256
- evidence/producer.txt
expire_in: 2 days
test_report:
stage: test
dependencies: []
script:
- mkdir -p test-results
- printf '%s\n' '<testsuite name="synthetic" tests="1" failures="0"><testcase classname="artifact" name="exists"/></testsuite>' > test-results/junit.xml
artifacts:
when: always
paths:
- test-results/junit.xml
reports:
junit: test-results/junit.xml
expire_in: 2 days
inspect_default_fetch:
stage: inspect
script:
- find dist evidence test-results -maxdepth 2 -type f -print | sort
- sha256sum -c evidence/app.sha256
- printf 'consumer_job=%s\nsha=%s\n' "$CI_JOB_ID" "$CI_COMMIT_SHA"
dependencies: [] on test_report is
deliberate: the test does not consume the build artifact, so
downloading it would be unnecessary data transfer. The inspect job,
however, intentionally demonstrates the default broad behavior.
4. Validate before pushing
Use CI Lint or the pipeline editor against the exact branch configuration. Validation proves syntax/logic, not that artifact paths will exist at runtime. Before pushing, predict:
- Which job uploads
dist/app.txt? - Which job uploads the typed JUnit report?
-
Which files should be visible to
inspect_default_fetch? - Which digest check should pass?
5. Run 1 — observe default artifact fetching
Push the branch and record pipeline ID, source, SHA, and each job
ID. In build_app, record the artifact archive/name and
its expiry. In the pipeline test UI, verify the synthetic JUnit
report is processed. Then open
inspect_default_fetch and confirm it sees both the
build artifact paths and the browsable report file.
ID + CI_PIPELINE_SOURCE + SHA
Job ID + artifact name
Job ID + JUnit ingestion
Job ID + downloaded paths
sha256sum -c succeeds
2-day declared expiry plus keep-latest caveat
6. Run 2 — narrow a stage-based consumer with dependencies
Now make the inspect job depend only on
build_app artifacts:
inspect_build_only:
stage: inspect
dependencies:
- build_app
script:
- test -f dist/app.txt
- test -f evidence/app.sha256
- test ! -e test-results/junit.xml
- sha256sum -c evidence/app.sha256
This preserves stage ordering but narrows transfer. The JUnit report producer can still exist and be ingested by GitLab, while this consumer avoids downloading its files. That is a dataflow change, not a scheduling change.
7. Run 3 — explicit DAG transfer with needs:artifacts
Next add a consumer that may start immediately after the build
producer finishes, regardless of unrelated jobs in the
test stage:
verify_early:
stage: inspect
needs:
- job: build_app
artifacts: true
script:
- cat evidence/producer.txt
- sha256sum -c evidence/app.sha256
- cmp src/message.txt dist/app.txt
The needs edge carries two meanings here: a control
dependency on build_app and permission to download that
producer’s artifacts. Because it does not need
test_report, it can run earlier than the stage barrier
would otherwise allow.
8. Negative experiment — keep ordering, remove artifact transfer
In a temporary commit, change the edge to
artifacts: false. The consumer still waits for the
build job, but dist/app.txt and
evidence/app.sha256 are not downloaded. Preserve the
failed consumer job ID and first missing-file error, then revert
only the flag.
verify_early:
stage: inspect
needs:
- job: build_app
artifacts: false
script:
- sha256sum -c evidence/app.sha256
9. Inspect artifact/report metadata without secrets
From the producer job page or Artifacts page, record:
- producer job ID/name and pipeline ID,
- artifact filename/paths visible in the archive,
- declared expiry and whether the project keeps latest-success artifacts,
- report type and pipeline test result display,
- consumer job IDs and digest verification result.
If you use the API for learning, target an exact job ID. Do not use a broad personal access token when the UI or an already-authorized lab context suffices.
10. Challenge — choose the layer, not the shortcut
A new job named package_notes needs only
dist/app.txt from build_app, should begin
as soon as build_app succeeds, and must not wait for
test_report. Which mechanism fits?
Expected design: use needs on
build_app with artifacts enabled. Do not rebuild
dist/app.txt, do not use cache as the source of truth,
and do not add a broad stage dependency simply because it is
familiar.
11. Cleanup
- Return the branch to the successful configuration if you keep it for the next lesson.
- Do not delete failed job evidence until diagnosis notes are complete.
- Delete only the exact disposable branch/project when finished.
- No credential cleanup is required because the lab created no real secrets or external identities.
Knowledge check
Why did test_report use dependencies: []?
It did not need previous-stage artifacts, so the empty list prevents unnecessary downloads while leaving stage scheduling unchanged.
What does dependencies change?
It narrows which earlier-stage jobs provide artifacts to a stage-based consumer; it does not create a DAG scheduling edge.
What does needs:artifacts add to a needs edge?
It controls whether the consumer fetches artifacts from that explicitly needed producer while the needs edge controls early scheduling.
Why preserve the failed artifacts:false run?
It is evidence that the control dependency was satisfied while the data-transfer contract was deliberately broken.
What proves the consumer got the same bytes?
A matching cryptographic digest checked against the producer-generated manifest, tied to producer job/pipeline/SHA identity.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. Artifact/report retention, access, API behavior, and cross-job transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.
- Job artifacts — creation, paths, expiry, download, artifact browsing, access, latest-success retention, and default previous-stage fetching.
-
CI/CD YAML syntax reference
— authoritative
artifacts,artifacts:access,artifacts:expire_in,dependencies, andneeds:artifactssemantics. - CI/CD artifacts report types — typed report ingestion and report-specific GitLab UI behavior.
- Unit test reports — JUnit report configuration and display.
- Job Artifacts API — artifact archive/file/report download, keep, and delete operations with exact job identity.
- Troubleshooting job artifacts — artifact expiry and upload/report problems.
- Caching in GitLab CI/CD — artifact-versus-cache boundary.
-
Make jobs start earlier with
needs— DAG dependency edges and artifact-transfer interaction.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.