Artifacts, Reports, Cache, Dependencies, Retention, and Data Flow Between Jobs: Guided Hands-On Workflow and Core Operations
Create and verify a tiny artifact, parsed JUnit report, disposable cache, and selective artifact downloads without secrets or external services.
Learning objectives
- Create a tiny deterministic artifact and bind it to the producing job/pipeline/SHA.
- Create a harmless JUnit report and distinguish report ingestion from job success.
- Create a rebuildable cache and verify miss/hit behavior from runner logs.
-
Narrow artifact downloads with
dependenciesandneeds:artifacts. - Inspect retention/access and clean up only disposable data.
dependencies, and needs:artifacts are
available in GitLab Free/Premium/Ultimate across GitLab.com,
Self-Managed, and Dedicated. If artifacts:expire_in is
omitted, the instance default controls expiry; GitLab also keeps
artifacts from the most recent successful pipeline on each ref by
default unless that behavior is disabled. Cache is a performance
optimization, not an authoritative release/evidence store. Protected
and non-protected refs use separate caches by default; disabling that
boundary or using cache:unprotect broadens who can
read/write the same cache and must be a deliberate trust decision.
Hosted storage/compute quotas and billing are volatile, so labs use
tiny files and a no-runner fixture path.
1. Disposable scenario and preflight
Use a disposable GitLab Free project or branch such as
ch16/data-flow-lab. The lab writes only tiny text/XML
files and a synthetic cache marker. It requires no package registry,
cloud account, external database, secret variable, or production
runner.
| Preflight item | Required state | Why |
|---|---|---|
| Project/ref |
Disposable project or branch
ch16/data-flow-lab.
|
Artifact/cache cleanup must not affect valuable evidence. |
| Role | Developer for commits/pipelines; Maintainer only if you choose artifact-deletion API cleanup. | Most learning does not require elevated access. |
| Runner | Any eligible runner with a basic shell, or CI Lint + supplied evidence fixture. | Hosted compute quota is not guaranteed. |
| Secrets | None. Do not add variables or credentials. | Artifacts/cache are intentionally inspectable in this lab. |
| Storage |
Files are a few bytes; use short expire_in.
|
Keeps storage/cost negligible. |
2. Create the three-job data-flow pipeline
The first job produces the authoritative build artifact and a harmless cache directory. The second creates a JUnit report while consuming only the build artifact. The third packages evidence and demonstrates that cache state is optional.
stages: [build, test, verify]
default:
cache:
key: "ch16-$CI_COMMIT_REF_SLUG"
paths:
- .chapter16-cache/
build_output:
stage: build
script:
- mkdir -p out .chapter16-cache
- printf 'commit=%s\n' "$CI_COMMIT_SHA" > out/build.txt
- printf 'cache-created-by=%s\n' "$CI_PIPELINE_ID" > .chapter16-cache/marker.txt
- sha256sum out/build.txt > out/build.txt.sha256
artifacts:
name: "ch16-build-$CI_COMMIT_SHORT_SHA"
paths:
- out/build.txt
- out/build.txt.sha256
expire_in: 2 days
access: developer
test_report:
stage: test
dependencies:
- build_output
script:
- test -f out/build.txt
- mkdir -p reports
- |
cat > reports/junit.xml <<'EOF'
<?xml version="1.0" encoding="UTF-8"?>
<testsuite name="chapter16" tests="1" failures="0">
<testcase classname="dataflow" name="artifact_received"/>
</testsuite>
EOF
artifacts:
when: always
reports:
junit: reports/junit.xml
paths:
- reports/junit.xml
expire_in: 2 days
verify_flow:
stage: verify
needs:
- job: build_output
artifacts: true
- job: test_report
artifacts: false
script:
- test -f out/build.txt
- sha256sum -c out/build.txt.sha256
- |
if test -f .chapter16-cache/marker.txt; then
echo "cache-state=present"
else
echo "cache-state=miss-rebuildable"
fi
3. Predict state before the first run
| Question | Prediction |
|---|---|
What must test_report receive? |
Only artifacts from build_output, because
dependencies narrows the automatic
previous-stage set.
|
Will verify_flow download JUnit XML? |
No; its needs edge to
test_report sets artifacts: false.
|
| Is cache marker authoritative? | No. A miss is acceptable and should not invalidate the build artifact. |
| What proves build identity? |
Artifact content/checksum + producer job/pipeline +
CI_COMMIT_SHA.
|
| What determines test job success? | Its script exit status, not the mere existence/ingestion of JUnit XML. |
4. Validate before committing
Use Pipeline Editor/CI Lint to validate the configuration. Then inspect the staged diff so artifact/cache paths are exactly the synthetic directories.
git switch -c ch16/data-flow-lab
git add .gitlab-ci.yml
git diff --cached --check
git diff --cached
git commit -m "ch16: add artifact report cache lab"
git push -u origin ch16/data-flow-lab
5. First run: capture artifact identity
Record pipeline ID and commit SHA. In Build → Jobs → build_output, verify the artifact archive exists and the job log contains the upload step. Do not infer success from the file being visible in a later job; bind it to the producer first.
PROJECT_ID="12345678"
PIPELINE_ID="123456789"
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
--jq '.[] | {id,name,status,stage,ref,artifacts_file,artifacts}'
6. Verify report interpretation independently
Open the pipeline details and its Tests tab/summary. The single
synthetic test case should be displayed as passed. Also inspect
test_report job metadata. This proves two different
facts: the XML was uploaded, and GitLab parsed it as JUnit.
If the Tests view is empty, inspect the report path and XML format before changing runner or permissions. A valid artifact archive is not proof of a valid report schema.
7. Observe cache miss and hit without depending on either
On the first run, runner logs may show a cache miss followed by cache creation. Push a trivial change that does not alter the cache key and run again. On compatible runner/cache topology, the second run may show cache restoration. If it still misses, the pipeline remains correct because no authoritative output depends on the cache.
8. Prove downstream file availability
In test_report, out/build.txt exists
because dependencies names the producer. In
verify_flow, build files exist because
needs:artifacts is enabled for
build_output. JUnit XML should not be present through
the test_report edge because that edge disables
artifact download.
Do not use ls -laR or an environment dump as evidence
in real pipelines; such broad output can reveal unrelated files.
Check only expected paths.
9. Inspect expiry and access
On the Build → Artifacts page or job details, verify the artifacts
are associated with the expected jobs and have the configured short
expiry. access: developer restricts UI/API downloads to
Developer-or-higher roles, but downstream CI semantics are separate.
Never describe artifact access as encryption or a secret manager.
10. Convert the test job to a DAG edge
Replace dependencies on test_report with
an explicit needs edge and validate again:
test_report:
stage: test
needs:
- job: build_output
artifacts: true
script:
- test -f out/build.txt
- ./write-safe-junit-fixture.sh
This change does two things: it permits earlier scheduling when the
producer completes, and it makes the artifact dependency explicit.
Do not combine needs and dependencies in
the same job as a normal design pattern; choose the model that
matches scheduling/data intent.
11. Challenge: choose the correct surface
Classify each requirement before writing YAML:
| Requirement | Correct surface | Reason |
|---|---|---|
| Keep a release binary for traceability. | Artifact now; release/package/registry for durable distribution. | Identity and retention matter. |
| Avoid downloading dependencies repeatedly. | Cache. | Rebuildable acceleration state. |
| Show test cases in GitLab. | JUnit report artifact. | GitLab parses a structured schema. |
| Prevent a downstream job from downloading unrelated artifacts. |
dependencies or needs:artifacts.
|
Defines explicit data flow. |
| Limit who can download a non-secret diagnostic bundle. | artifacts:access. |
UI/API download access control. |
12. Cleanup
The safest cleanup is to revert/remove the synthetic CI file on the disposable branch and then delete only that exact branch after verifying its identity. Individual artifact deletion is irreversible and normally requires Maintainer/Owner; it is optional in this lab because short expiry already bounds storage.
git switch main
git ls-remote --heads origin refs/heads/ch16/data-flow-lab
# Delete only after confirming this exact disposable branch:
git push origin --delete ch16/data-flow-lab
git branch -D ch16/data-flow-lab
Knowledge check
Why can the second pipeline be correct even if its cache still misses?
Because cache is optional acceleration state; the authoritative build artifact is produced independently.
What two facts must be verified for JUnit?
That the report artifact uploaded and that GitLab parsed/displayed the test cases.
What does artifacts: false on a needs edge
do?
It preserves the scheduling dependency but prevents artifact download from that needed job.
Why avoid broad recursive directory listings in real CI evidence?
They can disclose unrelated or sensitive workspace contents; verify only expected paths/metadata.
When is deleting artifacts during the lab unnecessary?
When short expiry and disposable scope already meet cleanup goals; irreversible deletion is optional.
Summary
You created three different data behaviors and verified them independently: an identified artifact, a parsed report, and a disposable cache. You then constrained downloads with explicit dependency edges and inspected retention/access without exposing sensitive values.
Official references
- GitLab Docs — Job artifacts
- GitLab Docs — Job artifacts troubleshooting
- GitLab Docs — Job Artifacts API
- GitLab Docs — CI/CD artifacts reports types
- GitLab Docs — Unit test reports
- GitLab Docs — Unit test report examples
- GitLab Docs — CI/CD caching
- GitLab Docs — CI/CD caching examples
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — Pass dotenv variables to specific jobs
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — Job artifacts administration
- GitLab Docs — Object storage
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.