Artifacts, Reports, Retention, Dependencies, needs:artifacts, and Cross-Job Data Flow: Diagnostics, Failure Modes, Security, and Performance
Artifact failures are often data-identity failures rather than build failures. This lesson diagnoses secret leakage, noisy paths, stale or rebuilt files, expiry surprises, report-vs-binary confusion, and incorrect transfer declarations while preserving producer job/SHA/digest evidence before repair.
Learning objectives
- Diagnose artifact leakage, wrong-path collection, stale/rebuilt consumers, premature expiry, and report-vs-deployable-output confusion.
- Separate config/rule/graph/runner/script failures from artifact upload, retention, authorization, download, and external deployment failures.
- Preserve first-failure job IDs, artifact metadata, digests, and consumer evidence before retrying or cleaning up.
- Apply the least destructive repair and avoid rebuilding an intended release artifact as a troubleshooting shortcut.
- Recognize when storage/transfer size, retention, or broad downloads create performance and cost problems.
1. Evidence-first diagnostic sequence
Do not start by rerunning the build. A rerun can create a new job ID and new bytes, destroying the evidence needed to explain why the original consumer failed. Preserve:
- pipeline source/ref/SHA and compiled configuration,
- producer and consumer job IDs/statuses,
- artifact name/paths/type/expiry/access and available metadata,
- producer digest/provenance manifest and consumer error,
- runner/executor context only if transfer/upload/download behavior depends on it,
- external deployment/target identity if the artifact already left GitLab.
Then locate the failing layer: creation → upload → retention → authorization → download → digest/format → external use.
flowchart TD
A[Source/config] --> B[Producer script]
B --> C[Artifact selection/upload]
C --> D[Retention + access]
D --> E[Consumer transfer]
E --> F[Digest/format verify]
F --> G[External use]
2. Failure: an artifact contains a secret
Symptom: a debug bundle contains a fake token, private key, cloud credential, or environment dump. Treat the same pattern as a real incident in production: preserve the artifact/job identity needed for response without copying the secret into tickets, rotate/revoke the affected credential, restrict/delete exposure according to incident procedure, and repair the job so secret material is never collected.
# BAD pattern — do not use in real pipelines
# tar -czf debug.tgz . # may sweep credentials and unrelated workspace files
# Safer pattern: explicit non-secret diagnostics only
artifacts:
paths:
- evidence/tool-version.txt
- evidence/test-summary.txt
- evidence/provenance.txt
expire_in: 1 day
3. Failure: wrong path uploads workspace noise
Symptom: artifact size suddenly grows, upload is
slow, or the archive contains dependencies/tmp files. Compare the
compiled artifacts:paths declaration with the artifact
browser. Narrow the paths rather than increasing artifact-size
limits or keeping the oversized archive longer.
Preserve the original artifact metadata as evidence of the faulty selection; do not delete first and then try to reconstruct what was uploaded.
4. Failure: downstream consumes a stale or rebuilt file
Symptom: consumer succeeds but its digest does not match the producer manifest, or a later job reruns the build and deploys a newly generated file. This is a correctness failure even if every job is green.
verify_release_candidate:
needs:
- job: build_app
artifacts: true
script:
- sha256sum -c evidence/app.sha256
- test "$CI_COMMIT_SHA" = "$(sed -n 's/^sha=//p' evidence/producer.txt)"
Repair by consuming the intended producer artifact and failing closed on digest/provenance mismatch. Do not “fix” a missing artifact by rebuilding in the deploy job.
5. Failure: expiry surprises a later consumer
Symptom: a consumer reports that needed artifacts
cannot be retrieved. Confirm the producer job still has artifacts
and record its expiry policy. Remember that keep-latest-success can
retain some artifacts beyond expire_in, while older
artifacts can become eligible for deletion after a newer successful
pipeline on the same ref.
Repair the retention horizon for future artifacts or promote the output to a durable distribution system when the use case exceeds ordinary CI retention. Do not depend on “Keep” clicks as an undocumented release process.
6. Failure: typed report mistaken for deployable binary
Symptom: deployment logic points at a JUnit/coverage/security report because “it is an artifact.” Inspect artifact type and contract. Reports are evidence inputs for GitLab features; a release binary/package has a different role and should have explicit producer/digest identity.
| Observed object | Correct interpretation |
|---|---|
| JUnit XML | Test evidence, not application payload |
| Coverage report | Coverage evidence/annotation input |
| dotenv report | Variable propagation format with strict safety/size constraints, not secret storage |
| Build archive/binary | Potential deployable candidate if identity and verification contract says so |
7. Intentionally broken example: the consumer gets no intended artifact
This example is syntactically valid but wrong for the data contract:
build_app:
stage: build
script:
- mkdir -p dist
- printf 'candidate\n' > dist/app.txt
artifacts:
paths: [dist/app.txt]
verify_app:
stage: test
needs:
- job: build_app
artifacts: false
script:
- test -f dist/app.txt
Interpret the evidence: pipeline creation succeeds;
build_app succeeds; the needs edge lets
verify_app wait for the producer; artifact transfer is
explicitly disabled; the consumer fails on a missing file. The least
destructive correction is artifacts: true, not a
rebuild.
8. Authorization failure: distinguish user access from job-token access
If a human cannot download an artifact, inspect project
visibility/membership and artifacts:access. If a CI job
cannot retrieve an artifact through an API or cross-project flow,
inspect CI_JOB_TOKEN permissions/allowlists and endpoint/tier
requirements separately. Do not “solve” either case with a broad PAT
in a variable.
9. Performance and cost: measure transfer, not just job duration
Large artifacts affect upload time, download time, runner bandwidth, object storage, and project storage quotas. Broad default fetching can make every later-stage job repeatedly download outputs it does not use.
| Evidence | Question |
|---|---|
| Artifact size / paths | Are we exporting only required outputs? |
| Consumer list | Does every later job really need this artifact? |
| Queue vs transfer time | Is the bottleneck runner capacity or artifact I/O? |
| Retention/storage | Are old outputs retained longer than needed? |
| Fan-out | Are many jobs downloading the same large archive when a smaller interface would suffice? |
Optimize by narrowing outputs/consumers, not by dropping verification or substituting a cache for a trusted artifact.
10. Compact recovery playbook
- Freeze the evidence tuple: pipeline/job/SHA/artifact/digest.
- Determine whether the failure is creation, upload, retention, access, transfer, digest/format, or external use.
- Correct only that layer in the disposable branch.
- Rerun the smallest safe scope; if a new producer run is required, label it as a new artifact identity.
- Never replace a missing release candidate by silently rebuilding during deployment.
- Verify consumer digest and downstream target identity after recovery.
Knowledge check
Why should you avoid immediately rerunning a failed producer?
A rerun creates new job/artifact identity and may destroy the evidence needed to diagnose the original failure.
A masked secret appears in an artifact. What failed?
Artifact content selection/secrets handling failed; log masking does not sanitize files. Rotate/revoke as needed and prevent the file from being collected.
A needs consumer waits correctly but has no file. What should you inspect first?
The need’s artifact-transfer flag and producer artifact existence, because ordering and data transfer are separate.
Why is rebuilding in deploy a dangerous “fix”?
It creates different bytes without proving equivalence to the verified release candidate, breaking provenance and promotion integrity.
What is the performance smell of broad default fetching?
Many later-stage jobs may repeatedly download large artifacts they do not use, increasing latency, bandwidth, and storage I/O.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-11. Artifact/report retention, access, API behavior, and cross-job transfer semantics are version-sensitive; verify the deployed GitLab version for Self-Managed/Dedicated installations.
- Job artifacts — creation, paths, expiry, download, artifact browsing, access, latest-success retention, and default previous-stage fetching.
-
CI/CD YAML syntax reference
— authoritative
artifacts,artifacts:access,artifacts:expire_in,dependencies, andneeds:artifactssemantics. - CI/CD artifacts report types — typed report ingestion and report-specific GitLab UI behavior.
- Unit test reports — JUnit report configuration and display.
- Job Artifacts API — artifact archive/file/report download, keep, and delete operations with exact job identity.
- Troubleshooting job artifacts — artifact expiry and upload/report problems.
- Caching in GitLab CI/CD — artifact-versus-cache boundary.
-
Make jobs start earlier with
needs— DAG dependency edges and artifact-transfer interaction.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.