Chapter 16Lesson 05~295 minutes

Checkpoint Lab — Artifacts, Reports, Cache, Dependencies, Retention, and Data Flow Between Jobs

Prove an end-to-end data contract with predictions, checksums, report ingestion, a safe miss/failure, retention review, and cleanup.

CheckpointArtifactReportCacheEvidenceCleanup

Checkpoint objectives

  • Predict artifact/report/cache availability before running a three-job pipeline.
  • Bind authoritative output to commit, pipeline, producer job, checksum, and retention.
  • Prove report ingestion independently from job success.
  • Cause and diagnose one safe cache miss or missing-artifact-style failure.
  • Clean up the disposable CI/ref and, optionally, synthetic artifacts through verified scope.
Availability baseline (verified 2026-08-21 against current GitLab documentation). Ordinary job artifacts, report artifacts such as JUnit, job caches, dependencies, and needs:artifacts are available in GitLab Free/Premium/Ultimate across GitLab.com, Self-Managed, and Dedicated. If artifacts:expire_in is omitted, the instance default controls expiry; GitLab also keeps artifacts from the most recent successful pipeline on each ref by default unless that behavior is disabled. Cache is a performance optimization, not an authoritative release/evidence store. Protected and non-protected refs use separate caches by default; disabling that boundary or using cache:unprotect broadens who can read/write the same cache and must be a deliberate trust decision. Hosted storage/compute quotas and billing are volatile, so labs use tiny files and a no-runner fixture path.

1. Mission and production-model connection

Your checkpoint is a small delivery data contract. One job creates an authoritative artifact, one creates a structured report, and one consumes only the data it needs while tolerating cache absence. You must write the expected file matrix before execution, compare it with evidence, and explain why each object has its chosen lifetime.

This adds the data plane to the operating model: Chapter 12 decides which jobs exist, Chapter 13 controls variables/secrets, Chapter 14 supplies execution, Chapter 15 controls scheduling/concurrency, and Chapter 16 defines what trustworthy data moves between those jobs.

2. Preflight and assumptions

Item Mandatory assumption
Tier/offering GitLab Free on GitLab.com, Self-Managed, or Dedicated is sufficient.
Project Disposable project/ref only.
Role Developer for normal lab; Maintainer only for optional artifact deletion.
Runner Tiny live jobs if available; CI Lint + evidence fixture otherwise.
Secrets None. No secret variables, tokens, customer data, environment dumps.
Storage Tiny text/XML/cache marker; expire_in: 2 days.

3. Write predictions before the run

Job Authoritative artifact expected? Report expected? Cache expected?
produce Creates out/product.txt + checksum. No. May create/update synthetic cache marker.
test Downloads product artifact. Creates JUnit report. Cache may be hit or miss; not required.
verify Downloads product artifact only. Does not download JUnit XML through its test edge. May be hit or miss; result must remain correct.

Also predict: producer SHA/pipeline identity, report test count, artifact expiry/access, and which runner log lines indicate cache restoration/upload.

4. Checkpoint configuration

stages: [build, test, verify]

default:
  cache:
    key: "ch16-checkpoint-$CI_COMMIT_REF_SLUG"
    paths: [.chapter16-cache/]

produce:
  stage: build
  script:
    - mkdir -p out .chapter16-cache
    - printf 'sha=%s\npipeline=%s\n' "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" > out/product.txt
    - sha256sum out/product.txt > out/product.txt.sha256
    - printf 'rebuildable-cache=%s\n' "$CI_COMMIT_REF_SLUG" > .chapter16-cache/marker.txt
  artifacts:
    name: "product-$CI_COMMIT_SHORT_SHA"
    paths: [out/product.txt, out/product.txt.sha256]
    expire_in: 2 days
    access: developer

test:
  stage: test
  needs:
    - job: produce
      artifacts: true
  script:
    - test -f out/product.txt
    - mkdir -p reports
    - |
      cat > reports/junit.xml <<'EOF'
      <?xml version="1.0" encoding="UTF-8"?>
      <testsuite name="chapter16-checkpoint" tests="1" failures="0">
        <testcase classname="dataflow" name="product_identity_verified"/>
      </testsuite>
      EOF
  artifacts:
    when: always
    reports:
      junit: reports/junit.xml
    expire_in: 2 days

verify:
  stage: verify
  needs:
    - job: produce
      artifacts: true
    - job: test
      artifacts: false
  script:
    - test -f out/product.txt
    - sha256sum -c out/product.txt.sha256
    - test ! -f reports/junit.xml
    - |
      if test -f .chapter16-cache/marker.txt; then
        echo "cache=present"
      else
        echo "cache=miss-acceptable"
      fi

5. Validate configuration and exact paths

Use CI Lint/Pipeline Editor. Inspect merged YAML and confirm no wildcard accidentally captures the repository, home directory, dotenv file, private key, or runner configuration. Then commit only on the disposable branch.

git switch -c ch16/checkpoint
git add .gitlab-ci.yml
git diff --cached --check
git diff --cached
git commit -m "ch16: data-flow checkpoint"
CHECKPOINT_SHA="$(git rev-parse HEAD)"
printf 'checkpoint_sha=%s\n' "$CHECKPOINT_SHA"
git push -u origin ch16/checkpoint

6. Capture pipeline/job evidence

Record pipeline ID/source/SHA and each job ID/status. Inspect artifact metadata without downloading broad content:

PROJECT_ID="12345678"
PIPELINE_ID="123456789"

glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID" \
  --jq '{id,sha,ref,source,status,web_url}'
glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs" --paginate \
  --jq '.[] | {id,name,status,stage,artifacts_file,artifacts,runner:(.runner.description // null)}'

7. Verify authoritative artifact identity

In verify, the checksum passes. Tie the artifact to the producer job ID, pipeline ID, and commit SHA. If you manually download it for the lab, inspect only the synthetic two-line file and checksum. In production, prefer digest/metadata verification over printing arbitrary artifact contents.

8. Verify report ingestion

Open pipeline Tests and confirm one test named product_identity_verified. This proves GitLab parsed the JUnit report. The test job also succeeded because its script exited zero. State these as separate observations.

9. Verify cache as disposable state

Inspect runner logs for cache restore/upload. If there is a miss, record it as a performance state and confirm all jobs still succeed. If there is a hit, record the key and runner/backend context but do not treat the marker as release evidence.

10. Cause one safe missing-artifact-style failure

On a second disposable commit, intentionally change verify to disable the artifact on the producer edge:

verify:
  needs:
    - job: produce
      artifacts: false
    - job: test
      artifacts: false
  script:
    - test -f out/product.txt

Predict that the pipeline graph remains valid but verify fails because the file is not downloaded. Run or use the supplied fixture, preserve the non-zero test failure, then repair artifacts: true. Do not use allow_failure, retry, or cache to hide the missing data edge.

11. Alternative failure when compute should be minimized: deliberate cache miss

Instead of the artifact failure, change the cache key to ch16-checkpoint-miss-$CI_PIPELINE_ID. Every pipeline uses a new key, producing a safe miss. Verify correctness remains unchanged. Restore the stable disposable key afterward. This is the preferred failure if you want to avoid a failed job.

12. Retention and access review

Confirm expire_in: 2 days and access: developer for the product artifact. Explain that latest-success retention may extend practical lifetime depending on project settings. The JUnit report uses the same short-lab retention. The cache has no evidence-retention promise and may be evicted independently.

13. Evidence packet

  • Checkpoint commit SHA and pipeline ID/source.
  • Producer/consumer job IDs and statuses.
  • Artifact archive metadata and checksum result.
  • JUnit Tests view showing one parsed case.
  • Cache hit/miss log lines with no secret/environment dump.
  • Prediction matrix versus observed files.
  • Preserved intentional failure/miss and repaired configuration.
  • Retention/access assumptions and cleanup result.

14. Cleanup / rollback

Restore/remove the checkpoint CI config, verify the exact disposable ref, then delete it. Short artifact expiry is sufficient for mandatory cleanup. If the entire project was created solely for this chapter, project deletion is acceptable only after confirming it contains no valuable data.

git switch main
git ls-remote --heads origin refs/heads/ch16/checkpoint
# Delete only the confirmed disposable ref:
git push origin --delete ch16/checkpoint
git branch -D ch16/checkpoint
Optional artifact API cleanup. Deleting job artifacts is irreversible and requires elevated project role. Use it only for a verified synthetic job. Do not bulk-delete artifacts in a shared project to “clean up the lab.”

15. Final verification checklist

  • Artifact is bound to the expected producer job/pipeline/SHA and checksum.
  • JUnit report is parsed separately from job status.
  • Cache presence/absence changes only performance, not correctness.
  • verify receives only the product artifact and no report artifact.
  • No secrets, dotenv credentials, private context, runner config, or broad environment listing is stored.
  • The intentional failure/miss has a preserved cause and deterministic repair.
  • The disposable branch/config is removed and artifact cleanup follows retention policy.

Knowledge check

Which checkpoint object is authoritative for the build result?

Why does verify depend on test with artifacts:false?

What should happen when the cache misses?

The intentionally broken verify job cannot find out/product.txt. What is the correct repair?

Why is short expiry sufficient for mandatory cleanup?

What must never appear in the evidence packet?

Checkpoint summary

You can now model GitLab delivery data explicitly: artifacts are identified outputs with access and retention, reports are structured platform inputs, caches are disposable acceleration state, and dependencies/needs define who receives what. That data-plane discipline is the prerequisite for Chapter 17, where environments and deployments attach pipeline outputs to real deployment state.

Official references

Next chapter

Environments, deployments, review apps, and deployment governance

Chapter 17 takes the traceable output you just produced and connects it to deployment/environment state, protection, approvals, and review-app lifecycles without confusing an artifact with a deployment.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.