Chapter 17Lesson 05~200 minutes

Checkpoint Lab — Parallel Jobs, parallel:matrix, Test Sharding, Fan-Out/Fan-In, and High-Throughput Pipeline Design

Refactor a serial synthetic suite into bounded shards, prove complete coverage and unique evidence, deliberately break fan-in evidence, and recover without hiding the first failure.

Checkpoint labCoverage proofFan-inEvidence packetCleanup

Learning objectives

  • Predict the job expansion and evidence layout before executing the checkpoint pipeline.
  • Convert a serial deterministic suite into bounded parallel shards while keeping the source SHA and test manifest fixed.
  • Prove complete and non-overlapping coverage from per-shard manifests.
  • Repair an intentionally broken artifact/fan-in design without rebuilding unrelated evidence.
  • Produce a compact evidence packet that another engineer can audit independently.

1. Checkpoint scenario and success criteria

You inherit a serial pipeline containing sixteen synthetic tests. Your task is to create four bounded shards, retain unique evidence for each, prove exact coverage, then reproduce and repair one deliberately broken artifact/fan-in design. The checkpoint is complete only when another engineer can reconstruct the source SHA, expanded jobs, shard ownership, queue/runtime behavior, and final coverage proof from retained evidence.

No production side effects: use a disposable project/branch, fake data, no secrets, and no privileged/cloud infrastructure.

2. Assumptions and preflight

Item Checkpoint assumption
GitLab feature tier Mandatory path uses Free syntax only.
Runner Any authorized runner capable of Alpine shell jobs; record actual Runner version/executor if visible.
Image alpine:3.22; record actual resolved image identity if your platform exposes it.
Source Disposable branch glci/ch17-checkpoint; record exact commit SHA.
Secrets None. Do not add project/group variables for this lab.
External systems None. Artifact storage is GitLab-owned evidence only.

3. Predict state changes before execution

Write down at least these predictions before pushing:

  1. The compiled graph changes from one serial test job to four concrete shard jobs plus one fan-in verification job.
  2. Each shard owns exactly four of sixteen test IDs under the modulo rule; no ID appears in two shard manifests.
  3. If effective runner capacity is lower than four, some shards will remain pending even though all are immediately runnable.
  4. Each shard uploads a unique directory and artifact name; the fan-in job receives all four.

Do not edit these predictions after seeing the result. The comparison between prediction and observation is part of the evidence.

4. Create the canonical test manifest and partition helper

git switch -c glci/ch17-checkpoint
mkdir -p tests ci
seq -w 1 16 | sed 's/^/case-/' > tests/checkpoint.txt
sha256sum tests/checkpoint.txt > tests/checkpoint.sha256

cat > ci/checkpoint-shard.sh <<'EOF'
#!/bin/sh
set -eu
idx="$1"; total="$2"; zero=$((idx - 1))
out="evidence/shard-$idx"
mkdir -p "$out"
awk -v z="$zero" -v n="$total" '((NR-1)%n)==z {print}' tests/checkpoint.txt > "$out/tests.txt"
count=$(wc -l < "$out/tests.txt" | tr -d ' ')
printf 'shard=%s/%s\ncount=%s\nsha=%s\n' "$idx" "$total" "$count" "$CI_COMMIT_SHA" > "$out/meta.txt"
while IFS= read -r t; do printf 'PASS %s\n' "$t"; sleep 1; done < "$out/tests.txt"
EOF
chmod +x ci/checkpoint-shard.sh

Before the pipeline, you can locally simulate indices 1–4 with a temporary CI_COMMIT_SHA=local-simulation environment if desired. The local simulation proves the partition function, not GitLab queue behavior.

5. Exact checkpoint pipeline

stages: [test, verify]

checkpoint:test:
  stage: test
  image: alpine:3.22
  parallel: 4
  script:
    - printf 'pipeline=%s job=%s source=%s sha=%s shard=%s/%s\n'         "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA" "$CI_NODE_INDEX" "$CI_NODE_TOTAL"
    - ./ci/checkpoint-shard.sh "$CI_NODE_INDEX" "$CI_NODE_TOTAL"
  artifacts:
    name: "ch17-checkpoint-$CI_NODE_INDEX-of-$CI_NODE_TOTAL"
    when: always
    expire_in: 1 week
    paths:
      - "evidence/shard-$CI_NODE_INDEX/"

checkpoint:verify:
  stage: verify
  image: alpine:3.22
  needs:
    - job: checkpoint:test
      artifacts: true
  script:
    - find evidence -maxdepth 2 -type f -print | sort
    - test "$(find evidence -mindepth 1 -maxdepth 1 -type d | wc -l)" -eq 4
    - cat evidence/shard-*/tests.txt | sort > observed.txt
    - sort tests/checkpoint.txt > expected.txt
    - test "$(wc -l < observed.txt)" -eq "$(sort -u observed.txt | wc -l)"
    - diff -u expected.txt observed.txt
    - printf 'coverage=complete\nsha=%s\n' "$CI_COMMIT_SHA" > verification.txt
  artifacts:
    when: always
    paths: [observed.txt, expected.txt, verification.txt]

6. Expected observations and independent verification

  1. Pipeline page shows four concrete checkpoint:test jobs and one verify job.
  2. Every shard trace prints the same pipeline/source/SHA but a distinct shard identity.
  3. Every shard artifact contains tests.txt and meta.txt under a unique directory.
  4. Each tests.txt contains four IDs. Concatenating and sorting all four exactly matches tests/checkpoint.txt.
  5. The verify job starts only after all four required shard jobs succeed because it needs the parallelized producer.
  6. Pending time may differ from execution time; record both rather than assuming simultaneous execution.

7. Deliberately break the evidence contract

Create a second disposable commit that changes the shard artifact path to a common location:

# Broken on purpose
artifacts:
  name: "shared-checkpoint"
  paths:
    - evidence/shared/

and change the script so every shard writes evidence/shared/tests.txt. Run the pipeline once. Preserve the pipeline/job IDs and the fan-in failure or incomplete evidence. Do not delete the broken pipeline.

Expected cause: independent jobs produced non-unique evidence identities. This is an artifact/dataflow defect, not a test failure or runner-capacity defect.

8. Repair without hiding the original cause

Restore evidence/shard-$CI_NODE_INDEX/ and the shard-qualified artifact name. Push a new commit. In the checkpoint note, link/record both the broken and repaired pipeline IDs and explain why the newer pipeline is a repair rather than a retry.

Do not rebuild or replace the old pipeline’s artifacts to make history look clean. The failed evidence is useful audit material.

9. Optional matrix extension

After the core checkpoint passes, optionally add a two-by-two synthetic matrix (MODE=[fast,strict], DATASET=[small,medium]) and predict four generated jobs. If your GitLab version supports matrix expressions and you want to explore 1:1 producer/consumer dependencies, do so on a separate disposable branch and record the version assumption. This extension is not required for completion.

10. Required evidence packet

Evidence Record
Source identity project path, branch, CI_PIPELINE_SOURCE, exact CI_COMMIT_SHA
Pipeline identity baseline/parallel/broken/repaired pipeline IDs and statuses
Graph expanded shard count, fan-in job, any matrix extension
Jobs/runners job IDs, shard identity, pending/start/finish times, runner/executor if visible
Partition canonical manifest hash, each shard tests.txt/meta.txt
Artifacts/reports artifact names, unique paths, expiry, fan-in download behavior
Coverage proof expected vs observed diff result; duplicate count; verification.txt
Assumptions GitLab/Runner/image versions known, unknown, or platform-managed

11. Verification checklist

  • Exactly four intended shards exist; no unexpected matrix or duplicate jobs.
  • All sixteen canonical test IDs appear exactly once across shard manifests.
  • All four artifact identities are unique.
  • The fan-in job proves completeness from retained files, not only upstream job status.
  • Queue time and execution time are recorded separately.
  • No credentials, protected resources, cloud subscriptions, or production data were used.
  • The broken pipeline remains preserved as first-failure evidence.

12. Cleanup / rollback

  1. Keep the evidence packet outside any secret store; it contains only synthetic metadata.
  2. Delete the disposable Git branch/project only if you no longer need the lab history. If preserving it, let the one-week artifacts expire naturally.
  3. Do not change or “reset” shared runner capacity as cleanup because this lab never required runner-setting mutations.
  4. If you created a local simulation directory, remove only that named temporary directory.

13. What Chapter 17 adds to the production operating model

Earlier chapters established source identity, rules, DAG dependencies, artifacts, reuse, downstream orchestration, and pre-merge trust. Chapter 17 adds a capacity-aware throughput model: every parallel job has an explicit partition identity, every shard emits collision-free evidence, runner queueing is measured separately from execution, and a fan-in gate proves the union of work before the pipeline claims completeness.

Chapter 18 builds directly on this: once many jobs can run concurrently, you need deliberate controls for resource_group, interruptibility, retry, timeout, and duplicate-pipeline behavior so concurrency does not create conflicting side effects or wasted work.

Knowledge check

What are the two strongest checkpoint proofs of complete sharding?

Why keep the broken artifact-collision pipeline?

If four shards exist but only two start immediately, is the pipeline incorrect?

What should a reviewer reject even if total duration improves?

Which Chapter 18 concern naturally follows this lab?

Next chapter

Chapter 18 — Concurrency controls

Move from safe fan-out to resource groups, interruptible jobs, retry/timeout policy, and duplicate-pipeline control.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Parallel-job limits, matrix expressions, runner concurrency controls, job-activity limits, and artifact-transfer behavior are version-sensitive. Re-check the deployed GitLab and GitLab Runner versions before using production capacity numbers or beta expression features.

Current assumptions used in this chapter: all mandatory examples use Free-tier CI/CD syntax and synthetic data. Numeric parallel accepts 1–200; parallel:matrix accepts at most 200 permutations. Matrix expressions are current but version-sensitive and are presented as an optional refinement, not a prerequisite. No cloud account, autoscaling fleet, protected environment, or administrator setting is required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.