Checkpoint Lab — Parallel Jobs, parallel:matrix, Test Sharding, Fan-Out/Fan-In, and High-Throughput Pipeline Design
Refactor a serial synthetic suite into bounded shards, prove complete coverage and unique evidence, deliberately break fan-in evidence, and recover without hiding the first failure.
Learning objectives
- Predict the job expansion and evidence layout before executing the checkpoint pipeline.
- Convert a serial deterministic suite into bounded parallel shards while keeping the source SHA and test manifest fixed.
- Prove complete and non-overlapping coverage from per-shard manifests.
- Repair an intentionally broken artifact/fan-in design without rebuilding unrelated evidence.
- Produce a compact evidence packet that another engineer can audit independently.
1. Checkpoint scenario and success criteria
You inherit a serial pipeline containing sixteen synthetic tests. Your task is to create four bounded shards, retain unique evidence for each, prove exact coverage, then reproduce and repair one deliberately broken artifact/fan-in design. The checkpoint is complete only when another engineer can reconstruct the source SHA, expanded jobs, shard ownership, queue/runtime behavior, and final coverage proof from retained evidence.
No production side effects: use a disposable project/branch, fake data, no secrets, and no privileged/cloud infrastructure.
2. Assumptions and preflight
| Item | Checkpoint assumption |
|---|---|
| GitLab feature tier | Mandatory path uses Free syntax only. |
| Runner | Any authorized runner capable of Alpine shell jobs; record actual Runner version/executor if visible. |
| Image |
alpine:3.22; record actual resolved image
identity if your platform exposes it.
|
| Source |
Disposable branch glci/ch17-checkpoint; record
exact commit SHA.
|
| Secrets | None. Do not add project/group variables for this lab. |
| External systems | None. Artifact storage is GitLab-owned evidence only. |
3. Predict state changes before execution
Write down at least these predictions before pushing:
- The compiled graph changes from one serial test job to four concrete shard jobs plus one fan-in verification job.
- Each shard owns exactly four of sixteen test IDs under the modulo rule; no ID appears in two shard manifests.
- If effective runner capacity is lower than four, some shards will remain pending even though all are immediately runnable.
- Each shard uploads a unique directory and artifact name; the fan-in job receives all four.
Do not edit these predictions after seeing the result. The comparison between prediction and observation is part of the evidence.
4. Create the canonical test manifest and partition helper
git switch -c glci/ch17-checkpoint
mkdir -p tests ci
seq -w 1 16 | sed 's/^/case-/' > tests/checkpoint.txt
sha256sum tests/checkpoint.txt > tests/checkpoint.sha256
cat > ci/checkpoint-shard.sh <<'EOF'
#!/bin/sh
set -eu
idx="$1"; total="$2"; zero=$((idx - 1))
out="evidence/shard-$idx"
mkdir -p "$out"
awk -v z="$zero" -v n="$total" '((NR-1)%n)==z {print}' tests/checkpoint.txt > "$out/tests.txt"
count=$(wc -l < "$out/tests.txt" | tr -d ' ')
printf 'shard=%s/%s\ncount=%s\nsha=%s\n' "$idx" "$total" "$count" "$CI_COMMIT_SHA" > "$out/meta.txt"
while IFS= read -r t; do printf 'PASS %s\n' "$t"; sleep 1; done < "$out/tests.txt"
EOF
chmod +x ci/checkpoint-shard.sh
Before the pipeline, you can locally simulate indices 1–4 with a
temporary CI_COMMIT_SHA=local-simulation environment if
desired. The local simulation proves the partition function, not
GitLab queue behavior.
5. Exact checkpoint pipeline
stages: [test, verify]
checkpoint:test:
stage: test
image: alpine:3.22
parallel: 4
script:
- printf 'pipeline=%s job=%s source=%s sha=%s shard=%s/%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA" "$CI_NODE_INDEX" "$CI_NODE_TOTAL"
- ./ci/checkpoint-shard.sh "$CI_NODE_INDEX" "$CI_NODE_TOTAL"
artifacts:
name: "ch17-checkpoint-$CI_NODE_INDEX-of-$CI_NODE_TOTAL"
when: always
expire_in: 1 week
paths:
- "evidence/shard-$CI_NODE_INDEX/"
checkpoint:verify:
stage: verify
image: alpine:3.22
needs:
- job: checkpoint:test
artifacts: true
script:
- find evidence -maxdepth 2 -type f -print | sort
- test "$(find evidence -mindepth 1 -maxdepth 1 -type d | wc -l)" -eq 4
- cat evidence/shard-*/tests.txt | sort > observed.txt
- sort tests/checkpoint.txt > expected.txt
- test "$(wc -l < observed.txt)" -eq "$(sort -u observed.txt | wc -l)"
- diff -u expected.txt observed.txt
- printf 'coverage=complete\nsha=%s\n' "$CI_COMMIT_SHA" > verification.txt
artifacts:
when: always
paths: [observed.txt, expected.txt, verification.txt]
6. Expected observations and independent verification
-
Pipeline page shows four concrete
checkpoint:testjobs and one verify job. - Every shard trace prints the same pipeline/source/SHA but a distinct shard identity.
-
Every shard artifact contains
tests.txtandmeta.txtunder a unique directory. -
Each
tests.txtcontains four IDs. Concatenating and sorting all four exactly matchestests/checkpoint.txt. - The verify job starts only after all four required shard jobs succeed because it needs the parallelized producer.
- Pending time may differ from execution time; record both rather than assuming simultaneous execution.
7. Deliberately break the evidence contract
Create a second disposable commit that changes the shard artifact path to a common location:
# Broken on purpose
artifacts:
name: "shared-checkpoint"
paths:
- evidence/shared/
and change the script so every shard writes
evidence/shared/tests.txt. Run the pipeline once.
Preserve the pipeline/job IDs and the fan-in failure or incomplete
evidence. Do not delete the broken pipeline.
Expected cause: independent jobs produced non-unique evidence identities. This is an artifact/dataflow defect, not a test failure or runner-capacity defect.
8. Repair without hiding the original cause
Restore evidence/shard-$CI_NODE_INDEX/ and the
shard-qualified artifact name. Push a new commit. In the checkpoint
note, link/record both the broken and repaired pipeline IDs and
explain why the newer pipeline is a repair rather than a retry.
Do not rebuild or replace the old pipeline’s artifacts to make history look clean. The failed evidence is useful audit material.
9. Optional matrix extension
After the core checkpoint passes, optionally add a two-by-two
synthetic matrix (MODE=[fast,strict],
DATASET=[small,medium]) and predict four generated
jobs. If your GitLab version supports matrix expressions and you
want to explore 1:1 producer/consumer dependencies, do so on a
separate disposable branch and record the version assumption. This
extension is not required for completion.
10. Required evidence packet
| Evidence | Record |
|---|---|
| Source identity |
project path, branch, CI_PIPELINE_SOURCE, exact
CI_COMMIT_SHA
|
| Pipeline identity | baseline/parallel/broken/repaired pipeline IDs and statuses |
| Graph | expanded shard count, fan-in job, any matrix extension |
| Jobs/runners | job IDs, shard identity, pending/start/finish times, runner/executor if visible |
| Partition |
canonical manifest hash, each shard
tests.txt/meta.txt
|
| Artifacts/reports | artifact names, unique paths, expiry, fan-in download behavior |
| Coverage proof |
expected vs observed diff result; duplicate count;
verification.txt
|
| Assumptions | GitLab/Runner/image versions known, unknown, or platform-managed |
11. Verification checklist
- Exactly four intended shards exist; no unexpected matrix or duplicate jobs.
- All sixteen canonical test IDs appear exactly once across shard manifests.
- All four artifact identities are unique.
- The fan-in job proves completeness from retained files, not only upstream job status.
- Queue time and execution time are recorded separately.
- No credentials, protected resources, cloud subscriptions, or production data were used.
- The broken pipeline remains preserved as first-failure evidence.
12. Cleanup / rollback
- Keep the evidence packet outside any secret store; it contains only synthetic metadata.
- Delete the disposable Git branch/project only if you no longer need the lab history. If preserving it, let the one-week artifacts expire naturally.
- Do not change or “reset” shared runner capacity as cleanup because this lab never required runner-setting mutations.
- If you created a local simulation directory, remove only that named temporary directory.
13. What Chapter 17 adds to the production operating model
Earlier chapters established source identity, rules, DAG dependencies, artifacts, reuse, downstream orchestration, and pre-merge trust. Chapter 17 adds a capacity-aware throughput model: every parallel job has an explicit partition identity, every shard emits collision-free evidence, runner queueing is measured separately from execution, and a fan-in gate proves the union of work before the pipeline claims completeness.
Chapter 18 builds directly on this: once many jobs can run
concurrently, you need deliberate controls for
resource_group, interruptibility, retry, timeout, and
duplicate-pipeline behavior so concurrency does not create
conflicting side effects or wasted work.
Knowledge check
What are the two strongest checkpoint proofs of complete sharding?
The observed union exactly matches the canonical manifest, and the observed count equals the unique observed count so no test is duplicated.
Why keep the broken artifact-collision pipeline?
It preserves the original causal evidence and proves the repair changed evidence identity rather than rewriting history.
If four shards exist but only two start immediately, is the pipeline incorrect?
Not necessarily. The graph is correct; available runner capacity may be two. Record pending time and continue validating coverage independently.
What should a reviewer reject even if total duration improves?
Any change that loses shard identity, permits artifact overwrites, omits coverage proof, or creates unsafe/unbounded runner demand.
Which Chapter 18 concern naturally follows this lab?
Controlling concurrent side effects and wasted work with resource groups, interruptibility, retries, timeouts, and duplicate-pipeline control.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. Parallel-job limits, matrix expressions, runner concurrency controls, job-activity limits, and artifact-transfer behavior are version-sensitive. Re-check the deployed GitLab and GitLab Runner versions before using production capacity numbers or beta expression features.
-
CI/CD YAML syntax reference
— current
parallel,parallel:matrix,needs, and artifact semantics. - Control how jobs run — sharding patterns, matrix jobs, and selecting parallelized dependencies.
-
Matrix expressions
— current compile-time
$[[ matrix.IDENTIFIER ]]behavior introduced in GitLab 18.6. -
Predefined variables
—
CI_NODE_INDEX,CI_NODE_TOTAL, pipeline/job IDs, source SHA, and timing evidence. -
GitLab Runner advanced configuration
—
concurrent, per-runnerlimit, andrequest_concurrency. - Runner fleet scaling — capacity planning, executor behavior, and the difference between requested fan-out and available workers.
Current assumptions used in this chapter: all
mandatory examples use Free-tier CI/CD syntax and synthetic data.
Numeric parallel accepts 1–200;
parallel:matrix accepts at most 200 permutations.
Matrix expressions are current but version-sensitive and are
presented as an optional refinement, not a prerequisite. No cloud
account, autoscaling fleet, protected environment, or administrator
setting is required.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.