Chapter 17Lesson 02~190 minutes

Parallel Jobs, parallel:matrix, Test Sharding, Fan-Out/Fan-In, and High-Throughput Pipeline Design: Guided Hands-On Workflow and Core Operations

Build a disposable deterministic test lab, expand it with parallel and matrix jobs, preserve unique shard evidence, and measure queue/runtime behavior without relying on paid infrastructure.

Hands-onCI_NODE_INDEXMatrixArtifactsQueue time

Learning objectives

  • Create a deterministic synthetic test list and partition it with `CI_NODE_INDEX` and `CI_NODE_TOTAL`.
  • Create a bounded `parallel:matrix` job and inspect the generated job identities and variables.
  • Publish unique per-shard evidence and aggregate it without depending on shared mutable files.
  • Measure execution time separately from pending/queue time so runner saturation is visible.
  • Prove that every synthetic test is executed exactly once and that no shard silently disappears.

1. Lab scope and preflight: one disposable repository, synthetic tests only

Create or reuse a throwaway GitLab project named, for example, glci-ch17-lab. The repository contains only generated text files and shell/Python helpers. Do not use production runners, proprietary tests, credentials, cloud accounts, or shared mutable environments.

Record before changing anything: GitLab offering/version if known, Runner version/executor if visible, project path, branch, CI_PIPELINE_SOURCE, source SHA, baseline pipeline/job IDs, and the baseline serial runtime. If using GitLab.com hosted runners, record that runner capacity is platform-managed rather than pretending you know its fleet settings.

2. Create a deterministic twelve-test manifest

The “tests” are deliberately synthetic so the partition rule—not a framework—is the subject.

mkdir -p tests ci evidence
cat > tests/list.txt <<'EOF'
test-01
test-02
test-03
test-04
test-05
test-06
test-07
test-08
test-09
test-10
test-11
test-12
EOF

sha256sum tests/list.txt > tests/list.sha256
git add tests/list.txt tests/list.sha256
git commit -m "ch17: add deterministic test manifest"

The SHA-256 file is evidence that serial and sharded runs started from the same manifest. The Git commit SHA remains the stronger repository identity; keep both.

3. Establish a serial baseline before parallelizing

stages: [test, verify]

serial:test:
  stage: test
  image: alpine:3.22
  script:
    - apk add --no-cache coreutils
    - mkdir -p evidence/serial
    - start=$(date +%s)
    - while IFS= read -r t; do printf '%s\n' "$t"; sleep 1; done < tests/list.txt | tee evidence/serial/tests.txt
    - end=$(date +%s)
    - printf 'seconds=%s\n' "$((end-start))" > evidence/serial/timing.txt
    - sha256sum tests/list.txt > evidence/serial/manifest.sha256
  artifacts:
    when: always
    paths: [evidence/serial/]

Run once and record the job ID, pending/start/finish timestamps, execution duration, artifact name, and source SHA. Do not compare a later sharded pipeline against a different commit and call the result a speedup.

4. Partition by index with an explicit, testable rule

CI_NODE_INDEX is one-based for parallel jobs. The script below converts it to zero-based arithmetic and assigns each line by modulo. Every input line has exactly one remainder, so the partition is deterministic.

cat > ci/run-shard.sh <<'EOF'
#!/bin/sh
set -eu
idx="$1"
total="$2"
zero=$((idx - 1))
out="evidence/shard-${idx}"
mkdir -p "$out"
awk -v z="$zero" -v n="$total" '((NR-1) % n) == z {print}' tests/list.txt > "$out/tests.txt"
count=$(wc -l < "$out/tests.txt" | tr -d ' ')
printf 'shard=%s/%s\ncount=%s\nsha=%s\n' "$idx" "$total" "$count" "$CI_COMMIT_SHA" > "$out/meta.txt"
while IFS= read -r t; do printf 'PASS %s\n' "$t"; sleep 1; done < "$out/tests.txt"
EOF
chmod +x ci/run-shard.sh

The script records no secrets and has no external side effect. An empty shard is observable because count=0 is retained rather than silently ignored.

5. Convert the serial job into four bounded shards

test:shard:
  stage: test
  image: alpine:3.22
  parallel: 4
  script:
    - apk add --no-cache coreutils
    - printf 'pipeline=%s job=%s source=%s sha=%s node=%s/%s\n'         "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA" "$CI_NODE_INDEX" "$CI_NODE_TOTAL"
    - ./ci/run-shard.sh "$CI_NODE_INDEX" "$CI_NODE_TOTAL"
  artifacts:
    name: "ch17-shard-$CI_NODE_INDEX-of-$CI_NODE_TOTAL"
    when: always
    paths:
      - "evidence/shard-$CI_NODE_INDEX/"
    expire_in: 1 week

Inspect the compiled job graph: exactly four concrete test:shard instances should exist. Then compare their pending times. Four jobs in the graph does not imply four workers started immediately.

6. Fan-in: prove all twelve test IDs exactly once

The fan-in job downloads artifacts from all instances when it needs the parallelized producer. Unique shard directories prevent file overwrite. The script reconstructs the observed union and compares it with the canonical manifest.

verify:coverage:
  stage: verify
  image: alpine:3.22
  needs:
    - job: test:shard
      artifacts: true
  script:
    - cat evidence/shard-*/tests.txt | sort > evidence/observed.txt
    - sort tests/list.txt > evidence/expected.txt
    - test "$(wc -l < evidence/observed.txt)" -eq "$(sort -u evidence/observed.txt | wc -l)"
    - diff -u evidence/expected.txt evidence/observed.txt
    - printf 'coverage=complete\nsha=%s\n' "$CI_COMMIT_SHA" > evidence/fan-in.txt
  artifacts:
    when: always
    paths:
      - evidence/observed.txt
      - evidence/expected.txt
      - evidence/fan-in.txt

The duplicate-count check and diff prove two different properties: no test was assigned twice, and no expected test was omitted.

7. Add a small compatibility matrix without exploding cardinality

test:profile:
  stage: test
  image: alpine:3.22
  parallel:
    matrix:
      - PROFILE: [fast, strict]
        DATASET: [small, medium]
  script:
    - mkdir -p "evidence/matrix-$PROFILE-$DATASET"
    - printf 'profile=%s\ndataset=%s\nsha=%s\n' "$PROFILE" "$DATASET" "$CI_COMMIT_SHA"         > "evidence/matrix-$PROFILE-$DATASET/meta.txt"
  artifacts:
    name: "matrix-$PROFILE-$DATASET"
    paths: ["evidence/matrix-$PROFILE-$DATASET/"]

Predict the cardinality before pushing: 2 profiles × 2 datasets = 4 jobs. This is small enough to inspect manually and large enough to teach explicit dimensions.

8. Measure queue time and runtime separately

Use the job UI or API metadata to record creation, queued/start, and finish timestamps for every shard. A simple evidence table is enough:

Job Shard Pending/queue Execution Runner/executor
test:shard 1/4 1/4 record observed record observed record observed
test:shard 2/4 2/4 record observed record observed record observed
test:shard 3/4 3/4 record observed record observed record observed
test:shard 4/4 4/4 record observed record observed record observed

If the fourth shard waits while the first three run, the bottleneck is capacity/queueing—not the partition script. Do not “fix” it by changing test allocation.

9. Deliberately create one evidence collision, then observe it

On a disposable branch only, temporarily change every shard to upload a common path such as evidence/report.txt. Let each shard write its own shard number there. The fan-in download can leave only one surviving file because identical paths overwrite one another.

Do not use this broken pattern for real evidence. Preserve the failed pipeline/job IDs and downloaded artifact state before fixing it.

Repair by restoring a shard-qualified directory or filename. The lesson is not “GitLab artifacts are unreliable”; it is “parallel producers require collision-free evidence identities.”

10. Challenge: choose the causal layer

Your pipeline graph shows four shards. All four jobs passed. The fan-in job reports only nine unique test IDs. Runner pending time was near zero. Which layer do you inspect first?

Start with the partition manifests and fan-in artifact layout, not runner capacity. The job graph and queue evidence already show that scheduling worked; the missing coverage is a dataflow/partition-evidence problem.

11. Cleanup and rollback

  1. Preserve the baseline and sharded pipeline IDs, source SHA, job IDs, timing table, and fan-in result.
  2. Delete only the disposable branch used for the deliberately broken collision if you no longer need it.
  3. Remove temporary synthetic artifact files from the working tree if they were accidentally committed; do not delete GitLab artifacts or caches broadly.
  4. Do not change runner or instance concurrency settings as part of cleanup.

Knowledge check

Why convert CI_NODE_INDEX before modulo sharding?

What independent checks does the fan-in example perform?

Four shards are pending for minutes. What layer is implicated first?

Why keep matrix cardinality at four in the lab?

What should you preserve before fixing an artifact collision?

Next lesson

Configuration, design choices, and tradeoffs

Choose shard/matrix architecture, balancing strategy, fan-in wiring, and capacity budgets with current GitLab limits.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Parallel-job limits, matrix expressions, runner concurrency controls, job-activity limits, and artifact-transfer behavior are version-sensitive. Re-check the deployed GitLab and GitLab Runner versions before using production capacity numbers or beta expression features.

Current assumptions used in this chapter: all mandatory examples use Free-tier CI/CD syntax and synthetic data. Numeric parallel accepts 1–200; parallel:matrix accepts at most 200 permutations. Matrix expressions are current but version-sensitive and are presented as an optional refinement, not a prerequisite. No cloud account, autoscaling fleet, protected environment, or administrator setting is required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.