Chapter 17Lesson 04~165 minutes

Parallel Jobs, parallel:matrix, Test Sharding, Fan-Out/Fan-In, and High-Throughput Pipeline Design: Diagnostics, Failure Modes, Security, and Performance

Diagnose matrix explosion, artifact collisions, empty shards, load imbalance, and runner saturation from preserved pipeline/job evidence before changing the configuration.

DiagnosticsCollisionEmpty shardImbalanceSaturation

Learning objectives

  • Preserve first-failure pipeline/job IDs and distinguish configuration expansion failures from runner-capacity and test-partition failures.
  • Diagnose matrix explosion, duplicate matrix identities, shared artifact names, empty shards, imbalance, and excessive queueing.
  • Interpret pending time, start/finish timestamps, shard manifests, and fan-in coverage evidence causally.
  • Repair only the smallest causal layer and rerun the smallest safe scope.
  • Avoid shortcuts such as unbounded autoscaling, blind retries, deleting all caches/artifacts, or treating a green subset as full coverage.

1. Evidence-first diagnostic sequence

  1. Preserve the pipeline ID, source/ref/SHA, failed job IDs, job traces, and artifact/report metadata.
  2. Confirm the compiled graph: expected shard/matrix count, concrete names, rules, and needs edges.
  3. Confirm effective non-secret shard inputs: CI_NODE_INDEX/CI_NODE_TOTAL or matrix variables.
  4. Inspect pending/start/finish timestamps and runner/executor identity.
  5. Inspect shard manifests and reports before touching caches or runners.
  6. Inspect fan-in artifact layout and completeness logic.
  7. Apply the smallest causal correction and rerun only the required scope.

This sequence prevents a common anti-pattern: increasing runner capacity to fix a missing-test partition bug, or rewriting the partition script to fix an infrastructure queue.

Start the incident record with CI_PIPELINE_SOURCE and CI_COMMIT_SHA; they anchor every shard, runner, artifact, and fan-in observation to the exact pipeline event and revision.

2. Failure mode: matrix explosion

Symptom: a small YAML edit unexpectedly creates dozens or hundreds of jobs, or pipeline creation fails at the matrix/active-job limit.

Cause: dimensions multiply. A new list of ten values added to an existing 5×4 matrix creates 200 combinations. Add another dimension and the configuration is invalid.

Evidence: compiled job count, matrix values in names, CI lint/creation error, and predicted cardinality calculation.

Repair: remove redundant dimensions, split hypotheses into staged selective pipelines, or gate expensive combinations by source/rules. Do not solve it with unbounded autoscaling.

3. Failure mode: duplicate matrix value combinations overwrite jobs

GitLab generates matrix job names from matrix values. If two matrix entries have different variable names but the same effective values, generated names can collide and one set overwrites the other during configuration expansion. This is a compilation-layer defect: no runner tuning can recover jobs that never exist in the compiled graph.

Inspect merged/compiled configuration and concrete job names first. Make matrix combinations semantically unique.

4. Failure mode: every shard writes the same artifact path

# Intentionally broken
shard:test:
  parallel: 4
  script:
    - printf '%s\n' "$CI_NODE_INDEX" > report.txt
  artifacts:
    paths: [report.txt]

fan-in:
  needs: [shard:test]
  script:
    - cat report.txt

Interpretation: fan-in needs every parallel producer, so artifacts from all instances are downloaded. Same-name files can overwrite one another; the surviving report.txt is not aggregate evidence.

Repair: use evidence/shard-$CI_NODE_INDEX/report.txt and assert the expected number of directories/files.

5. Failure mode: one shard is silently empty

Symptom: all shard jobs pass, total runtime looks excellent, but one job reports zero tests. If the fan-in gate only checks job status, the pipeline can be falsely green.

Evidence: shard count metadata, canonical manifest size, union/duplicate checks. An empty shard can be valid when shard count exceeds work units, but it must be intentional and visible.

Repair: reduce shard count or make the gate explicitly accept/flag empty shards according to policy. Do not add dummy tests merely to hide the symptom.

6. Failure mode: imbalance dominates the critical path

Suppose shard durations are 2, 2, 3, and 17 minutes. Parallel execution finishes near the slowest shard, not near the average. Queue time may be low, so runner capacity is not the problem. The partition function put expensive work together.

Preserve per-test/per-shard timing evidence, then rebalance deterministically. A history-aware splitter is a performance optimization; the canonical test manifest remains the correctness source.

7. Failure mode: more concurrency makes the pipeline slower

Symptom: after increasing shards from 8 to 40, median execution time is similar but pending time jumps. Other projects may also wait longer.

Cause: requested runnable jobs exceed available runner/fleet capacity, image pulls/services contend for shared resources, or runner manager concurrency settings are lower than expected.

Evidence: job creation/start timestamps, pending durations, runner assignment, executor metrics if authorized, and host/fleet saturation—not just total pipeline duration.

Repair: first reduce unnecessary fan-out. Capacity changes belong to the runner/fleet layer and require explicit cost/security review.

8. Failure mode map: repair the owning layer

Observed evidence Owning layer First repair
Expected jobs absent from graph YAML compilation/matrix/rules Fix matrix/rules; do not touch runners
Jobs exist but stay pending Runner/fleet/queue Bound fan-out or review authorized capacity
All jobs start, one has zero inputs Partition/dataflow Fix shard mapping/count policy
All shards pass, fan-in misses files Artifact naming/transfer Make evidence identities unique; verify downloads
Coverage complete but one shard takes 10× longer Partition balancing Rebalance workload, not credentials/runner trust

9. Security and disruption guardrails

  • Do not print tokens or full environments to distinguish shards; print non-secret shard metadata only.
  • Do not put untrusted fork code onto privileged runners to gain more capacity.
  • Do not disable TLS or broaden network access because many shards fail to fetch a dependency.
  • Do not create unbounded autoscaling as a troubleshooting shortcut.
  • Do not delete all caches/artifacts before preserving first-failure evidence.
  • Do not make a deployment or package-publish job parallel unless the side-effect model is explicitly designed for it.

10. Intentionally broken scenario: “green shards, incomplete suite”

Start from the Chapter 17 four-shard lab. Modify the partition script so it selects lines with ((NR-1) % total) == index while leaving index one-based. One remainder is never selected and another shard can become empty. Preserve the pipeline and job IDs.

Diagnosis should show: compiled graph correct → jobs assigned/running correctly → wrong effective partition arithmetic → shard manifests do not union to canonical manifest → fan-in fails. Repair only the zero-based conversion. Do not change runner capacity, artifact retention, or pipeline rules.

11. Retry and rerun policy

If one shard failed because its deterministic test failed, retrying may be useful only after preserving the original trace and deciding whether the failure is flaky. If the partition script or matrix configuration is wrong, retrying the same job cannot change the compiled configuration; create a corrected pipeline from a reviewed commit. If only fan-in logic is wrong and all shard artifacts remain valid, avoid rebuilding release-like evidence unnecessarily.

Knowledge check

A pipeline has 40 compiled shards but only four running. What evidence should you inspect first?

Why can a parallel artifact collision make a green pipeline unauditable?

Where do duplicate matrix-value job collisions occur?

What proves that a zero-test shard is harmless?

Why is blind retry a poor response to a bad partition rule?

Next lesson

Checkpoint lab

Refactor a serial suite into bounded shards, prove complete coverage, break one evidence contract, and recover without hiding the failure.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Parallel-job limits, matrix expressions, runner concurrency controls, job-activity limits, and artifact-transfer behavior are version-sensitive. Re-check the deployed GitLab and GitLab Runner versions before using production capacity numbers or beta expression features.

Current assumptions used in this chapter: all mandatory examples use Free-tier CI/CD syntax and synthetic data. Numeric parallel accepts 1–200; parallel:matrix accepts at most 200 permutations. Matrix expressions are current but version-sensitive and are presented as an optional refinement, not a prerequisite. No cloud account, autoscaling fleet, protected environment, or administrator setting is required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.