Parallel Jobs, parallel:matrix, Test Sharding, Fan-Out/Fan-In, and High-Throughput Pipeline Design: Diagnostics, Failure Modes, Security, and Performance
Diagnose matrix explosion, artifact collisions, empty shards, load imbalance, and runner saturation from preserved pipeline/job evidence before changing the configuration.
Learning objectives
- Preserve first-failure pipeline/job IDs and distinguish configuration expansion failures from runner-capacity and test-partition failures.
- Diagnose matrix explosion, duplicate matrix identities, shared artifact names, empty shards, imbalance, and excessive queueing.
- Interpret pending time, start/finish timestamps, shard manifests, and fan-in coverage evidence causally.
- Repair only the smallest causal layer and rerun the smallest safe scope.
- Avoid shortcuts such as unbounded autoscaling, blind retries, deleting all caches/artifacts, or treating a green subset as full coverage.
1. Evidence-first diagnostic sequence
- Preserve the pipeline ID, source/ref/SHA, failed job IDs, job traces, and artifact/report metadata.
-
Confirm the compiled graph: expected shard/matrix count, concrete
names, rules, and
needsedges. -
Confirm effective non-secret shard inputs:
CI_NODE_INDEX/CI_NODE_TOTALor matrix variables. - Inspect pending/start/finish timestamps and runner/executor identity.
- Inspect shard manifests and reports before touching caches or runners.
- Inspect fan-in artifact layout and completeness logic.
- Apply the smallest causal correction and rerun only the required scope.
This sequence prevents a common anti-pattern: increasing runner capacity to fix a missing-test partition bug, or rewriting the partition script to fix an infrastructure queue.
Start the incident record with CI_PIPELINE_SOURCE and
CI_COMMIT_SHA; they anchor every shard, runner,
artifact, and fan-in observation to the exact pipeline event and
revision.
2. Failure mode: matrix explosion
Symptom: a small YAML edit unexpectedly creates dozens or hundreds of jobs, or pipeline creation fails at the matrix/active-job limit.
Cause: dimensions multiply. A new list of ten values added to an existing 5×4 matrix creates 200 combinations. Add another dimension and the configuration is invalid.
Evidence: compiled job count, matrix values in names, CI lint/creation error, and predicted cardinality calculation.
Repair: remove redundant dimensions, split hypotheses into staged selective pipelines, or gate expensive combinations by source/rules. Do not solve it with unbounded autoscaling.
3. Failure mode: duplicate matrix value combinations overwrite jobs
GitLab generates matrix job names from matrix values. If two matrix entries have different variable names but the same effective values, generated names can collide and one set overwrites the other during configuration expansion. This is a compilation-layer defect: no runner tuning can recover jobs that never exist in the compiled graph.
Inspect merged/compiled configuration and concrete job names first. Make matrix combinations semantically unique.
4. Failure mode: every shard writes the same artifact path
# Intentionally broken
shard:test:
parallel: 4
script:
- printf '%s\n' "$CI_NODE_INDEX" > report.txt
artifacts:
paths: [report.txt]
fan-in:
needs: [shard:test]
script:
- cat report.txt
Interpretation: fan-in needs every
parallel producer, so artifacts from all instances are downloaded.
Same-name files can overwrite one another; the surviving
report.txt is not aggregate evidence.
Repair: use
evidence/shard-$CI_NODE_INDEX/report.txt and assert the
expected number of directories/files.
5. Failure mode: one shard is silently empty
Symptom: all shard jobs pass, total runtime looks excellent, but one job reports zero tests. If the fan-in gate only checks job status, the pipeline can be falsely green.
Evidence: shard count metadata,
canonical manifest size, union/duplicate checks. An empty shard can
be valid when shard count exceeds work units, but it must be
intentional and visible.
Repair: reduce shard count or make the gate explicitly accept/flag empty shards according to policy. Do not add dummy tests merely to hide the symptom.
6. Failure mode: imbalance dominates the critical path
Suppose shard durations are 2, 2, 3, and 17 minutes. Parallel execution finishes near the slowest shard, not near the average. Queue time may be low, so runner capacity is not the problem. The partition function put expensive work together.
Preserve per-test/per-shard timing evidence, then rebalance deterministically. A history-aware splitter is a performance optimization; the canonical test manifest remains the correctness source.
7. Failure mode: more concurrency makes the pipeline slower
Symptom: after increasing shards from 8 to 40, median execution time is similar but pending time jumps. Other projects may also wait longer.
Cause: requested runnable jobs exceed available runner/fleet capacity, image pulls/services contend for shared resources, or runner manager concurrency settings are lower than expected.
Evidence: job creation/start timestamps, pending durations, runner assignment, executor metrics if authorized, and host/fleet saturation—not just total pipeline duration.
Repair: first reduce unnecessary fan-out. Capacity changes belong to the runner/fleet layer and require explicit cost/security review.
8. Failure mode map: repair the owning layer
| Observed evidence | Owning layer | First repair |
|---|---|---|
| Expected jobs absent from graph | YAML compilation/matrix/rules | Fix matrix/rules; do not touch runners |
| Jobs exist but stay pending | Runner/fleet/queue | Bound fan-out or review authorized capacity |
| All jobs start, one has zero inputs | Partition/dataflow | Fix shard mapping/count policy |
| All shards pass, fan-in misses files | Artifact naming/transfer | Make evidence identities unique; verify downloads |
| Coverage complete but one shard takes 10× longer | Partition balancing | Rebalance workload, not credentials/runner trust |
9. Security and disruption guardrails
- Do not print tokens or full environments to distinguish shards; print non-secret shard metadata only.
- Do not put untrusted fork code onto privileged runners to gain more capacity.
- Do not disable TLS or broaden network access because many shards fail to fetch a dependency.
- Do not create unbounded autoscaling as a troubleshooting shortcut.
- Do not delete all caches/artifacts before preserving first-failure evidence.
- Do not make a deployment or package-publish job parallel unless the side-effect model is explicitly designed for it.
10. Intentionally broken scenario: “green shards, incomplete suite”
Start from the Chapter 17 four-shard lab. Modify the partition
script so it selects lines with
((NR-1) % total) == index while leaving
index one-based. One remainder is never selected and
another shard can become empty. Preserve the pipeline and job IDs.
Diagnosis should show: compiled graph correct → jobs assigned/running correctly → wrong effective partition arithmetic → shard manifests do not union to canonical manifest → fan-in fails. Repair only the zero-based conversion. Do not change runner capacity, artifact retention, or pipeline rules.
11. Retry and rerun policy
If one shard failed because its deterministic test failed, retrying may be useful only after preserving the original trace and deciding whether the failure is flaky. If the partition script or matrix configuration is wrong, retrying the same job cannot change the compiled configuration; create a corrected pipeline from a reviewed commit. If only fan-in logic is wrong and all shard artifacts remain valid, avoid rebuilding release-like evidence unnecessarily.
Knowledge check
A pipeline has 40 compiled shards but only four running. What evidence should you inspect first?
Pending/start timestamps and runner capacity/assignment. The graph exists; the bottleneck is queue/fleet capacity.
Why can a parallel artifact collision make a green pipeline unauditable?
Multiple successful producers can overwrite the same downloaded file, so the fan-in evidence no longer identifies every shard.
Where do duplicate matrix-value job collisions occur?
During configuration expansion/compilation, before runners execute anything.
What proves that a zero-test shard is harmless?
An explicit policy plus an independent union check showing every canonical test is still assigned exactly once.
Why is blind retry a poor response to a bad partition rule?
The same compiled/script logic will repeat the same defect and may obscure the first-failure evidence.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. Parallel-job limits, matrix expressions, runner concurrency controls, job-activity limits, and artifact-transfer behavior are version-sensitive. Re-check the deployed GitLab and GitLab Runner versions before using production capacity numbers or beta expression features.
-
CI/CD YAML syntax reference
— current
parallel,parallel:matrix,needs, and artifact semantics. - Control how jobs run — sharding patterns, matrix jobs, and selecting parallelized dependencies.
-
Matrix expressions
— current compile-time
$[[ matrix.IDENTIFIER ]]behavior introduced in GitLab 18.6. -
Predefined variables
—
CI_NODE_INDEX,CI_NODE_TOTAL, pipeline/job IDs, source SHA, and timing evidence. -
GitLab Runner advanced configuration
—
concurrent, per-runnerlimit, andrequest_concurrency. - Runner fleet scaling — capacity planning, executor behavior, and the difference between requested fan-out and available workers.
Current assumptions used in this chapter: all
mandatory examples use Free-tier CI/CD syntax and synthetic data.
Numeric parallel accepts 1–200;
parallel:matrix accepts at most 200 permutations.
Matrix expressions are current but version-sensitive and are
presented as an optional refinement, not a prerequisite. No cloud
account, autoscaling fleet, protected environment, or administrator
setting is required.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.