Parallel Jobs, parallel:matrix, Test Sharding, Fan-Out/Fan-In, and High-Throughput Pipeline Design: Configuration, Design Choices, and Tradeoffs
Choose between fixed parallel counts, matrices, balancing strategies, fan-in designs, and runner capacity using current GitLab limits and explicit correctness tradeoffs.
Learning objectives
- Choose fixed shard counts or matrices according to the shape of the work rather than YAML convenience.
- Compare simple modulo sharding with history-aware balancing and explain when imbalance dominates the critical path.
- Explain the current matrix-expression model and its compile-time, string-only scope.
- Bound fan-out against runner capacity, cost, active-job limits, and evidence-management overhead.
- Design artifact names, paths, and fan-in dependencies so parallel jobs cannot overwrite one another.
1. Design from the workload, not from the keyword
The configuration question is not “should we use
parallel or parallel:matrix?” It is “what
independent hypotheses or work partitions exist, what evidence must
each produce, and what runner capacity can execute them
economically?” Numeric parallelism is usually a good fit for
homogeneous partitions of one suite. A matrix is a good fit when
named dimensions represent meaningful compatibility cases.
For every design comparison, keep
CI_PIPELINE_SOURCE and the exact
CI_COMMIT_SHA fixed or explicitly recorded. Otherwise a
faster graph on a different source event or revision is not a
controlled comparison.
2. Fixed count versus matrix
| Design | Best fit | Primary risk | Evidence requirement |
|---|---|---|---|
Numeric parallel |
Homogeneous test shards | Empty/imbalanced shards | Index/total + shard manifest |
parallel:matrix |
Explicit compatibility dimensions | Cardinality explosion / duplicate values | Matrix variables + unique job/evidence identity |
| Manual named jobs | Very small stable cases | Duplication and drift | Explicit job name + purpose |
| Dynamic generated jobs | Highly variable workloads | Complexity/trust of generated config | Generator input/version + compiled graph |
3. Simple modulo versus history-aware balancing
Modulo sharding is deterministic, transparent, and excellent for a teaching lab. It assumes work units have roughly similar cost. Real test suites often violate that assumption: one integration test can take longer than fifty unit tests. Then the slowest shard defines the test stage’s critical path.
A history-aware splitter can use prior timing metadata to balance estimated duration, but that metadata becomes another correctness/performance input. Version it or record its provenance. If timing data is stale or missing, fall back safely and keep coverage independent of the balancing heuristic.
4. Fail-fast expectations versus complete evidence
Parallelism changes how failures arrive. One shard can fail while others are still running. Canceling remaining work may reduce cost for a developer-feedback pipeline, but it also forfeits complete failure evidence. For release qualification, teams often prefer all independent shards to finish so the next run is not a sequence of hidden failures.
Choose explicitly: optimize for first-failure latency, or optimize for complete diagnostic evidence. Do not let runner cancellation behavior accidentally define policy.
5. Throughput is bounded by runner capacity
GitLab compiles jobs; GitLab Runner executes them. A runner
manager’s global concurrent setting bounds total
simultaneously handled jobs, while each registered runner can have
its own limit. Job-request concurrency also affects how
quickly available work is acquired. On hosted infrastructure you may
not control these values at all.
Therefore a matrix of 64 jobs on a four-slot fleet is mostly a queue. Estimate startup overhead, average runtime, runner slots, and artifact traffic before increasing fan-out. Scaling runners can be valid, but it is infrastructure work with cost/security implications—not a YAML repair.
6. Cardinality budget: calculate before merge
Current GitLab caps a parallel:matrix expansion at 200
permutations. Treat that as an upper safety bound, not a target. For
dimensions OS=3, RUNTIME=4,
DB=3, PROFILE=2, the product is 72 jobs
before any other pipeline jobs.
3 OS × 4 runtimes × 3 databases × 2 profiles = 72 jobs
72 jobs × 90 seconds average execution = 108 runner-minutes
With 8 effective runner slots, idealized lower bound ≈ 13.5 minutes before queue/startup variance
The idealized bound is not a promise. Image pulls, services, cache misses, uneven tests, and platform queueing all add variance.
7. Matrix expressions and 1:1 dependencies
Current GitLab matrix expressions can wire matching
producer/consumer matrix jobs using
$[[ matrix.IDENTIFIER ]]. They are compile-time,
string-only, and scoped to matrix identifiers. This avoids an
all-to-all dependency graph.
build:
parallel:
matrix:
- OS: [linux, alpine]
ARCH: [amd64, arm64]
script: ./build.sh "$OS" "$ARCH"
verify:
parallel:
matrix:
- OS: [linux, alpine]
ARCH: [amd64, arm64]
needs:
- job: build
parallel:
matrix:
- OS: ['$[[ matrix.OS ]]']
ARCH: ['$[[ matrix.ARCH ]]']
script: ./verify.sh "$OS" "$ARCH"
Because matrix expressions were introduced in 18.6 and remain a version-sensitive capability, verify your instance before depending on them. The portable mental model is still “consumer needs exactly the producer partition it consumes.”
8. Artifact design: identity before convenience
Use a shard/matrix identity in both artifact name and path. A fan-in job that needs a parallelized producer can receive artifacts from every instance; same-name files can overwrite one another. Also remember that typed test reports and deployable binaries are different evidence classes. Aggregating JUnit results does not make those XML files release artifacts.
9. Parallelism multiplies trust exposure
Every additional job is another execution of repository-controlled scripts with whatever identity, variables, runner access, network reach, and services that job receives. Do not multiply privileged jobs, broad credentials, or mutable deployment side effects merely because a matrix is convenient. Keep test matrices read-only where possible and use the narrowest credentials only on the jobs that need them.
10. Decision table: choose a design and predict state
| Scenario | Recommended design | Prerequisites | State to verify |
|---|---|---|---|
| 12 deterministic tests, similar cost | parallel: 4 + modulo partition |
Free; ≥1 runner | 4 compiled shards; complete union; queue/runtime |
| 2 runtimes × 2 modes | 4-job matrix | Free | matrix job names/variables; 4 unique artifacts |
| Producer/consumer same matrix | Matrix + 1:1 needs:parallel:matrix |
Verify deployed GitLab; matrix expressions if used | each consumer waits only for matching producer |
| 100 uneven integration tests | History-aware balanced shards | Timing metadata + stable splitter | metadata provenance; balanced duration; coverage |
| 200+ combinations | Reduce dimensions / staged strategy | Architecture decision | bounded graph and explicit omitted coverage |
11. Upgrade and rollback strategy
Change one dimension at a time. Keep the previous shard count or
matrix definition in version control. Compare the same source/test
manifest when benchmarking. A rollback should be a configuration
reversion, not deletion of pipeline evidence or a runner-fleet
emergency change. If a new matrix expression is unsupported, fall
back to explicit needs:parallel:matrix entries or a
simpler stage/fan-in design.
Knowledge check
When is numeric parallel preferable to a
matrix?
When the work is homogeneous and only needs deterministic partitioning, such as splitting one test suite into equivalent shards.
Why can 72 matrix jobs be slower than 12?
Runner slots, startup/image-pull overhead, services, artifact traffic, and queueing can dominate after useful parallelism is exhausted.
What do matrix expressions change?
They simplify compile-time 1:1 dependency wiring; they do not add runtime logic or runner capacity.
Why is historical balancing not automatically more correct than modulo?
It can improve duration balance, but coverage still needs an independent manifest proof and the timing metadata itself can be stale.
What is the safest rollback for an over-aggressive matrix change?
Revert the CI configuration to the prior bounded graph while preserving the failed pipeline/job evidence.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. Parallel-job limits, matrix expressions, runner concurrency controls, job-activity limits, and artifact-transfer behavior are version-sensitive. Re-check the deployed GitLab and GitLab Runner versions before using production capacity numbers or beta expression features.
-
CI/CD YAML syntax reference
— current
parallel,parallel:matrix,needs, and artifact semantics. - Control how jobs run — sharding patterns, matrix jobs, and selecting parallelized dependencies.
-
Matrix expressions
— current compile-time
$[[ matrix.IDENTIFIER ]]behavior introduced in GitLab 18.6. -
Predefined variables
—
CI_NODE_INDEX,CI_NODE_TOTAL, pipeline/job IDs, source SHA, and timing evidence. -
GitLab Runner advanced configuration
—
concurrent, per-runnerlimit, andrequest_concurrency. - Runner fleet scaling — capacity planning, executor behavior, and the difference between requested fan-out and available workers.
Current assumptions used in this chapter: all
mandatory examples use Free-tier CI/CD syntax and synthetic data.
Numeric parallel accepts 1–200;
parallel:matrix accepts at most 200 permutations.
Matrix expressions are current but version-sensitive and are
presented as an optional refinement, not a prerequisite. No cloud
account, autoscaling fleet, protected environment, or administrator
setting is required.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.