Chapter 17Lesson 03~150 minutes

Parallel Jobs, parallel:matrix, Test Sharding, Fan-Out/Fan-In, and High-Throughput Pipeline Design: Configuration, Design Choices, and Tradeoffs

Choose between fixed parallel counts, matrices, balancing strategies, fan-in designs, and runner capacity using current GitLab limits and explicit correctness tradeoffs.

Design tradeoffsCapacityMatrix expressionsCostCorrectness

Learning objectives

  • Choose fixed shard counts or matrices according to the shape of the work rather than YAML convenience.
  • Compare simple modulo sharding with history-aware balancing and explain when imbalance dominates the critical path.
  • Explain the current matrix-expression model and its compile-time, string-only scope.
  • Bound fan-out against runner capacity, cost, active-job limits, and evidence-management overhead.
  • Design artifact names, paths, and fan-in dependencies so parallel jobs cannot overwrite one another.

1. Design from the workload, not from the keyword

The configuration question is not “should we use parallel or parallel:matrix?” It is “what independent hypotheses or work partitions exist, what evidence must each produce, and what runner capacity can execute them economically?” Numeric parallelism is usually a good fit for homogeneous partitions of one suite. A matrix is a good fit when named dimensions represent meaningful compatibility cases.

For every design comparison, keep CI_PIPELINE_SOURCE and the exact CI_COMMIT_SHA fixed or explicitly recorded. Otherwise a faster graph on a different source event or revision is not a controlled comparison.

2. Fixed count versus matrix

Design Best fit Primary risk Evidence requirement
Numeric parallel Homogeneous test shards Empty/imbalanced shards Index/total + shard manifest
parallel:matrix Explicit compatibility dimensions Cardinality explosion / duplicate values Matrix variables + unique job/evidence identity
Manual named jobs Very small stable cases Duplication and drift Explicit job name + purpose
Dynamic generated jobs Highly variable workloads Complexity/trust of generated config Generator input/version + compiled graph

3. Simple modulo versus history-aware balancing

Modulo sharding is deterministic, transparent, and excellent for a teaching lab. It assumes work units have roughly similar cost. Real test suites often violate that assumption: one integration test can take longer than fifty unit tests. Then the slowest shard defines the test stage’s critical path.

A history-aware splitter can use prior timing metadata to balance estimated duration, but that metadata becomes another correctness/performance input. Version it or record its provenance. If timing data is stale or missing, fall back safely and keep coverage independent of the balancing heuristic.

4. Fail-fast expectations versus complete evidence

Parallelism changes how failures arrive. One shard can fail while others are still running. Canceling remaining work may reduce cost for a developer-feedback pipeline, but it also forfeits complete failure evidence. For release qualification, teams often prefer all independent shards to finish so the next run is not a sequence of hidden failures.

Choose explicitly: optimize for first-failure latency, or optimize for complete diagnostic evidence. Do not let runner cancellation behavior accidentally define policy.

5. Throughput is bounded by runner capacity

GitLab compiles jobs; GitLab Runner executes them. A runner manager’s global concurrent setting bounds total simultaneously handled jobs, while each registered runner can have its own limit. Job-request concurrency also affects how quickly available work is acquired. On hosted infrastructure you may not control these values at all.

Therefore a matrix of 64 jobs on a four-slot fleet is mostly a queue. Estimate startup overhead, average runtime, runner slots, and artifact traffic before increasing fan-out. Scaling runners can be valid, but it is infrastructure work with cost/security implications—not a YAML repair.

6. Cardinality budget: calculate before merge

Current GitLab caps a parallel:matrix expansion at 200 permutations. Treat that as an upper safety bound, not a target. For dimensions OS=3, RUNTIME=4, DB=3, PROFILE=2, the product is 72 jobs before any other pipeline jobs.

3 OS × 4 runtimes × 3 databases × 2 profiles = 72 jobs
72 jobs × 90 seconds average execution = 108 runner-minutes
With 8 effective runner slots, idealized lower bound ≈ 13.5 minutes before queue/startup variance

The idealized bound is not a promise. Image pulls, services, cache misses, uneven tests, and platform queueing all add variance.

7. Matrix expressions and 1:1 dependencies

Current GitLab matrix expressions can wire matching producer/consumer matrix jobs using $[[ matrix.IDENTIFIER ]]. They are compile-time, string-only, and scoped to matrix identifiers. This avoids an all-to-all dependency graph.

build:
  parallel:
    matrix:
      - OS: [linux, alpine]
        ARCH: [amd64, arm64]
  script: ./build.sh "$OS" "$ARCH"

verify:
  parallel:
    matrix:
      - OS: [linux, alpine]
        ARCH: [amd64, arm64]
  needs:
    - job: build
      parallel:
        matrix:
          - OS: ['$[[ matrix.OS ]]']
            ARCH: ['$[[ matrix.ARCH ]]']
  script: ./verify.sh "$OS" "$ARCH"

Because matrix expressions were introduced in 18.6 and remain a version-sensitive capability, verify your instance before depending on them. The portable mental model is still “consumer needs exactly the producer partition it consumes.”

8. Artifact design: identity before convenience

Use a shard/matrix identity in both artifact name and path. A fan-in job that needs a parallelized producer can receive artifacts from every instance; same-name files can overwrite one another. Also remember that typed test reports and deployable binaries are different evidence classes. Aggregating JUnit results does not make those XML files release artifacts.

9. Parallelism multiplies trust exposure

Every additional job is another execution of repository-controlled scripts with whatever identity, variables, runner access, network reach, and services that job receives. Do not multiply privileged jobs, broad credentials, or mutable deployment side effects merely because a matrix is convenient. Keep test matrices read-only where possible and use the narrowest credentials only on the jobs that need them.

10. Decision table: choose a design and predict state

Scenario Recommended design Prerequisites State to verify
12 deterministic tests, similar cost parallel: 4 + modulo partition Free; ≥1 runner 4 compiled shards; complete union; queue/runtime
2 runtimes × 2 modes 4-job matrix Free matrix job names/variables; 4 unique artifacts
Producer/consumer same matrix Matrix + 1:1 needs:parallel:matrix Verify deployed GitLab; matrix expressions if used each consumer waits only for matching producer
100 uneven integration tests History-aware balanced shards Timing metadata + stable splitter metadata provenance; balanced duration; coverage
200+ combinations Reduce dimensions / staged strategy Architecture decision bounded graph and explicit omitted coverage

11. Upgrade and rollback strategy

Change one dimension at a time. Keep the previous shard count or matrix definition in version control. Compare the same source/test manifest when benchmarking. A rollback should be a configuration reversion, not deletion of pipeline evidence or a runner-fleet emergency change. If a new matrix expression is unsupported, fall back to explicit needs:parallel:matrix entries or a simpler stage/fan-in design.

Knowledge check

When is numeric parallel preferable to a matrix?

Why can 72 matrix jobs be slower than 12?

What do matrix expressions change?

Why is historical balancing not automatically more correct than modulo?

What is the safest rollback for an over-aggressive matrix change?

Next lesson

Diagnostics, failure modes, security, and performance

Preserve first-failure evidence and diagnose matrix explosion, artifact collisions, empty shards, imbalance, and runner saturation causally.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Parallel-job limits, matrix expressions, runner concurrency controls, job-activity limits, and artifact-transfer behavior are version-sensitive. Re-check the deployed GitLab and GitLab Runner versions before using production capacity numbers or beta expression features.

Current assumptions used in this chapter: all mandatory examples use Free-tier CI/CD syntax and synthetic data. Numeric parallel accepts 1–200; parallel:matrix accepts at most 200 permutations. Matrix expressions are current but version-sensitive and are presented as an optional refinement, not a prerequisite. No cloud account, autoscaling fleet, protected environment, or administrator setting is required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.