Chapter 17Lesson 01~155 minutes

Parallel Jobs, parallel:matrix, Test Sharding, Fan-Out/Fan-In, and High-Throughput Pipeline Design: Concepts, Architecture, and Mental Model

Learn why parallel execution is a partitioning and evidence problem, not merely a speed switch. Build a mental model from deterministic work partitions through runner queues to per-shard evidence and a completeness gate.

Parallel jobsShardingRunner queuesEvidenceFan-in

Learning objectives

  • Explain why parallelism is safe only when the work can be partitioned independently and the union of shard results is provably complete.
  • Distinguish numeric `parallel` expansion from `parallel:matrix` expansion and identify the evidence fields that make each shard reproducible.
  • Trace the lifecycle from compiled job expansion through the runner queue, shard execution, per-shard reports/artifacts, and a fan-in gate.
  • Explain current GitLab limits: 1–200 numeric parallel jobs, up to 200 matrix permutations, queueing when runner capacity is lower, and artifact-overwrite hazards.
  • Treat faster wall-clock time as a performance result, not as proof of test coverage or runner capacity health.

1. The practical problem: “more jobs” is not the same as “more throughput”

Chapter 16 established that pre-merge evidence is meaningful only when the exact ref and SHA under test are known. Chapter 17 keeps that identity fixed and asks a different question: how can one pipeline process more independent work without losing coverage, traceability, or control of runner demand?

Parallelism helps only when the work is partitionable. If four jobs all repeat the same tests, throughput does not improve. If four jobs each run a quarter of the suite but one quarter is empty, the pipeline can become fast and wrong. If 100 shards are requested but only four runner slots exist, most jobs simply wait in pending. The design unit is therefore not “number of jobs”; it is partition + capacity + evidence + fan-in proof.

Chapter invariant: a green parallel pipeline must let an auditor answer “which source SHA, which shard identity, which inputs, which tests, which runner/job, and which evidence contributed to the final gate?” without guessing.

2. Mental model: partition → expansion → queue → shards → fan-in

Start with a deterministic set of work—for example, twelve named tests. A partitioning rule maps every test to exactly one shard. GitLab compiles one job definition into several concrete jobs using parallel or parallel:matrix. Those concrete jobs enter the runner queue independently. Available runners execute some immediately; excess jobs remain pending. Each shard emits uniquely named evidence. A fan-in job checks that all required shards exist and that their union equals the original manifest.

High-throughput state model
            flowchart TD
              A[Test manifest / hypothesis] --> B[Partition rule]
              B --> C[parallel or matrix expansion]
              C --> D[Runnable jobs]
              D --> E[Runner queue]
              E --> F1[Shard 1]
              E --> F2[Shard 2]
              E --> F3[Shard N]
              F1 --> G[Unique reports / artifacts]
              F2 --> G
              F3 --> G
              G --> H[Fan-in completeness gate]
              H --> I[Coverage + timing evidence]
          

Notice the ownership boundaries. GitLab configuration compilation owns job expansion. Runner capacity owns how quickly expanded jobs leave the queue. The test script owns deterministic partitioning. Artifact/report storage owns retained shard evidence. The fan-in job owns completeness verification. None of those states can substitute for another.

3. State to record before changing concurrency

State Minimum evidence Why it matters
Source CI_PIPELINE_SOURCE, ref, CI_COMMIT_SHA Keeps serial and parallel comparisons on the same revision.
Compiled graph Concrete expanded job names and count Proves the requested fan-out really exists.
Shard identity CI_NODE_INDEX/CI_NODE_TOTAL or matrix values Lets evidence map back to one partition.
Runner/queue job ID, pending/start/finish timestamps, runner description/executor Separates queue delay from test runtime.
Evidence per-shard manifest/report/artifact name and digest Proves what each shard actually covered.
Fan-in expected shard set, observed shard set, union/duplicate/missing checks Prevents silent gaps.
External effects normally none in this chapter Parallel tests should stay side-effect bounded and disposable.

4. Numeric parallel: identical job definition, indexed instances

Current GitLab accepts a numeric parallel value from 1 through 200. A job with parallel: 4 becomes four concrete jobs. Each gets CI_NODE_INDEX and CI_NODE_TOTAL. The job name is also expanded, such as test 1/4 through test 4/4.

test:shard:
  stage: test
  parallel: 4
  script:
    - printf 'pipeline=%s sha=%s shard=%s/%s\n'         "$CI_PIPELINE_ID" "$CI_COMMIT_SHA" "$CI_NODE_INDEX" "$CI_NODE_TOTAL"
    - ./ci/run-shard.sh "$CI_NODE_INDEX" "$CI_NODE_TOTAL"

The predefined index is job identity, not coverage proof. Your script still has to map the deterministic test manifest to that index and prove the mapping.

5. parallel:matrix: fan-out by explicit dimensions

A matrix is better when each job represents a named combination such as operating system family, runtime version, or test profile. Current GitLab allows at most 200 permutations. Matrix identifiers become CI/CD variables in the generated jobs, and values are appended to job names.

test:matrix:
  stage: test
  parallel:
    matrix:
      - RUNTIME: ["3.12", "3.13"]
        PROFILE: [unit, integration]
  script:
    - printf 'runtime=%s profile=%s sha=%s\n' "$RUNTIME" "$PROFILE" "$CI_COMMIT_SHA"
    - ./ci/run-profile.sh "$RUNTIME" "$PROFILE"

This expands to four jobs. Matrix fan-out should correspond to a real compatibility hypothesis. Adding dimensions because YAML makes it easy creates cost without stronger evidence.

6. Current limits are correctness inputs, not trivia

Current behavior Operational consequence
Numeric parallel: 1–200 A requested count above the limit is a configuration problem, not a runner problem.
parallel:matrix: ≤200 permutations Multiply dimensions before committing; 5×5×10 already reaches 250 and is invalid.
Parallel jobs can exceed available runners Excess jobs remain pending; wall-clock gains flatten when capacity is saturated.
Active-job instance limits can reject pipeline creation A large fan-out can fail with job_activity_limit_exceeded.
Matrix values become part of job names Very long values can hit job-name limits; duplicate value combinations can overwrite generated jobs.

7. Runner queue is a separate performance state

Requested parallelism and achieved concurrency are different. A pipeline may contain 20 runnable shards while the runner fleet has capacity for only four. GitLab Runner exposes global concurrent, per-runner limit, and job-request concurrency controls; GitLab.com hosted runners have platform-managed capacity instead.

Measure at least two intervals: pending/queue time from job creation until start, and execution time from start until finish. If execution time falls but queue time rises more, the “optimization” made the user experience worse.

8. Per-shard evidence must be unique

A common failure is to let every shard upload report.json under the same artifact name. When a downstream job uses needs on a parallelized producer, GitLab can download artifacts from all parallel instances; if artifact files have the same name/path, later downloads can overwrite earlier ones. That destroys evidence even if every test passed.

artifacts:
  name: "shard-$CI_NODE_INDEX-of-$CI_NODE_TOTAL"
  when: always
  paths:
    - "evidence/shard-$CI_NODE_INDEX/"
  reports:
    junit: "evidence/shard-$CI_NODE_INDEX/junit.xml"

Unique names do not replace content validation. The fan-in step should also verify every expected shard manifest and test ID.

9. Fan-in is a completeness gate, not just another stage

A fan-in job should answer three questions independently: did every expected shard finish, did each shard publish the expected evidence, and does the union of shard manifests equal the source test manifest with no duplicates or omissions? Only then may the aggregate result claim full coverage.

Coverage proof
            flowchart TD
              M[Canonical test manifest] --> P[Partition function]
              P --> S1[Shard manifests]
              P --> S2[Shard reports]
              S1 --> G[Fan-in gate]
              S2 --> G
              G --> C1{All shard IDs present?}
              C1 --> C2{Union equals manifest?}
              C2 --> C3{No duplicate ownership?}
              C3 --> R[Complete evidence]
          

10. Matrix expressions: useful compile-time wiring, not runtime logic

GitLab introduced matrix expressions in 18.6. The $[[ matrix.IDENTIFIER ]] syntax can create 1:1 needs:parallel:matrix mappings without manually listing every combination. The expression is resolved when the pipeline is created, is string replacement only, and can reference matrix identifiers—not arbitrary CI/CD variables or inputs.

Because this feature is version-sensitive, this course treats it as an optional refinement. The core learning path works with ordinary parallel, explicit matrices, and conventional fan-in.

11. Read-only inspection first

Before editing YAML, inspect an existing pipeline and record: pipeline ID/source/ref/SHA; current job count and names; pending versus running jobs; runner/executor identity where visible; per-job start/finish timestamps; artifact/report names; and the final gate that claims completeness. If any field is absent, write that absence down instead of inventing it.

Do not change runner config.toml, autoscaling settings, project limits, or administrator settings just to make a demonstration faster. Chapter 17’s mandatory path changes only disposable repository configuration and synthetic data.

12. Misconceptions to remove

Misconception Correction
“parallel: 20 means 20 jobs run simultaneously.” It creates 20 job instances; actual concurrency depends on runner capacity and limits.
“A green shard means the suite is covered.” Only a coverage union check can prove the complete manifest was executed.
“Matrix dimensions are free.” They multiply job cardinality, queue pressure, artifact volume, and cost.
“The fan-in job can just depend on the parallel job name.” That may fetch all parallel artifacts; identical artifact names/paths can overwrite one another.
“More shards always reduce duration.” Small shards can become dominated by startup, queue, artifact, and scheduling overhead.

Knowledge check

What does parallel: 8 prove by itself?

Why record queue time separately from execution time?

What is the maximum current numeric parallel value and matrix permutation count?

Why is a shared report.json path dangerous across shards?

What must a fan-in coverage gate compare?

Next lesson

Guided hands-on workflow and core operations

Build a deterministic shard lab, a bounded matrix, unique evidence, queue-time measurements, and a fan-in coverage proof.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Parallel-job limits, matrix expressions, runner concurrency controls, job-activity limits, and artifact-transfer behavior are version-sensitive. Re-check the deployed GitLab and GitLab Runner versions before using production capacity numbers or beta expression features.

Current assumptions used in this chapter: all mandatory examples use Free-tier CI/CD syntax and synthetic data. Numeric parallel accepts 1–200; parallel:matrix accepts at most 200 permutations. Matrix expressions are current but version-sensitive and are presented as an optional refinement, not a prerequisite. No cloud account, autoscaling fleet, protected environment, or administrator setting is required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.