Chapter 18Lesson 01~155 minutes

Concurrency, resource_group, interruptible Jobs, Retry Policies, Timeouts, and Duplicate-Pipeline Control: Concepts, Architecture, and Mental Model

Build a concurrency mental model that separates pipeline duplication, safe cancellation, serialized resources, retries, timeouts, idempotency, and external reconciliation.

Concurrencyresource_groupInterruptibleRetryTimeout

Learning objectives

  • Explain why concurrency control is a side-effect correctness problem, not just a runner-capacity problem.
  • Distinguish pipeline suppression, job interruption, resource serialization, retry policy, timeout policy, and idempotent external reconciliation.
  • Trace competing pipeline/job IDs through cancellation state, resource-group queueing, retry attempts, execution timeout, and final side-effect evidence.
  • Explain the four current resource-group process modes and the safety assumptions behind newest-first modes.
  • Explain why retries and cancellations are safe only when the affected work is restartable or idempotent and first-failure evidence is preserved.

1. The practical problem: concurrency can repeat or interleave side effects

Chapter 17 deliberately increased throughput by running independent shards concurrently. Chapter 18 starts where that success becomes dangerous: two pipelines can now reach the same release, environment, device, namespace, package name, migration, or other mutable target at almost the same time.

The first instinct is often to “add retry” or “cancel old jobs.” Neither is a complete control. A canceled job may already have changed an external system. A retry may publish the same release twice. A timeout may terminate the runner while the external API continues processing. A branch and merge-request pipeline may duplicate the same validation workload before any job even starts.

Chapter invariant: every concurrency-sensitive side effect needs a stable identity, a serialization/idempotency strategy, bounded retry/timeout behavior, preserved first-failure evidence, and an explicit final reconciliation step.

2. Mental model: competing work → policy → one authorized side effect → reconciliation

First identify competing pipelines and jobs by pipeline ID, job ID, source, ref, and SHA. Pipeline-creation policy can suppress redundant pipelines before they consume runners. For pipelines that do exist, safe compute jobs can be marked interruptible. Jobs that touch one shared mutable resource can queue behind a resource_group. Retries and timeouts then operate on individual attempts, but they do not prove whether an external side effect happened. The final step is always reconciliation against the intended external state.

Concurrency-control causality
            flowchart TD
            A[Pipeline source/ref/SHA] --> B[workflow rules and auto-cancel]
            B --> C[Compiled jobs]
            C --> D[interruptible safe work]
            C --> E[resource_group queue]
            D --> F[Runner execution]
            E --> G[One side-effect owner]
            F --> H[Retry/timeout result]
            G --> H
            H --> I[External or simulated state]
            I --> J[Reconcile intended identity]
            J --> K[Evidence: IDs timestamps status receipt]
          

The important separation is that workflow controls pipeline creation, interruptible controls safe cancellation of running/pending work, resource_group controls ownership of a named resource, and retry/timeout controls an attempt. None of them independently guarantees the external system is correct.

3. State inventory before changing policy

State layer Evidence to record Why it matters
Source/revision CI_PIPELINE_SOURCE, ref, CI_COMMIT_SHA Explains why this pipeline exists and what revision owns the intent.
Compiled configuration workflow result, job inclusion, resource-group key Proves policy was compiled as intended.
Pipeline/job pipeline ID, job ID, status, queued/start/end time Separates cancellation, waiting, retry, and execution.
Interruption auto-cancel mode, interruptible, canceled job ID Shows whether stale work was intentionally stopped.
Retry failure reason/exit code, retry count, old/new job IDs Prevents a successful retry from hiding the first failure.
Timeout job timeout, runner ceiling, failure reason Explains whether GitLab stopped the attempt.
Resource lock resource-group key and process mode Defines which jobs must not overlap and their ordering policy.
External state idempotency key, release/deploy identity, reconciliation result Proves whether the side effect happened exactly as intended.

4. resource_group: serialize the resource, not the whole pipeline

A resource_group gives GitLab a mutex-like queue for one logical resource. Two jobs using the same key cannot run at the same time. Other jobs in their pipelines can still run concurrently, which is why the key should describe the smallest resource that truly requires serialization.

deploy:demo:
  stage: deploy
  resource_group: "training-environment"
  interruptible: false
  script:
    - printf 'pipeline=%s job=%s sha=%s start=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" "$(date -u +%FT%TZ)"
    - sleep 20
    - printf 'finish=%s\n' "$(date -u +%FT%TZ)"

This prevents overlap between deploy:demo jobs that use the same key. It does not stop test jobs from running in parallel and it does not guarantee that the deployment script itself is idempotent.

Process mode Ordering Primary use Safety condition
unordered Any ready job; default Ordering does not matter Side effect correct regardless of pipeline order
oldest_first Ascending pipeline ID Preserve chronological delivery Can trade latency for ordering; review child-pipeline deadlocks
newest_first Descending pipeline ID Prefer newest intended state Jobs must be idempotent
newest_ready_first Newest among jobs already ready Prefer fresh ready state with less starvation Jobs must be idempotent

5. interruptible and auto-cancel: cancel only restartable work

interruptible defaults to false. Mark a job true only when cancellation at any point is safe—for example, compilation or deterministic tests whose outputs are not externally committed. Deployment and publication jobs usually remain non-interruptible unless their external protocol was specifically designed for resumable cancellation.

workflow:
  auto_cancel:
    on_new_commit: interruptible

build:
  interruptible: true
  script: ./build.sh

deploy:
  interruptible: false
  resource_group: production
  script: ./deploy.sh

Current auto-cancel modes are conservative, interruptible, and none. Under the default conservative behavior, once a non-interruptible job has started, the pipeline is no longer eligible for redundant-pipeline cancellation. Under interruptible, only jobs explicitly marked interruptible are canceled.

6. Retry is attempt policy; idempotency is side-effect policy

A retry creates another processing attempt after failure. GitLab currently allows at most two retries. A broad retry: 2 repeats any supported failure type by default; a safer production design classifies the transient failures that are worth retrying.

integration-check:
  script: ./check-transient.sh
  retry:
    max: 1
    when:
      - runner_system_failure
      - runner_external_dependency_failure

If a job can create a release, deployment, payment-like operation, database mutation, or external ticket, retry cannot be the only safeguard. The side effect needs an idempotency key or create-or-update reconciliation based on a stable identity such as project + environment + artifact digest + intended version. Preserve the first failed job ID/trace before accepting the retry result.

Current-version note: GitLab 19.1 refined several runner failure reasons and deprecated older aggregate timeout categories. GitLab 19.3 documentation adds CI_JOB_RETRY_COUNT, which is useful for labs but must be version-checked before depending on it.

7. Timeouts bound waiting; they do not roll back external state

A timeout answers “how long may this job attempt occupy execution?” It does not answer “was the remote operation applied?” Current GitLab supports job-level timeout; project settings provide a default; a runner maximum timeout can impose a lower ceiling. Runner-side RUNNER_SCRIPT_TIMEOUT and RUNNER_AFTER_SCRIPT_TIMEOUT can reserve time for post-script evidence and artifact upload.

diagnostic-job:
  timeout: 20 minutes
  variables:
    RUNNER_SCRIPT_TIMEOUT: 15m
    RUNNER_AFTER_SCRIPT_TIMEOUT: 2m
  script:
    - ./bounded-operation.sh
  after_script:
    - ./write-nonsecret-diagnostic-summary.sh

If a deployment request times out locally, first query the deployment target with the intended identity. Blindly retrying can turn “unknown outcome” into duplicate side effects.

8. Duplicate-pipeline control belongs before runner execution

Chapter 8 separated pipeline creation from job inclusion, and Chapter 16 showed how the same push can otherwise produce both branch and merge-request pipelines. The most efficient concurrency control is often not creating redundant pipelines at all.

workflow:
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS && $CI_PIPELINE_SOURCE == "push"'
      when: never
    - if: '$CI_COMMIT_BRANCH' 

The explicit push guard is important in configurations that also use triggered/downstream pipelines: a triggered pipeline can have a branch variable but a different pipeline source and should not be accidentally suppressed.

9. Read-only inspection before modification

Before tuning concurrency, collect a small evidence table for two or more overlapping pipelines:

printf 'source=%s ref=%s sha=%s pipeline=%s job=%s timeout=%s\n'   "$CI_PIPELINE_SOURCE" "$CI_COMMIT_REF_NAME" "$CI_COMMIT_SHA"   "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_JOB_TIMEOUT"
printf 'retry_count=%s\n' "${CI_JOB_RETRY_COUNT:-version-not-exposed}"

From the GitLab UI or API, record which job is pending, running, canceled, failed, or waiting for resource. Do not print all environment variables and never expose tokens to discover this metadata.

10. What counts as proof

Claim Required evidence
Stale work was safely canceled Old/new pipeline IDs, job IDs, interruptible policy, canceled status, no committed side effect
Deployments did not overlap Same resource-group key plus non-overlapping start/finish timestamps
Retry was bounded and justified Original failure reason/exit code, retry count, new attempt/job identity, final result
Timeout was the cause Configured timeout, runner ceiling if relevant, trace/failure reason and timestamps
No duplicate branch+MR pipeline CI_PIPELINE_SOURCE, MR-open state/rule design, resulting pipeline list
External state is correct Stable side-effect identity plus read-back/reconciliation evidence

11. Common wrong mental models

  • “resource_group makes deployment idempotent.” It only prevents overlap for one key.
  • “interruptible means GitLab will clean up my external state.” Cancellation stops the job, not arbitrary remote systems.
  • “retry fixed it, so the first failure no longer matters.” Preserve the first attempt; the retry may have repeated side effects.
  • “a timeout means nothing happened.” The remote operation may have succeeded after the client stopped waiting.
  • “more resource-group scope is safer.” An unnecessarily broad key serializes unrelated work and increases latency without adding correctness.

12. Chapter trajectory

Lesson 2 turns the model into a disposable two-pipeline lab. Lesson 3 chooses modes and scope. Lesson 4 diagnoses causal failures without destructive resets. Lesson 5 combines the controls in a checkpoint and records final reconciliation evidence.

Knowledge check

What does a resource group guarantee?

Why are deployment jobs usually not marked interruptible?

How many retries can current GitLab retry policy request?

Why is a successful retry not enough evidence for a side effect?

Where should branch/MR duplicate suppression normally happen?

Next lesson

Guided hands-on workflow and core operations

Create overlapping pipelines, cancel only restartable work, serialize a fake deployment, inject a bounded retry, and suppress duplicate branch/MR pipelines.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Concurrency, auto-cancel, retry failure reasons, retry-count variables, and timeout behavior are version-sensitive. Re-check the GitLab and Runner versions used by production pipelines before applying the exact examples.

  • Resource groups — serialization, process modes, downstream-pipeline locking, waiting-for-resource diagnostics, and deadlock guidance.
  • CI/CD YAML syntax reference — resource_group, interruptible, retry, timeout, and workflow:auto_cancel.
  • workflow keyword — pipeline creation, duplicate branch/MR prevention, and CI_OPEN_MERGE_REQUESTS patterns.
  • Predefined variables — pipeline/job identity, job timeout, and current retry-attempt metadata.
  • Configure runners — runner maximum job timeout and script/after-script timeout controls.
  • Resource Groups API — reading/updating process mode for an existing resource group.

Current assumptions used in this chapter: mandatory examples use Free-tier CI/CD features and synthetic data. A resource group serializes one resource at a time. Current process modes are unordered (default), oldest_first, newest_first, and newest_ready_first; newest-first modes require idempotent jobs. workflow:auto_cancel:on_new_commit currently supports conservative (default), interruptible, and none. Job retry allows 0–2 retries; retry:exit_codes is generally available. Current GitLab 19.3 docs expose CI_JOB_RETRY_COUNT; older deployments need a different lab signal. Job-level timeout can override the project default but remains bounded by runner maximum timeout.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.