Chapter 18Lesson 03~150 minutes

Concurrency, resource_group, interruptible Jobs, Retry Policies, Timeouts, and Duplicate-Pipeline Control: Configuration, Design Choices, and Tradeoffs

Choose process modes, cancellation policy, retry scope, timeout layers, and resource-group granularity from correctness and evidence requirements rather than convenience.

Design tradeoffsProcess modesIdempotencyTimeout layersSafety

Learning objectives

  • Select resource-group granularity and process mode based on the actual external resource and ordering requirement.
  • Choose conservative, interruptible, or no auto-cancel behavior according to restartability and evidence needs.
  • Use retries only for classified transient failures and pair side effects with idempotency/reconciliation rather than assuming retries are harmless.
  • Place timeouts at the correct job/project/runner layer and reserve enough time for failure evidence and cleanup where required.
  • Balance stale-work cancellation and throughput against release/deployment safety and auditability.

1. Design from the mutable resource and failure contract

Concurrency controls should be selected from the thing that can be corrupted, not from whichever YAML keyword is familiar. Start with the external resource identity, ask whether operations commute, whether cancellation is safe, whether a retry can repeat the side effect, and what evidence proves the final state. Only then choose resource-group scope, process mode, interruption, retry, and timeout policy.

2. resource_group process-mode choices

Mode Strength Tradeoff Use when
unordered Serialization only No delivery ordering guarantee Any queued operation produces a valid state independent of order
oldest_first Serialize in pipeline order Can deploy stale work and increase latency Chronological delivery is required and dependencies cannot deadlock
newest_first Newest pipeline gets resource first Older jobs may be superseded/starved Desired state is latest and operation is idempotent
newest_ready_first Newest already-ready job first Still requires idempotency and careful semantics Prefer fresh ready work without letting not-yet-ready pipelines jump the queue

Changing process mode is API-managed state for an existing resource group, not a property you should casually mutate in a shared production lab. The mandatory course path teaches the modes read-only and keeps the default unless the learner owns an isolated project.

3. Resource-group key granularity: broad enough for correctness, narrow enough for throughput

A key such as production is correct only if every operation truly conflicts on the same production resource. If two regions are independent, production/$REGION can preserve concurrency—but only when $REGION is validated and maps one-to-one to the actual resource. A key that includes the commit SHA is usually too narrow because every pipeline gets a different lock and conflicting deployments can overlap.

Key Likely result Assessment
production One global production deployment at a time Safe but may be unnecessarily broad
production/$REGION One deployment per validated region Good if region is the real isolation boundary
production/$CI_COMMIT_SHA Different lock for every revision Usually unsafe for one shared target
$CI_ENVIRONMENT_NAME Serialize by environment identity Useful when environment names are deterministic and bounded

4. Cancel stale work versus preserve evidence

Work type Interruptible? Reason
Compile/lint/unit test Usually yes Restarting on a newer SHA is safe; no external commit
Report generation from immutable inputs Often yes Can be regenerated if canceled before upload
Release/package publication Usually no May create a durable external identity
Deployment/migration Usually no by default Cancellation can leave partial state
Read-only health check Often yes No mutation, but preserve evidence if used for a gate

workflow:auto_cancel:on_new_commit: conservative is the current default. Choose interruptible when you deliberately want stale safe jobs canceled while non-interruptible jobs continue. Use none where preserving every pipeline is more important than saving compute.

5. Retries versus idempotency and compensation

Retry policy should answer two independent questions: is the failure plausibly transient? and can repeating this operation be safe? The first chooses retry:when or retry:exit_codes; the second belongs to the application/external API contract.

Failure Retry? Required safeguard
Runner host reclaimed before script starts Often yes No external mutation occurred; preserve system-failure evidence
Temporary registry/network lookup failure Often yes Bound attempts; immutable dependency identity
Deterministic test assertion Usually no Fix code/test; retry can hide flakiness
Release create returned unknown outcome Not until reconciled Query by version/idempotency key first
Database migration partially applied Not blindly Migration tool must expose idempotent/resumable state

6. Timeout layers and evidence budget

A project timeout is a default policy. A job-level timeout specializes one job and can exceed the project default, but the runner maximum timeout is an upper bound. If you need failure artifacts or cleanup, set RUNNER_SCRIPT_TIMEOUT shorter than the job timeout and reserve enough time for after_script and uploads.

Do not use extremely short timeouts to simulate circuit breakers around remote deployments unless the deployment API has an explicit cancellation/reconciliation contract. A client timeout is often an unknown outcome, not a rollback.

7. Duplicate pipelines: allowlist creation rather than clean up later

Pipeline suppression is cheaper and easier to audit than canceling duplicate pipelines after jobs start. Prefer a positive workflow model that names accepted pipeline sources. When switching between branch and MR pipelines, use CI_OPEN_MERGE_REQUESTS together with CI_PIPELINE_SOURCE == "push" so API/trigger/downstream sources are not accidentally blocked.

workflow:
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS && $CI_PIPELINE_SOURCE == "push"'
      when: never
    - if: '$CI_COMMIT_BRANCH'
    - if: '$CI_COMMIT_TAG' 

Record CI_PIPELINE_SOURCE and CI_COMMIT_SHA in evidence jobs so reviewers can prove which rule path created the pipeline.

8. Resource groups with child/downstream pipelines

A resource group can guard a trigger job that starts a child/downstream deployment pipeline. When the resource must remain locked until the downstream work finishes, current GitLab guidance requires a waiting trigger strategy such as strategy: mirror. Review resource-group ordering carefully: oldest_first combined with parent/child pipelines that compete for the same key can deadlock if the parent waits on a child while an older parent job owns the next queue position.

Design principle: place the lock at the orchestration boundary that owns the entire side effect. Do not acquire the same resource independently in parent and child in a way that creates cyclic waiting.

9. Worked decision table

Scenario Choice Prerequisite/evidence
Fast MR validation superseded by new commits interruptible: true + auto-cancel interruptible Jobs have no durable side effects; preserve old/new pipeline IDs
One training environment shared by many pipelines One resource-group key, default unordered initially Start/finish timestamps prove non-overlap
Production desired-state deployment where newest wins Consider newest-ready-first Operation is idempotent; old desired states can be safely superseded
Transient runner/registry failure Retry max 1 for classified failure Original attempt retained; dependency identity immutable
Release publication No blind retry; serialize and reconcile version Release identity/version queried before create
Branch/MR duplicate workload workflow suppression Pipeline source and open-MR evidence

10. Performance and cost are downstream of correctness

Overly broad serialization raises wait time and increases stale-work pressure. Overly narrow serialization can corrupt shared state. Aggressive auto-cancel saves runner capacity but can discard evidence you intended to retain. Broad retry multiplies compute and can amplify outages. Tune these controls from measured queue time, retry rate, timeout frequency, and external reconciliation outcomes—not from anecdotal “pipelines feel slow.”

11. Review checklist for a production change

  • What exact external resource does the resource-group key represent?
  • Could two different keys still mutate the same target?
  • Which jobs are safe to cancel at any instruction boundary?
  • Which failure reasons are truly transient, and what is the retry maximum?
  • What stable idempotency/reconciliation key proves whether the side effect already happened?
  • What job/project/runner timeout actually applies?
  • Which pipeline sources are allowed to exist, and can a branch+MR duplicate still be created?
  • Are first failed/canceled attempts retained long enough for audit?

Knowledge check

When is newest_first unsafe?

Why is a resource-group key containing CI_COMMIT_SHA often wrong for one shared environment?

What is the difference between retry policy and idempotency?

What caps a job-level timeout?

Why should duplicate branch/MR pipelines be prevented with workflow rules?

Next lesson

Diagnostics, failure modes, security, and performance

Preserve first-failure evidence and diagnose canceled side effects, duplicate releases, lock contention, timeouts, and duplicate pipelines causally.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Concurrency, auto-cancel, retry failure reasons, retry-count variables, and timeout behavior are version-sensitive. Re-check the GitLab and Runner versions used by production pipelines before applying the exact examples.

  • Resource groups — serialization, process modes, downstream-pipeline locking, waiting-for-resource diagnostics, and deadlock guidance.
  • CI/CD YAML syntax reference — resource_group, interruptible, retry, timeout, and workflow:auto_cancel.
  • workflow keyword — pipeline creation, duplicate branch/MR prevention, and CI_OPEN_MERGE_REQUESTS patterns.
  • Predefined variables — pipeline/job identity, job timeout, and current retry-attempt metadata.
  • Configure runners — runner maximum job timeout and script/after-script timeout controls.
  • Resource Groups API — reading/updating process mode for an existing resource group.

Current assumptions used in this chapter: mandatory examples use Free-tier CI/CD features and synthetic data. A resource group serializes one resource at a time. Current process modes are unordered (default), oldest_first, newest_first, and newest_ready_first; newest-first modes require idempotent jobs. workflow:auto_cancel:on_new_commit currently supports conservative (default), interruptible, and none. Job retry allows 0–2 retries; retry:exit_codes is generally available. Current GitLab 19.3 docs expose CI_JOB_RETRY_COUNT; older deployments need a different lab signal. Job-level timeout can override the project default but remains bounded by runner maximum timeout.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.