Chapter 18Lesson 02~190 minutes

Concurrency, resource_group, interruptible Jobs, Retry Policies, Timeouts, and Duplicate-Pipeline Control: Guided Hands-On Workflow and Core Operations

Run a disposable two-pipeline lab that demonstrates interruptible work, resource-group serialization, targeted transient retries, timeout evidence, and source-aware duplicate suppression.

Hands-onAuto-cancelSerializationRetry evidenceDuplicate control

Learning objectives

  • Create two overlapping disposable pipelines and capture source/ref/SHA plus pipeline/job IDs before changing policy.
  • Mark only safe build/test work interruptible and observe auto-cancel behavior without making the fake deployment interruptible.
  • Serialize a fake deployment with a stable resource-group key and prove non-overlap from timestamps.
  • Inject a transient exit-code failure, use bounded retry, and distinguish the original failed attempt from the successful retry.
  • Prevent duplicate branch/MR pipelines with workflow rules and record the reason each pipeline exists or is suppressed.

1. Guided lab: observe overlap before adding controls

This lab uses one disposable GitLab project/branch family, synthetic sleeps, fake deployment receipts, and no credentials. You will intentionally create two close-together pipelines, capture their identities, then add controls one layer at a time. The objective is not “make the pipeline green”; it is to explain exactly why a job was canceled, retried, timed out, or serialized.

Safety boundary: the fake deployment only sleeps and writes an artifact receipt. It does not call a real environment, package registry, cloud API, or production service.

2. Preflight and assumptions

Item Mandatory path
GitLab tier Free-compatible syntax only
Runner Any authorized runner able to execute Alpine shell jobs
Image alpine:3.22; record resolved identity if visible
Branch glci/ch18-concurrency
Secrets None
Admin settings None; process-mode API changes are optional only
Current-version feature CI_JOB_RETRY_COUNT example assumes GitLab 19.3+; fallback is documented

3. Baseline pipeline: let the overlap be visible

stages: [build, verify, deploy]

workflow:
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH'

build:slow:
  stage: build
  image: alpine:3.22
  script:
    - printf 'source=%s ref=%s sha=%s pipeline=%s job=%s start=%s\n'         "$CI_PIPELINE_SOURCE" "$CI_COMMIT_REF_NAME" "$CI_COMMIT_SHA"         "$CI_PIPELINE_ID" "$CI_JOB_ID" "$(date -u +%FT%TZ)"
    - sleep 45
    - printf 'finish=%s\n' "$(date -u +%FT%TZ)"

verify:quick:
  stage: verify
  image: alpine:3.22
  script:
    - sleep 5
    - echo "verified $CI_COMMIT_SHA"

deploy:fake:
  stage: deploy
  image: alpine:3.22
  script:
    - mkdir -p evidence
    - printf 'pipeline=%s job=%s sha=%s start=%s\n'         "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" "$(date -u +%FT%TZ)" | tee evidence/deploy.txt
    - sleep 25
    - printf 'finish=%s\n' "$(date -u +%FT%TZ)" | tee -a evidence/deploy.txt
  artifacts:
    when: always
    expire_in: 1 week
    paths: [evidence/deploy.txt]

Push commit A, then make a harmless text change and push commit B while A is still running. Preserve both pipeline IDs and the three job IDs from each. At this point, concurrent fake deploy jobs are possible because no resource lock exists.

4. Mark only safe stale work interruptible

Change build:slow and verify:quick to interruptible: true, keep deploy:fake non-interruptible, and add current explicit auto-cancel behavior:

workflow:
  auto_cancel:
    on_new_commit: interruptible
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH'

build:slow:
  interruptible: true
  # ...

verify:quick:
  interruptible: true
  # ...

deploy:fake:
  interruptible: false
  # ...

Again push two commits close together. If the old safe jobs are still pending/running, the newer commit can cancel those interruptible jobs. If the non-interruptible deployment has already started, it is not canceled by this mode. Record canceled and surviving job IDs rather than only the final pipeline color.

5. Serialize the fake environment

Add one resource key to the side-effect job:

deploy:fake:
  stage: deploy
  image: alpine:3.22
  interruptible: false
  resource_group: "ch18-training-environment"
  script:
    - mkdir -p evidence
    - printf 'pipeline=%s job=%s sha=%s start=%s\n'         "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" "$(date -u +%FT%TZ)" | tee evidence/deploy.txt
    - sleep 25
    - printf 'finish=%s\n' "$(date -u +%FT%TZ)" | tee -a evidence/deploy.txt
  artifacts:
    when: always
    paths: [evidence/deploy.txt]

Create two pipelines that both reach deployment. One deployment can be waiting for resource while the other owns the key. Compare the first job’s finish timestamp with the second job’s start timestamp; non-overlap is your direct serialization evidence.

Do not widen the key to the whole project unless every mutable target truly conflicts. A key such as environment/$TARGET is often safer and faster when the target value is validated and bounded.

6. Inject a classified transient failure and observe retry identity

Current GitLab 19.3 exposes CI_JOB_RETRY_COUNT. Use it only in this synthetic lab to create a deterministic first-attempt failure:

verify:transient:
  stage: verify
  image: alpine:3.22
  retry:
    max: 1
    exit_codes: 75
  script:
    - printf 'pipeline=%s job=%s retry=%s sha=%s\n'         "$CI_PIPELINE_ID" "$CI_JOB_ID" "${CI_JOB_RETRY_COUNT:-unsupported}" "$CI_COMMIT_SHA"
    - |
      if [ "${CI_JOB_RETRY_COUNT:-unsupported}" = "0" ]; then
        echo "synthetic temporary failure before any side effect"
        exit 75
      fi
    - echo "second attempt succeeds without external mutation"

Preserve the first failed job ID and trace, then the retried attempt’s job ID/count. If your GitLab is older than 19.3 and the variable is not exposed, replace this job with a local shell simulation; do not invent a retry counter.

7. Compare timeout evidence without touching external systems

verify:timeout-demo:
  stage: verify
  image: alpine:3.22
  timeout: 30 seconds
  allow_failure: true
  script:
    - printf 'job=%s timeout=%s start=%s\n' "$CI_JOB_ID" "$CI_JOB_TIMEOUT" "$(date -u +%FT%TZ)"
    - sleep 45
    - echo "this line should not be reached if the timeout is enforced"

Run this only in the disposable project. Record the configured timeout, trace termination reason, and timestamps. allow_failure: true keeps the teaching pipeline moving; it does not redefine the timeout as success. In production, do not hide a real deployment timeout behind allow_failure.

8. Suppress duplicate branch/MR pipelines at workflow creation

Once you create a disposable merge request from the lab branch, the broad baseline workflow can produce both branch and MR pipelines. Replace it with the documented switch pattern:

workflow:
  auto_cancel:
    on_new_commit: interruptible
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS && $CI_PIPELINE_SOURCE == "push"'
      when: never
    - if: '$CI_COMMIT_BRANCH' 

Push another commit. The evidence should show one intended MR pipeline rather than both MR and push branch pipelines. This is a pipeline-creation fix; adding job-level rules alone would still pay the cost of creating the duplicate pipeline.

9. Local faithful simulation of an idempotent side effect

GitLab runner workspaces are intentionally not a reliable shared external database, so do not pretend a workspace file proves cross-pipeline idempotency. Instead, model the external contract locally in a temporary directory:

tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT

apply_once() {
  state="$1/state.tsv"; key="$2"; payload="$3"
  mkdir -p "$1"
  if [ -f "$state" ] && awk -F '\t' -v k="$key" '$1==k {found=1} END{exit !found}' "$state"; then
    printf 'already-applied key=%s\n' "$key"
    return 0
  fi
  printf '%s\t%s\n' "$key" "$payload" >> "$state"
  printf 'applied key=%s\n' "$key"
}

key='project-123:env-training:artifact-sha256-demo'
apply_once "$tmp" "$key" 'release-v1'
apply_once "$tmp" "$key" 'release-v1'
cat "$tmp/state.tsv"

The second call becomes a no-op because the operation identity is stable. A real deployment system should provide an equivalent create-or-update/idempotency/read-back mechanism; a resource group alone cannot replace it.

10. Evidence to retain from the workflow

Evidence What to capture
Pipeline identity source, ref, SHA, both overlapping pipeline IDs
Interruption old canceled job ID, new job ID, auto-cancel mode
Serialization resource-group key, waiting/running statuses, deploy start/finish timestamps
Retry first failed job ID, exit code/reason, retry count, successful attempt ID
Timeout timeout value, job ID, failure/allow-failure presentation, timestamps
Duplicate suppression MR-open state and pipeline list/source after workflow repair
Reconciliation local idempotency key and exactly one state row

11. Challenge: choose the correct layer

Your fake deployment never overlaps, but every push to an MR branch still creates two test pipelines. Which layer should you change? Pipeline creation, using source-aware workflow:rules. Making test jobs share a resource group would serialize duplicated work but would not remove the duplicate pipeline.

Second challenge: the deployment API returns a timeout after 20 seconds, but a later read shows the desired artifact digest already active. Do you retry? Not automatically. Reconcile by stable deployment identity first; if the intended state exists, record success/recovery rather than repeat the side effect.

12. Cleanup

  1. Close/delete the disposable MR if you created one.
  2. Delete the glci/ch18-concurrency branch only after recording the pipeline/job IDs you need.
  3. Let one-week synthetic artifacts expire or delete only those lab artifacts through the disposable project UI.
  4. Do not change shared runner settings or resource-group process mode as part of cleanup; the mandatory lab did not require those mutations.
  5. Local mktemp state is removed by the trap.

Knowledge check

Why does the lab keep deploy:fake non-interruptible?

What proves resource-group serialization?

Why is CI_JOB_RETRY_COUNT not a universal assumption?

Why is allow_failure used only for the timeout teaching job?

What is the correct fix for duplicate branch and MR pipelines?

Next lesson

Configuration, design choices, and tradeoffs

Choose resource-group scope/process mode, interruption, retry, timeout, and workflow policy from the actual side-effect contract.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Concurrency, auto-cancel, retry failure reasons, retry-count variables, and timeout behavior are version-sensitive. Re-check the GitLab and Runner versions used by production pipelines before applying the exact examples.

  • Resource groups — serialization, process modes, downstream-pipeline locking, waiting-for-resource diagnostics, and deadlock guidance.
  • CI/CD YAML syntax reference — resource_group, interruptible, retry, timeout, and workflow:auto_cancel.
  • workflow keyword — pipeline creation, duplicate branch/MR prevention, and CI_OPEN_MERGE_REQUESTS patterns.
  • Predefined variables — pipeline/job identity, job timeout, and current retry-attempt metadata.
  • Configure runners — runner maximum job timeout and script/after-script timeout controls.
  • Resource Groups API — reading/updating process mode for an existing resource group.

Current assumptions used in this chapter: mandatory examples use Free-tier CI/CD features and synthetic data. A resource group serializes one resource at a time. Current process modes are unordered (default), oldest_first, newest_first, and newest_ready_first; newest-first modes require idempotent jobs. workflow:auto_cancel:on_new_commit currently supports conservative (default), interruptible, and none. Job retry allows 0–2 retries; retry:exit_codes is generally available. Current GitLab 19.3 docs expose CI_JOB_RETRY_COUNT; older deployments need a different lab signal. Job-level timeout can override the project default but remains bounded by runner maximum timeout.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.