Chapter 18Lesson 04~170 minutes

Concurrency, resource_group, interruptible Jobs, Retry Policies, Timeouts, and Duplicate-Pipeline Control: Diagnostics, Failure Modes, Security, and Performance

Diagnose canceled side effects, duplicate releases, resource-group contention, misleading timeouts, and duplicate branch/MR pipelines without erasing first-failure evidence.

DiagnosticsCancellationRetry hazardsWaiting for resourceDuplicates

Learning objectives

  • Preserve pipeline/job IDs, first-failure traces, resource-group state, and external-side-effect evidence before retrying or canceling anything.
  • Diagnose a canceled job that partially changed external state as a reconciliation problem rather than blindly restarting it.
  • Diagnose duplicate releases caused by retry, over-broad or under-broad resource-group keys, timeout masking, and duplicate pipelines.
  • Separate configuration/rules defects from queue/resource locking, runner/executor, shell/tool/network, and external-system defects.
  • Apply the least destructive repair and rerun only the smallest safe scope.

1. Evidence-first diagnostic sequence

Concurrency failures are easy to “fix” by canceling, retrying, or deleting the visible state. That destroys the timeline needed to understand the race. Preserve evidence first, then move through layers in order:

  1. Record pipeline IDs, job IDs, source/ref/CI_COMMIT_SHA, timestamps, and first-failure/cancellation traces.
  2. Inspect compiled configuration and workflow/job-rule decisions.
  3. Inspect job graph and whether the job is pending, running, canceled, retrying, or waiting for a resource.
  4. Confirm runner/executor/image/toolchain only if the job reached execution.
  5. Inspect script/tool/network evidence.
  6. Inspect artifacts/reports/caches without deleting them.
  7. Inspect the external or simulated target by stable identity.
  8. Apply the smallest causal correction and rerun only the smallest safe scope.

2. Failure: cancellation leaves an external side effect

Symptom: an old pipeline job is canceled, but the environment shows the old artifact active.

Wrong reaction: immediately retry the canceled job or redeploy the newest pipeline without reading the target.

Diagnosis: confirm the old job was interruptible: true; identify the external operation; query its current deployment/version/digest; compare it with the old and new pipeline SHAs. The root cause is not “GitLab failed to cancel”; the contract allowed a mutable side effect in a cancelable job.

Repair: make side-effect jobs non-interruptible by default, add idempotent create-or-update/reconciliation semantics, and move cancellation to pre-deployment compute jobs.

3. Failure: retry duplicates a release or publication

Symptom: attempt 1 times out after sending a publish request; attempt 2 tries to create the same version and receives “already exists,” or worse, creates a second mutable object.

Preserve attempt 1’s job ID and timeout evidence. Query the release/package system by immutable version/idempotency key. If the intended object already exists with the expected digest, mark the recovery as reconciliation rather than another publish. Retry should be bounded to failures known to occur before the side effect or supported by the target’s idempotency contract.

4. Failure: resource-group key is too broad or too narrow

Observation Likely cause Repair
Unrelated regional deployments wait behind each other Key too broad Split by validated region/resource identity
Two jobs mutate same environment concurrently Key too narrow or includes per-pipeline value Use one stable key for the real shared target
Job waits forever while unrelated orchestration also needs key Dependency/lock cycle Move lock to owning orchestration boundary; inspect process mode
New deployments queue behind obviously obsolete work Ordering policy mismatch Consider safe newer-first mode only after proving idempotency

5. “Waiting for resource” is a queue state, not automatically a defect

When a job is waiting for a resource, first identify the job currently using the resource and its status. Under ordered process modes, a created job in an older pipeline can be the next queue owner even if another job looks ready. Current GitLab documentation recommends inspecting current/upcoming resource-group jobs and pipeline state before canceling anything.

With child/downstream pipelines, be especially careful with oldest_first: a parent waiting for a child that needs the same resource as a later parent job can form a deadlock. Preserve the whole pipeline relationship before changing process mode or rerunning jobs.

6. Failure: timeout hides the actual layer

A job timeout can be caused by slow script work, waiting on a network call, a runner maximum that is lower than expected, or an external API that never returns. Record CI_JOB_TIMEOUT, project/runner policy if available, last trace timestamp, and external request identity. If the remote target can be queried, reconcile it before rerun. Do not “fix” the symptom by multiplying timeout values without measuring where time is spent.

7. Failure: duplicate branch and MR pipelines waste capacity

Evidence pattern: the same commit produces one pipeline with CI_PIPELINE_SOURCE=push and another with merge_request_event. Job-level rules may still create expensive subsets in both.

Repair: switch pipeline creation at workflow:rules using CI_OPEN_MERGE_REQUESTS and a push guard. Preserve the duplicate pipeline IDs as proof that the repair removed creation rather than merely canceling later.

8. Intentionally broken example: unsafe retry + interruptible side effect

# BROKEN ON PURPOSE
deploy:broken:
  stage: deploy
  image: alpine:3.22
  interruptible: true
  retry: 2
  timeout: 20 seconds
  script:
    - echo "pretend to create release $CI_COMMIT_SHA"
    - sleep 30
  resource_group: "release-$CI_COMMIT_SHA"

This configuration stacks several defects:

  • The job can be canceled after “creating” its side effect.
  • Retry is broad and can repeat the side effect.
  • The timeout creates an unknown outcome.
  • The resource-group key includes the SHA, so two revisions do not serialize against one shared release namespace.

A repair uses a stable release-target key, non-interruptible behavior for the publication boundary, narrow retry only before publication or against an idempotent API, and a reconciliation query after any unknown outcome.

9. Layer-by-layer failure matrix

Symptom Layer to inspect first Do not jump to
Two pipelines exist for one MR push workflow/pipeline creation Runner scaling or resource_group
Safe test job from old pipeline still running auto-cancel + interruptible policy Manual kill of unrelated jobs
Deploy job says waiting for resource resource-group owner/upcoming queue Increasing runner count
Retry succeeded after timeout attempt evidence + external reconciliation Declaring success from green status alone
Job times out before artifacts upload job/script/after-script timeout budget Deleting artifacts/caches
Two deployments overlap resource key identity Broad PAT or privileged runner changes

10. Security and disruption guardrails

  • Never print CI_JOB_TOKEN, PATs, runner tokens, OIDC tokens, or full environment dumps while diagnosing retries.
  • Do not put untrusted MR code on privileged runners to reduce waiting time.
  • Do not change a shared resource-group process mode with a broad PAT from a teaching pipeline.
  • Do not disable TLS to “fix” transient network failures.
  • Do not mark production deployment jobs interruptible merely to save compute.
  • Do not retry releases/packages/deployments until the target has been reconciled by stable identity.
  • Do not delete and recreate external resources simply to make a blocked pipeline move.

11. Performance diagnosis before tuning

Measure how much time is spent in runner pending state versus resource-group waiting versus actual execution. If resource waiting dominates, more runners cannot help. If runner pending dominates, a narrower correct resource key will not help. If retries dominate, classify failure reasons and fix the upstream instability. If duplicate pipelines dominate, prevent creation. Optimization follows the measured bottleneck and must preserve the side-effect contract.

12. Recovery rule: reconcile before retrying side effects

The recovery question is: “what state exists now, and does it match the intended identity?” For a fake lab, that may be a receipt keyed by environment + artifact digest. For production, it could be a deployment record, package version/digest, migration version, or cloud resource tag. If the intended state exists, continue from that state. If it does not, apply one bounded idempotent correction. Preserve both the original failure and the recovery evidence.

Knowledge check

A canceled deployment job changed the target before cancellation. What is the next action?

Why can release-$CI_COMMIT_SHA be an unsafe resource-group key?

Does adding more runners fix Waiting for resource?

What evidence proves duplicate pipeline creation?

Why preserve the first failed retry attempt?

Next lesson

Checkpoint lab

Run two competing pipelines, prove serialization/idempotency intent, recover from a transient attempt, and retain an auditable evidence packet.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Concurrency, auto-cancel, retry failure reasons, retry-count variables, and timeout behavior are version-sensitive. Re-check the GitLab and Runner versions used by production pipelines before applying the exact examples.

  • Resource groups — serialization, process modes, downstream-pipeline locking, waiting-for-resource diagnostics, and deadlock guidance.
  • CI/CD YAML syntax reference — resource_group, interruptible, retry, timeout, and workflow:auto_cancel.
  • workflow keyword — pipeline creation, duplicate branch/MR prevention, and CI_OPEN_MERGE_REQUESTS patterns.
  • Predefined variables — pipeline/job identity, job timeout, and current retry-attempt metadata.
  • Configure runners — runner maximum job timeout and script/after-script timeout controls.
  • Resource Groups API — reading/updating process mode for an existing resource group.

Current assumptions used in this chapter: mandatory examples use Free-tier CI/CD features and synthetic data. A resource group serializes one resource at a time. Current process modes are unordered (default), oldest_first, newest_first, and newest_ready_first; newest-first modes require idempotent jobs. workflow:auto_cancel:on_new_commit currently supports conservative (default), interruptible, and none. Job retry allows 0–2 retries; retry:exit_codes is generally available. Current GitLab 19.3 docs expose CI_JOB_RETRY_COUNT; older deployments need a different lab signal. Job-level timeout can override the project default but remains bounded by runner maximum timeout.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.