Concurrency, resource_group, interruptible Jobs, Retry Policies, Timeouts, and Duplicate-Pipeline Control: Diagnostics, Failure Modes, Security, and Performance
Diagnose canceled side effects, duplicate releases, resource-group contention, misleading timeouts, and duplicate branch/MR pipelines without erasing first-failure evidence.
Learning objectives
- Preserve pipeline/job IDs, first-failure traces, resource-group state, and external-side-effect evidence before retrying or canceling anything.
- Diagnose a canceled job that partially changed external state as a reconciliation problem rather than blindly restarting it.
- Diagnose duplicate releases caused by retry, over-broad or under-broad resource-group keys, timeout masking, and duplicate pipelines.
- Separate configuration/rules defects from queue/resource locking, runner/executor, shell/tool/network, and external-system defects.
- Apply the least destructive repair and rerun only the smallest safe scope.
1. Evidence-first diagnostic sequence
Concurrency failures are easy to “fix” by canceling, retrying, or deleting the visible state. That destroys the timeline needed to understand the race. Preserve evidence first, then move through layers in order:
-
Record pipeline IDs, job IDs,
source/ref/
CI_COMMIT_SHA, timestamps, and first-failure/cancellation traces. -
Inspect compiled configuration and
workflow/job-rule decisions. - Inspect job graph and whether the job is pending, running, canceled, retrying, or waiting for a resource.
- Confirm runner/executor/image/toolchain only if the job reached execution.
- Inspect script/tool/network evidence.
- Inspect artifacts/reports/caches without deleting them.
- Inspect the external or simulated target by stable identity.
- Apply the smallest causal correction and rerun only the smallest safe scope.
2. Failure: cancellation leaves an external side effect
Symptom: an old pipeline job is canceled, but the environment shows the old artifact active.
Wrong reaction: immediately retry the canceled job or redeploy the newest pipeline without reading the target.
Diagnosis: confirm the old job was
interruptible: true; identify the external operation;
query its current deployment/version/digest; compare it with the old
and new pipeline SHAs. The root cause is not “GitLab failed to
cancel”; the contract allowed a mutable side effect in a cancelable
job.
Repair: make side-effect jobs non-interruptible by default, add idempotent create-or-update/reconciliation semantics, and move cancellation to pre-deployment compute jobs.
3. Failure: retry duplicates a release or publication
Symptom: attempt 1 times out after sending a publish request; attempt 2 tries to create the same version and receives “already exists,” or worse, creates a second mutable object.
Preserve attempt 1’s job ID and timeout evidence. Query the release/package system by immutable version/idempotency key. If the intended object already exists with the expected digest, mark the recovery as reconciliation rather than another publish. Retry should be bounded to failures known to occur before the side effect or supported by the target’s idempotency contract.
4. Failure: resource-group key is too broad or too narrow
| Observation | Likely cause | Repair |
|---|---|---|
| Unrelated regional deployments wait behind each other | Key too broad | Split by validated region/resource identity |
| Two jobs mutate same environment concurrently | Key too narrow or includes per-pipeline value | Use one stable key for the real shared target |
| Job waits forever while unrelated orchestration also needs key | Dependency/lock cycle | Move lock to owning orchestration boundary; inspect process mode |
| New deployments queue behind obviously obsolete work | Ordering policy mismatch | Consider safe newer-first mode only after proving idempotency |
5. “Waiting for resource” is a queue state, not automatically a defect
When a job is waiting for a resource, first identify the job
currently using the resource and its status. Under ordered process
modes, a created job in an older pipeline can be the
next queue owner even if another job looks ready. Current GitLab
documentation recommends inspecting current/upcoming resource-group
jobs and pipeline state before canceling anything.
With child/downstream pipelines, be especially careful with
oldest_first: a parent waiting for a child that needs
the same resource as a later parent job can form a deadlock.
Preserve the whole pipeline relationship before changing process
mode or rerunning jobs.
6. Failure: timeout hides the actual layer
A job timeout can be caused by slow script work, waiting on a
network call, a runner maximum that is lower than expected, or an
external API that never returns. Record CI_JOB_TIMEOUT,
project/runner policy if available, last trace timestamp, and
external request identity. If the remote target can be queried,
reconcile it before rerun. Do not “fix” the symptom by multiplying
timeout values without measuring where time is spent.
7. Failure: duplicate branch and MR pipelines waste capacity
Evidence pattern: the same commit produces one
pipeline with CI_PIPELINE_SOURCE=push and another with
merge_request_event. Job-level rules may still create
expensive subsets in both.
Repair: switch pipeline creation at
workflow:rules using
CI_OPEN_MERGE_REQUESTS and a push guard.
Preserve the duplicate pipeline IDs as proof that the repair removed
creation rather than merely canceling later.
8. Intentionally broken example: unsafe retry + interruptible side effect
# BROKEN ON PURPOSE
deploy:broken:
stage: deploy
image: alpine:3.22
interruptible: true
retry: 2
timeout: 20 seconds
script:
- echo "pretend to create release $CI_COMMIT_SHA"
- sleep 30
resource_group: "release-$CI_COMMIT_SHA"
This configuration stacks several defects:
- The job can be canceled after “creating” its side effect.
- Retry is broad and can repeat the side effect.
- The timeout creates an unknown outcome.
- The resource-group key includes the SHA, so two revisions do not serialize against one shared release namespace.
A repair uses a stable release-target key, non-interruptible behavior for the publication boundary, narrow retry only before publication or against an idempotent API, and a reconciliation query after any unknown outcome.
9. Layer-by-layer failure matrix
| Symptom | Layer to inspect first | Do not jump to |
|---|---|---|
| Two pipelines exist for one MR push | workflow/pipeline creation | Runner scaling or resource_group |
| Safe test job from old pipeline still running | auto-cancel + interruptible policy | Manual kill of unrelated jobs |
| Deploy job says waiting for resource | resource-group owner/upcoming queue | Increasing runner count |
| Retry succeeded after timeout | attempt evidence + external reconciliation | Declaring success from green status alone |
| Job times out before artifacts upload | job/script/after-script timeout budget | Deleting artifacts/caches |
| Two deployments overlap | resource key identity | Broad PAT or privileged runner changes |
10. Security and disruption guardrails
-
Never print
CI_JOB_TOKEN, PATs, runner tokens, OIDC tokens, or full environment dumps while diagnosing retries. - Do not put untrusted MR code on privileged runners to reduce waiting time.
- Do not change a shared resource-group process mode with a broad PAT from a teaching pipeline.
- Do not disable TLS to “fix” transient network failures.
- Do not mark production deployment jobs interruptible merely to save compute.
- Do not retry releases/packages/deployments until the target has been reconciled by stable identity.
- Do not delete and recreate external resources simply to make a blocked pipeline move.
11. Performance diagnosis before tuning
Measure how much time is spent in runner pending state versus resource-group waiting versus actual execution. If resource waiting dominates, more runners cannot help. If runner pending dominates, a narrower correct resource key will not help. If retries dominate, classify failure reasons and fix the upstream instability. If duplicate pipelines dominate, prevent creation. Optimization follows the measured bottleneck and must preserve the side-effect contract.
12. Recovery rule: reconcile before retrying side effects
The recovery question is: “what state exists now, and does it match the intended identity?” For a fake lab, that may be a receipt keyed by environment + artifact digest. For production, it could be a deployment record, package version/digest, migration version, or cloud resource tag. If the intended state exists, continue from that state. If it does not, apply one bounded idempotent correction. Preserve both the original failure and the recovery evidence.
Knowledge check
A canceled deployment job changed the target before cancellation. What is the next action?
Query/reconcile the target by stable identity before any retry; cancellation alone does not roll back external state.
Why can release-$CI_COMMIT_SHA be an unsafe resource-group key?
Different revisions receive different locks, so operations against one shared release target can overlap.
Does adding more runners fix Waiting for resource?
No. Resource-group waiting is a serialization/ordering state, not runner-capacity scarcity.
What evidence proves duplicate pipeline creation?
Two pipeline IDs for the same revision/context with sources such as push and merge_request_event.
Why preserve the first failed retry attempt?
It contains the original failure timing/reason and may be the only evidence of whether a side effect began before the successful retry.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. Concurrency, auto-cancel, retry failure reasons, retry-count variables, and timeout behavior are version-sensitive. Re-check the GitLab and Runner versions used by production pipelines before applying the exact examples.
- Resource groups — serialization, process modes, downstream-pipeline locking, waiting-for-resource diagnostics, and deadlock guidance.
-
CI/CD YAML syntax reference
—
resource_group,interruptible,retry,timeout, andworkflow:auto_cancel. -
workflow keyword
— pipeline creation, duplicate branch/MR prevention, and
CI_OPEN_MERGE_REQUESTSpatterns. - Predefined variables — pipeline/job identity, job timeout, and current retry-attempt metadata.
- Configure runners — runner maximum job timeout and script/after-script timeout controls.
- Resource Groups API — reading/updating process mode for an existing resource group.
Current assumptions used in this chapter: mandatory
examples use Free-tier CI/CD features and synthetic data. A resource
group serializes one resource at a time. Current process modes are
unordered (default), oldest_first,
newest_first, and newest_ready_first;
newest-first modes require idempotent jobs.
workflow:auto_cancel:on_new_commit currently supports
conservative (default), interruptible, and
none. Job retry allows 0–2 retries;
retry:exit_codes is generally available. Current GitLab
19.3 docs expose CI_JOB_RETRY_COUNT; older deployments
need a different lab signal. Job-level timeout can
override the project default but remains bounded by runner maximum
timeout.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.