Concurrency, resource_group, interruptible Jobs, Retry Policies, Timeouts, and Duplicate-Pipeline Control: Guided Hands-On Workflow and Core Operations
Run a disposable two-pipeline lab that demonstrates interruptible work, resource-group serialization, targeted transient retries, timeout evidence, and source-aware duplicate suppression.
Learning objectives
- Create two overlapping disposable pipelines and capture source/ref/SHA plus pipeline/job IDs before changing policy.
- Mark only safe build/test work interruptible and observe auto-cancel behavior without making the fake deployment interruptible.
- Serialize a fake deployment with a stable resource-group key and prove non-overlap from timestamps.
- Inject a transient exit-code failure, use bounded retry, and distinguish the original failed attempt from the successful retry.
- Prevent duplicate branch/MR pipelines with workflow rules and record the reason each pipeline exists or is suppressed.
1. Guided lab: observe overlap before adding controls
This lab uses one disposable GitLab project/branch family, synthetic sleeps, fake deployment receipts, and no credentials. You will intentionally create two close-together pipelines, capture their identities, then add controls one layer at a time. The objective is not “make the pipeline green”; it is to explain exactly why a job was canceled, retried, timed out, or serialized.
2. Preflight and assumptions
| Item | Mandatory path |
|---|---|
| GitLab tier | Free-compatible syntax only |
| Runner | Any authorized runner able to execute Alpine shell jobs |
| Image |
alpine:3.22; record resolved identity if
visible
|
| Branch | glci/ch18-concurrency |
| Secrets | None |
| Admin settings | None; process-mode API changes are optional only |
| Current-version feature |
CI_JOB_RETRY_COUNT example assumes GitLab
19.3+; fallback is documented
|
3. Baseline pipeline: let the overlap be visible
stages: [build, verify, deploy]
workflow:
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- if: '$CI_COMMIT_BRANCH'
build:slow:
stage: build
image: alpine:3.22
script:
- printf 'source=%s ref=%s sha=%s pipeline=%s job=%s start=%s\n' "$CI_PIPELINE_SOURCE" "$CI_COMMIT_REF_NAME" "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID" "$(date -u +%FT%TZ)"
- sleep 45
- printf 'finish=%s\n' "$(date -u +%FT%TZ)"
verify:quick:
stage: verify
image: alpine:3.22
script:
- sleep 5
- echo "verified $CI_COMMIT_SHA"
deploy:fake:
stage: deploy
image: alpine:3.22
script:
- mkdir -p evidence
- printf 'pipeline=%s job=%s sha=%s start=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" "$(date -u +%FT%TZ)" | tee evidence/deploy.txt
- sleep 25
- printf 'finish=%s\n' "$(date -u +%FT%TZ)" | tee -a evidence/deploy.txt
artifacts:
when: always
expire_in: 1 week
paths: [evidence/deploy.txt]
Push commit A, then make a harmless text change and push commit B while A is still running. Preserve both pipeline IDs and the three job IDs from each. At this point, concurrent fake deploy jobs are possible because no resource lock exists.
4. Mark only safe stale work interruptible
Change build:slow and verify:quick to
interruptible: true, keep
deploy:fake non-interruptible, and add current explicit
auto-cancel behavior:
workflow:
auto_cancel:
on_new_commit: interruptible
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- if: '$CI_COMMIT_BRANCH'
build:slow:
interruptible: true
# ...
verify:quick:
interruptible: true
# ...
deploy:fake:
interruptible: false
# ...
Again push two commits close together. If the old safe jobs are still pending/running, the newer commit can cancel those interruptible jobs. If the non-interruptible deployment has already started, it is not canceled by this mode. Record canceled and surviving job IDs rather than only the final pipeline color.
5. Serialize the fake environment
Add one resource key to the side-effect job:
deploy:fake:
stage: deploy
image: alpine:3.22
interruptible: false
resource_group: "ch18-training-environment"
script:
- mkdir -p evidence
- printf 'pipeline=%s job=%s sha=%s start=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" "$(date -u +%FT%TZ)" | tee evidence/deploy.txt
- sleep 25
- printf 'finish=%s\n' "$(date -u +%FT%TZ)" | tee -a evidence/deploy.txt
artifacts:
when: always
paths: [evidence/deploy.txt]
Create two pipelines that both reach deployment. One deployment can be waiting for resource while the other owns the key. Compare the first job’s finish timestamp with the second job’s start timestamp; non-overlap is your direct serialization evidence.
Do not widen the key to the whole project unless
every mutable target truly conflicts. A key such as
environment/$TARGET is often safer and faster when the
target value is validated and bounded.
6. Inject a classified transient failure and observe retry identity
Current GitLab 19.3 exposes CI_JOB_RETRY_COUNT. Use it
only in this synthetic lab to create a deterministic first-attempt
failure:
verify:transient:
stage: verify
image: alpine:3.22
retry:
max: 1
exit_codes: 75
script:
- printf 'pipeline=%s job=%s retry=%s sha=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "${CI_JOB_RETRY_COUNT:-unsupported}" "$CI_COMMIT_SHA"
- |
if [ "${CI_JOB_RETRY_COUNT:-unsupported}" = "0" ]; then
echo "synthetic temporary failure before any side effect"
exit 75
fi
- echo "second attempt succeeds without external mutation"
Preserve the first failed job ID and trace, then the retried attempt’s job ID/count. If your GitLab is older than 19.3 and the variable is not exposed, replace this job with a local shell simulation; do not invent a retry counter.
7. Compare timeout evidence without touching external systems
verify:timeout-demo:
stage: verify
image: alpine:3.22
timeout: 30 seconds
allow_failure: true
script:
- printf 'job=%s timeout=%s start=%s\n' "$CI_JOB_ID" "$CI_JOB_TIMEOUT" "$(date -u +%FT%TZ)"
- sleep 45
- echo "this line should not be reached if the timeout is enforced"
Run this only in the disposable project. Record the configured
timeout, trace termination reason, and timestamps.
allow_failure: true keeps the teaching pipeline moving;
it does not redefine the timeout as success. In production, do not
hide a real deployment timeout behind allow_failure.
8. Suppress duplicate branch/MR pipelines at workflow creation
Once you create a disposable merge request from the lab branch, the broad baseline workflow can produce both branch and MR pipelines. Replace it with the documented switch pattern:
workflow:
auto_cancel:
on_new_commit: interruptible
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- if: '$CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS && $CI_PIPELINE_SOURCE == "push"'
when: never
- if: '$CI_COMMIT_BRANCH'
Push another commit. The evidence should show one intended MR
pipeline rather than both MR and push branch pipelines. This is a
pipeline-creation fix; adding job-level rules alone
would still pay the cost of creating the duplicate pipeline.
9. Local faithful simulation of an idempotent side effect
GitLab runner workspaces are intentionally not a reliable shared external database, so do not pretend a workspace file proves cross-pipeline idempotency. Instead, model the external contract locally in a temporary directory:
tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT
apply_once() {
state="$1/state.tsv"; key="$2"; payload="$3"
mkdir -p "$1"
if [ -f "$state" ] && awk -F '\t' -v k="$key" '$1==k {found=1} END{exit !found}' "$state"; then
printf 'already-applied key=%s\n' "$key"
return 0
fi
printf '%s\t%s\n' "$key" "$payload" >> "$state"
printf 'applied key=%s\n' "$key"
}
key='project-123:env-training:artifact-sha256-demo'
apply_once "$tmp" "$key" 'release-v1'
apply_once "$tmp" "$key" 'release-v1'
cat "$tmp/state.tsv"
The second call becomes a no-op because the operation identity is stable. A real deployment system should provide an equivalent create-or-update/idempotency/read-back mechanism; a resource group alone cannot replace it.
10. Evidence to retain from the workflow
| Evidence | What to capture |
|---|---|
| Pipeline identity | source, ref, SHA, both overlapping pipeline IDs |
| Interruption | old canceled job ID, new job ID, auto-cancel mode |
| Serialization | resource-group key, waiting/running statuses, deploy start/finish timestamps |
| Retry | first failed job ID, exit code/reason, retry count, successful attempt ID |
| Timeout | timeout value, job ID, failure/allow-failure presentation, timestamps |
| Duplicate suppression | MR-open state and pipeline list/source after workflow repair |
| Reconciliation | local idempotency key and exactly one state row |
11. Challenge: choose the correct layer
Your fake deployment never overlaps, but every push to an MR branch
still creates two test pipelines. Which layer should you change?
Pipeline creation, using source-aware
workflow:rules. Making test jobs share a resource group
would serialize duplicated work but would not remove the duplicate
pipeline.
Second challenge: the deployment API returns a timeout after 20 seconds, but a later read shows the desired artifact digest already active. Do you retry? Not automatically. Reconcile by stable deployment identity first; if the intended state exists, record success/recovery rather than repeat the side effect.
12. Cleanup
- Close/delete the disposable MR if you created one.
-
Delete the
glci/ch18-concurrencybranch only after recording the pipeline/job IDs you need. - Let one-week synthetic artifacts expire or delete only those lab artifacts through the disposable project UI.
- Do not change shared runner settings or resource-group process mode as part of cleanup; the mandatory lab did not require those mutations.
- Local
mktempstate is removed by the trap.
Knowledge check
Why does the lab keep deploy:fake non-interruptible?
It represents a side-effect boundary where cancellation could create unknown external state; safe compute jobs are the appropriate cancellation target.
What proves resource-group serialization?
The same resource-group key plus job status/timestamps showing the second side-effect job starts only after the first releases it.
Why is CI_JOB_RETRY_COUNT not a universal assumption?
It is current GitLab 19.3 behavior; older deployments do not expose it, so the lab must version-check or simulate.
Why is allow_failure used only for the timeout teaching job?
It lets the disposable pipeline continue while preserving a visible timeout failure; it is not a production recovery mechanism.
What is the correct fix for duplicate branch and MR pipelines?
Source-aware workflow rules that prevent redundant pipeline creation.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. Concurrency, auto-cancel, retry failure reasons, retry-count variables, and timeout behavior are version-sensitive. Re-check the GitLab and Runner versions used by production pipelines before applying the exact examples.
- Resource groups — serialization, process modes, downstream-pipeline locking, waiting-for-resource diagnostics, and deadlock guidance.
-
CI/CD YAML syntax reference
—
resource_group,interruptible,retry,timeout, andworkflow:auto_cancel. -
workflow keyword
— pipeline creation, duplicate branch/MR prevention, and
CI_OPEN_MERGE_REQUESTSpatterns. - Predefined variables — pipeline/job identity, job timeout, and current retry-attempt metadata.
- Configure runners — runner maximum job timeout and script/after-script timeout controls.
- Resource Groups API — reading/updating process mode for an existing resource group.
Current assumptions used in this chapter: mandatory
examples use Free-tier CI/CD features and synthetic data. A resource
group serializes one resource at a time. Current process modes are
unordered (default), oldest_first,
newest_first, and newest_ready_first;
newest-first modes require idempotent jobs.
workflow:auto_cancel:on_new_commit currently supports
conservative (default), interruptible, and
none. Job retry allows 0–2 retries;
retry:exit_codes is generally available. Current GitLab
19.3 docs expose CI_JOB_RETRY_COUNT; older deployments
need a different lab signal. Job-level timeout can
override the project default but remains bounded by runner maximum
timeout.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.