Checkpoint Lab — Concurrency, resource_group, interruptible Jobs, Retry Policies, Timeouts, and Duplicate-Pipeline Control
Run two competing pipelines against one simulated environment, prove serialization and idempotent intent, inject a transient retry/interruption condition, and produce an auditable evidence packet.
Learning objectives
- Predict how two competing pipelines will be canceled, queued, retried, and serialized before executing them.
- Prove that only one fake deployment job owns the shared resource at a time and that safe stale work can be interrupted.
- Recover from a deliberately transient failure without repeating the simulated side effect.
- Demonstrate an idempotency-key/reconciliation model in a local faithful simulation that requires no production service.
- Produce an evidence packet that ties source SHA, pipeline/job attempts, resource-group behavior, and final reconciled state together.
1. Checkpoint scenario and success criteria
You own a disposable “training environment” that must receive at most one fake deployment at a time. Two commits are pushed close together. Safe validation work should be cancelable when stale; the deployment boundary must not be canceled blindly. A synthetic transient check must retry exactly once, and the evidence packet must prove what happened without relying on the final pipeline color.
The checkpoint is complete when you can prove: (1) which pipelines/jobs competed, (2) which stale work was canceled, (3) that deployment jobs using one resource key did not overlap, (4) that a transient failure was retried in a bounded way, (5) that duplicate branch/MR pipelines are prevented, and (6) that the side-effect model is idempotent/reconcilable in the local faithful simulation.
2. Assumptions and preflight
| Item | Checkpoint assumption |
|---|---|
| GitLab tier | Mandatory path uses Free features |
| GitLab behavior | Verified against current docs on 2026-09-12; retry-count variable assumes 19.3+ |
| Runner | Any authorized runner; record actual Runner/executor if visible |
| Image |
alpine:3.22; production should pin verified
immutable image identity where practical
|
| Branch | glci/ch18-checkpoint |
| MR | Optional but recommended for duplicate-pipeline proof |
| Credentials | None |
| External systems | None; side effect is synthetic plus a local idempotency simulation |
3. Predict at least four state changes before execution
- Commit B creates a newer pipeline; safe interruptible work in pipeline A may be canceled, while a started non-interruptible fake deployment is preserved.
-
Both fake deployment jobs compile with the same
resource_groupkey, so only one can own the resource at a time. - The transient verification job fails with exit code 75 on retry count 0 and succeeds on retry count 1, producing two distinct job-attempt records.
- After the MR-aware workflow rule is present, an MR push creates the MR pipeline while the redundant push branch pipeline is suppressed.
- The local idempotency simulation records one state row even when the same operation key is applied twice.
Write these predictions in your checkpoint notes before pushing. Do not rewrite them after observing the results.
4. Create the synthetic project files
git switch -c glci/ch18-checkpoint
mkdir -p ci
cat > ci/idempotency-demo.sh <<'EOF'
#!/bin/sh
set -eu
state_dir="$1"; key="$2"; payload="$3"
mkdir -p "$state_dir"
state="$state_dir/state.tsv"
if [ -f "$state" ] && awk -F '\t' -v k="$key" '$1==k {found=1} END{exit !found}' "$state"; then
printf 'already-applied key=%s\n' "$key"
exit 0
fi
printf '%s\t%s\n' "$key" "$payload" >> "$state"
printf 'applied key=%s payload=%s\n' "$key" "$payload"
EOF
chmod +x ci/idempotency-demo.sh
The helper is not the GitLab resource lock. It is a local faithful model of the external API property you need in addition to GitLab serialization.
5. Exact checkpoint pipeline
stages: [build, verify, deploy, evidence]
workflow:
auto_cancel:
on_new_commit: interruptible
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- if: '$CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS && $CI_PIPELINE_SOURCE == "push"'
when: never
- if: '$CI_COMMIT_BRANCH'
build:safe:
stage: build
image: alpine:3.22
interruptible: true
script:
- printf 'source=%s ref=%s sha=%s pipeline=%s job=%s start=%s\n' "$CI_PIPELINE_SOURCE" "$CI_COMMIT_REF_NAME" "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID" "$(date -u +%FT%TZ)"
- sleep 40
- echo "build complete"
verify:transient:
stage: verify
image: alpine:3.22
interruptible: true
retry:
max: 1
exit_codes: 75
script:
- printf 'pipeline=%s job=%s retry=%s sha=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "${CI_JOB_RETRY_COUNT:-unsupported}" "$CI_COMMIT_SHA"
- |
if [ "${CI_JOB_RETRY_COUNT:-unsupported}" = "0" ]; then
echo "synthetic transient failure before side effect"
exit 75
fi
- echo "transient check recovered"
deploy:fake:
stage: deploy
image: alpine:3.22
interruptible: false
resource_group: "ch18-training-environment"
script:
- mkdir -p evidence
- printf 'resource=ch18-training-environment\npipeline=%s\njob=%s\nsha=%s\nstart=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" "$(date -u +%FT%TZ)" | tee evidence/deploy.txt
- sleep 30
- printf 'finish=%s\n' "$(date -u +%FT%TZ)" | tee -a evidence/deploy.txt
artifacts:
when: always
expire_in: 1 week
paths: [evidence/deploy.txt]
evidence:identity:
stage: evidence
image: alpine:3.22
script:
- printf 'source=%s ref=%s sha=%s pipeline=%s job=%s\n' "$CI_PIPELINE_SOURCE" "$CI_COMMIT_REF_NAME" "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID"
If CI_JOB_RETRY_COUNT is not available on your GitLab
version, do not force this exact retry job. Replace it with the
local retry simulation in the next section and mark the platform
limitation in the evidence packet.
6. Execute two competing pipelines deliberately
- Commit and push the checkpoint pipeline as commit A.
-
While
build:safeis still running, make a harmless documentation change and push commit B. - Record both pipeline IDs immediately.
- Observe whether the stale interruptible build/verify work in pipeline A is canceled when B arrives.
-
Allow both pipelines that reach
deploy:faketo contend for the same resource key. Record which job runs and which waits. -
Download or inspect each deployment receipt. The later
deployment’s
startmust not precede the previous owner’sfinish.
If auto-cancel prevents pipeline A from reaching deploy, that is
valid evidence too. To demonstrate pure serialization separately,
create two manual/otherwise non-redundant disposable pipelines that
both reach deployment, or temporarily set
on_new_commit: none in a separate lab commit and record
that assumption. Do not weaken a production policy just to produce a
screenshot.
7. Prove retry behavior without hiding attempt 1
For GitLab 19.3+, the transient job should first fail with retry count 0 and exit 75, then GitLab processes one retry with retry count 1. Record both job attempt IDs/statuses and the exact source SHA. The deployment stage should only execute after verification succeeds, so the synthetic transient failure occurs before the side-effect boundary.
If the retry fails for a deterministic reason other than the synthetic exit code, stop and diagnose that first; do not keep rerunning until green.
8. Prove the external-side-effect contract locally
tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT
key='project-demo:training-env:artifact-sha256-demo'
./ci/idempotency-demo.sh "$tmp" "$key" 'desired-v1'
./ci/idempotency-demo.sh "$tmp" "$key" 'desired-v1'
printf 'rows=%s\n' "$(wc -l < "$tmp/state.tsv")"
cat "$tmp/state.tsv"
test "$(wc -l < "$tmp/state.tsv")" -eq 1
This verifies a key invariant: a repeated request with the same operation identity does not duplicate the intended state. In a real target, replace the file with a supported idempotency key or read-before/create-or-update API and verify the artifact/deployment digest.
9. Deliberate failure: make the resource key too narrow
On a disposable branch, temporarily change:
resource_group: "ch18-training-environment-$CI_COMMIT_SHA"
Create two pipelines. The keys differ, so GitLab is free to run both fake deployments concurrently. Preserve the pipeline/job IDs and overlapping timestamps. This is the intentional first failure.
Repair by restoring the stable
ch18-training-environment key. Create a new pipeline
and prove non-overlap. Do not delete the broken pipeline; it
demonstrates that the key identifies the real resource, not merely a
job.
10. Optional timeout recovery exercise
Add a separate verify:timeout-demo job with a 30-second
job timeout and a 45-second sleep, marked
allow_failure: true. Preserve its timeout evidence.
Then answer: would this be safe around a real deployment API? Only
if the target offers a reliable cancellation/reconciliation
contract. Otherwise the correct recovery is to query the target
before retrying.
11. Required evidence packet
| Category | Required evidence |
|---|---|
| Source/config |
branch/MR context, CI_PIPELINE_SOURCE, ref,
CI_COMMIT_SHA, exact YAML commit
|
| Pipelines | both competing pipeline IDs and statuses |
| Jobs | build/verify/deploy job IDs, canceled/failed/retried/waiting/running states |
| Retry |
attempt IDs, CI_JOB_RETRY_COUNT if supported,
exit code/reason
|
| Resource lock | key, process-mode assumption, wait/ownership evidence, deploy timestamps |
| Artifacts | both deployment receipt artifacts and retention |
| Duplicate control | MR/push pipeline list showing intended source only after workflow repair |
| Reconciliation | local idempotency key, one-row final state, limitations note |
| Runtime | Runner/executor/image identity if visible; unknowns explicitly marked |
12. Verification checklist
- Every recorded pipeline/job can be tied to one exact SHA.
- Only safe compute jobs are interruptible.
- The fake side-effect job is non-interruptible and uses one stable resource-group key.
- Two jobs with that key never overlap in observed execution time.
- The transient failure is retried no more than configured and the first failure remains visible.
- The MR workflow does not create a redundant push pipeline for the same branch when the MR pipeline exists.
- The local side-effect simulation applies the same idempotency key twice but stores one state row.
- No token, real secret, cloud resource, production environment, release, or package was touched.
13. Cleanup / rollback
- Restore the correct stable resource-group key if you ran the deliberate broken branch.
- Close the disposable MR and delete only the lab branches/projects you created.
- Let synthetic one-week artifacts expire or remove them from the disposable project only.
- Do not cancel/delete unrelated pipelines, change shared runner settings, or mutate resource-group process modes as cleanup.
-
The local temporary reconciliation directory is removed
automatically by
trap.
14. What Chapter 18 adds to the production operating model
Chapter 17 proved that parallel work can scale safely when every shard is bounded and auditable. Chapter 18 adds the controls needed when concurrent pipelines approach mutable state: redundant pipeline creation is suppressed early, stale safe work is interruptible, one logical resource is serialized, retries are classified and bounded, timeouts preserve an evidence budget, and every uncertain side effect is reconciled by stable identity before another attempt.
Chapter 19 moves from these concurrency controls to GitLab
environments and deployment records: environment tiers, URLs,
deployment history, and operational traceability. The
resource_group key introduced here will become one part
of a larger deployment identity rather than the whole deployment
model.
Knowledge check
What is the strongest proof that the two fake deployments were serialized?
They use the same resource-group key and their recorded execution intervals do not overlap.
Why does the checkpoint inject retry failure before the deploy stage?
It demonstrates retry mechanics without risking repetition of the simulated side effect.
What deliberate defect makes the resource lock ineffective?
Appending CI_COMMIT_SHA to the key gives competing revisions different locks even though they target the same environment.
What should happen after an unknown-outcome timeout around a real deployment?
Reconcile/query the target by stable identity before deciding whether another attempt is needed.
What concept bridges to Chapter 19?
A serialized side effect must also become a traceable GitLab environment/deployment record tied to exact source and artifact identity.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. Concurrency, auto-cancel, retry failure reasons, retry-count variables, and timeout behavior are version-sensitive. Re-check the GitLab and Runner versions used by production pipelines before applying the exact examples.
- Resource groups — serialization, process modes, downstream-pipeline locking, waiting-for-resource diagnostics, and deadlock guidance.
-
CI/CD YAML syntax reference
—
resource_group,interruptible,retry,timeout, andworkflow:auto_cancel. -
workflow keyword
— pipeline creation, duplicate branch/MR prevention, and
CI_OPEN_MERGE_REQUESTSpatterns. - Predefined variables — pipeline/job identity, job timeout, and current retry-attempt metadata.
- Configure runners — runner maximum job timeout and script/after-script timeout controls.
- Resource Groups API — reading/updating process mode for an existing resource group.
Current assumptions used in this chapter: mandatory
examples use Free-tier CI/CD features and synthetic data. A resource
group serializes one resource at a time. Current process modes are
unordered (default), oldest_first,
newest_first, and newest_ready_first;
newest-first modes require idempotent jobs.
workflow:auto_cancel:on_new_commit currently supports
conservative (default), interruptible, and
none. Job retry allows 0–2 retries;
retry:exit_codes is generally available. Current GitLab
19.3 docs expose CI_JOB_RETRY_COUNT; older deployments
need a different lab signal. Job-level timeout can
override the project default but remains bounded by runner maximum
timeout.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.