Chapter 18Lesson 05~205 minutes

Checkpoint Lab — Concurrency, resource_group, interruptible Jobs, Retry Policies, Timeouts, and Duplicate-Pipeline Control

Run two competing pipelines against one simulated environment, prove serialization and idempotent intent, inject a transient retry/interruption condition, and produce an auditable evidence packet.

Checkpoint labCompeting pipelinesReconciliationEvidence packetCleanup

Learning objectives

  • Predict how two competing pipelines will be canceled, queued, retried, and serialized before executing them.
  • Prove that only one fake deployment job owns the shared resource at a time and that safe stale work can be interrupted.
  • Recover from a deliberately transient failure without repeating the simulated side effect.
  • Demonstrate an idempotency-key/reconciliation model in a local faithful simulation that requires no production service.
  • Produce an evidence packet that ties source SHA, pipeline/job attempts, resource-group behavior, and final reconciled state together.

1. Checkpoint scenario and success criteria

You own a disposable “training environment” that must receive at most one fake deployment at a time. Two commits are pushed close together. Safe validation work should be cancelable when stale; the deployment boundary must not be canceled blindly. A synthetic transient check must retry exactly once, and the evidence packet must prove what happened without relying on the final pipeline color.

The checkpoint is complete when you can prove: (1) which pipelines/jobs competed, (2) which stale work was canceled, (3) that deployment jobs using one resource key did not overlap, (4) that a transient failure was retried in a bounded way, (5) that duplicate branch/MR pipelines are prevented, and (6) that the side-effect model is idempotent/reconcilable in the local faithful simulation.

2. Assumptions and preflight

Item Checkpoint assumption
GitLab tier Mandatory path uses Free features
GitLab behavior Verified against current docs on 2026-09-12; retry-count variable assumes 19.3+
Runner Any authorized runner; record actual Runner/executor if visible
Image alpine:3.22; production should pin verified immutable image identity where practical
Branch glci/ch18-checkpoint
MR Optional but recommended for duplicate-pipeline proof
Credentials None
External systems None; side effect is synthetic plus a local idempotency simulation

3. Predict at least four state changes before execution

  1. Commit B creates a newer pipeline; safe interruptible work in pipeline A may be canceled, while a started non-interruptible fake deployment is preserved.
  2. Both fake deployment jobs compile with the same resource_group key, so only one can own the resource at a time.
  3. The transient verification job fails with exit code 75 on retry count 0 and succeeds on retry count 1, producing two distinct job-attempt records.
  4. After the MR-aware workflow rule is present, an MR push creates the MR pipeline while the redundant push branch pipeline is suppressed.
  5. The local idempotency simulation records one state row even when the same operation key is applied twice.

Write these predictions in your checkpoint notes before pushing. Do not rewrite them after observing the results.

4. Create the synthetic project files

git switch -c glci/ch18-checkpoint
mkdir -p ci
cat > ci/idempotency-demo.sh <<'EOF'
#!/bin/sh
set -eu
state_dir="$1"; key="$2"; payload="$3"
mkdir -p "$state_dir"
state="$state_dir/state.tsv"
if [ -f "$state" ] && awk -F '\t' -v k="$key" '$1==k {found=1} END{exit !found}' "$state"; then
  printf 'already-applied key=%s\n' "$key"
  exit 0
fi
printf '%s\t%s\n' "$key" "$payload" >> "$state"
printf 'applied key=%s payload=%s\n' "$key" "$payload"
EOF
chmod +x ci/idempotency-demo.sh

The helper is not the GitLab resource lock. It is a local faithful model of the external API property you need in addition to GitLab serialization.

5. Exact checkpoint pipeline

stages: [build, verify, deploy, evidence]

workflow:
  auto_cancel:
    on_new_commit: interruptible
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS && $CI_PIPELINE_SOURCE == "push"'
      when: never
    - if: '$CI_COMMIT_BRANCH'

build:safe:
  stage: build
  image: alpine:3.22
  interruptible: true
  script:
    - printf 'source=%s ref=%s sha=%s pipeline=%s job=%s start=%s\n'         "$CI_PIPELINE_SOURCE" "$CI_COMMIT_REF_NAME" "$CI_COMMIT_SHA"         "$CI_PIPELINE_ID" "$CI_JOB_ID" "$(date -u +%FT%TZ)"
    - sleep 40
    - echo "build complete"

verify:transient:
  stage: verify
  image: alpine:3.22
  interruptible: true
  retry:
    max: 1
    exit_codes: 75
  script:
    - printf 'pipeline=%s job=%s retry=%s sha=%s\n'         "$CI_PIPELINE_ID" "$CI_JOB_ID" "${CI_JOB_RETRY_COUNT:-unsupported}" "$CI_COMMIT_SHA"
    - |
      if [ "${CI_JOB_RETRY_COUNT:-unsupported}" = "0" ]; then
        echo "synthetic transient failure before side effect"
        exit 75
      fi
    - echo "transient check recovered"

deploy:fake:
  stage: deploy
  image: alpine:3.22
  interruptible: false
  resource_group: "ch18-training-environment"
  script:
    - mkdir -p evidence
    - printf 'resource=ch18-training-environment\npipeline=%s\njob=%s\nsha=%s\nstart=%s\n'         "$CI_PIPELINE_ID" "$CI_JOB_ID" "$CI_COMMIT_SHA" "$(date -u +%FT%TZ)" | tee evidence/deploy.txt
    - sleep 30
    - printf 'finish=%s\n' "$(date -u +%FT%TZ)" | tee -a evidence/deploy.txt
  artifacts:
    when: always
    expire_in: 1 week
    paths: [evidence/deploy.txt]

evidence:identity:
  stage: evidence
  image: alpine:3.22
  script:
    - printf 'source=%s ref=%s sha=%s pipeline=%s job=%s\n'         "$CI_PIPELINE_SOURCE" "$CI_COMMIT_REF_NAME" "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID"

If CI_JOB_RETRY_COUNT is not available on your GitLab version, do not force this exact retry job. Replace it with the local retry simulation in the next section and mark the platform limitation in the evidence packet.

6. Execute two competing pipelines deliberately

  1. Commit and push the checkpoint pipeline as commit A.
  2. While build:safe is still running, make a harmless documentation change and push commit B.
  3. Record both pipeline IDs immediately.
  4. Observe whether the stale interruptible build/verify work in pipeline A is canceled when B arrives.
  5. Allow both pipelines that reach deploy:fake to contend for the same resource key. Record which job runs and which waits.
  6. Download or inspect each deployment receipt. The later deployment’s start must not precede the previous owner’s finish.

If auto-cancel prevents pipeline A from reaching deploy, that is valid evidence too. To demonstrate pure serialization separately, create two manual/otherwise non-redundant disposable pipelines that both reach deployment, or temporarily set on_new_commit: none in a separate lab commit and record that assumption. Do not weaken a production policy just to produce a screenshot.

7. Prove retry behavior without hiding attempt 1

For GitLab 19.3+, the transient job should first fail with retry count 0 and exit 75, then GitLab processes one retry with retry count 1. Record both job attempt IDs/statuses and the exact source SHA. The deployment stage should only execute after verification succeeds, so the synthetic transient failure occurs before the side-effect boundary.

If the retry fails for a deterministic reason other than the synthetic exit code, stop and diagnose that first; do not keep rerunning until green.

8. Prove the external-side-effect contract locally

tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT
key='project-demo:training-env:artifact-sha256-demo'
./ci/idempotency-demo.sh "$tmp" "$key" 'desired-v1'
./ci/idempotency-demo.sh "$tmp" "$key" 'desired-v1'
printf 'rows=%s\n' "$(wc -l < "$tmp/state.tsv")"
cat "$tmp/state.tsv"
test "$(wc -l < "$tmp/state.tsv")" -eq 1

This verifies a key invariant: a repeated request with the same operation identity does not duplicate the intended state. In a real target, replace the file with a supported idempotency key or read-before/create-or-update API and verify the artifact/deployment digest.

9. Deliberate failure: make the resource key too narrow

On a disposable branch, temporarily change:

resource_group: "ch18-training-environment-$CI_COMMIT_SHA"

Create two pipelines. The keys differ, so GitLab is free to run both fake deployments concurrently. Preserve the pipeline/job IDs and overlapping timestamps. This is the intentional first failure.

Repair by restoring the stable ch18-training-environment key. Create a new pipeline and prove non-overlap. Do not delete the broken pipeline; it demonstrates that the key identifies the real resource, not merely a job.

10. Optional timeout recovery exercise

Add a separate verify:timeout-demo job with a 30-second job timeout and a 45-second sleep, marked allow_failure: true. Preserve its timeout evidence. Then answer: would this be safe around a real deployment API? Only if the target offers a reliable cancellation/reconciliation contract. Otherwise the correct recovery is to query the target before retrying.

11. Required evidence packet

Category Required evidence
Source/config branch/MR context, CI_PIPELINE_SOURCE, ref, CI_COMMIT_SHA, exact YAML commit
Pipelines both competing pipeline IDs and statuses
Jobs build/verify/deploy job IDs, canceled/failed/retried/waiting/running states
Retry attempt IDs, CI_JOB_RETRY_COUNT if supported, exit code/reason
Resource lock key, process-mode assumption, wait/ownership evidence, deploy timestamps
Artifacts both deployment receipt artifacts and retention
Duplicate control MR/push pipeline list showing intended source only after workflow repair
Reconciliation local idempotency key, one-row final state, limitations note
Runtime Runner/executor/image identity if visible; unknowns explicitly marked

12. Verification checklist

  • Every recorded pipeline/job can be tied to one exact SHA.
  • Only safe compute jobs are interruptible.
  • The fake side-effect job is non-interruptible and uses one stable resource-group key.
  • Two jobs with that key never overlap in observed execution time.
  • The transient failure is retried no more than configured and the first failure remains visible.
  • The MR workflow does not create a redundant push pipeline for the same branch when the MR pipeline exists.
  • The local side-effect simulation applies the same idempotency key twice but stores one state row.
  • No token, real secret, cloud resource, production environment, release, or package was touched.

13. Cleanup / rollback

  1. Restore the correct stable resource-group key if you ran the deliberate broken branch.
  2. Close the disposable MR and delete only the lab branches/projects you created.
  3. Let synthetic one-week artifacts expire or remove them from the disposable project only.
  4. Do not cancel/delete unrelated pipelines, change shared runner settings, or mutate resource-group process modes as cleanup.
  5. The local temporary reconciliation directory is removed automatically by trap.

14. What Chapter 18 adds to the production operating model

Chapter 17 proved that parallel work can scale safely when every shard is bounded and auditable. Chapter 18 adds the controls needed when concurrent pipelines approach mutable state: redundant pipeline creation is suppressed early, stale safe work is interruptible, one logical resource is serialized, retries are classified and bounded, timeouts preserve an evidence budget, and every uncertain side effect is reconciled by stable identity before another attempt.

Chapter 19 moves from these concurrency controls to GitLab environments and deployment records: environment tiers, URLs, deployment history, and operational traceability. The resource_group key introduced here will become one part of a larger deployment identity rather than the whole deployment model.

Knowledge check

What is the strongest proof that the two fake deployments were serialized?

Why does the checkpoint inject retry failure before the deploy stage?

What deliberate defect makes the resource lock ineffective?

What should happen after an unknown-outcome timeout around a real deployment?

What concept bridges to Chapter 19?

Next chapter

Chapter 19 — Environments and deployments

Move from concurrency-safe side effects to environment identity, deployment history, URLs, tiers, and operational traceability.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Concurrency, auto-cancel, retry failure reasons, retry-count variables, and timeout behavior are version-sensitive. Re-check the GitLab and Runner versions used by production pipelines before applying the exact examples.

  • Resource groups — serialization, process modes, downstream-pipeline locking, waiting-for-resource diagnostics, and deadlock guidance.
  • CI/CD YAML syntax reference — resource_group, interruptible, retry, timeout, and workflow:auto_cancel.
  • workflow keyword — pipeline creation, duplicate branch/MR prevention, and CI_OPEN_MERGE_REQUESTS patterns.
  • Predefined variables — pipeline/job identity, job timeout, and current retry-attempt metadata.
  • Configure runners — runner maximum job timeout and script/after-script timeout controls.
  • Resource Groups API — reading/updating process mode for an existing resource group.

Current assumptions used in this chapter: mandatory examples use Free-tier CI/CD features and synthetic data. A resource group serializes one resource at a time. Current process modes are unordered (default), oldest_first, newest_first, and newest_ready_first; newest-first modes require idempotent jobs. workflow:auto_cancel:on_new_commit currently supports conservative (default), interruptible, and none. Job retry allows 0–2 retries; retry:exit_codes is generally available. Current GitLab 19.3 docs expose CI_JOB_RETRY_COUNT; older deployments need a different lab signal. Job-level timeout can override the project default but remains bounded by runner maximum timeout.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.