Chapter 37Lesson 05~190 minutes

Checkpoint Lab — Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks

Complete a three-layer incident drill, preserve a hashed evidence packet, restore configuration/routing/target state with minimal changes, and hand off a reusable playbook.

Checkpoint labIncident drillEvidence packetRecovery playbook

Learning objectives

  • Predict configuration, runner-routing, deployment, and evidence-preservation state changes before repair.
  • Preserve and repair three independent failure layers without deleting original evidence.
  • Justify whether the real GitLab equivalent should retry a job or create a new pipeline.
  • Build a before/after evidence packet with hashes, runner/target state, assumptions, and limitations.
  • Close the course troubleshooting loop and bridge to the Chapter 38 production capstone.

Checkpoint boundary: this timed drill is fully local and disposable. It uses synthetic YAML, runner inventory, job traces, and target state. The optional GitLab mapping may be performed only in an authorized disposable project. No paid tier, administrator change, cloud subscription, managed Kubernetes cluster, protected production environment, or real credential is required.

1. Checkpoint scenario: three-layer incident drill

Atlas Relay has a release candidate blocked by three independent failures. Layer A is configuration compilation. Layer B is runner scheduling. Layer C is deployment verification. Your objective is not speed alone: preserve each first failure, write predictions before changes, restore a valid pipeline/target with the least destructive changes, and produce an evidence packet another engineer can audit.

Use a 45-minute target if you want a timed exercise: 10 minutes triage/predictions, 20 minutes repairs, 10 minutes verification/evidence, 5 minutes playbook review. The time target is optional; evidence quality is mandatory.

2. Assumptions and current-platform mapping

  • Documentation checked: 2026-09-13.
  • Mandatory path: Python 3.10+, Git, PyYAML in a temporary virtual environment.
  • Real GitLab optional path: CI Lint works across Free/Premium/Ultimate; record actual GitLab, glab, and Runner versions.
  • Optional Runner diagnostics: use read-only list/verify/status first; never unregister/reconfigure a shared runner for this lab.
  • No secret/debug trace is required. If you choose to test debug tracing in a disposable project, use fake variables only.

3. Create the disposable incident workspace

Create all checkpoint state under one guarded path.

LAB="${TMPDIR:-/tmp}/gitlab-ch37-checkpoint"
rm -rf "$LAB"
mkdir -p "$LAB"/{ci,runner,evidence,target,recovery}
cd "$LAB"
python -m venv .venv
. .venv/bin/activate 2>/dev/null || . .venv/Scripts/activate
python -m pip install --quiet 'PyYAML>=6,<7'
git init -q
git config user.name 'Chapter 37 Learner'
git config user.email 'learner@example.invalid'

cat > ci/.gitlab-ci.yml <<'YAML'
stages: [test, deploy]
test:
  stage: test
  script: echo test
   tags: [linux, docker]
deploy:
  stage: deploy
  script: echo deploy
  tags: [linux, docker, gpu]
YAML
cat > runner/inventory.json <<'JSON'
{"runner_id":88,"version":"record-real-version-in-production","executor":"docker","tags":["linux","docker"],"online":true}
JSON
cat > target/current.json <<'JSON'
{"source_sha":"sha-old","artifact_digest":"sha256:old","healthy":true}
JSON
printf 'expected_source_sha=sha-good\nexpected_artifact_digest=sha256:good\n' > recovery/expected.env
git add . && git commit -qm 'checkpoint: establish broken incident state'
printf 'checkpoint_source=%s\n' "$(git rev-parse HEAD)" | tee evidence/source.txt

4. Write predictions before repairing anything

Record at least these predictions in evidence/predictions.txt.

Prediction Expected observation before repair Independent proof after repair
P1 — configuration YAML parser/lint fails; no runner diagnosis is yet causal YAML parses; optional GitLab CI Lint valid
P2 — scheduling After config repair, deploy requires gpu but inventory lacks it Chosen job tag set is fully contained in trusted runner tags
P3 — deployment Target reports old SHA/digest despite pipeline-side “deploy” intent Target file reports expected SHA/digest and healthy=true
P4 — evidence preservation Broken files/traces remain available after repair Before/after files have separate hashes in manifest

5. Layer A — preserve and repair configuration failure

Capture the original file and parser error before fixing indentation.

cp ci/.gitlab-ci.yml evidence/A-original.gitlab-ci.yml
set +e
python - <<'PY' > evidence/A-yaml-before.json 2>&1
import json,yaml
p='ci/.gitlab-ci.yml'
try:
    d=yaml.safe_load(open(p)); print(json.dumps({'valid_yaml':True,'keys':list(d)}))
except Exception as e:
    print(json.dumps({'valid_yaml':False,'error':str(e)})); raise
PY
printf '%s\n' "$?" > evidence/A-yaml-before.exit
set -e

python - <<'PY'
p='ci/.gitlab-ci.yml'
s=open(p).read().replace('   tags: [linux, docker]','  tags: [linux, docker]')
open(p,'w').write(s)
PY
python - <<'PY' | tee evidence/A-yaml-after.json
import json,yaml
d=yaml.safe_load(open('ci/.gitlab-ci.yml'))
print(json.dumps({'valid_yaml':True,'jobs':[k for k in d if k!='stages']}))
PY

Optional real GitLab proof: run glab ci lint --dry-run --include-jobs --ref main in an authorized disposable project because PyYAML proves only YAML syntax, not GitLab CI semantics.

6. Layer B — preserve and repair runner-routing failure

Do not provision an imaginary GPU runner. First decide whether the deployment job actually needs the gpu capability. In this synthetic scenario it does not, so remove only that unjustified tag in the repaired configuration while preserving the original.

python - <<'PY' | tee evidence/B-runner-before.json
import json,yaml
cfg=yaml.safe_load(open('ci/.gitlab-ci.yml'))
r=json.load(open('runner/inventory.json'))
need=set(cfg['deploy']['tags']); have=set(r['tags'])
print(json.dumps({'job':'deploy','required':sorted(need),'runner_tags':sorted(have),'missing':sorted(need-have)}))
assert need-have, 'expected a scheduling mismatch'
PY
cp ci/.gitlab-ci.yml evidence/B-before-tag-fix.gitlab-ci.yml
python - <<'PY'
p='ci/.gitlab-ci.yml'
s=open(p).read().replace('tags: [linux, docker, gpu]','tags: [linux, docker]')
open(p,'w').write(s)
PY
python - <<'PY' | tee evidence/B-runner-after.json
import json,yaml
cfg=yaml.safe_load(open('ci/.gitlab-ci.yml')); r=json.load(open('runner/inventory.json'))
need=set(cfg['deploy']['tags']); have=set(r['tags']); missing=need-have
print(json.dumps({'required':sorted(need),'runner_tags':sorted(have),'missing':sorted(missing)}))
assert not missing
PY

7. Layer C — deployment verification and rollback/roll-forward proof

Preserve the old target state. The safest repair in this simulation is a controlled roll-forward to the expected immutable identity, followed by an independent read.

cp target/current.json evidence/C-target-before.json
python - <<'PY' | tee evidence/C-verify-before.json
import json
x=json.load(open('target/current.json'))
ok=(x['source_sha']=='sha-good' and x['artifact_digest']=='sha256:good' and x['healthy'])
print(json.dumps({'verified':ok,'target':x})); assert not ok
PY
cat > target/current.json <<'JSON'
{"source_sha":"sha-good","artifact_digest":"sha256:good","healthy":true}
JSON
python - <<'PY' | tee evidence/C-verify-after.json
import json
x=json.load(open('target/current.json'))
ok=(x['source_sha']=='sha-good' and x['artifact_digest']=='sha256:good' and x['healthy'])
print(json.dumps({'verified':ok,'target':x})); assert ok
PY

In production, the equivalent action might be rollback or roll-forward. The invariant is the same: do not stop at “deployment command returned zero.” Verify the actual target version/digest and health.

8. Decide what you would rerun

Because checkpoint layers A and B changed CI configuration, the real GitLab equivalent should be a new pipeline tied to the repaired configuration/SHA, not a retry of the old job expecting it to see new YAML. Layer C changes external target state and must be verified independently. If only a transient network failure had occurred with unchanged config/source and idempotent behavior, a single job retry could be appropriate.

9. Build the evidence packet and hash manifest

Keep before/after artifacts separate. The manifest makes accidental editing visible.

git add ci/.gitlab-ci.yml target/current.json
GIT_EDITOR=true git commit -qm 'fix: repair checkpoint configuration routing and target state' || true
git rev-parse HEAD | tee evidence/repaired-source.txt
python - <<'PY'
import hashlib,json,platform
from pathlib import Path
files=sorted(str(p) for p in Path('evidence').glob('*') if p.is_file())
def h(p): return hashlib.sha256(Path(p).read_bytes()).hexdigest()
manifest={
 'chapter':37,
 'generated_at_assumption':'2026-09-13',
 'initial_source':Path('evidence/source.txt').read_text().strip(),
 'repaired_source':Path('evidence/repaired-source.txt').read_text().strip(),
 'runner':json.load(open('runner/inventory.json')),
 'target_final':json.load(open('target/current.json')),
 'python':platform.python_version(),
 'files':[{'path':p,'sha256':h(p)} for p in files],
 'limitations':[
   'PyYAML validates YAML syntax, not GitLab CI semantics.',
   'Runner matching is a local model; real GitLab scheduling also depends on scope/protection/online/capacity.',
   'Target deployment state is a local simulation; production requires provider/application verification.'
 ]
}
Path('evidence/manifest.json').write_text(json.dumps(manifest,indent=2)+'\n')
print(json.dumps(manifest,indent=2))
PY
sha256sum evidence/* 2>/dev/null | tee evidence/all-sha256.txt

10. Write the reusable playbook

Your final recovery/playbook.md should include: symptom, evidence-freeze rule, source/config checks, rules/job graph checks, runner checks, runtime/network checks, artifact/deployment checks, forbidden shortcuts, retry-vs-new-pipeline decision, rollback/roll-forward verification, escalation owner, and evidence retention requirements.

# GitLab CI/CD incident playbook
1. Preserve pipeline/job IDs, SHA, first trace/report, merged config, runner and target state.
2. Classify the first failed state transition.
3. Reject unrelated fixes explicitly.
4. Apply the least-destructive causal correction.
5. Retry one safe job only when source/config are unchanged and side effects are idempotent.
6. Create a new pipeline when source/config/include/policy changed.
7. Never disable TLS, print secrets, broaden PATs, or erase evidence as a shortcut.
8. Verify GitLab state and external target independently.
9. Record permanent remediation owner/date and any temporary mitigation expiry.

11. Verification checklist and cleanup

  • Original malformed YAML and parser error remain in evidence.
  • Repaired YAML parses and the real-GitLab path, if used, passes CI Lint simulation.
  • Original runner-tag mismatch and repaired match are both recorded.
  • Original target SHA/digest and verified final target are both recorded.
  • No debug trace or secret value was needed.
  • The retry/new-pipeline decision is justified from changed state.
  • Evidence manifest contains hashes and limitations.
cd /tmp 2>/dev/null || cd /
printf 'about_to_remove=%s\n' "$LAB"
case "$LAB" in
  */gitlab-ch37-checkpoint) rm -rf "$LAB" ;;
  *) echo 'Refusing unexpected cleanup path' >&2; exit 2 ;;
esac

12. Knowledge check

A job is pending and requires linux,docker,gpu; the online runner has linux,docker. Which layer should you investigate first?

Why can retrying one job be the wrong action after a referenced include changed?

A first test attempt fails and a retry passes with no code change. What has been proven?

Why is CI_DEBUG_TRACE=true security-sensitive?

A rollback job is green, but the target still serves the old digest. Is recovery complete?

When should you prefer a new pipeline over a job retry?

13. What Chapter 37 adds to the production operating model

Chapter 37 adds a repeatable incident loop to the course: preserve first failure, locate the failed state transition, reject unrelated hypotheses, repair the smallest causal layer, choose retry versus new pipeline deliberately, and verify both GitLab and external state. This prevents recovery from becoming a sequence of destructive guesses.

Chapter 38 now has the final prerequisite it needs. The capstone can build a complete delivery platform and deliberately inject failures with a known evidence, diagnosis, recovery, and handoff discipline.

Next chapter

Next: Production Capstone: Build, Secure, Scale, Observe, and Govern a Complete GitLab Delivery Platform

Integrate the course into one production-grade GitLab delivery platform, then prove its security, governance, observability, failure recovery, and operational handoff.

Knowledge check

What evidence should you preserve before retrying a failed GitLab job?

A job remains pending with no trace. Should you start by debugging the application script?

Why is blind retry a poor default for a flaky failure?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Further reading — current official GitLab sources

Behavior and version-sensitive assumptions in this chapter were checked against current GitLab documentation on 2026-09-13. Re-check your exact GitLab and GitLab Runner versions before production recovery automation.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.