Checkpoint Lab — Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks
Complete a three-layer incident drill, preserve a hashed evidence packet, restore configuration/routing/target state with minimal changes, and hand off a reusable playbook.
Learning objectives
- Predict configuration, runner-routing, deployment, and evidence-preservation state changes before repair.
- Preserve and repair three independent failure layers without deleting original evidence.
- Justify whether the real GitLab equivalent should retry a job or create a new pipeline.
- Build a before/after evidence packet with hashes, runner/target state, assumptions, and limitations.
- Close the course troubleshooting loop and bridge to the Chapter 38 production capstone.
Checkpoint boundary: this timed drill is fully local and disposable. It uses synthetic YAML, runner inventory, job traces, and target state. The optional GitLab mapping may be performed only in an authorized disposable project. No paid tier, administrator change, cloud subscription, managed Kubernetes cluster, protected production environment, or real credential is required.
1. Checkpoint scenario: three-layer incident drill
Atlas Relay has a release candidate blocked by three independent failures. Layer A is configuration compilation. Layer B is runner scheduling. Layer C is deployment verification. Your objective is not speed alone: preserve each first failure, write predictions before changes, restore a valid pipeline/target with the least destructive changes, and produce an evidence packet another engineer can audit.
Use a 45-minute target if you want a timed exercise: 10 minutes triage/predictions, 20 minutes repairs, 10 minutes verification/evidence, 5 minutes playbook review. The time target is optional; evidence quality is mandatory.
2. Assumptions and current-platform mapping
- Documentation checked: 2026-09-13.
- Mandatory path: Python 3.10+, Git, PyYAML in a temporary virtual environment.
-
Real GitLab optional path: CI Lint works across
Free/Premium/Ultimate; record actual GitLab,
glab, and Runner versions. -
Optional Runner diagnostics: use read-only
list/verify/statusfirst; never unregister/reconfigure a shared runner for this lab. - No secret/debug trace is required. If you choose to test debug tracing in a disposable project, use fake variables only.
3. Create the disposable incident workspace
Create all checkpoint state under one guarded path.
LAB="${TMPDIR:-/tmp}/gitlab-ch37-checkpoint"
rm -rf "$LAB"
mkdir -p "$LAB"/{ci,runner,evidence,target,recovery}
cd "$LAB"
python -m venv .venv
. .venv/bin/activate 2>/dev/null || . .venv/Scripts/activate
python -m pip install --quiet 'PyYAML>=6,<7'
git init -q
git config user.name 'Chapter 37 Learner'
git config user.email 'learner@example.invalid'
cat > ci/.gitlab-ci.yml <<'YAML'
stages: [test, deploy]
test:
stage: test
script: echo test
tags: [linux, docker]
deploy:
stage: deploy
script: echo deploy
tags: [linux, docker, gpu]
YAML
cat > runner/inventory.json <<'JSON'
{"runner_id":88,"version":"record-real-version-in-production","executor":"docker","tags":["linux","docker"],"online":true}
JSON
cat > target/current.json <<'JSON'
{"source_sha":"sha-old","artifact_digest":"sha256:old","healthy":true}
JSON
printf 'expected_source_sha=sha-good\nexpected_artifact_digest=sha256:good\n' > recovery/expected.env
git add . && git commit -qm 'checkpoint: establish broken incident state'
printf 'checkpoint_source=%s\n' "$(git rev-parse HEAD)" | tee evidence/source.txt
4. Write predictions before repairing anything
Record at least these predictions in
evidence/predictions.txt.
| Prediction | Expected observation before repair | Independent proof after repair |
|---|---|---|
| P1 — configuration | YAML parser/lint fails; no runner diagnosis is yet causal | YAML parses; optional GitLab CI Lint valid |
| P2 — scheduling |
After config repair, deploy requires gpu but
inventory lacks it
|
Chosen job tag set is fully contained in trusted runner tags |
| P3 — deployment | Target reports old SHA/digest despite pipeline-side “deploy” intent | Target file reports expected SHA/digest and healthy=true |
| P4 — evidence preservation | Broken files/traces remain available after repair | Before/after files have separate hashes in manifest |
5. Layer A — preserve and repair configuration failure
Capture the original file and parser error before fixing indentation.
cp ci/.gitlab-ci.yml evidence/A-original.gitlab-ci.yml
set +e
python - <<'PY' > evidence/A-yaml-before.json 2>&1
import json,yaml
p='ci/.gitlab-ci.yml'
try:
d=yaml.safe_load(open(p)); print(json.dumps({'valid_yaml':True,'keys':list(d)}))
except Exception as e:
print(json.dumps({'valid_yaml':False,'error':str(e)})); raise
PY
printf '%s\n' "$?" > evidence/A-yaml-before.exit
set -e
python - <<'PY'
p='ci/.gitlab-ci.yml'
s=open(p).read().replace(' tags: [linux, docker]',' tags: [linux, docker]')
open(p,'w').write(s)
PY
python - <<'PY' | tee evidence/A-yaml-after.json
import json,yaml
d=yaml.safe_load(open('ci/.gitlab-ci.yml'))
print(json.dumps({'valid_yaml':True,'jobs':[k for k in d if k!='stages']}))
PY
Optional real GitLab proof: run
glab ci lint --dry-run --include-jobs --ref main in an
authorized disposable project because PyYAML proves only YAML
syntax, not GitLab CI semantics.
6. Layer B — preserve and repair runner-routing failure
Do not provision an imaginary GPU runner. First decide whether the
deployment job actually needs the gpu capability. In
this synthetic scenario it does not, so remove only that unjustified
tag in the repaired configuration while preserving the original.
python - <<'PY' | tee evidence/B-runner-before.json
import json,yaml
cfg=yaml.safe_load(open('ci/.gitlab-ci.yml'))
r=json.load(open('runner/inventory.json'))
need=set(cfg['deploy']['tags']); have=set(r['tags'])
print(json.dumps({'job':'deploy','required':sorted(need),'runner_tags':sorted(have),'missing':sorted(need-have)}))
assert need-have, 'expected a scheduling mismatch'
PY
cp ci/.gitlab-ci.yml evidence/B-before-tag-fix.gitlab-ci.yml
python - <<'PY'
p='ci/.gitlab-ci.yml'
s=open(p).read().replace('tags: [linux, docker, gpu]','tags: [linux, docker]')
open(p,'w').write(s)
PY
python - <<'PY' | tee evidence/B-runner-after.json
import json,yaml
cfg=yaml.safe_load(open('ci/.gitlab-ci.yml')); r=json.load(open('runner/inventory.json'))
need=set(cfg['deploy']['tags']); have=set(r['tags']); missing=need-have
print(json.dumps({'required':sorted(need),'runner_tags':sorted(have),'missing':sorted(missing)}))
assert not missing
PY
7. Layer C — deployment verification and rollback/roll-forward proof
Preserve the old target state. The safest repair in this simulation is a controlled roll-forward to the expected immutable identity, followed by an independent read.
cp target/current.json evidence/C-target-before.json
python - <<'PY' | tee evidence/C-verify-before.json
import json
x=json.load(open('target/current.json'))
ok=(x['source_sha']=='sha-good' and x['artifact_digest']=='sha256:good' and x['healthy'])
print(json.dumps({'verified':ok,'target':x})); assert not ok
PY
cat > target/current.json <<'JSON'
{"source_sha":"sha-good","artifact_digest":"sha256:good","healthy":true}
JSON
python - <<'PY' | tee evidence/C-verify-after.json
import json
x=json.load(open('target/current.json'))
ok=(x['source_sha']=='sha-good' and x['artifact_digest']=='sha256:good' and x['healthy'])
print(json.dumps({'verified':ok,'target':x})); assert ok
PY
In production, the equivalent action might be rollback or roll-forward. The invariant is the same: do not stop at “deployment command returned zero.” Verify the actual target version/digest and health.
8. Decide what you would rerun
Because checkpoint layers A and B changed CI configuration, the real GitLab equivalent should be a new pipeline tied to the repaired configuration/SHA, not a retry of the old job expecting it to see new YAML. Layer C changes external target state and must be verified independently. If only a transient network failure had occurred with unchanged config/source and idempotent behavior, a single job retry could be appropriate.
9. Build the evidence packet and hash manifest
Keep before/after artifacts separate. The manifest makes accidental editing visible.
git add ci/.gitlab-ci.yml target/current.json
GIT_EDITOR=true git commit -qm 'fix: repair checkpoint configuration routing and target state' || true
git rev-parse HEAD | tee evidence/repaired-source.txt
python - <<'PY'
import hashlib,json,platform
from pathlib import Path
files=sorted(str(p) for p in Path('evidence').glob('*') if p.is_file())
def h(p): return hashlib.sha256(Path(p).read_bytes()).hexdigest()
manifest={
'chapter':37,
'generated_at_assumption':'2026-09-13',
'initial_source':Path('evidence/source.txt').read_text().strip(),
'repaired_source':Path('evidence/repaired-source.txt').read_text().strip(),
'runner':json.load(open('runner/inventory.json')),
'target_final':json.load(open('target/current.json')),
'python':platform.python_version(),
'files':[{'path':p,'sha256':h(p)} for p in files],
'limitations':[
'PyYAML validates YAML syntax, not GitLab CI semantics.',
'Runner matching is a local model; real GitLab scheduling also depends on scope/protection/online/capacity.',
'Target deployment state is a local simulation; production requires provider/application verification.'
]
}
Path('evidence/manifest.json').write_text(json.dumps(manifest,indent=2)+'\n')
print(json.dumps(manifest,indent=2))
PY
sha256sum evidence/* 2>/dev/null | tee evidence/all-sha256.txt
10. Write the reusable playbook
Your final recovery/playbook.md should include:
symptom, evidence-freeze rule, source/config checks, rules/job graph
checks, runner checks, runtime/network checks, artifact/deployment
checks, forbidden shortcuts, retry-vs-new-pipeline decision,
rollback/roll-forward verification, escalation owner, and evidence
retention requirements.
# GitLab CI/CD incident playbook
1. Preserve pipeline/job IDs, SHA, first trace/report, merged config, runner and target state.
2. Classify the first failed state transition.
3. Reject unrelated fixes explicitly.
4. Apply the least-destructive causal correction.
5. Retry one safe job only when source/config are unchanged and side effects are idempotent.
6. Create a new pipeline when source/config/include/policy changed.
7. Never disable TLS, print secrets, broaden PATs, or erase evidence as a shortcut.
8. Verify GitLab state and external target independently.
9. Record permanent remediation owner/date and any temporary mitigation expiry.
11. Verification checklist and cleanup
- Original malformed YAML and parser error remain in evidence.
- Repaired YAML parses and the real-GitLab path, if used, passes CI Lint simulation.
- Original runner-tag mismatch and repaired match are both recorded.
- Original target SHA/digest and verified final target are both recorded.
- No debug trace or secret value was needed.
- The retry/new-pipeline decision is justified from changed state.
- Evidence manifest contains hashes and limitations.
cd /tmp 2>/dev/null || cd /
printf 'about_to_remove=%s\n' "$LAB"
case "$LAB" in
*/gitlab-ch37-checkpoint) rm -rf "$LAB" ;;
*) echo 'Refusing unexpected cleanup path' >&2; exit 2 ;;
esac
12. Knowledge check
A job is pending and requires
linux,docker,gpu; the online runner has
linux,docker. Which layer should you investigate
first?
Scheduling/runner matching. The runner lacks one required tag, so the job has not reached executor or script execution.
Why can retrying one job be the wrong action after a referenced include changed?
A job retry stays with the original pipeline configuration snapshot. If you intend to evaluate the changed include, create a new pipeline and compare the merged configuration/evidence.
A first test attempt fails and a retry passes with no code change. What has been proven?
Nondeterminism/transience, not root-cause remediation. Preserve attempt 1 and investigate the differing dependency/environment/state.
Why is
CI_DEBUG_TRACE=true security-sensitive?
GitLab warns that debug trace exposes all variables and secrets available to the job. Use narrow diagnostics first and control access/rotation if trace is unavoidable.
A rollback job is green, but the target still serves the old digest. Is recovery complete?
No. Job success is GitLab execution evidence; external target identity/health must be independently verified.
When should you prefer a new pipeline over a job retry?
When source, CI configuration, includes, policy, or intended rule evaluation changed—or when you need a clean experiment tied to the repaired state.
13. What Chapter 37 adds to the production operating model
Chapter 37 adds a repeatable incident loop to the course: preserve first failure, locate the failed state transition, reject unrelated hypotheses, repair the smallest causal layer, choose retry versus new pipeline deliberately, and verify both GitLab and external state. This prevents recovery from becoming a sequence of destructive guesses.
Chapter 38 now has the final prerequisite it needs. The capstone can build a complete delivery platform and deliberately inject failures with a known evidence, diagnosis, recovery, and handoff discipline.
Knowledge check
What evidence should you preserve before retrying a failed GitLab job?
Preserve pipeline/job IDs, source ref and SHA, compiled configuration/rule outcome, first-failure trace, runner/executor/image identity, and relevant artifact/deployment/external evidence.
A job remains pending with no trace. Should you start by debugging the application script?
No. A pending job has not executed the script. Inspect job eligibility, requested tags, runner scope/protection/status, manager connectivity, and executor capacity first.
Why is blind retry a poor default for a flaky failure?
It can erase the timing and first-failure evidence needed to distinguish a real flake from runner, network, configuration, or external-system failure.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Further reading — current official GitLab sources
Behavior and version-sensitive assumptions in this chapter were checked against current GitLab documentation on 2026-09-13. Re-check your exact GitLab and GitLab Runner versions before production recovery automation.
- GitLab Docs — Validate GitLab CI/CD configuration
- GitLab Docs — CI Lint API
- GitLab Docs — Pipeline editor / full configuration
- GitLab Docs — CI/CD YAML syntax and retry
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — CI/CD jobs and statuses
- GitLab Docs — Troubleshooting CI/CD variables / debug trace
- GitLab Docs — Services and CI_DEBUG_SERVICES
- GitLab Docs — GitLab Runner commands
- GitLab Docs — Troubleshooting GitLab Runner
- GitLab Docs — Runner advanced configuration
- GitLab Docs — Docker executor image pull errors
- GitLab Docs — Docker build troubleshooting (DNS/TLS)
- GitLab Docs — Job artifacts and destructive deletion
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.