Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks: Guided Hands-On Workflow and Core Operations
Practice a progressive six-fault workflow in a disposable local lab: invalid YAML, no pipeline, runner mismatch, DNS failure, flaky test, and failed deployment verification.
Learning objectives
- Create a free disposable troubleshooting workspace with synthetic data and no real credentials.
- Preserve each first failure before making a repair or rerunning anything.
- Diagnose invalid configuration, rule selection, runner tag mismatch, DNS failure, deterministic flake, and target-state mismatch.
- Map local evidence to optional GitLab CI Lint, pipeline/job, runner and trace inspection.
- Choose the causal layer for a new symptom instead of copying a repair sequence.
Lab boundary: the mandatory path is local, free,
and disposable. It creates files only under a temporary
gitlab-ch37-lab directory and uses a Python virtual
environment plus synthetic target state. Optional real-GitLab
commands are inspection-first and require an authorized disposable
project. No production runner, credential, registry, cloud account,
or deployment target is required.
1. Scenario: Atlas Relay has six different “pipeline failures”
You are given one small service whose team reports that “GitLab CI is flaky.” Instead of retrying until green, you will classify six faults: malformed configuration, valid configuration that creates no pipeline for the event, a job that cannot match a runner, a DNS/image-style dependency failure, a deterministic flaky test, and a deployment verification failure. Every fault gets a before-state, preserved trace, minimal repair, and final proof.
2. Preflight and tool assumptions
The local lab uses Python 3.10+ and Git. PyYAML is installed only
inside a temporary virtual environment. If you also have
glab, you may lint the final YAML against a disposable
GitLab project, but the chapter does not require it.
python --version
git --version
LAB="${TMPDIR:-/tmp}/gitlab-ch37-lab"
rm -rf "$LAB"
mkdir -p "$LAB"/{ci,tools,evidence,target,state}
cd "$LAB"
python -m venv .venv
# Git Bash / Linux / macOS:
. .venv/bin/activate 2>/dev/null || . .venv/Scripts/activate
python -m pip install --quiet 'PyYAML>=6,<7'
python - <<'PY'
import yaml,sys
print('pyyaml',yaml.__version__,'python',sys.version.split()[0])
PY
3. Create a tiny diagnostic harness
The harness intentionally models only the state boundaries needed for this chapter. It does not claim to emulate GitLab. GitLab remains authoritative for real CI compilation; the harness gives you deterministic local faults so everyone can practice the same causal sequence.
# Save as tools/diagnose.py
import argparse, json, socket, sys
from pathlib import Path
import yaml
p=argparse.ArgumentParser()
p.add_argument('mode',choices=['yaml','rules','runner','dns','flaky','deploy'])
p.add_argument('--file')
p.add_argument('--source',default='push')
p.add_argument('--job-tags',default='linux,docker')
p.add_argument('--runner-tags',default='linux,docker')
p.add_argument('--host',default='registry.invalid')
p.add_argument('--state-dir',default='state')
p.add_argument('--target',default='target/current.json')
a=p.parse_args()
def emit(kind,status,**kw):
out={'layer':kind,'status':status,**kw}
print(json.dumps(out,sort_keys=True))
return out
if a.mode=='yaml':
try:
data=yaml.safe_load(Path(a.file).read_text())
assert isinstance(data,dict)
emit('configuration','parsed',top_level=list(data))
except Exception as e:
emit('configuration','invalid',error=str(e)); sys.exit(2)
elif a.mode=='rules':
data=yaml.safe_load(Path(a.file).read_text()) or {}
rules=(data.get('workflow') or {}).get('rules') or []
allowed=False
for r in rules:
expr=r.get('if','')
if 'CI_PIPELINE_SOURCE' in expr and f'"{a.source}"' in expr:
allowed=r.get('when','on_success')!='never'; break
if r.get('when')=='always': allowed=True; break
emit('rules','pipeline-created' if allowed else 'no-pipeline',source=a.source)
sys.exit(0 if allowed else 3)
elif a.mode=='runner':
job={x for x in a.job_tags.split(',') if x}
runner={x for x in a.runner_tags.split(',') if x}
missing=sorted(job-runner)
emit('runner','matched' if not missing else 'pending',missing_tags=missing)
sys.exit(0 if not missing else 4)
elif a.mode=='dns':
try:
addrs=socket.getaddrinfo(a.host,443)
emit('network','resolved',host=a.host,answers=len(addrs))
except Exception as e:
emit('network','dns-failure',host=a.host,error=type(e).__name__); sys.exit(5)
elif a.mode=='flaky':
d=Path(a.state_dir); d.mkdir(exist_ok=True)
c=d/'flaky-count'; n=int(c.read_text() or '0') if c.exists() else 0
c.write_text(str(n+1))
if n==0: emit('script','failed-first-attempt',attempt=1); sys.exit(6)
emit('script','passed',attempt=n+1)
elif a.mode=='deploy':
target=json.loads(Path(a.target).read_text())
expected='sha-demo-good'
ok=target.get('source_sha')==expected and target.get('healthy') is True
emit('deployment','verified' if ok else 'verification-failed',expected_sha=expected,target=target)
sys.exit(0 if ok else 7)
cat > tools/diagnose.py <<'PYFILE'
# Paste the Python program from this section exactly here.
PYFILE
chmod +x tools/diagnose.py
4. Fault 1 — invalid YAML: stop before runner debugging
Create a malformed file and preserve the parser failure. Local YAML parsing catches syntax, but real GitLab CI Lint must still validate GitLab-specific keywords and logic.
cat > ci/broken-syntax.yml <<'YAML'
stages: [test]
test:
script:
- echo hello
tags: [linux]
YAML
set +e
python tools/diagnose.py yaml --file ci/broken-syntax.yml \
> evidence/01-invalid-yaml.json 2>&1
printf '%s\n' "$?" > evidence/01-invalid-yaml.exit
set -e
cat evidence/01-invalid-yaml.json
Interpretation: no pipeline/job/runner evidence exists yet. The smallest repair is indentation/configuration, not runner restart, token changes, or network debugging.
5. Fault 2 — valid YAML, but this event creates no pipeline
Now the YAML parses. The problem is activation logic: the workflow allows merge-request pipelines only, while you simulate a push.
cat > ci/rules.yml <<'YAML'
workflow:
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- when: never
stages: [test]
test:
script: echo test
YAML
python tools/diagnose.py yaml --file ci/rules.yml | tee evidence/02-yaml-valid.json
set +e
python tools/diagnose.py rules --file ci/rules.yml --source push \
> evidence/02-no-pipeline.json
printf '%s\n' "$?" > evidence/02-no-pipeline.exit
set -e
cat evidence/02-no-pipeline.json
python tools/diagnose.py rules --file ci/rules.yml --source merge_request_event \
| tee evidence/02-mr-created.json
On real GitLab, use CI Lint pipeline simulation and
include_jobs to test the relevant ref/source
assumptions. Do not add when: always merely to make a
pipeline appear unless that is truly the intended activation
contract.
6. Fault 3 — a stuck/pending job caused by runner tags
A runner must have every tag requested by a job. Model a job asking
for linux,docker,gpu while the available runner has
only linux,docker.
set +e
python tools/diagnose.py runner \
--job-tags linux,docker,gpu --runner-tags linux,docker \
> evidence/03-runner-pending.json
printf '%s\n' "$?" > evidence/03-runner-pending.exit
set -e
cat evidence/03-runner-pending.json
python tools/diagnose.py runner \
--job-tags linux,docker --runner-tags linux,docker \
| tee evidence/03-runner-matched.json
The fix is not automatically “remove the tag.” In production you must decide whether the job truly requires a specialized trust/capability boundary. If it does, repair runner availability/routing; if it does not, remove the unjustified requirement through review.
7. Fault 4 — image/network-style failure: prove DNS before blaming the app
The reserved .invalid top-level domain is used to
produce a deterministic DNS failure without touching a real
registry.
set +e
python tools/diagnose.py dns --host registry.invalid \
> evidence/04-dns-failure.json
printf '%s\n' "$?" > evidence/04-dns-failure.exit
set -e
cat evidence/04-dns-failure.json
A real image-pull failure can be caused by image name, registry availability, DNS, proxy, certificate trust, credentials, pull policy, or executor configuration. Never “fix” an x509 error with TLS verification disabled. Capture the failing hostname, resolver result, certificate chain/issuer and runner/executor path, then repair trust correctly.
8. Fault 5 — a deterministic flaky test: preserve attempt 1
This synthetic test fails the first time and passes the second. It demonstrates why a green retry is evidence of nondeterminism, not evidence that nothing was wrong.
rm -f state/flaky-count
set +e
python tools/diagnose.py flaky --state-dir state \
> evidence/05-flake-attempt1.json
printf '%s\n' "$?" > evidence/05-flake-attempt1.exit
set -e
python tools/diagnose.py flaky --state-dir state \
| tee evidence/05-flake-attempt2.json
cat evidence/05-flake-attempt1.json
cat evidence/05-flake-attempt2.json
Production response: retain the first failed job ID/log/report,
classify whether a retry is safe, then file/route remediation for
the nondeterministic dependency. Do not silently turn every
script_failure into retry: 2.
9. Fault 6 — deployment job success is not target success
Create a local “deployment target” record whose source SHA is wrong. The deployment verifier fails until the target state is corrected and independently re-read.
cat > target/current.json <<'JSON'
{"source_sha":"sha-demo-old","healthy":true,"release":"2026.09.13-old"}
JSON
set +e
python tools/diagnose.py deploy --target target/current.json \
> evidence/06-deploy-failed.json
printf '%s\n' "$?" > evidence/06-deploy-failed.exit
set -e
cat evidence/06-deploy-failed.json
cat > target/current.json <<'JSON'
{"source_sha":"sha-demo-good","healthy":true,"release":"2026.09.13-fixed"}
JSON
python tools/diagnose.py deploy --target target/current.json \
| tee evidence/06-deploy-verified.json
10. Optional GitLab inspection path
With an authorized disposable project, reproduce the same evidence chain using GitLab itself. CI Lint is Free-tier. Job/pipeline APIs are also available broadly, but authentication and project permissions still apply.
# Example inspection only; adapt to your authenticated disposable project.
glab ci lint --dry-run --include-jobs --ref main
# Record pipeline/job JSON without printing credentials.
# glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID"
# glab api "projects/$PROJECT_ID/pipelines/$PIPELINE_ID/jobs?include_retried=true"
# glab api "projects/$PROJECT_ID/jobs/$JOB_ID/trace" > "evidence/job-$JOB_ID.trace"
11. Challenge — choose the layer before the command
You see a deployment job marked success, but the
service still reports the previous artifact digest. Which layer do
you investigate first?
Expected reasoning: not YAML compilation, not runner tags, and not automatic retry. Preserve the deployment/job identifiers, inspect environment/deployment metadata, verify the promoted artifact/digest, then inspect the external target. Repair only the handoff or target reconciliation that failed.
Knowledge check
What evidence should you preserve before retrying a failed GitLab job?
Preserve pipeline/job IDs, source ref and SHA, compiled configuration/rule outcome, first-failure trace, runner/executor/image identity, and relevant artifact/deployment/external evidence.
A job remains pending with no trace. Should you start by debugging the application script?
No. A pending job has not executed the script. Inspect job eligibility, requested tags, runner scope/protection/status, manager connectivity, and executor capacity first.
Why is blind retry a poor default for a flaky failure?
It can erase the timing and first-failure evidence needed to distinguish a real flake from runner, network, configuration, or external-system failure.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Further reading — current official GitLab sources
Behavior and version-sensitive assumptions in this chapter were checked against current GitLab documentation on 2026-09-13. Re-check your exact GitLab and GitLab Runner versions before production recovery automation.
- GitLab Docs — Validate GitLab CI/CD configuration
- GitLab Docs — CI Lint API
- GitLab Docs — Pipeline editor / full configuration
- GitLab Docs — CI/CD YAML syntax and retry
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — CI/CD jobs and statuses
- GitLab Docs — Troubleshooting CI/CD variables / debug trace
- GitLab Docs — Services and CI_DEBUG_SERVICES
- GitLab Docs — GitLab Runner commands
- GitLab Docs — Troubleshooting GitLab Runner
- GitLab Docs — Runner advanced configuration
- GitLab Docs — Docker executor image pull errors
- GitLab Docs — Docker build troubleshooting (DNS/TLS)
- GitLab Docs — Job artifacts and destructive deletion
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.