Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks: Concepts, Architecture, and Mental Model
Build an evidence-first mental model for diagnosing GitLab CI/CD from pipeline creation and compiled configuration through runner execution, artifacts, deployment, and external state.
Learning objectives
- Classify failures by pipeline creation, compilation/rules, scheduling, runner/executor, runtime/network, evidence-transfer, deployment, or external state.
- Preserve source SHA, pipeline/job identity, merged configuration, runner context, and first-failure evidence before mutation.
- Use CI Lint and job statuses to rule layers in or out without confusing configuration validation with runtime success.
- Explain why pending, preparing, running, failed, and external-unhealthy states require different diagnostics.
- Connect troubleshooting evidence to reproducible DevOps operation and governance.
Current-platform note: GitLab documentation checked 2026-09-13. CI Lint is available on Free, Premium, and Ultimate across GitLab.com, Self-Managed, and Dedicated. Exact job failure-reason names evolve; the current YAML reference already documents GitLab 19.1 changes to several retry/failure classifications, so always record the instance and Runner versions before interpreting a production incident.
1. Troubleshooting starts by refusing to guess
A failed pipeline is only a symptom. The causal defect can live before a pipeline exists, during configuration compilation, while rules decide whether jobs exist, in the scheduler/runner boundary, inside the executor, in DNS/TLS or another dependency, in the job script, in artifact/report transfer, or after GitLab reports a successful deployment job while the external target is unhealthy.
Chapter 36 added policy origin and governance evidence. Chapter 37 turns that discipline into an incident method: preserve the first observable failure, identify which state transition did not happen as expected, change the smallest responsible layer, and prove recovery against both GitLab state and any external target.
2. Name the states before touching Retry
Source / pipeline intent
Pipeline source, project/ref, exact source SHA, user/API intent, and policy revision. A retry without this identity can mix evidence from different experiments.
Compiled configuration
Lint result, expanded/merged YAML, includes, workflow/job rules,
stages, needs, variables visible to evaluation, and
the jobs that should exist.
Queue / runner state
Job ID/status, required tags, protected-runner conditions, runner availability, runner manager/version, executor and image preparation.
Runtime / network state
Workspace, shell, tool versions, DNS resolution, TLS trust, registry/package endpoint behavior, and script exit status.
Evidence-transfer state
Reports, artifacts, caches, registry objects and their IDs/digests. Job success does not prove a report was ingested or an artifact was retained.
Deployment / external state
Environment/deployment record plus independent target health, version/digest, rollback result, and any provider-side event.
Recovery state
What changed, who changed it, whether the old failure remains preserved, what was rerun, and which evidence proves the final state.
3. The evidence-first causal chain
The diagnostic sequence is deliberately ordered from cheapest and most stable evidence toward more invasive inspection. First ask whether GitLab created the intended pipeline for the intended SHA. Then prove what configuration GitLab evaluated. Only after a job exists should you investigate queueing and runner selection. Only after a runner accepts the job should you blame the script, network, image, or tool. Finally, a successful deployment job still requires target-state verification.
Every arrow is a hypothesis boundary. If no pipeline was created,
runner debugging is wasted effort. If a job is
pending with unmatched tags, application logs are
irrelevant. If the script succeeds but the target still serves the
old version, the failure is downstream of the job.
flowchart TD
A[Symptom or alert] --> B[Pipeline source / ref / SHA]
B --> C[CI Lint + full/merged config + rules]
C --> D[Pipeline/job graph + status]
D --> E[Runner tags / manager / executor / image]
E --> F[Script / tool / DNS / TLS / network]
F --> G[Reports / artifacts / cache / registry]
G --> H[Environment / deployment / external target]
H --> I[Least-destructive repair]
I --> J[Smallest safe retry or new pipeline]
J --> K[Independent recovery verification]
4. Separate “invalid configuration” from “no matching pipeline”
These failures look similar to a beginner because both can produce
“nothing useful ran,” but they occur at different stages. An invalid
YAML/configuration error blocks compilation. A valid configuration
can still create no pipeline because
workflow:rules deliberately reject the event. A valid
pipeline can then omit a specific job because that job’s
rules do not match.
| Observation | Likely layer | First evidence |
|---|---|---|
yaml_errors / lint invalid |
Configuration syntax or GitLab CI schema | CI Lint errors and exact submitted YAML |
| Lint valid; simulation has no expected job | Rules / pipeline-source logic |
dry_run + include_jobs, source/ref
assumptions
|
| Pipeline exists; job absent | Job-rule or dependency selection | Full/merged config + job list |
Job exists and is pending |
Scheduling / runner match | Job tags, protected status, runner availability |
Job reaches preparing then fails |
Executor/image/helper/network setup | Runner trace + executor/image details |
Job is running then fails |
Script/tool/network/identity | First failing command, exit code, safe logs |
| Deploy job succeeds; target unhealthy | Deployment/provider/verification | Environment record + external target health |
5. Read-only inspection before mutation
On a real disposable project, collect identifiers before rerunning anything. Use the UI or API/CLI available to you; the exact commands below are examples, not permission to inspect production systems you do not own.
# Record tool versions first.
git --version
glab --version 2>/dev/null || true
gitlab-runner --version 2>/dev/null || true
# Local source identity.
git rev-parse HEAD
git status --short
# Free CI Lint path when glab is authenticated to an authorized project.
glab ci lint --dry-run --include-jobs --ref main 2>/dev/null || true
# Optional runner-host read-only checks on an authorized runner manager.
gitlab-runner list 2>/dev/null || true
gitlab-runner verify 2>/dev/null || true
gitlab-runner status 2>/dev/null || true
gitlab-runner verify proves that registered runners can
connect to GitLab; current Runner documentation explicitly says it
does not prove that the Runner service is using them. Treat
service state, runner registration, job routing, and actual job
execution as separate evidence.
6. CI Lint is a compiler-side diagnostic, not a runner test
Current GitLab CI Lint can do more than parse YAML. Project-scoped
lint can resolve local includes and project variables, and
dry_run can simulate pipeline creation. The API can
return the merged YAML and, with include_jobs, the jobs
that would exist. A successful lint does not prove a runner exists,
an image can be pulled, DNS works, or a deployment target is
healthy.
# Example only: use a disposable authorized project and never put a token in the command history.
# Prefer an authenticated CLI or a header sourced from a protected local environment.
jq --null-input --arg yaml "$(cat .gitlab-ci.yml)" \
'{content:$yaml,dry_run:true,include_jobs:true,ref:"main"}' \
> /tmp/ch37-lint-request.json
# POST /projects/:id/ci/lint with the JSON body above.
# Preserve response.valid, response.errors, response.warnings,
# response.merged_yaml, response.includes, and response.jobs.
Version-sensitive API detail: for validating an
existing configuration, current docs use
content_ref and dry_run_ref; older
sha/ref forms are deprecated in that GET
flow. Check your instance docs before automating it.
7. A pending job is evidence, not a generic “runner problem”
GitLab job status narrows the search. pending means the
job is queued waiting for a runner. Runner tags are conjunctive: a
runner must have every tag listed by the job.
preparing means a runner is preparing the execution
environment. waiting_for_resource points to
resource-group serialization rather than missing capacity. These
statuses prevent a blind “restart the runner” reflex.
| Job state | Do not assume | Inspect instead |
|---|---|---|
created |
The runner is broken | Dependencies/stage/DAG and whether scheduling has begun |
pending |
More CPU will fix it | Runner tags/scope/protection/online status/capacity |
preparing |
The test failed | Executor, helper/image pull, volume/network setup |
running |
GitLab scheduling failed | Script/tool/network/identity/resource behavior |
waiting_for_resource |
Runner starvation | Resource group owner and upstream blockers |
failed |
Retry is safe | Failure reason + first failing evidence + side effects |
8. DevOps connection: reproducibility includes failure reproduction
A delivery pipeline is reproducible only if another engineer can reconstruct the same source SHA, compiled configuration, rule outcome, runner/executor context, toolchain, evidence outputs, and external state. Troubleshooting therefore produces an evidence packet, not just a fixed pipeline. The packet is what lets a team distinguish a real remediation from an accidental green rerun.
Knowledge check
What evidence should you preserve before retrying a failed GitLab job?
Preserve pipeline/job IDs, source ref and SHA, compiled configuration/rule outcome, first-failure trace, runner/executor/image identity, and relevant artifact/deployment/external evidence.
A job remains pending with no trace. Should you start by debugging the application script?
No. A pending job has not executed the script. Inspect job eligibility, requested tags, runner scope/protection/status, manager connectivity, and executor capacity first.
Why is blind retry a poor default for a flaky failure?
It can erase the timing and first-failure evidence needed to distinguish a real flake from runner, network, configuration, or external-system failure.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Further reading — current official GitLab sources
Behavior and version-sensitive assumptions in this chapter were checked against current GitLab documentation on 2026-09-13. Re-check your exact GitLab and GitLab Runner versions before production recovery automation.
- GitLab Docs — Validate GitLab CI/CD configuration
- GitLab Docs — CI Lint API
- GitLab Docs — Pipeline editor / full configuration
- GitLab Docs — CI/CD YAML syntax and retry
- GitLab Docs — Jobs API
- GitLab Docs — Pipelines API
- GitLab Docs — CI/CD jobs and statuses
- GitLab Docs — Troubleshooting CI/CD variables / debug trace
- GitLab Docs — Services and CI_DEBUG_SERVICES
- GitLab Docs — GitLab Runner commands
- GitLab Docs — Troubleshooting GitLab Runner
- GitLab Docs — Runner advanced configuration
- GitLab Docs — Docker executor image pull errors
- GitLab Docs — Docker build troubleshooting (DNS/TLS)
- GitLab Docs — Job artifacts and destructive deletion
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.