Chapter 37Lesson 01~105 minutes

Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks: Concepts, Architecture, and Mental Model

Build an evidence-first mental model for diagnosing GitLab CI/CD from pipeline creation and compiled configuration through runner execution, artifacts, deployment, and external state.

TroubleshootingCI LintRunner stateEvidence first

Learning objectives

  • Classify failures by pipeline creation, compilation/rules, scheduling, runner/executor, runtime/network, evidence-transfer, deployment, or external state.
  • Preserve source SHA, pipeline/job identity, merged configuration, runner context, and first-failure evidence before mutation.
  • Use CI Lint and job statuses to rule layers in or out without confusing configuration validation with runtime success.
  • Explain why pending, preparing, running, failed, and external-unhealthy states require different diagnostics.
  • Connect troubleshooting evidence to reproducible DevOps operation and governance.

Current-platform note: GitLab documentation checked 2026-09-13. CI Lint is available on Free, Premium, and Ultimate across GitLab.com, Self-Managed, and Dedicated. Exact job failure-reason names evolve; the current YAML reference already documents GitLab 19.1 changes to several retry/failure classifications, so always record the instance and Runner versions before interpreting a production incident.

1. Troubleshooting starts by refusing to guess

A failed pipeline is only a symptom. The causal defect can live before a pipeline exists, during configuration compilation, while rules decide whether jobs exist, in the scheduler/runner boundary, inside the executor, in DNS/TLS or another dependency, in the job script, in artifact/report transfer, or after GitLab reports a successful deployment job while the external target is unhealthy.

Chapter 36 added policy origin and governance evidence. Chapter 37 turns that discipline into an incident method: preserve the first observable failure, identify which state transition did not happen as expected, change the smallest responsible layer, and prove recovery against both GitLab state and any external target.

2. Name the states before touching Retry

Source / pipeline intent

Pipeline source, project/ref, exact source SHA, user/API intent, and policy revision. A retry without this identity can mix evidence from different experiments.

Compiled configuration

Lint result, expanded/merged YAML, includes, workflow/job rules, stages, needs, variables visible to evaluation, and the jobs that should exist.

Queue / runner state

Job ID/status, required tags, protected-runner conditions, runner availability, runner manager/version, executor and image preparation.

Runtime / network state

Workspace, shell, tool versions, DNS resolution, TLS trust, registry/package endpoint behavior, and script exit status.

Evidence-transfer state

Reports, artifacts, caches, registry objects and their IDs/digests. Job success does not prove a report was ingested or an artifact was retained.

Deployment / external state

Environment/deployment record plus independent target health, version/digest, rollback result, and any provider-side event.

Recovery state

What changed, who changed it, whether the old failure remains preserved, what was rerun, and which evidence proves the final state.

3. The evidence-first causal chain

The diagnostic sequence is deliberately ordered from cheapest and most stable evidence toward more invasive inspection. First ask whether GitLab created the intended pipeline for the intended SHA. Then prove what configuration GitLab evaluated. Only after a job exists should you investigate queueing and runner selection. Only after a runner accepts the job should you blame the script, network, image, or tool. Finally, a successful deployment job still requires target-state verification.

Every arrow is a hypothesis boundary. If no pipeline was created, runner debugging is wasted effort. If a job is pending with unmatched tags, application logs are irrelevant. If the script succeeds but the target still serves the old version, the failure is downstream of the job.

Symptom → evidence → causal layer → minimal repair → verified recovery
            flowchart TD
             A[Symptom or alert] --> B[Pipeline source / ref / SHA]
             B --> C[CI Lint + full/merged config + rules]
             C --> D[Pipeline/job graph + status]
             D --> E[Runner tags / manager / executor / image]
             E --> F[Script / tool / DNS / TLS / network]
             F --> G[Reports / artifacts / cache / registry]
             G --> H[Environment / deployment / external target]
             H --> I[Least-destructive repair]
             I --> J[Smallest safe retry or new pipeline]
             J --> K[Independent recovery verification]
          

4. Separate “invalid configuration” from “no matching pipeline”

These failures look similar to a beginner because both can produce “nothing useful ran,” but they occur at different stages. An invalid YAML/configuration error blocks compilation. A valid configuration can still create no pipeline because workflow:rules deliberately reject the event. A valid pipeline can then omit a specific job because that job’s rules do not match.

Observation Likely layer First evidence
yaml_errors / lint invalid Configuration syntax or GitLab CI schema CI Lint errors and exact submitted YAML
Lint valid; simulation has no expected job Rules / pipeline-source logic dry_run + include_jobs, source/ref assumptions
Pipeline exists; job absent Job-rule or dependency selection Full/merged config + job list
Job exists and is pending Scheduling / runner match Job tags, protected status, runner availability
Job reaches preparing then fails Executor/image/helper/network setup Runner trace + executor/image details
Job is running then fails Script/tool/network/identity First failing command, exit code, safe logs
Deploy job succeeds; target unhealthy Deployment/provider/verification Environment record + external target health

5. Read-only inspection before mutation

On a real disposable project, collect identifiers before rerunning anything. Use the UI or API/CLI available to you; the exact commands below are examples, not permission to inspect production systems you do not own.

# Record tool versions first.
git --version
glab --version 2>/dev/null || true
gitlab-runner --version 2>/dev/null || true

# Local source identity.
git rev-parse HEAD
git status --short

# Free CI Lint path when glab is authenticated to an authorized project.
glab ci lint --dry-run --include-jobs --ref main 2>/dev/null || true

# Optional runner-host read-only checks on an authorized runner manager.
gitlab-runner list 2>/dev/null || true
gitlab-runner verify 2>/dev/null || true
gitlab-runner status 2>/dev/null || true

gitlab-runner verify proves that registered runners can connect to GitLab; current Runner documentation explicitly says it does not prove that the Runner service is using them. Treat service state, runner registration, job routing, and actual job execution as separate evidence.

6. CI Lint is a compiler-side diagnostic, not a runner test

Current GitLab CI Lint can do more than parse YAML. Project-scoped lint can resolve local includes and project variables, and dry_run can simulate pipeline creation. The API can return the merged YAML and, with include_jobs, the jobs that would exist. A successful lint does not prove a runner exists, an image can be pulled, DNS works, or a deployment target is healthy.

# Example only: use a disposable authorized project and never put a token in the command history.
# Prefer an authenticated CLI or a header sourced from a protected local environment.
jq --null-input --arg yaml "$(cat .gitlab-ci.yml)" \
  '{content:$yaml,dry_run:true,include_jobs:true,ref:"main"}' \
  > /tmp/ch37-lint-request.json

# POST /projects/:id/ci/lint with the JSON body above.
# Preserve response.valid, response.errors, response.warnings,
# response.merged_yaml, response.includes, and response.jobs.

Version-sensitive API detail: for validating an existing configuration, current docs use content_ref and dry_run_ref; older sha/ref forms are deprecated in that GET flow. Check your instance docs before automating it.

7. A pending job is evidence, not a generic “runner problem”

GitLab job status narrows the search. pending means the job is queued waiting for a runner. Runner tags are conjunctive: a runner must have every tag listed by the job. preparing means a runner is preparing the execution environment. waiting_for_resource points to resource-group serialization rather than missing capacity. These statuses prevent a blind “restart the runner” reflex.

Job state Do not assume Inspect instead
created The runner is broken Dependencies/stage/DAG and whether scheduling has begun
pending More CPU will fix it Runner tags/scope/protection/online status/capacity
preparing The test failed Executor, helper/image pull, volume/network setup
running GitLab scheduling failed Script/tool/network/identity/resource behavior
waiting_for_resource Runner starvation Resource group owner and upstream blockers
failed Retry is safe Failure reason + first failing evidence + side effects

8. DevOps connection: reproducibility includes failure reproduction

A delivery pipeline is reproducible only if another engineer can reconstruct the same source SHA, compiled configuration, rule outcome, runner/executor context, toolchain, evidence outputs, and external state. Troubleshooting therefore produces an evidence packet, not just a fixed pipeline. The packet is what lets a team distinguish a real remediation from an accidental green rerun.

Next lesson

Next: Pipeline Troubleshooting, YAML Debugging, Runner Failures, Network Issues, Flaky Jobs, and Recovery Playbooks: Guided Hands-On Workflow and Core Operations

Build a disposable fault set and walk it from invalid YAML through deployment verification without losing the first failure.

Knowledge check

What evidence should you preserve before retrying a failed GitLab job?

A job remains pending with no trace. Should you start by debugging the application script?

Why is blind retry a poor default for a flaky failure?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Further reading — current official GitLab sources

Behavior and version-sensitive assumptions in this chapter were checked against current GitLab documentation on 2026-09-13. Re-check your exact GitLab and GitLab Runner versions before production recovery automation.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.