Chapter 10Lesson 04~225 minutes

GitLab CI/CD Foundations: .gitlab-ci.yml, Pipelines, Jobs, Stages, and Execution Model: Diagnostics, Failure Modes, Security, and Performance

Diagnose invalid CI configuration, pipelines suppressed at creation time, pending jobs with no eligible runner, wrong-SHA assumptions, and over-trusted job tokens using evidence from the correct control or execution plane.

DiagnosticsYAML vs CI schemaworkflow rulesPending jobsWrong SHACI_JOB_TOKEN

Learning objectives

  • Distinguish YAML syntax failure, GitLab CI schema/logic failure, pipeline-creation suppression, runner scheduling failure, and job execution failure.
  • Diagnose a pending job by runner eligibility rather than changing unrelated scripts.
  • Prove which commit/ref/source a successful or failed job actually used before interpreting its result.
  • Treat CI_JOB_TOKEN as restricted temporary authority and diagnose authorization failures without broadening scope blindly.
  • Use preserve evidence → scope → inspect → least-destructive correction → verify as the standard CI/CD diagnostic sequence.
Availability baseline (verified 2026-08-21). Core GitLab CI/CD, .gitlab-ci.yml, Pipeline Editor/CI Lint, jobs, stages, ordinary branch and merge-request pipelines, predefined variables, artifacts, runner scheduling, and CI_JOB_TOKEN are available on GitLab Free across GitLab.com, Self-Managed, and Dedicated. GitLab-hosted runners are provided for GitLab.com and GitLab Dedicated; Self-Managed installations provide and operate their own runner capacity. Hosted compute quotas, credits, machine types, and billing can change, so this chapter never hard-codes them and provides a validation-only path whenever a runner is unavailable. Current glab uses the glab ci command family; older pipe/pipeline aliases are deprecated.

1. Diagnostic sequence: identify the failed phase first

Use the same disciplined sequence from earlier chapters, now expanded for CI/CD:

  1. Preserve evidence: pipeline URL/ID, job ID, source, ref, SHA, status, log excerpt, runner ID/tags, and configuration commit.
  2. Identify scope: GitLab offering/instance → namespace/project → ref/MR → pipeline source/ID → job/stage → runner/executor → artifact/token/resource.
  3. Inspect the owning layer: CI Lint/config, pipeline graph, runner list/tags/protection, job log, API status/body, or job-token policy.
  4. Choose the least destructive correction: repair one rule/tag/command/permission rather than disabling governance or adding broad credentials.
  5. Verify: prove the new pipeline/job corresponds to the intended SHA and that the original cause no longer exists.

2. Symptom-to-layer map

Observed state Likely owning layer Inspect first Wrong first reaction
CI Lint says invalid GitLab CI config/schema Lint error + merged/full configuration. Register or restart runners.
Push occurs but no pipeline exists Pipeline creation workflow:rules, source, ref, config location. Edit runner tags.
Pipeline exists, job is pending Runner scheduling Runner scope/status/tags/protection and job tags. Rewrite shell command.
Job runs then fails Executor/script Job log, exit code, checkout SHA, environment. Change workflow rules.
Job succeeds but tests wrong revision Identity/assumption Pipeline SHA, CI_COMMIT_SHA, checkout. Retry repeatedly without comparing SHAs.
API request returns 401/403 Authentication/authorization Token type/lifetime, endpoint support, allowlist, triggering user permission. Replace job token with broad PAT.

3. Broken example A: YAML parses, GitLab rejects the CI model

This file is syntactically valid YAML, but the job names a stage that is not declared. Generic YAML validation can pass while GitLab CI Lint rejects it.

stages:
  - test

compile:
  stage: build
  script:
    - echo "compile"
Interpret the error. The failure belongs to GitLab configuration validation: build is not in the declared stages. The repair is to add the intended stage or assign the job to test, then lint again. No runner should be involved yet.
stages:
  - build
  - test

compile:
  stage: build
  script:
    - echo "compile"

4. Broken example B: valid configuration deliberately creates no pipeline

A configuration can be valid and still suppress pipeline creation. This is a creation-time result, not a failed pipeline.

workflow:
  rules:
    - when: never

never_created:
  script:
    - echo "A job definition exists, but no pipeline is created."
If the push produces no pipeline, inspect the evaluated workflow and CI_PIPELINE_SOURCE assumptions. Do not wait for a runner: no job entered the queue.

5. Broken example C: no runner satisfies the job tags

The following job can create a pipeline successfully yet remain pending if no available runner carries every declared tag.

needs_impossible_runner:
  tags:
    - ch10-lab-no-such-runner
  script:
    - echo "This should not execute in the lab."
Evidence Interpretation
Pipeline exists Configuration was valid enough to create the graph.
Job status = pending The job is queued, not executing.
UI reports no runner matching tags / runner list lacks tag Scheduling mismatch is causal.
Removing the synthetic tag makes a normal untagged runner eligible, if policy permits Least-destructive repair restores eligibility without changing the command.
Lab hygiene: if you test this live, use only a disposable branch/project, capture the pending evidence promptly, then cancel the pipeline or repair/remove the synthetic tag. Do not leave deliberately stuck jobs accumulating.

6. Broken assumption D: “green” for the wrong commit

Imagine pipeline P1 passed for SHA a1a1…. A new commit moves the branch to b2b2…. P1 remains a valid historical result for a1a1…, but it is not evidence for b2b2…. The repair is not “rerun P1 until green”; create/identify the pipeline whose recorded SHA equals the candidate commit.

BRANCH="feature/example"
git fetch origin "$BRANCH"
REMOTE_SHA="$(git rev-parse FETCH_HEAD)"

glab ci list --ref "$BRANCH" --output json --per-page 10   --jq '.[] | {id,status,source,ref,sha}'

printf 'remote_tip=%s\n' "$REMOTE_SHA"

7. Broken assumption E: CI_JOB_TOKEN should access everything the user can

GitLab documents a limited set of resources/endpoints for job-token authentication. The token also expires when the job ends. Cross-project access normally requires the target’s job-token allowlist plus the triggering user’s underlying permission. A syntactically correct JOB-TOKEN header can therefore receive 401 or 403 for a legitimate policy reason.

Check Question
Lifetime Is the request happening inside the running job, before token revocation?
Endpoint support Is this REST resource documented for CI_JOB_TOKEN? Job tokens cannot authenticate GraphQL requests.
Target scope For cross-project access, is the source project/group on the target allowlist where required?
Human permission Does the user who triggered the job have the required target-project permission?
Runner trust Could another untrusted job steal the token from an insecure shared execution environment?
Do not “fix” a 403 by echoing the token, placing it in an artifact, embedding it in a URL, or replacing it with an unscoped long-lived PAT. Preserve the status/body without secret material, verify the documented endpoint and authorization chain, then grant only the missing capability if it is actually required.

8. Runner security is causally relevant to token safety

GitLab explicitly warns that insecure runner configuration can let jobs steal tokens from other jobs. Reusing a shell executor host for mutually untrusted projects, or using privileged containers on a reused machine without a deliberate isolation design, expands the blast radius. This is why Chapter 14 treats runner executors/isolation as security architecture.

At this chapter’s level, the rule is simple: do not send untrusted CI configuration to a runner merely because the runner is powerful and available. Match the runner trust zone to the source/configuration trust level.

9. Performance and cost only after correctness

A pending job caused by no matching runner is not solved by buying faster compute. A pipeline never created does not consume job runtime. A slow job that is actually running may justify profiling command duration, image/setup time, network, or runner capacity. Keep these causal distinctions so cost optimization does not mask correctness failures.

Hosted-runner compute limits and billing are volatile. Record the current account/project policy when it materially affects a lab, but do not encode current minute/credit numbers into operational logic.

10. Minimal evidence pack for a CI incident

# Example read-only evidence collection; replace project/ref values.
REPO="GROUP/PROJECT"
REF="feature/example"
mkdir -p ci-evidence

git fetch origin "$REF"
git rev-parse FETCH_HEAD > ci-evidence/remote-sha.txt

glab ci list -R "$REPO" --ref "$REF" --output json --per-page 20   > ci-evidence/pipelines.json

glab runner list -R "$REPO" --output json   > ci-evidence/runners.json

# Copy only non-secret log excerpts/IDs needed for diagnosis.
# Never archive CI_JOB_TOKEN or runner authentication tokens.

Knowledge check

A pipeline does not exist after a push. Why is restarting a runner a poor first action?

What makes “valid YAML” insufficient evidence?

A job is pending with tags [linux, gpu], but the only runner has [linux]. What is the cause?

Why might a 403 from a CI_JOB_TOKEN be correct even if the pipeline-triggering user is a Maintainer somewhere?

What should be compared before trusting a green job for release?

Summary

CI/CD diagnosis becomes tractable when you identify the phase: parse/validate → create pipeline → materialize job → schedule runner → execute checkout/script → authorize resource → upload evidence. Preserve IDs and SHAs, repair the owning layer, and resist broad-credential or governance-bypass shortcuts.

Official references

Next lesson

Checkpoint: predict, run, break, repair, prove

Lesson 5 integrates the chapter into one disposable pipeline exercise with explicit prediction, one safe failure, independent verification, and complete cleanup.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.