Chapter 01Lesson 04~115 minutes

Continuous Integration and Delivery Foundations, GitLab Pipeline Architecture, and Delivery Flow: Diagnostics, Failure Modes, Security, and Performance

Troubleshoot GitLab CI/CD by following evidence from pipeline source and compiled configuration through job creation, queueing, runner/executor assignment, scripts, artifacts, and external targets. You will diagnose realistic failures without leaking secrets, weakening runner isolation, destroying first-failure evidence, or using blind retries as a substitute for root-cause analysis.

DiagnosticsFailure modesRunner securityEvidence firstPerformance

Learning objectives

  • Apply an evidence-first diagnostic sequence from pipeline creation through runner execution and external target state.
  • Diagnose valid YAML that creates no desired pipeline or job, and distinguish that from runner or script failure.
  • Recognize pending-job runner problems, ref/SHA confusion, missing artifacts, and green-job/failed-target mismatches.
  • Avoid unsafe troubleshooting shortcuts such as broad tokens, secret dumps, privileged runners, TLS disablement, and blind retries.
  • Separate queue time, execution time, and external wait time before attempting performance optimization.

1. Preserve first-failure evidence before changing anything

The fastest-looking troubleshooting action—rerun, edit, delete, retry—can destroy the evidence that distinguishes one cause from another. Before changing configuration, preserve the pipeline ID, source, ref, SHA, compiled configuration or lint result, relevant job IDs/statuses, the first failing trace, runner identity if assigned, and any artifact/deployment/external resource identifiers already created.

Then move through the system in causal order:

  1. Source/ref/SHA: what actually attempted to create the pipeline?
  2. Configuration: which CI configuration was resolved and was it valid?
  3. Pipeline decision: did workflow logic allow a pipeline?
  4. Job graph: which jobs exist, and which were omitted/skipped?
  5. Queue/runner: did an eligible runner claim the job?
  6. Executor/runtime: was the environment prepared as expected?
  7. Script/tool/network: where is the first non-zero or unexpected result?
  8. Artifact/report/cache: did the expected durable evidence get created and uploaded?
  9. Environment/deployment/external target: did the side effect really complete and become healthy?
Diagnostic principle: start at the last state you can prove, then inspect the next expected transition. Do not begin three layers downstream because that is where the most familiar tool lives.

2. Failure mode: valid YAML, but no push pipeline

The following configuration is syntactically valid, but it deliberately allows only scheduled pipelines. A developer can push it, see no new push pipeline, and incorrectly blame the runner. The runner is not involved because no eligible push pipeline is created.

workflow:
  rules:
    - if: '$CI_PIPELINE_SOURCE == "schedule"'
    - when: never

observe:
  script:
    - echo "This job can exist only in an allowed pipeline."

Use CI Lint and pipeline simulation to inspect this before touching runner settings. Current GitLab CI Lint can simulate pipeline creation and expose rule/graph problems. For a push-oriented lab, the minimal repair could explicitly allow push and web sources:

workflow:
  rules:
    - if: '$CI_PIPELINE_SOURCE == "push"'
    - if: '$CI_PIPELINE_SOURCE == "web"'
    - when: never

This example teaches a layer boundary: YAML validity, pipeline creation, and runner execution are three different states. A green linter is useful but does not guarantee that the current source matches a rule.

3. Failure mode: pipeline exists, but the expected job does not

Pipeline-level logic may allow creation while job-level rules omit one particular job. If the pipeline exists but the job is absent from the graph, inspect job rules and the actual CI_PIPELINE_SOURCE/ref context. Do not wait for a runner to execute a job that does not exist.

pipeline_identity:
  script:
    - echo "pipeline exists"

manual_only_check:
  script:
    - echo "included only for a web-created pipeline"
  rules:
    - if: '$CI_PIPELINE_SOURCE == "web"'

In a push pipeline, manual_only_check is absent by design. That is not “skipped after execution”; it was not included in the graph. Precise status language avoids needless runner debugging.

4. Failure mode: job exists but stays pending

A pending job has advanced farther: GitLab created the pipeline and the job. Now inspect matching runners. Runner tags, scope, protected status, pause/offline state, executor availability, and capacity can all affect whether a job starts. A deliberately impossible tag demonstrates the symptom without changing any runner:

pending_demo:
  tags:
    - course-runner-tag-that-does-not-exist
  script:
    - echo "This line will not run without a matching runner."

In a disposable test, the correct repair is to remove/correct the unnecessary job tag or use an explicitly authorized runner whose tags match the intended workload. The wrong repair is to register a new privileged production runner, make a protected runner unprotected, or broaden runner access without understanding the trust consequence.

Pending is not failed script execution. Preserve the job ID and runner eligibility evidence. There is no first failing command because no command has run yet.

5. Failure mode: “GitLab ran the wrong commit”

A common investigation compares an old pipeline to the current head of a branch. The branch may have moved after the pipeline was created. That does not change the historical pipeline SHA. Compare the pipeline's recorded commit with the job's CI_COMMIT_SHA and repository history, not with an assumption that main still points where it did hours ago.

# In a local clone, compare immutable identities.
git rev-parse HEAD
git rev-parse origin/main
git show --no-patch --oneline YOUR_RECORDED_PIPELINE_SHA

Also remember that advanced merge-request pipeline types can have source/revision semantics that differ from a simple branch pipeline. Chapter 16 handles those details. In Chapter 01, the lesson is simply: preserve the SHA before reasoning from a moving ref.

6. Failure mode: runner starts, then the script fails

Once the job is running, execution evidence exists. Read from the first failing command rather than the last generic error. Check working directory, tool availability, quoting, exit codes, network dependencies, and the exact runtime image/executor. Do not hide a meaningful failure with an unconditional || true.

set -eu
printf 'sha=%s\n' "$CI_COMMIT_SHA"
printf 'pwd=%s\n' "$PWD"
git --version
git rev-parse HEAD
test "$CI_COMMIT_SHA" = "$(git rev-parse HEAD)"

This narrow probe is safer than dumping the environment. It verifies assumptions directly and lets the original non-zero exit status remain visible.

7. Failure mode: green script, missing artifact

A script can exit successfully and still fail to produce the file that the artifact configuration expects. GitLab Runner reports messages such as “No files to upload” when paths do not match actual output. Treat artifact upload as a separate transition after script execution.

make_evidence:
  script:
    - mkdir -p evidence
    - printf 'sha=%s\n' "$CI_COMMIT_SHA" > evidence/identity.txt
  artifacts:
    paths:
      # Intentional typo: the script created identity.txt.
      - evidence/identty.txt

The causal repair is to make the producer path and artifact path agree, then verify the uploaded artifact from the job UI. Do not interpret a green command as proof that GitLab stored the intended file.

8. Failure mode: deployment job is green, target is unhealthy

This chapter does not deploy anything, but the diagnostic pattern belongs in the foundation. A shell command can receive HTTP 202/200 from a deployment API and exit 0 while rollout continues asynchronously. GitLab then knows what the job reported; it does not magically know whether users can reach a healthy application.

Do not stop evidence at request acceptance
            flowchart TD
              A[Deploy job] --> B[Provider accepts request]
              B --> C[GitLab job succeeds]
              B --> D[Provider rollout continues]
              D --> E[Target health checks]
              E --> F[Healthy]
              E --> G[Unhealthy / rollback / incident]
          

Production pipelines should correlate deployment request/resource identity with target health and a defined recovery procedure. “Green CI” and “healthy service” are different assertions.

9. Unsafe troubleshooting shortcuts increase blast radius

Shortcut Why it is unsafe Safer diagnostic pattern
Print all variables Can expose tokens, credentials, and internal context. Print an allowlist of non-secret IDs and metadata.
Use a broad PAT because job identity is denied Hides the authorization design problem and expands privilege/lifetime. Read the denial, identify required resource/scope, choose narrow identity.
Disable TLS verification Turns trust validation into silent interception risk. Fix CA/trust/network configuration.
Run untrusted code on Shell/privileged runner Can expose host/network/other-job state. Use isolated disposable execution appropriate to trust level.
Delete/recreate resource blindly Destroys evidence and may lose unrelated state. Inspect exact identity, apply smallest reversible correction.
Retry until green Can repeat side effects and hide intermittent cause. Preserve first failure, classify retry safety, retry smallest scope.

10. Runner failures are also security signals

Current GitLab documentation warns that Shell executor jobs have high security risk for untrusted builds because they execute with the runner user's host permissions and can potentially access other project state. Non-privileged Docker can provide stronger isolation, but privileged containers, host socket mounts, broad capabilities, or a powerful Kubernetes service account can collapse that boundary.

Therefore “fix the runner” should always include a trust question: what authority would this job gain if we made it run? Do not resolve a scheduling problem by accidentally granting untrusted branch or fork code access to a production-capable execution path.

11. Performance: measure queue time separately from execution time

A pipeline may be slow because jobs wait for runners, because one job has a long critical path, because artifact transfer is large, because an external dependency is slow, or because a test itself is expensive. These causes require different fixes.

Time Likely owner Useful evidence
Pipeline creation / graph delay Configuration/control plane Creation timestamps, configuration complexity, rules/includes.
Queued → started Runner scheduling/capacity Job queued/start timestamps, runner availability.
Job runtime Executor + script/tool workload Trace timestamps, tool profiling, runtime image/tool versions.
Artifact transfer Runner/network/storage Upload/download trace and artifact size.
Deployment wait External provider/target Provider resource IDs and target timestamps.

More runners do not fix a CPU-bound test; more cache does not fix a policy wait; more parallel jobs can increase queue pressure. Measure before tuning. Chapter 35 owns systematic optimization.

12. Retries and reruns can repeat side effects

An observation or test job is often safe to retry because it is read-only or produces pipeline-scoped evidence. A release, package publish, infrastructure apply, or deployment may not be safe to repeat blindly. Before retrying a side-effecting job, determine what already completed, whether the operation is idempotent, and how the external resource is identified.

Preserve the first failure before retrying. A successful second attempt does not erase the operational question of why attempt one failed or whether it partially changed an external system.

13. Compact troubleshooting playbook

Symptom First evidence Do not start with
No pipeline Source/ref/SHA + config + workflow logic. Runner logs.
Pipeline, missing job Compiled graph + job rules. Shell command changes.
Pending job Runner matching/capacity/protection. Application code.
Running job fails First failing command + runtime assumptions. Deleting pipeline history.
Missing artifact Producer path + upload trace + artifact config. Unrelated runner scaling.
Green deploy, bad service Provider rollout and target health. Rebuilding the artifact.

Summary

GitLab CI/CD diagnostics becomes tractable when you keep pipeline creation, job inclusion, scheduling, execution, evidence storage, deployment recording, and target health separate. Preserve the first failure, identify the last proven state, and investigate the next transition. Security controls are part of troubleshooting: never trade runner isolation, credential scope, TLS, or audit evidence for a quick green result.

Next lesson

Checkpoint Lab

Build a two-stage disposable pipeline, predict its state changes, trigger it from push and the web UI, preserve an evidence packet, verify artifact transfer across jobs, inject one controlled failure, and explain exactly what each green status proves.

Knowledge check

The YAML is valid, but a push creates no pipeline. Is the runner your first suspect?

A job exists but is pending because it requests a tag no runner has. What is the causal fix?

A job script succeeded, but the artifact path has a typo. What does the green script prove?

Why is printenv a poor default debugging command in CI?

Why can a blind retry be dangerous for deployment or release jobs?

Official references and version notes

Version and compatibility note

Diagnostic UI details and runner capabilities evolve. Failure layers and current runner-security statements in this lesson were checked against primary GitLab documentation on 2026-09-11. Revalidate executor and pipeline-rule behavior before applying the playbook to a later platform version.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.