Continuous Integration and Delivery Foundations, GitLab Pipeline Architecture, and Delivery Flow: Diagnostics, Failure Modes, Security, and Performance
Troubleshoot GitLab CI/CD by following evidence from pipeline source and compiled configuration through job creation, queueing, runner/executor assignment, scripts, artifacts, and external targets. You will diagnose realistic failures without leaking secrets, weakening runner isolation, destroying first-failure evidence, or using blind retries as a substitute for root-cause analysis.
Learning objectives
- Apply an evidence-first diagnostic sequence from pipeline creation through runner execution and external target state.
- Diagnose valid YAML that creates no desired pipeline or job, and distinguish that from runner or script failure.
- Recognize pending-job runner problems, ref/SHA confusion, missing artifacts, and green-job/failed-target mismatches.
- Avoid unsafe troubleshooting shortcuts such as broad tokens, secret dumps, privileged runners, TLS disablement, and blind retries.
- Separate queue time, execution time, and external wait time before attempting performance optimization.
1. Preserve first-failure evidence before changing anything
The fastest-looking troubleshooting action—rerun, edit, delete, retry—can destroy the evidence that distinguishes one cause from another. Before changing configuration, preserve the pipeline ID, source, ref, SHA, compiled configuration or lint result, relevant job IDs/statuses, the first failing trace, runner identity if assigned, and any artifact/deployment/external resource identifiers already created.
Then move through the system in causal order:
- Source/ref/SHA: what actually attempted to create the pipeline?
- Configuration: which CI configuration was resolved and was it valid?
-
Pipeline decision: did
workflowlogic allow a pipeline? - Job graph: which jobs exist, and which were omitted/skipped?
- Queue/runner: did an eligible runner claim the job?
- Executor/runtime: was the environment prepared as expected?
- Script/tool/network: where is the first non-zero or unexpected result?
- Artifact/report/cache: did the expected durable evidence get created and uploaded?
- Environment/deployment/external target: did the side effect really complete and become healthy?
2. Failure mode: valid YAML, but no push pipeline
The following configuration is syntactically valid, but it deliberately allows only scheduled pipelines. A developer can push it, see no new push pipeline, and incorrectly blame the runner. The runner is not involved because no eligible push pipeline is created.
workflow:
rules:
- if: '$CI_PIPELINE_SOURCE == "schedule"'
- when: never
observe:
script:
- echo "This job can exist only in an allowed pipeline."
Use CI Lint and pipeline simulation to inspect this before touching runner settings. Current GitLab CI Lint can simulate pipeline creation and expose rule/graph problems. For a push-oriented lab, the minimal repair could explicitly allow push and web sources:
workflow:
rules:
- if: '$CI_PIPELINE_SOURCE == "push"'
- if: '$CI_PIPELINE_SOURCE == "web"'
- when: never
This example teaches a layer boundary: YAML validity, pipeline creation, and runner execution are three different states. A green linter is useful but does not guarantee that the current source matches a rule.
3. Failure mode: pipeline exists, but the expected job does not
Pipeline-level logic may allow creation while job-level
rules omit one particular job. If the pipeline exists
but the job is absent from the graph, inspect job rules and the
actual CI_PIPELINE_SOURCE/ref context. Do not wait for
a runner to execute a job that does not exist.
pipeline_identity:
script:
- echo "pipeline exists"
manual_only_check:
script:
- echo "included only for a web-created pipeline"
rules:
- if: '$CI_PIPELINE_SOURCE == "web"'
In a push pipeline, manual_only_check is absent by
design. That is not “skipped after execution”; it was not included
in the graph. Precise status language avoids needless runner
debugging.
4. Failure mode: job exists but stays pending
A pending job has advanced farther: GitLab created the pipeline and the job. Now inspect matching runners. Runner tags, scope, protected status, pause/offline state, executor availability, and capacity can all affect whether a job starts. A deliberately impossible tag demonstrates the symptom without changing any runner:
pending_demo:
tags:
- course-runner-tag-that-does-not-exist
script:
- echo "This line will not run without a matching runner."
In a disposable test, the correct repair is to remove/correct the unnecessary job tag or use an explicitly authorized runner whose tags match the intended workload. The wrong repair is to register a new privileged production runner, make a protected runner unprotected, or broaden runner access without understanding the trust consequence.
5. Failure mode: “GitLab ran the wrong commit”
A common investigation compares an old pipeline to the
current head of a branch. The branch may have moved after
the pipeline was created. That does not change the historical
pipeline SHA. Compare the pipeline's recorded commit with the job's
CI_COMMIT_SHA and repository history, not with an
assumption that main still points where it did hours
ago.
# In a local clone, compare immutable identities.
git rev-parse HEAD
git rev-parse origin/main
git show --no-patch --oneline YOUR_RECORDED_PIPELINE_SHA
Also remember that advanced merge-request pipeline types can have source/revision semantics that differ from a simple branch pipeline. Chapter 16 handles those details. In Chapter 01, the lesson is simply: preserve the SHA before reasoning from a moving ref.
6. Failure mode: runner starts, then the script fails
Once the job is running, execution evidence exists. Read from the
first failing command rather than the last generic error. Check
working directory, tool availability, quoting, exit codes, network
dependencies, and the exact runtime image/executor. Do not hide a
meaningful failure with an unconditional
|| true.
set -eu
printf 'sha=%s\n' "$CI_COMMIT_SHA"
printf 'pwd=%s\n' "$PWD"
git --version
git rev-parse HEAD
test "$CI_COMMIT_SHA" = "$(git rev-parse HEAD)"
This narrow probe is safer than dumping the environment. It verifies assumptions directly and lets the original non-zero exit status remain visible.
7. Failure mode: green script, missing artifact
A script can exit successfully and still fail to produce the file that the artifact configuration expects. GitLab Runner reports messages such as “No files to upload” when paths do not match actual output. Treat artifact upload as a separate transition after script execution.
make_evidence:
script:
- mkdir -p evidence
- printf 'sha=%s\n' "$CI_COMMIT_SHA" > evidence/identity.txt
artifacts:
paths:
# Intentional typo: the script created identity.txt.
- evidence/identty.txt
The causal repair is to make the producer path and artifact path agree, then verify the uploaded artifact from the job UI. Do not interpret a green command as proof that GitLab stored the intended file.
8. Failure mode: deployment job is green, target is unhealthy
This chapter does not deploy anything, but the diagnostic pattern belongs in the foundation. A shell command can receive HTTP 202/200 from a deployment API and exit 0 while rollout continues asynchronously. GitLab then knows what the job reported; it does not magically know whether users can reach a healthy application.
flowchart TD
A[Deploy job] --> B[Provider accepts request]
B --> C[GitLab job succeeds]
B --> D[Provider rollout continues]
D --> E[Target health checks]
E --> F[Healthy]
E --> G[Unhealthy / rollback / incident]
Production pipelines should correlate deployment request/resource identity with target health and a defined recovery procedure. “Green CI” and “healthy service” are different assertions.
9. Unsafe troubleshooting shortcuts increase blast radius
| Shortcut | Why it is unsafe | Safer diagnostic pattern |
|---|---|---|
| Print all variables | Can expose tokens, credentials, and internal context. | Print an allowlist of non-secret IDs and metadata. |
| Use a broad PAT because job identity is denied | Hides the authorization design problem and expands privilege/lifetime. | Read the denial, identify required resource/scope, choose narrow identity. |
| Disable TLS verification | Turns trust validation into silent interception risk. | Fix CA/trust/network configuration. |
| Run untrusted code on Shell/privileged runner | Can expose host/network/other-job state. | Use isolated disposable execution appropriate to trust level. |
| Delete/recreate resource blindly | Destroys evidence and may lose unrelated state. | Inspect exact identity, apply smallest reversible correction. |
| Retry until green | Can repeat side effects and hide intermittent cause. | Preserve first failure, classify retry safety, retry smallest scope. |
10. Runner failures are also security signals
Current GitLab documentation warns that Shell executor jobs have high security risk for untrusted builds because they execute with the runner user's host permissions and can potentially access other project state. Non-privileged Docker can provide stronger isolation, but privileged containers, host socket mounts, broad capabilities, or a powerful Kubernetes service account can collapse that boundary.
Therefore “fix the runner” should always include a trust question: what authority would this job gain if we made it run? Do not resolve a scheduling problem by accidentally granting untrusted branch or fork code access to a production-capable execution path.
11. Performance: measure queue time separately from execution time
A pipeline may be slow because jobs wait for runners, because one job has a long critical path, because artifact transfer is large, because an external dependency is slow, or because a test itself is expensive. These causes require different fixes.
| Time | Likely owner | Useful evidence |
|---|---|---|
| Pipeline creation / graph delay | Configuration/control plane | Creation timestamps, configuration complexity, rules/includes. |
| Queued → started | Runner scheduling/capacity | Job queued/start timestamps, runner availability. |
| Job runtime | Executor + script/tool workload | Trace timestamps, tool profiling, runtime image/tool versions. |
| Artifact transfer | Runner/network/storage | Upload/download trace and artifact size. |
| Deployment wait | External provider/target | Provider resource IDs and target timestamps. |
More runners do not fix a CPU-bound test; more cache does not fix a policy wait; more parallel jobs can increase queue pressure. Measure before tuning. Chapter 35 owns systematic optimization.
12. Retries and reruns can repeat side effects
An observation or test job is often safe to retry because it is read-only or produces pipeline-scoped evidence. A release, package publish, infrastructure apply, or deployment may not be safe to repeat blindly. Before retrying a side-effecting job, determine what already completed, whether the operation is idempotent, and how the external resource is identified.
Preserve the first failure before retrying. A successful second attempt does not erase the operational question of why attempt one failed or whether it partially changed an external system.
13. Compact troubleshooting playbook
| Symptom | First evidence | Do not start with |
|---|---|---|
| No pipeline | Source/ref/SHA + config + workflow logic. | Runner logs. |
| Pipeline, missing job | Compiled graph + job rules. | Shell command changes. |
| Pending job | Runner matching/capacity/protection. | Application code. |
| Running job fails | First failing command + runtime assumptions. | Deleting pipeline history. |
| Missing artifact | Producer path + upload trace + artifact config. | Unrelated runner scaling. |
| Green deploy, bad service | Provider rollout and target health. | Rebuilding the artifact. |
Summary
GitLab CI/CD diagnostics becomes tractable when you keep pipeline creation, job inclusion, scheduling, execution, evidence storage, deployment recording, and target health separate. Preserve the first failure, identify the last proven state, and investigate the next transition. Security controls are part of troubleshooting: never trade runner isolation, credential scope, TLS, or audit evidence for a quick green result.
Knowledge check
The YAML is valid, but a push creates no pipeline. Is the runner your first suspect?
No. Inspect pipeline source/ref/SHA and
workflow:rules first. A runner cannot execute a
pipeline that was never created.
A job exists but is pending because it requests a tag no runner has. What is the causal fix?
Correct the job's intended runner/tag relationship or provide an authorized matching runner. Do not broaden privileges or expose a protected production runner merely to make it start.
A job script succeeded, but the artifact path has a typo. What does the green script prove?
Only that the script commands completed successfully. Artifact creation/upload is a later state and must be verified separately.
Why is printenv a poor default debugging command
in CI?
It can print credentials and sensitive variables. Use a narrow allowlist of non-secret metadata and direct hypothesis tests.
Why can a blind retry be dangerous for deployment or release jobs?
The first attempt may have partially completed external side effects. Without idempotency and exact resource identity, retrying can duplicate or corrupt state.
Official references and version notes
- Get started with GitLab CI/CD — current first-principles overview of pipeline configuration, jobs, stages, and runners.
- CI/CD pipelines — pipeline sources, basic/stage/DAG pipeline forms, and current pipeline behavior.
- CI/CD YAML syntax reference — authoritative keyword semantics and compatibility notes.
- Predefined CI/CD variables — availability phases and runtime identity variables such as source, SHA, pipeline, job, and runner metadata.
- Runners — current runner scheduling and execution model.
- Validate GitLab CI/CD configuration — validation and simulation used to separate configuration problems from execution problems.
- Specify when jobs run with rules — current ordered rule evaluation and pipeline-source behavior.
- Troubleshooting job artifacts — current artifact retrieval/upload failure guidance.
- Security for self-managed runners — runner trust and executor risks.
- Shell executor — current maintenance-mode status and untrusted-build warning.
Diagnostic UI details and runner capabilities evolve. Failure layers and current runner-security statements in this lesson were checked against primary GitLab documentation on 2026-09-11. Revalidate executor and pipeline-rule behavior before applying the playbook to a later platform version.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.