GitLab CI/CD Foundations: .gitlab-ci.yml, Pipelines, Jobs, Stages, and Execution Model: Configuration, Design Choices, and Tradeoffs
Choose stages, pipeline surfaces, and runner ownership deliberately by balancing clarity, trust, portability, isolation, capacity, compatibility, maintainability, and cost.
Learning objectives
- Choose a stages-first pipeline when clarity is more valuable than premature DAG optimization.
- Compare GitLab-hosted and self-managed runners across trust, isolation, maintenance, compatibility, capacity, and cost.
- Choose branch/push versus merge-request pipeline context based on the evidence the workflow actually needs.
- Separate GitLab platform CI/CD concepts taught here from deeper syntax/ecosystem material delegated to the dedicated GitLab CI/CD course.
- Use a decision table that explicitly addresses maintainability, security, governance, reliability, compatibility, performance, and cost.
.gitlab-ci.yml, Pipeline Editor/CI Lint,
jobs, stages, ordinary branch and merge-request pipelines, predefined
variables, artifacts, runner scheduling, and
CI_JOB_TOKEN are available on GitLab Free across
GitLab.com, Self-Managed, and Dedicated. GitLab-hosted runners are
provided for GitLab.com and GitLab Dedicated; Self-Managed
installations provide and operate their own runner capacity. Hosted
compute quotas, credits, machine types, and billing can change, so
this chapter never hard-codes them and provides a validation-only path
whenever a runner is unavailable. Current glab uses the
glab ci command family; older pipe/pipeline
aliases are deprecated.
1. CI/CD design starts with evidence, not YAML cleverness
A pipeline should answer a delivery question with the least complex trustworthy execution model. “Can this commit build?” “Does this merge request satisfy pre-merge checks?” and “Can this release be deployed?” are different questions. Choosing stages, pipeline source, runner ownership, and identity scope follows from the question and trust boundary.
In Chapter 10, prefer simple, visible causal chains. Chapter 15 introduces DAG optimization; Chapters 12–13 introduce richer creation rules and inputs; Chapter 14 focuses on runners; Chapters 16–21 deepen artifacts, environments, composition, reuse, performance, and cost. Do not import those mechanisms prematurely merely because GitLab supports them.
2. Stages first; DAG later when dependency evidence justifies it
Stages are easy to reason about: jobs within a stage can run in
parallel and the next stage waits for the prior stage to succeed.
This may serialize independent work, but it gives beginners and
operators a clear failure boundary. A DAG using
needs can shorten critical paths by expressing
job-to-job dependencies directly, but it adds graph complexity and
artifact/data-flow considerations.
| Choice | Use when | Benefit | Cost/risk |
|---|---|---|---|
| Stages-first | Small pipelines, onboarding, clear gates, modest duration. | Simple ordering and diagnosis. | Can wait unnecessarily for unrelated jobs. |
| DAG / needs | Measured latency matters and true dependencies are understood. | Jobs can start as soon as dependencies complete. | More graph/data-flow complexity; taught in Chapter 15. |
3. GitLab-hosted versus self-managed runners
A runner decision determines where untrusted repository code executes and what network/filesystem/credentials it can reach. GitLab-hosted runners reduce infrastructure operations; self-managed runners provide environment/network/customization control but make your organization responsible for hardening, patching, isolation, scaling, and credential hygiene.
| Dimension | GitLab-hosted runner | Self-managed runner |
|---|---|---|
| Maintainability | Provider operates runner infrastructure. | You own installation, updates, autoscaling, retirement, and recovery. |
| Security/isolation | Provider-defined ephemeral/isolation model; verify current runner type. | You choose executor/isolation; a poor shell/privileged/shared-host design can leak job data. |
| Governance | Project/group policy still controls eligibility; provider infra is external trust. | Can place runners in controlled networks but must govern who may send jobs there. |
| Reliability | Depends on service availability and current capacity/quota. | Depends on your fleet capacity, health, networking, and operations. |
| Compatibility | Standard published environments/machine types. | Can support custom OS, tools, hardware, or private network access. |
| Performance | Elastic options may exist; startup/runtime characteristics vary. | Can tune hardware/cache/network locality but must measure and scale it. |
| Cost | Hosted compute policy/credits/billing can change; verify current plan. | Infrastructure + operational labor + idle capacity are your responsibility. |
4. Push/branch pipeline versus merge-request pipeline
A push pipeline commonly has
CI_PIPELINE_SOURCE=push and evaluates the pushed ref. A
merge-request pipeline uses merge_request_event and
exposes MR-specific context. Merely opening an MR does not magically
transform every branch pipeline into an MR pipeline; the
configuration must opt into MR pipeline behavior through matching
rules or equivalent current configuration.
| Question | Prefer | Reason |
|---|---|---|
| Does every pushed branch remain buildable? | Push/branch pipeline | Direct evidence tied to the branch commit. |
| Should checks run in MR-specific context and use MR metadata? | Merge-request pipeline | Provides MR source/event context and supports MR-specific rules. |
| Do I need the source combined with latest target? | Not answered by ordinary MR pipeline alone | That moves toward merged-results pipelines/merge trains, covered around integration strategy and tier-dependent behavior. |
5. Pipeline-source policy belongs at creation time
workflow:rules controls whether a pipeline is created.
Job rules decides whether a job is added to an
already-created pipeline. Keep those layers separate. A common
design bug is enabling both push and MR pipelines unintentionally
for the same branch and then interpreting duplicate results as
independent evidence.
# Illustration only; Chapter 12 develops rules systematically.
workflow:
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
- if: '$CI_PIPELINE_SOURCE == "push"'
show_source:
script:
- printf 'source=%s sha=%s\n' "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA"
6. Human-triggered context versus automation identity
The pipeline source and triggering actor influence authorization. A
CI_JOB_TOKEN has a restricted resource set and inherits
effective capability from the user who triggered the pipeline,
subject to job-token policy. That is safer than blindly reusing a
long-lived personal token, but it is not a permission bypass.
Design the job so the minimum required capability is obvious. A test job should not need repository push or deployment authority. Later chapters introduce variables/secrets and deployment environments; this chapter’s safest job has no external write access at all.
7. Integrated GitLab CI/CD versus the dedicated CI/CD course
This GitLab course needs substantial CI/CD knowledge because pipelines are woven into merge governance, registries, releases, security, environments, and automation. Chapters 10–21 therefore teach the execution and platform integration model needed to operate GitLab safely.
The separate GitLab CI/CD course can go deeper into specialized syntax patterns, ecosystem integrations, advanced reusable architectures, runner fleets, performance tuning, and pipeline engineering without forcing every GitLab platform learner through all niche details. The boundary is deliberate: learn enough here to reason about GitLab as a delivery platform, then deepen CI/CD specialization separately.
8. Worked decision: a small service team
Scenario: six developers maintain an ordinary web service. They need fast pre-merge confidence, have no private-network test dependency, and do not want to administer runner infrastructure yet.
| Decision | Selected approach | Justification |
|---|---|---|
| Pipeline graph | Two or three simple stages | Maintainable and easy to diagnose; no measured DAG need yet. |
| Pipeline context | Push for branch smoke checks + MR-specific checks only when needed | Evidence is explicit by source; avoid accidental duplicate pipelines. |
| Runner | GitLab-hosted when available | No private-network requirement; minimizes infrastructure toil. |
| Identity | CI_JOB_TOKEN only where documented; otherwise no credential | Short-lived and scoped; test jobs should not acquire write authority. |
| Artifact | Small, bounded test report/evidence only | Keeps traceability without turning artifacts into bulk storage. |
| Cost control | Tiny jobs + current quota check + validation fallback | Avoids hard-coded compute entitlements and unnecessary usage. |
9. Design anti-patterns to reject
- “Put every job in one stage for speed.” This erases useful gates and dependency meaning.
- “Use our production shell runner because it is already online.” Availability is not a trust justification.
- “Give CI a broad PAT so every integration works.” Long-lived ambient authority expands blast radius.
- “A branch pipeline is enough for all MR decisions.” It may lack the MR source/context your policy expects.
- “Optimize DAG/cache before measuring.” Complexity without bottleneck evidence reduces maintainability.
Knowledge check
When is a stages-first design usually preferable to needs-based DAG execution?
When the pipeline is small, sequencing is understandable, and there is no measured latency problem that justifies extra graph complexity.
Why might a self-managed runner be justified despite higher operational cost?
It can provide controlled private-network access, custom hardware/tools, or environment requirements that hosted runners do not satisfy.
Why can enabling both push and merge_request_event indiscriminately create confusion?
The same branch/MR activity can create multiple pipelines whose results look redundant unless their purpose is deliberately separated.
What is the security flaw in giving a test job a broad long-lived PAT “for convenience”?
The job gains authority unrelated to testing, increasing the impact of malicious or compromised CI configuration.
What should trigger a move from stages to a DAG?
Measured critical-path or concurrency needs plus a clear model of true job dependencies, not stylistic preference.
Summary
CI/CD architecture is a set of policy choices: graph complexity, pipeline source, runner ownership, and job identity. Start with a simple pipeline whose evidence is easy to explain; add DAGs, specialized runners, or broader integration only when requirements justify their trust, operational, and cost consequences.
Official references
- GitLab Docs — CI/CD
- GitLab Docs — Get started with GitLab CI/CD
- GitLab Docs — CI/CD YAML syntax reference
- GitLab Docs — Validate CI/CD configuration
- GitLab Docs — Pipeline editor
- GitLab Docs — Pipelines
- GitLab Docs — CI/CD jobs
- GitLab Docs — Job execution flow
- GitLab Docs — Predefined CI/CD variables
- GitLab Docs — CI/CD variables
- GitLab Docs — CI/CD job token
- GitLab Docs — Runners
- GitLab Docs — Configure runners
- GitLab Docs — GitLab-hosted runners
- GitLab Docs — Pipelines API
- GitLab Docs — Jobs API
- GitLab Docs — CI Lint API
- GitLab Docs — Job rules and CI_PIPELINE_SOURCE
- GitLab CLI — ci
- GitLab CLI — ci list
- GitLab CLI — ci get
- GitLab CLI — ci status
- GitLab CLI — runner list
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.