GitLab Runner Architecture, Registration, Runner Managers, Job Execution, and Lifecycle: Diagnostics, Failure Modes, Security, and Performance
A stuck job is not automatically a broken runner, and an offline runner is not automatically a network outage. This lesson diagnoses tag/scope/protection mismatches, stale or unreachable managers, exposed runner authentication tokens, over-broad runner assignment, and workspace residue by preserving pipeline/job/runner evidence before changing infrastructure.
Learning objectives
- Diagnose a job that is stuck because no runner satisfies all tags, scope, or protection requirements.
- Separate a paused runner configuration from an offline/stale runner manager and from a job-level execution failure.
- Respond to a suspected runner authentication-token exposure by preserving evidence and rotating/revoking the token instead of merely editing job YAML.
- Identify how over-broad scope and persistent host residue can leak code or credentials across otherwise unrelated jobs.
- Repair the smallest causal layer while preserving original pipeline/job/runner-manager evidence and avoiding blind re-registration.
1. Diagnostic sequence from job graph to runner manager
flowchart TD
A[Preserve pipeline/job/SHA] --> B[Does job exist and remain pending?]
B --> C[Compare requested tags / scope / protection / paused]
C --> D[Inspect runner managers: online/offline/stale/version/system_id]
D --> E[Inspect manager service/network/logs]
E --> F[Inspect executor preparation]
F --> G[Inspect script/tool/network failure]
G --> H[Apply smallest correction]
H --> I[Retry only the safe scope and compare evidence]
If the job is not in the pipeline, return to Chapters 01–02: no runner can execute a job that configuration/rules never created. If the job exists but is pending, runner eligibility/capacity becomes the next layer.
2. Failure mode: no runner satisfies all job tags
# Runner configuration tags:
# ch04-lab, linux
stuck_probe:
tags:
- ch04-lab
- gpu
script:
- echo "this should remain pending"
The YAML is valid. The runner may be online. The job still remains
pending because the runner lacks gpu. Preserve the
pending job ID and both tag sets. Repair the causal intent: either
remove an accidental job tag or provision an authorized runner that
actually has the capability. Do not relabel arbitrary hardware as
gpu merely to make the pipeline green.
3. Failure mode: stale or offline manager
A runner can be unpaused with correct tags while its only manager is offline. Check manager contact time/status, then the service/container/VM and network path to GitLab. For a containerized lab manager, inspect:
docker ps -a --filter "name=^/ch04-runner-manager$"
docker logs --tail 120 ch04-runner-manager
# If the container is intentionally stopped, record that fact before starting it.
Do not immediately run register again. Re-registration
can create duplicate managers/config entries and obscure the
original outage. First determine whether the expected manager
already exists and why it stopped contacting GitLab.
4. Failure mode: paused runner mistaken for an outage
Pause is a GitLab-side administrative state. The manager may be
healthy and contacting GitLab, but new jobs are intentionally
ignored. This is useful during maintenance or retirement. Compare
the runner’s paused field/UI state with manager
online/offline state before touching the host.
Current API guidance deprecates the older active field
in favor of paused. New automation should model the
current field rather than perpetuating legacy terminology.
5. Failure mode: runner authentication token exposed
Suppose a troubleshooting transcript accidentally includes the runner authentication token. The incident is not fixed by deleting the transcript alone. The credential may have been copied and could be used to clone the runner identity.
- Preserve incident metadata without redistributing the secret.
- Restrict/disable the affected runner if needed to stop new work.
- Reset/rotate the runner authentication token through current GitLab UI/API guidance.
- Update authorized manager configuration securely.
- Verify managers reconnect with expected system IDs/versions and old credentials no longer authenticate.
- Review jobs and manager activity for unexpected execution.
6. Failure mode: runner scope is broader than intended
A project-specific signing/test runner that was accidentally made available to a whole group can execute code from more repositories than its credentials/network were designed for. The fix is authorization design, not job retries.
Before narrowing scope, inventory jobs/projects that legitimately use the runner. Then create/migrate capacity deliberately and remove assignments without breaking unrelated pipelines. For high-value capabilities, prefer dedicated tags, protected refs, narrow scope, and isolated workers.
7. Failure mode: shared host residue between jobs
Persistent runners can retain Git checkouts, caches, temporary files, language package state, container layers, credentials accidentally written by scripts, and host-level changes. Shell executor is especially risky because jobs run directly with the runner user’s host permissions.
Symptoms can be subtle: a build succeeds only because a prior job installed a dependency; a secret file exists from an earlier pipeline; one project reads another project’s workspace. The safest response for untrusted/high-risk workloads is usually stronger isolation or disposable workers, not an ever-growing cleanup script.
git clean -fdx /, blanket home-directory
deletion, or similarly broad cleanup commands.
Cleanup must be scoped to known runner build/cache paths and tested
against the executor model.
8. Failure mode: version drift or incompatible features
Record GitLab Runner version before blaming YAML. New GitLab features can require matching Runner support, and Runner updates can change executor/helper behavior. The first troubleshooting step in GitLab Runner documentation is to confirm GitLab and Runner version compatibility.
gitlab-runner --version
# Or for the disposable container:
docker exec ch04-runner-manager gitlab-runner --version
Upgrade through a controlled canary/staging process; do not update every production runner during an incident without a rollback plan and evidence.
9. Performance symptom: queue delay is not necessarily execution slowness
Separate queue time from job runtime. A job can wait because no eligible manager is idle even when its script takes seconds. Adding more tags can make eligibility narrower and increase queueing; enabling untagged work can increase contention and trust exposure. Measure queue depth, manager utilization and job runtime before tuning.
10. Intentionally broken diagnostic exercise
runner_diagnostic_probe:
tags:
- ch04-lab
- nonexistent-capability
script:
- printf 'runner=%s
' "$CI_RUNNER_ID"
Expected evidence: pipeline exists; job exists; job remains pending; no job trace begins; the known runner can be online and unpaused; runner tags do not satisfy the job. Repair only the fake tag, create a new pipeline, and preserve the old pending job/pipeline as evidence of the causal mismatch.
11. Recovery acceptance criteria
- The original failing/pending pipeline and job IDs are preserved.
- The cause is assigned to configuration creation, runner eligibility, manager availability, executor, or job runtime—not vaguely “CI.”
- The repair changes only the responsible layer.
- No new broad token, runner assignment, privileged mode, untagged access, or network reachability was introduced as a shortcut.
- The repaired run records exact SHA, runner ID/version and relevant manager identity.
- Old exposed credentials, duplicate managers, or disposable resources are actually retired.
Knowledge check
A job is pending and the matching runner is online. What should you compare before restarting Runner?
Requested job tags, runner tags, scope/assignment, protected-ref
eligibility, run_untagged, and paused state.
Why is re-registering an offline runner a poor first diagnostic step?
It can create duplicate manager/configuration state and hide the original service/network cause.
A runner token was printed once but the log was deleted. Is rotation still necessary?
Yes. Deleting the visible copy does not prove the secret was never captured; treat the credential as exposed and rotate/revoke it.
What is the security problem with a persistent shared Shell runner?
Jobs execute with host user permissions and can leave/read cross-job state, so malicious or buggy code can compromise other workloads and the host/network.
Why separate queue time from runtime?
Runner eligibility/capacity can dominate waiting even when the job command itself is fast; the fixes are different.
Official references and version notes
- Runners — runner categories, job scheduling, GitLab-hosted versus self-managed runners, and execution flow.
- Manage runners — project/group/instance scope, creation workflow, ownership and pause/resume operations.
-
Configure runners
— tags,
run_untagged, protected runners, authentication-token rotation, and routing behavior. - Registering runners and new runner creation workflow — current runner authentication-token registration and deprecated legacy registration-token behavior.
-
GitLab Runner commands
—
register,list,verify,run,stop, andunregisterlifecycle commands. -
Runner fleet planning
— manager
system_ididentity and modern unregister/delete distinctions. - Runners API — runner details, managers, status, pause, job history, authentication-token reset, and deletion semantics.
- Security for self-managed runners — remote-code-execution trust, Shell executor risk, persistent-runner residue, isolation, and credential exposure.
- GitLab Runner documentation — current compatibility guidance recommends keeping Runner major.minor aligned with GitLab; GitLab.com users should keep self-managed runners current.
Version-sensitive statements were rechecked against current
primary GitLab documentation and the GitLab Runner release history
on 2026-09-11. The latest stable Runner tag
visible in the upstream release history at that verification point
is v19.3.1 (2026-08-24); GitLab 19.4 is scheduled
after this guide-authoring date, so executable examples pin
gitlab/gitlab-runner:v19.3.1 instead of a moving
latest tag. The legacy runner-registration-token
workflow is deprecated and scheduled for removal in GitLab 20.0;
this chapter uses runner authentication tokens and the modern
creation workflow.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.