Chapter 04Lesson 04~125 minutes

GitLab Runner Architecture, Registration, Runner Managers, Job Execution, and Lifecycle: Diagnostics, Failure Modes, Security, and Performance

A stuck job is not automatically a broken runner, and an offline runner is not automatically a network outage. This lesson diagnoses tag/scope/protection mismatches, stale or unreachable managers, exposed runner authentication tokens, over-broad runner assignment, and workspace residue by preserving pipeline/job/runner evidence before changing infrastructure.

DiagnosticsStuck jobsOffline runnersToken exposureWorkspace residue

Learning objectives

  • Diagnose a job that is stuck because no runner satisfies all tags, scope, or protection requirements.
  • Separate a paused runner configuration from an offline/stale runner manager and from a job-level execution failure.
  • Respond to a suspected runner authentication-token exposure by preserving evidence and rotating/revoking the token instead of merely editing job YAML.
  • Identify how over-broad scope and persistent host residue can leak code or credentials across otherwise unrelated jobs.
  • Repair the smallest causal layer while preserving original pipeline/job/runner-manager evidence and avoiding blind re-registration.
Evidence-first rule. Before restarting, re-registering, rotating, deleting, or broadening a runner, preserve the pipeline/job ID, exact SHA, requested tags, runner ID, manager system ID/version/status where available, and the first runner/job error. Infrastructure retries can erase the state that explains the incident.

1. Diagnostic sequence from job graph to runner manager

Runner diagnosis — do not skip eligibility
            flowchart TD
              A[Preserve pipeline/job/SHA] --> B[Does job exist and remain pending?]
              B --> C[Compare requested tags / scope / protection / paused]
              C --> D[Inspect runner managers: online/offline/stale/version/system_id]
              D --> E[Inspect manager service/network/logs]
              E --> F[Inspect executor preparation]
              F --> G[Inspect script/tool/network failure]
              G --> H[Apply smallest correction]
              H --> I[Retry only the safe scope and compare evidence]
          

If the job is not in the pipeline, return to Chapters 01–02: no runner can execute a job that configuration/rules never created. If the job exists but is pending, runner eligibility/capacity becomes the next layer.

2. Failure mode: no runner satisfies all job tags

# Runner configuration tags:
#   ch04-lab, linux

stuck_probe:
  tags:
    - ch04-lab
    - gpu
  script:
    - echo "this should remain pending"

The YAML is valid. The runner may be online. The job still remains pending because the runner lacks gpu. Preserve the pending job ID and both tag sets. Repair the causal intent: either remove an accidental job tag or provision an authorized runner that actually has the capability. Do not relabel arbitrary hardware as gpu merely to make the pipeline green.

3. Failure mode: stale or offline manager

A runner can be unpaused with correct tags while its only manager is offline. Check manager contact time/status, then the service/container/VM and network path to GitLab. For a containerized lab manager, inspect:

docker ps -a --filter "name=^/ch04-runner-manager$"
docker logs --tail 120 ch04-runner-manager
# If the container is intentionally stopped, record that fact before starting it.

Do not immediately run register again. Re-registration can create duplicate managers/config entries and obscure the original outage. First determine whether the expected manager already exists and why it stopped contacting GitLab.

4. Failure mode: paused runner mistaken for an outage

Pause is a GitLab-side administrative state. The manager may be healthy and contacting GitLab, but new jobs are intentionally ignored. This is useful during maintenance or retirement. Compare the runner’s paused field/UI state with manager online/offline state before touching the host.

Current API guidance deprecates the older active field in favor of paused. New automation should model the current field rather than perpetuating legacy terminology.

5. Failure mode: runner authentication token exposed

Suppose a troubleshooting transcript accidentally includes the runner authentication token. The incident is not fixed by deleting the transcript alone. The credential may have been copied and could be used to clone the runner identity.

  1. Preserve incident metadata without redistributing the secret.
  2. Restrict/disable the affected runner if needed to stop new work.
  3. Reset/rotate the runner authentication token through current GitLab UI/API guidance.
  4. Update authorized manager configuration securely.
  5. Verify managers reconnect with expected system IDs/versions and old credentials no longer authenticate.
  6. Review jobs and manager activity for unexpected execution.
Never place a real runner authentication token in a CI variable just so a job can “fix” its own runner. That crosses the infrastructure control plane into repository-executed code and creates a credential-exfiltration path.

6. Failure mode: runner scope is broader than intended

A project-specific signing/test runner that was accidentally made available to a whole group can execute code from more repositories than its credentials/network were designed for. The fix is authorization design, not job retries.

Before narrowing scope, inventory jobs/projects that legitimately use the runner. Then create/migrate capacity deliberately and remove assignments without breaking unrelated pipelines. For high-value capabilities, prefer dedicated tags, protected refs, narrow scope, and isolated workers.

7. Failure mode: shared host residue between jobs

Persistent runners can retain Git checkouts, caches, temporary files, language package state, container layers, credentials accidentally written by scripts, and host-level changes. Shell executor is especially risky because jobs run directly with the runner user’s host permissions.

Symptoms can be subtle: a build succeeds only because a prior job installed a dependency; a secret file exists from an earlier pipeline; one project reads another project’s workspace. The safest response for untrusted/high-risk workloads is usually stronger isolation or disposable workers, not an ever-growing cleanup script.

Do not use git clean -fdx /, blanket home-directory deletion, or similarly broad cleanup commands. Cleanup must be scoped to known runner build/cache paths and tested against the executor model.

8. Failure mode: version drift or incompatible features

Record GitLab Runner version before blaming YAML. New GitLab features can require matching Runner support, and Runner updates can change executor/helper behavior. The first troubleshooting step in GitLab Runner documentation is to confirm GitLab and Runner version compatibility.

gitlab-runner --version
# Or for the disposable container:
docker exec ch04-runner-manager gitlab-runner --version

Upgrade through a controlled canary/staging process; do not update every production runner during an incident without a rollback plan and evidence.

9. Performance symptom: queue delay is not necessarily execution slowness

Separate queue time from job runtime. A job can wait because no eligible manager is idle even when its script takes seconds. Adding more tags can make eligibility narrower and increase queueing; enabling untagged work can increase contention and trust exposure. Measure queue depth, manager utilization and job runtime before tuning.

10. Intentionally broken diagnostic exercise

runner_diagnostic_probe:
  tags:
    - ch04-lab
    - nonexistent-capability
  script:
    - printf 'runner=%s
' "$CI_RUNNER_ID"

Expected evidence: pipeline exists; job exists; job remains pending; no job trace begins; the known runner can be online and unpaused; runner tags do not satisfy the job. Repair only the fake tag, create a new pipeline, and preserve the old pending job/pipeline as evidence of the causal mismatch.

11. Recovery acceptance criteria

  • The original failing/pending pipeline and job IDs are preserved.
  • The cause is assigned to configuration creation, runner eligibility, manager availability, executor, or job runtime—not vaguely “CI.”
  • The repair changes only the responsible layer.
  • No new broad token, runner assignment, privileged mode, untagged access, or network reachability was introduced as a shortcut.
  • The repaired run records exact SHA, runner ID/version and relevant manager identity.
  • Old exposed credentials, duplicate managers, or disposable resources are actually retired.
Next lesson

Checkpoint Lab — Runner Architecture and Lifecycle

Provision and retire one disposable runner end to end, proving exact routing and documenting the controls required before organizational rollout.

Knowledge check

A job is pending and the matching runner is online. What should you compare before restarting Runner?

Why is re-registering an offline runner a poor first diagnostic step?

A runner token was printed once but the log was deleted. Is rotation still necessary?

What is the security problem with a persistent shared Shell runner?

Why separate queue time from runtime?

Official references and version notes

  • Runners — runner categories, job scheduling, GitLab-hosted versus self-managed runners, and execution flow.
  • Manage runners — project/group/instance scope, creation workflow, ownership and pause/resume operations.
  • Configure runners — tags, run_untagged, protected runners, authentication-token rotation, and routing behavior.
  • Registering runners and new runner creation workflow — current runner authentication-token registration and deprecated legacy registration-token behavior.
  • GitLab Runner commands — register, list, verify, run, stop, and unregister lifecycle commands.
  • Runner fleet planning — manager system_id identity and modern unregister/delete distinctions.
  • Runners API — runner details, managers, status, pause, job history, authentication-token reset, and deletion semantics.
  • Security for self-managed runners — remote-code-execution trust, Shell executor risk, persistent-runner residue, isolation, and credential exposure.
  • GitLab Runner documentation — current compatibility guidance recommends keeping Runner major.minor aligned with GitLab; GitLab.com users should keep self-managed runners current.
Version and compatibility note

Version-sensitive statements were rechecked against current primary GitLab documentation and the GitLab Runner release history on 2026-09-11. The latest stable Runner tag visible in the upstream release history at that verification point is v19.3.1 (2026-08-24); GitLab 19.4 is scheduled after this guide-authoring date, so executable examples pin gitlab/gitlab-runner:v19.3.1 instead of a moving latest tag. The legacy runner-registration-token workflow is deprecated and scheduled for removal in GitLab 20.0; this chapter uses runner authentication tokens and the modern creation workflow.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.