Chapter 14Lesson 04~255 minutes

GitLab Runners, Executors, Tags, Registration, Autoscaling, Isolation, and Security: Diagnostics, Failure Modes, Security, and Performance

Diagnose pending jobs, protected-runner mismatches, obsolete registration workflows, stale runner state, privileged executor risk, and unsafe cross-project trust without weakening controls.

Diagnostics410 GoneProtected runnersPrivileged modeStale runnersSecurity

Learning objectives

  • Apply a preserve → scope → inspect → correct → verify diagnostic sequence to runner failures.
  • Distinguish tag, scope, protection, pause/status, capacity, and configuration failures from CI YAML failures.
  • Interpret obsolete legacy registration failures and migrate to runner authentication tokens instead of weakening instance policy.
  • Recognize privileged/persistent runner compromise paths and contain them before restoring capacity.
  • Detect and remove stale runner records and local runner-manager configuration during decommissioning.
Availability baseline (verified 2026-08-21 against current GitLab documentation). GitLab Runner, project/group/instance runner scopes, tags, protected runners, the supported executor framework, and self-managed runners are available across Free/Premium/Ultimate. GitLab-hosted runners are a GitLab.com service and are also available for GitLab Dedicated under a separately provisioned Limited Availability offering; Self-Managed installations provide their own runner infrastructure. Hosted-runner compute can be quota/billing constrained, so every mandatory exercise has a no-runner/static or evidence-fixture path. Legacy runner registration tokens are deprecated and scheduled for removal in GitLab 20.0; this chapter teaches the modern runner creation workflow with runner authentication tokens.

1. Diagnostic sequence: preserve evidence before changing the runner

  1. Preserve: pipeline ID, job ID/status, pipeline source, ref/SHA, requested tags, job trace if one exists, and runner inventory.
  2. Scope: offering, project/group, runner type, protected-ref state, membership/role, executor, host/network, and whether the job ever started.
  3. Inspect: runner visibility, pause/status, tags, run-untagged, protection, project enablement, host service logs/version, and capacity.
  4. Correct minimally: fix the specific mismatched dimension; do not disable protection or broaden scope as a generic repair.
  5. Verify: rerun only the synthetic job and correlate the new job with the expected runner ID/ref/SHA.

2. Failure A: job pending because tags do not match

Broken job:

gpu_test:
  tags: [linux, gpu]
  script:
    - ./run-gpu-test.sh

Runner inventory contains [linux,docker] but no gpu. The job never reaches an executor. There is no shell exit code because no script has run. GitLab's matching rule requires one runner to have every job tag.

Repair: either route the job to genuine GPU capacity or remove the gpu requirement if the workload does not need it. Do not rename a generic runner tag to fake hardware capability.

3. Failure B: protected runner cannot execute an unprotected branch

A deployment runner is project-scoped, online, and has the correct tags, yet a feature-branch job remains pending. Inspection shows the runner is Protected. This is expected behavior: protected runners are intended for protected refs (with current MR-specific conditions).

Wrong repair: uncheck Protected just to make the job run. Correct repair: keep deployment capacity protected and run feature-branch validation on a separate low-privilege runner. If the job genuinely belongs on the deployment runner, first prove the ref/merge-request satisfies the protected-ref policy.

4. Failure C: tags match, but runner is paused/offline/stale

Observed state Meaning Repair direction
paused=true Runner intentionally ignores new jobs. Confirm maintenance owner; resume only when the host is safe.
offline Runner has not contacted GitLab recently and is unavailable. Check runner service, network/DNS/TLS, host health, then verify.
stale Runner has not contacted GitLab for an extended period. Determine whether host still exists; decommission stale record/config if not.
never_contacted Runner object exists but manager has never connected. Complete registration/startup or delete the unused runner object.

5. Failure D: a runner exists, but not for this project

A group may show inherited runners while a project has instance runners disabled; a project runner may not be enabled for a second project; a locked project runner cannot be enabled for additional projects. Verify the runner's actual assignment path before editing job tags.

When sharing a project runner, remember that runner configuration changes affect every project using that runner. That is a governance reason to prefer narrow ownership for sensitive workloads.

6. Failure E: obsolete registration instructions return 410 Gone

An old automation script runs:

gitlab-runner register \
  --registration-token "$OLD_REGISTRATION_TOKEN" \
  --executor shell \
  --tag-list legacy

On a current instance where legacy registration is disabled, registration can fail with 410 Gone - runner registration disallowed. This is not an executor failure.

Repair: create the runner in GitLab (UI or supported API), capture the runner authentication token securely, register the manager with --token, and move tags/protection/run-untagged metadata to the GitLab-side runner configuration. Legacy registration tokens are scheduled for removal in GitLab 20.0.

7. Failure F: untrusted MR code reaches privileged/persistent infrastructure

Contain first. If you suspect a privileged or persistent runner executed malicious/untrusted code, pause the runner and isolate the host/network before collecting further evidence. Rotate/revoke credentials that may have been exposed. Do not simply retry the pipeline on the same host.

GitLab warns that self-managed runners are a remote-code-execution service. A privileged Docker job can potentially gain root-equivalent host capability; a Shell job directly shares the host. If a compromised runner had internal network access, mounted credentials, SSH keys, cloud metadata access, or shared caches, treat those as potential exposure paths.

Recovery should rebuild or reimage the disposable runner host rather than “cleaning” an untrusted persistent host in place when compromise cannot be ruled out.

8. Fork/MR trust: repository ownership does not equal execution trust

Fork and merge-request code is user-controlled input. Never route it to a persistent privileged runner with production secrets or internal-network access. Protected variables and protected runners are useful controls, but they must be combined with project/ref policy and careful pipeline-source design from Chapters 7, 12, and 13.

9. Runner authentication token exposure

If the runner host filesystem or config.toml is exposed, treat the runner authentication token as compromised. Pause/isolate the runner, rotate or replace the runner authentication token using supported runner/API controls, and investigate whether unauthorized runner managers appeared. Never attach config.toml to an issue or CI artifact.

10. Stale runner after host decommissioning

Deleting a VM does not automatically prove the GitLab runner object and local configuration lifecycle are clean. Conversely, deleting a runner object in GitLab can leave a local config.toml entry that continues contacting the instance.

For runners created with authentication tokens, GitLab documents runner objects and runner managers separately: gitlab-runner unregister can remove the manager association while the runner object remains. Your decommission checklist must intentionally remove both sides and verify the result.

11. Performance diagnosis without disabling security controls

If a job is eligible but queues for a long time, separate matching from capacity. Check runner online state, concurrent job limits, fleet utilization, autoscaler provisioning health, cloud quota, image pulls, and executor startup time. Increasing runner scope, enabling untagged jobs, or turning off protection can hide a capacity problem by letting the wrong runners accept work.

For autoscalers, watch provisioning errors and cloud resource limits. For persistent runners, watch CPU/memory/disk pressure and cleanup. For hosted runners, verify current service capacity and account compute entitlements rather than hard-coding minute quotas into runbooks.

12. Intentionally broken example: repair one tag mismatch only

Start from a disposable pending job tagged ch14-no-such-runner. Capture the job page/API status and visible runner tag lists. Then repair by editing only the synthetic job:

# Before: impossible selector
tags: [ch14-no-such-runner]

# After: either no tags for an available untagged hosted runner,
# or the exact tag of your disposable isolated runner.
# Do not edit production runner policy to satisfy the lab.

Commit the repair, record the new SHA, and verify the new job either runs on the predicted runner or remains pending for a different, now-visible reason. Preserve the original failed job as evidence; do not rewrite its history.

13. Runner failure decision tree

job pending?
  -> no visible runner for project?       inspect scope/assignment
  -> runner paused/offline/stale?         repair runner/host availability
  -> job has tags runner lacks?           repair capability/routing
  -> job untagged but runner disallows?   choose correct runner/policy
  -> protected mismatch?                  keep trust boundary; route elsewhere
  -> all eligibility matches?             inspect capacity/autoscaler

job started and failed?
  -> now inspect executor preparation, image, shell, network, script, artifacts
     (this is no longer a runner-selection failure)

14. Safe rollback and evidence checklist

  • Cancel only synthetic pending jobs after capturing their state.
  • Restore/remove synthetic CI tags rather than weakening runner protection.
  • Pause suspect runners before forensic work; isolate compromised hosts.
  • Rotate/revoke exposed runner/CI credentials before cleaning logs/history.
  • For disposable runner removal, unregister the manager and delete the GitLab runner object.
  • Verify no stale runner record, local config volume, VM, cache bucket/object, or firewall exception remains from the lab.

Knowledge check

A job is pending and no job log exists. What layer should you diagnose first?

What does 410 Gone during legacy runner registration usually indicate on a current instance?

Why is unprotecting a deployment runner a poor fix for a feature-branch pending job?

What should you do first if untrusted code may have run on a privileged runner?

Why can deleting the VM alone leave runner governance incomplete?

If tags, scope, protection, and status all match but jobs queue, what class of problem remains?

Summary

Runner diagnosis is most effective when it respects layer boundaries. A pending job is a selection/capacity problem until proven otherwise. Preserve the failed state, identify exactly which eligibility predicate is false, make the smallest correction, and verify on a new commit/job without weakening protected or privileged infrastructure.

Official references

Next lesson

Checkpoint: prove the runner control plane end to end

Lesson 5 combines inventory, prediction, a normal route, a deliberately impossible route, evidence capture, optional isolated registration, and complete cleanup into one production-style runner governance exercise.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.