GitLab Runners, Executors, Tags, Registration, Autoscaling, Isolation, and Security: Diagnostics, Failure Modes, Security, and Performance
Diagnose pending jobs, protected-runner mismatches, obsolete registration workflows, stale runner state, privileged executor risk, and unsafe cross-project trust without weakening controls.
Learning objectives
- Apply a preserve → scope → inspect → correct → verify diagnostic sequence to runner failures.
- Distinguish tag, scope, protection, pause/status, capacity, and configuration failures from CI YAML failures.
- Interpret obsolete legacy registration failures and migrate to runner authentication tokens instead of weakening instance policy.
- Recognize privileged/persistent runner compromise paths and contain them before restoring capacity.
- Detect and remove stale runner records and local runner-manager configuration during decommissioning.
1. Diagnostic sequence: preserve evidence before changing the runner
- Preserve: pipeline ID, job ID/status, pipeline source, ref/SHA, requested tags, job trace if one exists, and runner inventory.
- Scope: offering, project/group, runner type, protected-ref state, membership/role, executor, host/network, and whether the job ever started.
- Inspect: runner visibility, pause/status, tags, run-untagged, protection, project enablement, host service logs/version, and capacity.
- Correct minimally: fix the specific mismatched dimension; do not disable protection or broaden scope as a generic repair.
- Verify: rerun only the synthetic job and correlate the new job with the expected runner ID/ref/SHA.
3. Failure B: protected runner cannot execute an unprotected branch
A deployment runner is project-scoped, online, and has the correct tags, yet a feature-branch job remains pending. Inspection shows the runner is Protected. This is expected behavior: protected runners are intended for protected refs (with current MR-specific conditions).
Wrong repair: uncheck Protected just to make the job run. Correct repair: keep deployment capacity protected and run feature-branch validation on a separate low-privilege runner. If the job genuinely belongs on the deployment runner, first prove the ref/merge-request satisfies the protected-ref policy.
4. Failure C: tags match, but runner is paused/offline/stale
| Observed state | Meaning | Repair direction |
|---|---|---|
| paused=true | Runner intentionally ignores new jobs. | Confirm maintenance owner; resume only when the host is safe. |
| offline | Runner has not contacted GitLab recently and is unavailable. | Check runner service, network/DNS/TLS, host health, then verify. |
| stale | Runner has not contacted GitLab for an extended period. | Determine whether host still exists; decommission stale record/config if not. |
| never_contacted | Runner object exists but manager has never connected. | Complete registration/startup or delete the unused runner object. |
5. Failure D: a runner exists, but not for this project
A group may show inherited runners while a project has instance runners disabled; a project runner may not be enabled for a second project; a locked project runner cannot be enabled for additional projects. Verify the runner's actual assignment path before editing job tags.
When sharing a project runner, remember that runner configuration changes affect every project using that runner. That is a governance reason to prefer narrow ownership for sensitive workloads.
6. Failure E: obsolete registration instructions return 410 Gone
An old automation script runs:
gitlab-runner register \
--registration-token "$OLD_REGISTRATION_TOKEN" \
--executor shell \
--tag-list legacy
On a current instance where legacy registration is disabled,
registration can fail with
410 Gone - runner registration disallowed. This is not
an executor failure.
Repair: create the runner in GitLab (UI or
supported API), capture the runner authentication token securely,
register the manager with --token, and move
tags/protection/run-untagged metadata to the GitLab-side runner
configuration. Legacy registration tokens are scheduled for removal
in GitLab 20.0.
7. Failure F: untrusted MR code reaches privileged/persistent infrastructure
GitLab warns that self-managed runners are a remote-code-execution service. A privileged Docker job can potentially gain root-equivalent host capability; a Shell job directly shares the host. If a compromised runner had internal network access, mounted credentials, SSH keys, cloud metadata access, or shared caches, treat those as potential exposure paths.
Recovery should rebuild or reimage the disposable runner host rather than “cleaning” an untrusted persistent host in place when compromise cannot be ruled out.
8. Fork/MR trust: repository ownership does not equal execution trust
Fork and merge-request code is user-controlled input. Never route it to a persistent privileged runner with production secrets or internal-network access. Protected variables and protected runners are useful controls, but they must be combined with project/ref policy and careful pipeline-source design from Chapters 7, 12, and 13.
9. Runner authentication token exposure
If the runner host filesystem or config.toml is
exposed, treat the runner authentication token as compromised.
Pause/isolate the runner, rotate or replace the runner
authentication token using supported runner/API controls, and
investigate whether unauthorized runner managers appeared. Never
attach config.toml to an issue or CI artifact.
10. Stale runner after host decommissioning
Deleting a VM does not automatically prove the GitLab runner object
and local configuration lifecycle are clean. Conversely, deleting a
runner object in GitLab can leave a local
config.toml entry that continues contacting the
instance.
For runners created with authentication tokens, GitLab documents
runner objects and runner managers separately:
gitlab-runner unregister can remove the manager
association while the runner object remains. Your decommission
checklist must intentionally remove both sides and verify the
result.
11. Performance diagnosis without disabling security controls
If a job is eligible but queues for a long time, separate matching from capacity. Check runner online state, concurrent job limits, fleet utilization, autoscaler provisioning health, cloud quota, image pulls, and executor startup time. Increasing runner scope, enabling untagged jobs, or turning off protection can hide a capacity problem by letting the wrong runners accept work.
For autoscalers, watch provisioning errors and cloud resource limits. For persistent runners, watch CPU/memory/disk pressure and cleanup. For hosted runners, verify current service capacity and account compute entitlements rather than hard-coding minute quotas into runbooks.
12. Intentionally broken example: repair one tag mismatch only
Start from a disposable pending job tagged
ch14-no-such-runner. Capture the job page/API status
and visible runner tag lists. Then repair by editing only the
synthetic job:
# Before: impossible selector
tags: [ch14-no-such-runner]
# After: either no tags for an available untagged hosted runner,
# or the exact tag of your disposable isolated runner.
# Do not edit production runner policy to satisfy the lab.
Commit the repair, record the new SHA, and verify the new job either runs on the predicted runner or remains pending for a different, now-visible reason. Preserve the original failed job as evidence; do not rewrite its history.
13. Runner failure decision tree
job pending?
-> no visible runner for project? inspect scope/assignment
-> runner paused/offline/stale? repair runner/host availability
-> job has tags runner lacks? repair capability/routing
-> job untagged but runner disallows? choose correct runner/policy
-> protected mismatch? keep trust boundary; route elsewhere
-> all eligibility matches? inspect capacity/autoscaler
job started and failed?
-> now inspect executor preparation, image, shell, network, script, artifacts
(this is no longer a runner-selection failure)
14. Safe rollback and evidence checklist
- Cancel only synthetic pending jobs after capturing their state.
- Restore/remove synthetic CI tags rather than weakening runner protection.
- Pause suspect runners before forensic work; isolate compromised hosts.
- Rotate/revoke exposed runner/CI credentials before cleaning logs/history.
- For disposable runner removal, unregister the manager and delete the GitLab runner object.
- Verify no stale runner record, local config volume, VM, cache bucket/object, or firewall exception remains from the lab.
Knowledge check
A job is pending and no job log exists. What layer should you diagnose first?
Runner eligibility and scheduling: scope, pause/status, tags, untagged policy, protection, and capacity. The executor script has not started.
What does 410 Gone during legacy runner registration usually indicate on a current instance?
The deprecated registration-token workflow is disabled. Migrate to runner creation plus runner authentication token registration.
Why is unprotecting a deployment runner a poor fix for a feature-branch pending job?
It weakens a trust boundary and lets less-trusted refs reach sensitive capacity. Route the feature job to separate low-privilege capacity instead.
What should you do first if untrusted code may have run on a privileged runner?
Pause/isolate the runner and host, then rotate/revoke potentially exposed credentials and investigate. Do not retry on the same host.
Why can deleting the VM alone leave runner governance incomplete?
The GitLab runner object/manager association and local configuration lifecycle may remain; both sides require explicit cleanup/verification.
If tags, scope, protection, and status all match but jobs queue, what class of problem remains?
Capacity/performance: runner concurrency, busy workers, autoscaler provisioning, cloud quota, startup/image/network delays, or hosted-service entitlement/capacity.
Summary
Runner diagnosis is most effective when it respects layer boundaries. A pending job is a selection/capacity problem until proven otherwise. Preserve the failed state, identify exactly which eligibility predicate is false, make the smallest correction, and verify on a new commit/job without weakening protected or privileged infrastructure.
Official references
- GitLab Docs — Get started with GitLab Runner
- GitLab Docs — Manage runners and runner scope
- GitLab Docs — Configure runners, tags, and protected runners
- GitLab Docs — New runner creation/registration workflow
- GitLab Docs — Register runners
- GitLab Docs — Runner commands and unregister behavior
- GitLab Docs — Runner executors
- GitLab Docs — Docker executor
- GitLab Docs — Shell executor
- GitLab Docs — Kubernetes executor
- GitLab Docs — Docker Autoscaler executor
- GitLab Docs — Instance executor
- GitLab Docs — Self-managed runner security
- GitLab Docs — Runner fleet scaling
- GitLab Docs — Instance-group autoscaler
- GitLab Docs — GitLab-hosted runners
- GitLab Docs — Hosted runners on Linux for GitLab.com
- GitLab Docs — Hosted runners for GitLab Dedicated
- GitLab Docs — Runners API
- GitLab Docs — Token overview / runner authentication tokens
- GitLab Docs — CI/CD YAML tags
- GitLab Docs — Protected branches and CI/CD
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.