Chapter 05Lesson 04~140 minutes

Runner Executors: Shell, Docker, Docker Autoscaler, Kubernetes, SSH, and Custom Execution Models: Diagnostics, Failure Modes, Security, and Performance

Executor failures often look like application failures until you identify the actual boundary. This lesson diagnoses unsafe Shell use, privileged Docker/container breakout, Docker-socket exposure, Kubernetes over-privilege, image/shell mismatches, autoscaler provisioning failures, and the false assumption that a container is automatically a hard multi-tenant security boundary.

DiagnosticsPrivileged modeDocker socketKubernetes RBACImage failures

Learning objectives

  • Diagnose whether a failure belongs to job routing, executor preparation, image/shell startup, service/networking, job script, or teardown rather than changing unrelated pipeline YAML.
  • Recognize privileged Docker, host PID/network/socket exposure, and broad Kubernetes RBAC as boundary collapses rather than convenience settings.
  • Preserve executor/worker evidence before retries so ephemeral container/pod/instance failures remain diagnosable.
  • Repair image and routing failures by correcting the smallest causal configuration instead of weakening isolation or granting broad privilege.
  • Explain why “it ran in a container” is insufficient evidence for safe multi-tenant execution.
Diagnostic rule: preserve the first failed job’s pipeline/job ID, runner/manager identity, executor, worker/image identity, timestamps, trace, and executor-side logs/events before retrying. Ephemeral workers can disappear, and a retry can hide the original infrastructure state.

1. Failure taxonomy: locate the boundary before changing configuration

Symptom First layer to inspect Do not start by
Job remains pending Runner scope/tags/protection/pause/capacity. Changing image or script commands.
Job assigned, image cannot be pulled Executor preparation, registry/network/auth/image reference. Adding retries everywhere.
Container starts but shell/entrypoint fails Image shell/entrypoint compatibility. Granting privileged mode.
Service hostname exists but app cannot connect Service readiness/port/protocol/DNS/network. Assuming helper health check proves application readiness.
Kubernetes pod unschedulable Events, quota/resources/node selectors/taints/admission. Editing test code.
Autoscaler never provisions worker Manager/Fleeting/provider IAM/quota/policy/capacity. Re-registering the runner blindly.
Script completes, cleanup/artifact phase fails Post-build transfer/teardown/storage. Declaring the test itself failed.

2. Evidence-first executor diagnostic sequence

  1. Record pipeline source/ref/SHA, pipeline ID and failing job ID.
  2. Confirm the compiled job exists and is eligible; preserve its requested tags and timeout.
  3. Record runner ID, manager/system identity, Runner version and executor.
  4. Identify the concrete worker: host/container/pod/instance and image/OS.
  5. Preserve executor preparation logs, image-pull output, Kubernetes events or Fleeting/provider errors.
  6. Only after the worker is healthy, inspect script exit/status and service readiness.
  7. Preserve artifact/cache/report upload failures separately from script failures.
  8. Apply the least-privilege correction to the smallest causal layer.
  9. Retry only the smallest safe scope, noting that retries can create new workers and new external conditions.

3. Failure mode: untrusted code reaches a Shell executor

This is a routing/trust-policy failure even if every command “works.” A public fork or broadly writable repository should not be able to schedule arbitrary scripts onto a persistent host that contains sensitive code, credentials, sockets, devices or internal network reachability.

Preserve: pipeline/job/source identity, runner ID/tags/scope/protection, which ref/user triggered it, and host/manager logs. Contain: pause the affected runner if necessary, rotate credentials reachable from that host according to exposure assessment, and isolate/rebuild a potentially compromised worker. Repair: narrow scope/tags/protected refs and move untrusted work to isolated container/pod/ephemeral workers.

Do not “prove” Shell risk by exfiltrating local files or scanning internal services. The architecture and documented host-local execution model are sufficient to justify the control.

4. Failure mode: privileged Docker collapses the intended boundary

GitLab warns that privileged Docker gives the container host-level capabilities and can enable container breakout. Host PID namespace and similarly broad configuration also undermine isolation. Treat a privileged build runner as special-purpose infrastructure exposed to repository-controlled code.

If privileged execution is genuinely required—for example, a narrowly controlled nested-container workflow—separate it from general jobs, restrict it to protected/trusted refs, use isolated ephemeral VMs when practical, limit network/credentials, and destroy/reimage the worker after use. Never turn on privileged mode simply because a tool returned “permission denied.”

# Anti-pattern: do not use as a troubleshooting shortcut.
[runners.docker]
  privileged = true

# Safer diagnostic questions first:
# - What exact syscall/device/socket does the tool require?
# - Can the build use a rootless/daemonless builder?
# - Can capability/volume access be narrowed?
# - Can privileged work move to a dedicated one-use worker?

5. Failure mode: host Docker socket exposure is mistaken for ordinary container isolation

Binding /var/run/docker.sock into a job container lets commands talk to the host Docker daemon. A job may then start sibling containers, mount host files, or otherwise control resources far beyond its own container. The fact that the job container itself is not marked privileged does not make the host-daemon access low privilege.

Diagnose what the workload actually needs. Prefer a purpose-built image builder or a dedicated isolated builder architecture rather than giving general test jobs host-daemon authority. If socket access is unavoidable, treat the runner host as fully within the job’s trust boundary and restrict who can schedule work.

6. Failure mode: Kubernetes runner has more cluster authority than the job model needs

The runner manager needs Kubernetes API permissions to create/watch/delete job resources. Those permissions should be scoped to the runner’s namespace and required resource types. Broad cluster-admin credentials or hostPath/privileged pod settings can turn an ordinary CI workload into cluster-level risk.

When a pod fails, preserve Kubernetes events and the runner-manager log before deletion. If the error is Forbidden, identify the exact API verb/resource/namespace being denied; grant only the permission required by the documented executor feature instead of replacing the service account with cluster-admin.

# Read-only diagnostics in an authorized runner namespace.
kubectl -n "$RUNNER_NS" get pod "$JOB_POD" -o wide
kubectl -n "$RUNNER_NS" describe pod "$JOB_POD"
kubectl -n "$RUNNER_NS" get events --sort-by=.lastTimestamp

# Inspect authorization without granting it.
kubectl auth can-i create pods   --as=system:serviceaccount:"$RUNNER_NS":"$RUNNER_SA"   -n "$RUNNER_NS"

7. Failure mode: wrong image entrypoint or missing shell

GitLab Runner expects job images to satisfy executor-specific shell/entrypoint requirements. A minimal image may contain the application binary but no compatible shell. A custom entrypoint can also prevent Runner from starting the generated job script correctly.

Preserve the image reference/digest and exact executor error. Reproduce locally with the same image and inspect its configured entrypoint/shell. Repair the image or executor-compatible entrypoint rather than broadening privileges.

IMAGE="registry.example.invalid/lab/tool@sha256:<digest>"
# In a real authorized environment, inspect—not mutate—the image metadata.
docker image inspect "$IMAGE"   --format 'entrypoint={{json .Config.Entrypoint}} cmd={{json .Config.Cmd}} user={{.Config.User}}'

# Verify the expected shell exists if the executor requires it.
docker run --rm --entrypoint /bin/sh "$IMAGE" -c 'printf "shell-ok\n"' 

8. Failure mode: autoscaler/provider state is hidden behind “stuck job”

With Docker Autoscaler or Instance, there is an extra control chain: Runner manager → task scaling policy → Fleeting plugin → cloud provider/resource → instance connectivity → job environment. Capacity can fail because of provider IAM, quota, image availability, networking, SSH/connectivity, instance startup, plugin mismatch, or competing management of the same autoscaling resource.

Preserve runner-manager and Fleeting/plugin logs, instance/group events, exact worker image/version, provider error code, and capacity policy. Do not respond by making max_instances unbounded or sharing one autoscaling group across multiple runner configurations; current GitLab guidance requires dedicated autoscaling resources per Docker Autoscaler configuration.

9. Performance: isolation, reuse, and queue time trade against one another

Executor performance includes more than job CPU time. Measure queue time, worker provisioning time, image-pull time, source/artifact/cache transfer, service readiness, script runtime, teardown, and idle capacity. A persistent Shell host may start instantly but accumulate drift/residue. A one-use VM may isolate better but add boot latency. Kubernetes may schedule quickly when nodes are warm but queue under quota or node pressure.

Optimize from evidence: pre-bake stable toolchains into verified worker/images, use caches only for performance, bound autoscaler idle capacity, separate large/specialized workloads by routing tags, and retain enough metrics to see whether bottlenecks are runner capacity, image pulls, network, or the job itself.

10. Intentionally broken example: image failure, not script failure

broken_image_probe:
  tags: [ch05-docker]
  image: alpine:does-not-exist-ch05
  script:
    - echo "You should never see this line"

Expected evidence: the pipeline and job are created; an eligible Docker runner accepts the job; executor preparation attempts to pull the image; image resolution/pull fails; the script never starts. Preserve job ID, runner ID/executor, image reference and first error.

fixed_image_probe:
  tags: [ch05-docker]
  image: alpine:3.22
  script:
    - printf 'sha=%s job=%s runner=%s\n' "$CI_COMMIT_SHA" "$CI_JOB_ID" "$CI_RUNNER_ID"
    - cat /etc/os-release | sed -n '1,4p' 

The repair changes only the invalid image reference. It does not add privileged mode, change runner scope, ignore failures, or retry indefinitely.

11. Recovery checklist

  • Preserve first-failure executor/worker evidence before retry.
  • Correct routing before changing executor configuration.
  • Correct image/shell compatibility before granting privilege.
  • Correct specific Kubernetes RBAC/resources before broadening cluster authority.
  • Correct Fleeting/provider capacity/authentication before increasing limits.
  • Destroy/reimage workers that may have executed hostile privileged code.
  • Verify teardown: no orphan container/pod/instance, unexpected volume, or stale credential remains.
  • Record the residual risk; “green after retry” is not evidence that the original boundary was safe.
Next lesson

Checkpoint Lab

Run one trusted workload through two disposable execution models, inject an executor-specific failure, preserve the evidence, and justify the production choice.

Knowledge check

A Docker job fails “image not found” before script:. Which layer failed?

Why is mounting the host Docker socket a major trust decision even without privileged=true?

A Kubernetes executor gets Forbidden: cannot create pods. What is the safe first fix?

Why preserve autoscaler logs before retrying a failed job?

A privileged job may have been compromised. Is deleting only the container sufficient?

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

  • GitLab Runner executors — current executor selection guidance, compatibility matrix, actively developed paths, and maintenance-mode executors.
  • Shell executor — host-local execution, shell selection, process termination, maintenance-mode status, and security warnings.
  • Docker executor — image/service model, container lifecycle, volumes, image pull behavior, and executor configuration.
  • Docker Autoscaler executor — Fleeting-based autoscaling, Docker feature compatibility, autoscaling resources, capacity and ephemeral-worker patterns.
  • Instance executor — Fleeting-based instance provisioning, full host access, worker images, capacity, and native-step considerations.
  • Kubernetes executor — per-job Pod creation, build/helper/service containers, RBAC requirements, entrypoint behavior, and executor configuration.
  • SSH executor and Custom executor — exceptional/maintenance-mode execution models and operational constraints.
  • Security for self-managed runners — Shell risk, Docker privilege/capability guidance, network segmentation, and trust-boundary recommendations.
  • Use Docker to build Docker images — security and executor implications for Docker-in-Docker and socket/pipe binding.
Verification note

Executor behavior and status were rechecked against current primary GitLab documentation on 2026-09-11. At verification time, Docker, Docker Autoscaler, Instance, and Kubernetes are actively developed executor paths; Shell, SSH, VirtualBox, Parallels, and Custom are maintenance mode, and Docker Machine is deprecated. Docker Autoscaler is documented as generally available since GitLab Runner 17.1 and uses Fleeting plugins. Re-check these facts before future course revisions.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.