Checkpoint Lab — Runner Executors: Shell, Docker, Docker Autoscaler, Kubernetes, SSH, and Custom Execution Models
The checkpoint runs the same synthetic trusted test through two disposable execution models, captures comparable evidence, injects one routing/image failure, repairs the causal layer, and produces a production recommendation. The goal is not to crown one executor universally “best,” but to justify the safer choice for a stated workload and trust model.
Learning objectives
- Execute the same trusted synthetic test through two disposable execution models or an equivalent faithful local simulation.
- Predict and verify worker identity, filesystem persistence, image/runtime identity, privilege, network, and teardown differences.
- Inject one executor-specific routing or image failure, preserve first-failure evidence, and repair only the causal layer.
- Produce a comparable evidence packet tied to source SHA, pipeline/job IDs, runner/executor identity, worker/image identity, logs, and cleanup.
- Write a production recommendation that explicitly states workload trust, isolation requirement, operational ownership, elasticity, cost, and residual risk.
1. Mission and success criteria
Your team has one small deterministic test that can run in more than one environment. You must compare two execution models, not merely obtain two green statuses. The preferred live pair is a narrowly scoped disposable Shell runner and a non-privileged Docker runner; if registering runners is not appropriate, the mandatory local path compares host-style execution with a disposable Docker container and uses supplied GitLab evidence fixtures.
Success means you can answer: what source ran, which runner/manager accepted it, which executor/worker environment executed it, what persisted after completion, what privilege/network/mount boundary existed, how failure appeared, how it was repaired, and which model is safer for the stated workload.
2. Preflight and predictions before execution
Record the GitLab offering/tier, project path, lab branch/SHA, Runner version(s), executor names, and image identity. Then write predictions before running:
- The Shell-style run will use a persistent host/process boundary and can leave workspace state unless cleanup removes it.
- The Docker run will use the configured image and a disposable container; container deletion will not erase deliberately mounted/persisted storage.
- A job requesting a nonexistent runner tag will remain pending before executor preparation begins.
- A Docker job with a nonexistent image tag will be assigned but fail before the user script emits output.
# Local simulation preflight
LAB_ROOT="${TMPDIR:-/tmp}/ch05-checkpoint"
mkdir -p "$LAB_ROOT/evidence" "$LAB_ROOT/docker-output"
git rev-parse HEAD 2>/dev/null || true
docker version --format 'server={{.Server.Version}}'
docker ps -a --filter 'name=^/ch05-checkpoint-'
3. Live GitLab path: same trusted probe, two tagged executors
stages: [probe]
.executor_probe:
stage: probe
script:
- set -eu
- printf 'source=%s sha=%s\n' "$CI_PIPELINE_SOURCE" "$CI_COMMIT_SHA"
- printf 'pipeline=%s job=%s\n' "$CI_PIPELINE_ID" "$CI_JOB_ID"
- printf 'runner_id=%s runner=%s\n' "$CI_RUNNER_ID" "$CI_RUNNER_DESCRIPTION"
- printf 'runner_version=%s arch=%s\n' "$CI_RUNNER_VERSION" "$CI_RUNNER_EXECUTABLE_ARCH"
- printf 'host=%s uid=%s pwd=%s\n' "$(hostname)" "$(id -u)" "$PWD"
- uname -a
- printf 'created_by_job=%s\n' "$CI_JOB_ID" > ch05-boundary.txt
artifacts:
when: always
paths: [ch05-boundary.txt]
expire_in: 1 day
probe_shell:
extends: .executor_probe
tags: [ch05-shell]
probe_docker:
extends: .executor_probe
tags: [ch05-docker]
image: alpine:3.22
Both runner configurations must belong only to the disposable lab and execute trusted source. The Docker runner must be non-privileged with no host Docker socket or sensitive mounts. If those conditions cannot be guaranteed, use the local simulation instead.
4. Free/local fallback: compare without registering runners
cat > "$LAB_ROOT/probe.sh" <<'EOF'
#!/bin/sh
set -eu
printf 'host=%s uid=%s pwd=%s\n' "$(hostname)" "$(id -u)" "$PWD"
uname -a
printf 'marker=%s\n' "$(hostname)" > boundary-marker.txt
EOF
chmod +x "$LAB_ROOT/probe.sh"
# Host-style execution.
(
cd "$LAB_ROOT"
./probe.sh
) | tee "$LAB_ROOT/evidence/host-style.txt"
# Container execution.
cp "$LAB_ROOT/probe.sh" "$LAB_ROOT/docker-output/probe.sh"
docker run --rm --name ch05-checkpoint-docker --read-only --tmpfs /tmp:rw,noexec,nosuid,size=32m --mount type=bind,src="$LAB_ROOT/docker-output",dst=/work --workdir /work alpine:3.22 /bin/sh ./probe.sh | tee "$LAB_ROOT/evidence/docker.txt"
docker ps -a --filter 'name=^/ch05-checkpoint-docker$' | tee "$LAB_ROOT/evidence/docker-after.txt"
Save image inspection output/digest as part of the evidence. The fallback demonstrates boundary mechanics but cannot prove GitLab runner scheduling or manager identity; record that limitation explicitly.
5. Capture comparable evidence before failure injection
| Evidence | Shell / host style | Docker |
|---|---|---|
| Source | Pipeline/source/ref/SHA or local Git SHA. | Same. |
| Runner/manager | Runner ID/description/version/system identity if live. | Runner ID/description/version/system identity if live. |
| Executor | shell. |
docker. |
| Worker | Host name/OS/user. | Runner host + container ID/hostname/image. |
| Image/toolchain | Host OS/packages/tool versions. | Container image release + digest/tool versions. |
| Filesystem | Workspace location and residue after job. | Container layer + mounts/volumes and their cleanup. |
| Privilege | Runner host user/groups/devices/network. | Container user/capabilities/privileged flag/mounts/socket access. |
| Teardown | Host persists. | Container should be removed; persistent mounts handled separately. |
6. Failure injection A: routing failure
routing_failure:
stage: probe
tags: [ch05-no-such-runner-tag]
script:
- echo "This line should not execute while no matching runner exists"
Prediction: pipeline/job creation succeeds, but the job remains pending because no runner satisfies the tag. Preserve pipeline/job ID and requested tags. Repair by changing only the job to the intended existing lab tag or by restoring the intended lab runner—do not create a broad untagged runner merely to make the queue move.
7. Failure injection B: Docker image resolution failure
image_failure:
stage: probe
tags: [ch05-docker]
image: alpine:does-not-exist-ch05
script:
- echo "This line must not be reached"
Prediction: the Docker runner accepts the job,
executor preparation attempts image resolution/pull, and the job
fails before script:. Preserve the first image error,
runner/executor identity and job ID. Repair only the image
reference:
image_repaired:
stage: probe
tags: [ch05-docker]
image: alpine:3.22
script:
- printf 'repaired sha=%s job=%s runner=%s\n' "$CI_COMMIT_SHA" "$CI_JOB_ID" "$CI_RUNNER_ID"
- cat /etc/os-release | sed -n '1,4p'
Do not respond with privileged mode, host socket mounts,
allow_failure, or blind retries. None of those fixes a
nonexistent image.
8. Verify predictions independently
After host-style/Shell execution, verify the workspace/host persists and identify any marker residue. Remove only the lab marker during cleanup.
After Docker execution with --rm or a normal
executor job, verify the job container is gone; separately
inspect deliberate persistent mounts/artifacts.
The no-such-tag job should never reach executor preparation while no eligible runner exists.
The invalid-image job should have runner/executor preparation evidence but no user-script line.
The corrected image should run without changes to privilege, runner scope or unrelated pipeline rules.
9. Make the production recommendation from constraints
Assume the target workload is ordinary Linux build/test code written by project developers, may eventually include merge-request contributions, requires no host devices, and benefits from a repeatable toolchain. A reasonable recommendation is a non-privileged Docker executor or an ephemeral Docker Autoscaler variant when stronger cross-job worker destruction/elasticity is required.
Document why Shell is rejected for the general pool: persistent host exposure and weaker isolation for potentially untrusted code. Also document Docker’s residual risks: shared host kernel, daemon/runner-host compromise potential, sensitive mounts/capabilities, network reachability and persistent cache/volume services. If privileged image builds are introduced later, route them to a distinct protected/ephemeral runner rather than weakening the general pool.
recommendation:
workload: "ordinary Linux build/test"
code_trust: "developer + future MR code"
selected_executor: "docker"
required_controls:
- "non-privileged job containers"
- "no host Docker socket"
- "no sensitive host mounts"
- "narrow runner scope/tags"
- "protected credentials unavailable to untrusted refs"
- "verified image release and digest"
- "network segmentation appropriate to CI"
scale_path: "docker-autoscaler with one-use workers when justified"
rejected_default: "shell for untrusted/general shared workloads"
residual_risk: "shared kernel and persistent runner-manager/host services"
10. Cleanup and rollback
# Local simulation: remove only the exact lab directory after preserving evidence.
docker ps -a --filter 'name=^/ch05-checkpoint-'
# If the exact lab container exists because a previous command was interrupted:
docker rm -f ch05-checkpoint-docker 2>/dev/null || true
# Review before deleting the local lab files.
find "$LAB_ROOT" -maxdepth 2 -type f -print
rm -rf -- "$LAB_ROOT"
For live GitLab runners, follow Chapter 04 retirement discipline: stop accepting work, preserve manager/job evidence, unregister disposable managers, delete only the lab runner configurations, rotate/delete temporary credentials, and remove exact disposable Docker/VM resources. Do not use global prune/delete commands.
11. Required evidence packet
- Source: project/path, branch/ref, immutable SHA, pipeline source, pipeline ID.
- Jobs: two successful comparison job IDs plus routing-failure and image-failure job IDs where live.
- Runner: runner ID/description, manager/system identity where available, Runner version and executor.
- Worker: host/container/pod/instance identity and lifecycle/reuse statement.
- Image/runtime: Docker image release/digest or host OS/tool versions.
- Boundary: user/UID, privilege, capabilities, mounts/socket exposure, network assumptions, Kubernetes RBAC if applicable.
- Failure evidence: first pending/routing or image-pull evidence before repair.
- Cleanup: proof that disposable container/runner/worker resources and temporary tokens are retired as intended.
- Assumptions: GitLab offering/tier, executor availability, local/cloud/Kubernetes simulation boundaries and anything not proven by the lab.
12. What Chapter 05 adds to the operating model
You can now separate runner-manager identity from executor and worker identity, compare isolation and lifecycle rather than merely syntax, preserve evidence across ephemeral environments, diagnose routing versus preparation versus script failures, and justify executor choice from trust and operational constraints.
Chapter 06 moves up one layer into CI/CD variables: predefined/project/group/job/file variables, masking/protection, expansion and precedence. Executor boundaries remain crucial because every variable ultimately reaches some execution environment, and secret safety depends on which code and worker can access it.
Knowledge check
What proves the two checkpoint jobs actually used different execution models?
Runner/executor metadata plus worker evidence: host identity for Shell and container/image identity for Docker, correlated with the pipeline/job IDs and same source SHA.
Why does the routing-failure job stay pending rather than fail its script?
No runner satisfies the requested tag, so executor preparation and the user script never begin.
What is the correct repair for the invalid Docker image?
Use a verified valid image reference/digest. Do not broaden privilege or hide the failure.
Why can Docker be recommended while still documenting residual risk?
Executor selection is risk reduction, not perfect isolation. Docker improves reproducibility/separation but still shares a host kernel and depends on daemon/mount/network configuration.
When would you prefer Docker Autoscaler over a static Docker runner?
When burst elasticity or one-use ephemeral worker isolation justifies Fleeting/cloud complexity, startup latency and cost controls.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
- GitLab Runner executors — current executor selection guidance, compatibility matrix, actively developed paths, and maintenance-mode executors.
- Shell executor — host-local execution, shell selection, process termination, maintenance-mode status, and security warnings.
- Docker executor — image/service model, container lifecycle, volumes, image pull behavior, and executor configuration.
- Docker Autoscaler executor — Fleeting-based autoscaling, Docker feature compatibility, autoscaling resources, capacity and ephemeral-worker patterns.
- Instance executor — Fleeting-based instance provisioning, full host access, worker images, capacity, and native-step considerations.
- Kubernetes executor — per-job Pod creation, build/helper/service containers, RBAC requirements, entrypoint behavior, and executor configuration.
- SSH executor and Custom executor — exceptional/maintenance-mode execution models and operational constraints.
- Security for self-managed runners — Shell risk, Docker privilege/capability guidance, network segmentation, and trust-boundary recommendations.
- Use Docker to build Docker images — security and executor implications for Docker-in-Docker and socket/pipe binding.
Executor behavior and status were rechecked against current primary GitLab documentation on 2026-09-11. At verification time, Docker, Docker Autoscaler, Instance, and Kubernetes are actively developed executor paths; Shell, SSH, VirtualBox, Parallels, and Custom are maintenance mode, and Docker Machine is deprecated. Docker Autoscaler is documented as generally available since GitLab Runner 17.1 and uses Fleeting plugins. Re-check these facts before future course revisions.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.