Chapter 03Lesson 04~120 minutes

Docker Engine Components: dockerd, containerd, runc, BuildKit, APIs, and Container Lifecycle: Diagnostics, Failure Modes, Security, and Performance

Diagnose Engine-internal failures by preserving first evidence and locating the failing layer before changing anything; avoid destructive runtime-directory edits, blind restarts, privileged shortcuts, and wrong-layer fixes.

DiagnosticsFailure localizationAPI versionsRuntime stateEvidence first

Learning objectives

  • Use a causal diagnostic sequence from context/API through runtime, storage/security, and external dependencies.
  • Diagnose a wrong context and forced API mismatch without touching runtime state.
  • Explain why a successful custom-builder build may not create a local Engine image.
  • Reject destructive runtime-directory edits, direct managed-containerd mutation, and privilege escalation as generic fixes.
  • Measure the failing phase before tuning performance.
Chapter 03 platform baseline — verified 2026-09-20. Docker Engine 29.8.1 is current. The 29.8 API matrix lists API 1.55 maximum / 1.40 minimum. Engine 29.8 packages BuildKit 0.33.0 and runc 1.5.1; the 29.8.1 static-binary packaging update carries containerd 2.3.5. Treat these as dated reference points only: capture docker version, docker info, docker buildx version, and platform-specific runtime evidence on the machine that actually runs the lab.

1. Preserve evidence before touching the layer you suspect

The fastest-looking fix—restart Docker, delete runtime files, run a lower-level tool, add privilege—often erases the information that would identify the actual failing layer. Start with a timestamped evidence packet.

date -u
docker context show
docker version
docker info
docker buildx version
docker buildx ls
docker ps -a --no-trunc
docker events --since 10m --until "$(date -u +%Y-%m-%dT%H:%M:%SZ)" 2>/dev/null | tail -n 100

2. Diagnostic sequence by causal layer

  1. Client/context/endpoint: did the request reach the intended daemon?
  2. Engine API/dockerd: did the daemon accept and validate the operation?
  3. Image/build input: did the exact content exist and resolve?
  4. Runtime creation: did containerd/shim/runc create the task/process?
  5. Process/health: did PID 1 start, stay running, and behave correctly?
  6. Storage/resource/security: did mounts, cgroups, capabilities, LSM policy, or permissions block the process?
  7. Network/external dependency: did post-start communication fail?

Only move downward when the previous layer has evidence of success.

3. Intentionally broken example: wrong context, not broken runtime

This safe exercise creates a local Docker context that points to a nonexistent Unix socket. It does not change the daemon.

docker context create da-ch03-broken   --docker 'host=unix:///tmp/da-ch03-missing.sock'

docker --context da-ch03-broken version || true

docker context show
docker --context default version 2>/dev/null || true

docker context rm da-ch03-broken

The expected failure occurs before dockerd, containerd, or runc on the real host receives a request. The repair is to select/correct the endpoint, not to restart runtimes.

4. Intentionally broken example: forced obsolete API version

Engine 29.8 supports API negotiation, with the current matrix listing minimum API 1.40. Forcing a much older API version is a reversible way to demonstrate an API-layer failure.

DOCKER_API_VERSION=1.24 docker version || true
unset DOCKER_API_VERSION
docker version

Preserve the exact error. If normal negotiation works immediately afterward, there is no evidence that runc or containerd was the problem.

5. “Build succeeded but image missing”: builder/exporter diagnosis

A custom builder can successfully solve a build while leaving the result only in BuildKit cache because no local-load or push exporter was requested.

WORK="${TMPDIR:-/tmp}/da-ch03-missing-image"
mkdir -p "$WORK"
printf 'FROM busybox:1.36.1
RUN echo ok > /ok
' > "$WORK/Dockerfile"

docker buildx create --name da-ch03-isolated --driver docker-container --use --bootstrap
docker buildx build --progress=plain -t da-ch03-cacheonly:local "$WORK"
docker image inspect da-ch03-cacheonly:local || true

# Correct the output contract instead of blaming Engine image storage.
docker buildx build --progress=plain --load -t da-ch03-cacheonly:local "$WORK"
docker image inspect da-ch03-cacheonly:local --format '{{.Id}}'

docker buildx use default 2>/dev/null || true
docker buildx rm da-ch03-isolated
docker image rm da-ch03-cacheonly:local 2>/dev/null || true
rm -rf "$WORK"
Expected diagnosis: first build success proves BuildKit completion, not Engine image-store presence. --load changes the exporter outcome. No direct containerd repair is appropriate.

6. Daemon restart is not container restart

If dockerd restarts, the control plane changes. If a workload process restarts, application runtime state changes. Live-restore can intentionally keep certain standalone containers running across daemon unavailability, but only within documented constraints. Therefore a daemon PID change does not by itself prove the application process restarted, and an application PID change does not prove the daemon restarted.

7. Unsafe shortcuts to reject

Shortcut Why it is wrong Safer evidence-first alternative
Delete /run/containerd or Docker runtime directories Destroys live state/evidence and can orphan ownership. Preserve logs/processes; use supported service recovery.
Run ctr mutations in Docker's namespace Bypasses Engine object ownership. Use Docker API/CLI; lower-level inspection read-only if needed.
Add --privileged Collapses security boundaries without proving cause. Inspect the exact denied capability/device/mount/LSM rule.
Blind daemon restart Erases timing evidence and can disrupt workloads. Capture logs/events/config first; restart only with a hypothesis.
Broad prune Deletes unrelated content/cache and evidence. Remove exact lab-owned IDs/tags/builders.

8. Performance: locate the bottleneck before tuning

Slow docker build can be context transfer, cache miss, BuildKit worker CPU/I/O, registry latency, snapshotter/storage behavior, or exporter time. Slow docker run can be image pull/unpack, runtime setup, mount/network initialization, or application startup. Measure the phase and preserve timestamps before increasing resources or concurrency.

Next lesson

Next: Checkpoint Lab — Docker Engine Components: dockerd, containerd, runc, BuildKit, APIs, and Container Lifecycle

Continue through the same Engine evidence chain while adding the next layer of operational reasoning.

Knowledge check

An invalid daemon option prevents dockerd from starting. Would changing runc permissions be a justified first fix?

A developer says “the build succeeded, so the image must be in docker image ls.” What missing fact should you ask for?

Why is deleting /run/containerd or Docker runtime directories a poor troubleshooting shortcut?

What evidence distinguishes a wrong context from a daemon failure?

Official references and version notes

Current baseline, not a frozen requirement

Verified 2026-09-20: Docker Engine 29.8.1 is current; the Engine 29.8 API matrix lists maximum API 1.55 and minimum API 1.40. Engine 29.8 packaging includes BuildKit 0.33.0 and runc 1.5.1; 29.8.1 updates the static-binary containerd package to 2.3.5. Package-managed distributions and Docker Desktop can bundle or expose components differently. Record the actual client/server/API/component/builder versions on the learner's environment before diagnosing compatibility.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.