Chapter 01Lesson 04~100 minutes

Containers, Virtual Machines, Linux Isolation, OCI Standards, and Docker Architecture: Diagnostics, Failure Modes, Security, and Performance

Diagnose Docker from evidence instead of guesswork: preserve first failure, localize symptoms to client/daemon/image/runtime/security/storage/network layers, reject privilege-escalation shortcuts, and repair only the smallest justified scope.

DiagnosticsFailure modesSecurityEvidencePerformance

Learning objectives

  • Use a causal diagnostic sequence from client/context through daemon, image, runtime/process, security/storage/network, and external dependencies.
  • Explain why container-as-VM thinking, broad daemon access, and --privileged troubleshooting are unsafe or misleading.
  • Distinguish image success, container-object existence, process state, writable-layer state, and application health.
  • Preserve IDs, digests, logs, exit codes, inspect state, namespace/cgroup evidence, and measurements before retry/recreate.
  • Diagnose one intentionally denied container operation without weakening the isolation boundary.
Chapter 01 technical baseline — verified 2026-09-20. Docker Engine 29.8.1 is the current Engine 29 patch release in Docker's release notes. Its packages update the bundled containerd static binaries to v2.3.5. Docker and OCI internals are version-sensitive, so every lab starts by recording the actual client, server, API, kernel, context, cgroup, runtime, and image identity instead of assuming this baseline.

1. Preserve first-failure evidence before “fixing Docker”

Docker failures become expensive when every symptom triggers a restart, recreate, prune, privilege escalation, or daemon change. Those actions can erase the state that identifies the failing layer. The Chapter 01 diagnostic rule is therefore: capture the current context, versions, object IDs/digests, container state, logs, and inspect data before changing anything.

date -u +%Y-%m-%dT%H:%M:%SZ
docker context show
docker version
docker info
docker ps -a --no-trunc

From there, move layer by layer: client/context → daemon/API → image identity → container/runtime/process → mounts/resources/security → network/DNS/exposure → registry/external dependency.

2. Use a causal failure ladder

Stop at the first layer where evidence diverges from expectation
flowchart TD
  A[Client and selected context] --> B[Daemon and Engine API]
  B --> C[Image reference and immutable digest]
  C --> D[Container config and runtime creation]
  D --> E[Main process and exit code]
  E --> F[Mounts, cgroups and security controls]
  F --> G[Network, DNS and publication]
  G --> H[Registry or external dependency]
  H --> I[Application-level health]
            

This sequence prevents a network symptom from causing an image rebuild, or a process permission problem from causing a firewall change.

3. Failure mode — treating a container as a VM with its own kernel

Suppose a learner expects alpine:3.22 to run an “Alpine kernel” because /etc/os-release says Alpine. Capture the two observations instead:

LAB_IMAGE_TAG="alpine:3.22"
docker pull "$LAB_IMAGE_TAG"
ALPINE_REF="$(docker image inspect "$LAB_IMAGE_TAG" --format '{{index .RepoDigests 0}}')"

uname -srmo
docker run --rm "$ALPINE_REF" uname -srmo
docker run --rm "$ALPINE_REF" cat /etc/os-release

On native Linux Engine, the kernel release comes from the same host kernel while userspace files come from the image. On Desktop, compare against the Desktop VM boundary, not the physical host kernel. The repair is a corrected model—not a container setting.

4. Failure mode — “operation not permitted” does not mean “use --privileged”

The following disposable command intentionally tries to mount a tmpfs inside an ordinary container. It should fail on a normal hardened configuration because mounting generally requires privileges such as CAP_SYS_ADMIN that Docker does not grant by default.

docker run --rm "$ALPINE_REF" sh -c '
  mkdir -p /mnt/test
  mount -t tmpfs tmpfs /mnt/test
'

Preserve the error. Then inspect the default capability/security posture rather than re-running with --privileged:

docker run --rm "$ALPINE_REF" sh -c 'grep -E "Cap(Prm|Eff|Bnd)" /proc/1/status'
docker info --format '{{json .SecurityOptions}}' 2>/dev/null || true
Do not execute a privileged retry for this lesson. The purpose of the broken example is to understand that the denial is an expected protection boundary. If a real workload needs a mount operation, redesign around host-managed mounts or justify the narrowest specific capability and security policy in a controlled environment.

5. Failure mode — Docker socket permission denied

If docker version prints client information but fails to connect to the Unix socket, first distinguish “daemon not running,” “wrong context/endpoint,” and “client lacks permission.” On Linux, inspect without mutating:

docker context show
docker context inspect "$(docker context show)"
ls -l /var/run/docker.sock 2>/dev/null || true
id
systemctl is-active docker 2>/dev/null || true

The unsafe reflex is chmod 777 /var/run/docker.sock or exposing port 2375. Docker documents daemon access as security-sensitive; group membership on a rootful host grants root-level privileges. Fix ownership/access according to the host's administration policy, or use rootless Docker—not world-writable daemon control.

6. Failure mode — confusing image, container, and process state

A successful pull can coexist with a failed container. A container object can persist after its main process exits. Diagnose each state explicitly.

BROKEN="da-ch01-exit-demo"

docker run --name "$BROKEN" "$ALPINE_REF" sh -c 'echo before-exit; exit 23' || true

docker image inspect "$ALPINE_REF" --format 'image-id={{.Id}}'
docker inspect "$BROKEN" --format \
'container={{.Id}} status={{.State.Status}} exit={{.State.ExitCode}} finished={{.State.FinishedAt}}'
docker logs "$BROKEN"

docker rm "$BROKEN"

The image remains valid; the container's main process deliberately returned exit code 23. Re-pulling the image or restarting the daemon would not explain that application/process result.

7. Failure mode — assuming container writable state is durable application data

A container can create files in its writable layer, but that layer belongs to the container lifecycle. If you remove and recreate the container, those changes are not a durable data strategy.

TMP="da-ch01-writable-demo"
docker run --name "$TMP" "$ALPINE_REF" sh -c 'echo ephemeral > /evidence.txt; cat /evidence.txt'
docker diff "$TMP"
docker rm "$TMP"
# The container writable layer has now been removed with the container.

Do not “recover” this by editing a stopped container or committing it into an image. Later storage chapters teach explicit volumes, bind mounts, backups, and restore verification.

8. Failure mode — assuming namespaces are a complete security boundary

Namespaces isolate views of kernel resources, but container processes still interact with the same underlying Linux kernel on native Engine. Security is layered: capability bounding, seccomp, AppArmor/SELinux, user namespaces/rootless operation, mount choices, device exposure, network controls, patching, and workload trust all matter.

A container escape vulnerability is a kernel/runtime/security event, not evidence that namespaces are useless. Likewise, namespace separation does not justify placing mutually hostile code on one privileged Docker host without additional isolation.

9. Performance — measure the layer before tuning it

Container startup delay can come from image pull/verification, daemon load, filesystem extraction, runtime creation, application initialization, DNS, or external services. CPU/memory symptoms can be application behavior, cgroup limits, host contention, Desktop VM limits, or storage/network I/O. “Containers are lightweight” does not identify a bottleneck.

docker stats --no-stream
docker inspect "$LAB_CONTAINER" 2>/dev/null || true
docker info

Record measurements on the environment that matters. Docker Desktop VM performance is not automatically representative of a Linux production host.

10. Chapter 01 incident playbook

  1. Record UTC time, host/platform, context, client/server/API versions.
  2. Record exact image reference/digest and container ID before recreation.
  3. Inspect .State, exit code, PID, command, labels, mounts, networks, and security settings.
  4. Preserve console/application logs and relevant daemon events/logs.
  5. On native Linux, inspect the host PID, namespace links, cgroup path, and kernel/LSM evidence when relevant.
  6. Identify the first layer whose state differs from expectation.
  7. Apply the smallest reversible correction; do not change multiple layers at once.
  8. Re-run only the smallest safe scope and verify the original symptom plus side effects.

11. Intentionally broken lab — explain, do not bypass

Run the denied tmpfs mount from section 4 and capture three items: the exact command, its exit status/error text, and the container security context you can inspect without elevating privilege. Write a one-paragraph diagnosis that distinguishes a Docker client failure, image failure, runtime failure, and kernel permission denial.

set +e
docker run --name da-ch01-denied \
  --label devops-academy.lab=chapter01 \
  "$ALPINE_REF" \
  sh -c 'mkdir -p /mnt/test && mount -t tmpfs tmpfs /mnt/test'
STATUS=$?
set -e
printf 'exit=%s
' "$STATUS"

docker inspect da-ch01-denied --format \
'privileged={{.HostConfig.Privileged}} cap-add={{json .HostConfig.CapAdd}} security-opt={{json .HostConfig.SecurityOpt}} status={{.State.Status}} exit={{.State.ExitCode}}'
docker logs da-ch01-denied

docker rm da-ch01-denied

Expected reasoning: the client reached the daemon, the image started, and the process attempted an operation the default privilege boundary does not allow. The correct lesson is not to disable that boundary.

Knowledge check

A container exits with code 127. Should the first response be to restart the Docker daemon?

Why is chmod 777 /var/run/docker.sock an unsafe socket-permission fix?

The tmpfs mount fails inside the lab container. What important thing did the failure prove?

Why preserve a container ID and image digest before recreate?

Why can Desktop performance measurements differ from Linux production?

Summary

Evidence-first diagnosis prevents destructive guesswork. A container is not a VM; “operation not permitted” is not a request for --privileged; daemon access is a security boundary; images, container objects, and processes have separate identities; writable layers are not durable data; and namespaces are only one security layer. Preserve the first failure, localize it to the owning layer, and apply the narrowest correction.

Next lesson

Next: Checkpoint Lab — Containers, Virtual Machines, Linux Isolation, OCI Standards, and Docker Architecture

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version-sensitive statements in this chapter were rechecked against primary documentation on 2026-09-20. Docker Engine, Docker Desktop, containerd, runc, OCI specifications, security defaults, and platform integration continue to evolve. Record the versions actually reported by your own environment before treating an example as production policy.

Current-version note

Docker Engine 29.8.1 is the current Engine 29 patch release as of this lesson's verification date. The current Engine API page contains an example for 29.8.1 reporting API 1.56 while its version matrix lists Docker 29.8 with maximum API 1.55. This chapter therefore treats the output of docker version on the actual client/daemon pair as authoritative lab evidence and does not hard-code an expected negotiated API number.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.