Chapter 41Lesson 02~245 minutes

Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents: Guided Hands-On Workflow and Core Operations

Diagnose controlled Docker incidents involving context selection, crash loops, DNS, bind mounts, bounded ENOSPC, and build inputs while preserving first-failure evidence.

Hands-onCrash loopsDNSMountsBuildKit

Learning objectives

  • Run six bounded incidents that each fail at a different Docker layer without touching a production daemon.
  • Capture context, object IDs, inspect state, logs/events, mount/network evidence, and BuildKit progress before repair.
  • Diagnose a restart loop, DNS typo, read-only bind, bounded tmpfs ENOSPC, and missing build input.
  • Apply one minimal correction per incident and repeat the failed check.
  • Clean only exact lab objects and retain selected evidence files deliberately.

1. Lab contract and preflight

This lab uses a recorded Alpine image digest, exact ch41-* names, one user-defined bridge, one small bind directory, one 4 MiB tmpfs, and one tiny build context. It does not restart Docker, alter daemon JSON, change host firewall/DNS policy, broaden permissions, or prune shared resources.

Safety boundary. Confirm the active Docker context is an authorized local/disposable daemon. The exercises intentionally create failures; they must not run against a production or shared context.
set -eu
printf 'timestamp=%s
' "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
printf 'context=%s
' "$(docker context show)"
docker version
docker info --format 'Server={{.ServerVersion}} Driver={{.Driver}} Cgroup={{.CgroupVersion}}'
docker buildx version

docker pull alpine:3.22
BASE_REF="$(docker image inspect alpine:3.22 --format '{{index .RepoDigests 0}}')"
printf 'base=%s
' "$BASE_REF"

2. Incident A — client/context failure

Before debugging the daemon, prove that the client can select the requested context. This deliberately references a context that should not exist; no daemon object is mutated.

docker --context ch41-context-does-not-exist ps 2>&1 | tee ch41-context-error.log || true
printf 'actual_context=%s
' "$(docker context show)"
docker context inspect "$(docker context show)" > ch41-context-inspect.json

Interpretation: a “context not found” result is client/context state. The smallest correction is to use the intended named context explicitly or repair context configuration—not restart containers.

3. Incident B — controlled crash loop

The process exits with a synthetic application code and a bounded on-failure:3 restart policy. Preserve state after the retries instead of deleting the container immediately.

docker rm -f ch41-crash 2>/dev/null || true
docker run -d \
  --name ch41-crash \
  --label devops.academy.lab=ch41 \
  --restart on-failure:3   "$BASE_REF" sh -c 'echo "synthetic startup failure" >&2; exit 17'
sleep 5

docker ps -a --filter name=ch41-crash --no-trunc
docker inspect ch41-crash --format 'Id={{.Id}} Status={{.State.Status}} Exit={{.State.ExitCode}} RestartCount={{.RestartCount}} Image={{.Image}}'
docker logs --timestamps ch41-crash | tee ch41-crash.log
END="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
docker events --since 10m --until "$END" --filter container=ch41-crash > ch41-crash.events || true

The daemon has successfully executed the policy; the application process is the failing layer. Repair the process contract only after saving the original ID/log/event packet.

OLD_ID="$(docker inspect ch41-crash --format '{{.Id}}')"
docker rm ch41-crash
docker run -d --name ch41-crash --label devops.academy.lab=ch41 "$BASE_REF" sleep 300
NEW_ID="$(docker inspect ch41-crash --format '{{.Id}}')"
printf 'old=%s
new=%s
' "$OLD_ID" "$NEW_ID"
docker inspect ch41-crash --format 'Status={{.State.Status}} Exit={{.State.ExitCode}}' 

4. Incident C — user-defined-network DNS typo

Create a scoped network and a sleeping peer. The first query uses the wrong service name; the second repeats the same resolver check with the correct name.

docker network rm ch41-net 2>/dev/null || true
docker network create --label devops.academy.lab=ch41 ch41-net
docker rm -f ch41-api 2>/dev/null || true
docker run -d --name ch41-api --network ch41-net --label devops.academy.lab=ch41 "$BASE_REF" sleep 300

docker run --rm --network ch41-net "$BASE_REF" nslookup ch41-api-typo 2>&1 | tee ch41-dns-fail.log || true
docker network inspect ch41-net > ch41-network.json
docker inspect ch41-api --format '{{json .NetworkSettings.Networks}}'

docker run --rm --network ch41-net "$BASE_REF" nslookup ch41-api | tee ch41-dns-fixed.log

The fix changes the queried identity, not host DNS. On a user-defined bridge, Docker’s embedded DNS supplies container-name resolution. If external names fail instead, inspect the container resolver/upstream path as a different branch.

5. Incident D — deterministic read-only bind failure

This incident avoids host-specific UID/GID behavior by making the mount contract explicitly read-only. The error should be interpreted as mount configuration, not “Docker permissions are broken.”

rm -rf ch41-host
mkdir ch41-host
printf 'seed
' > ch41-host/input.txt

docker run \
  --rm \
  --name ch41-ro \
  --mount type=bind,source="$PWD/ch41-host",target=/work,readonly   "$BASE_REF" sh -c 'echo attempt > /work/output.txt'   2>&1 | tee ch41-bind-fail.log || true

docker create --name ch41-ro-inspect   --mount type=bind,source="$PWD/ch41-host",target=/work,readonly   "$BASE_REF" true >/dev/null
docker inspect ch41-ro-inspect --format '{{json .Mounts}}' | tee ch41-bind-inspect.json
docker rm ch41-ro-inspect

docker run \
  --rm \
  --name ch41-rw \
  --mount type=bind,source="$PWD/ch41-host",target=/work   "$BASE_REF" sh -c 'echo fixed > /work/output.txt && cat /work/output.txt'
cat ch41-host/output.txt

The correction is narrow: only this disposable mount becomes writable because the workload contract requires a write. In a real incident, also verify host ownership, SELinux/AppArmor policy, user namespaces, and the runtime user before changing permissions.

6. Incident E — bounded “disk full” using tmpfs

Never fill the host filesystem just to practice ENOSPC. A small tmpfs gives a local, self-cleaning capacity boundary.

docker run \
  --rm \
  --name ch41-enospc \
  --mount type=tmpfs,destination=/scratch,tmpfs-size=4m   "$BASE_REF" sh -c 'df -h /scratch; dd if=/dev/zero of=/scratch/blob bs=1M count=8; df -h /scratch'   2>&1 | tee ch41-enospc.log || true

If the failure is “No space left on device,” the evidence points to that mount capacity. Do not respond by pruning unrelated images or volumes. In a real host-pressure incident, separately inspect bytes, inodes, Docker-wide usage, build cache, container writable size, volumes, and logs.

7. Incident F — BuildKit input failure

The Dockerfile asks for a source file that does not exist in the context. Preserve plain BuildKit progress before creating the missing file.

rm -rf ch41-build
mkdir ch41-build
cat > ch41-build/Dockerfile <<'EOF'
# syntax=docker/dockerfile:1
FROM alpine:3.22
WORKDIR /app
COPY required.txt ./
CMD ["cat","/app/required.txt"]
EOF

docker buildx build --progress=plain --load -t ch41-build:broken ch41-build   > ch41-build-fail.log 2>&1 || true
cat ch41-build-fail.log

printf 'present
' > ch41-build/required.txt
docker buildx build --progress=plain --load -t ch41-build:fixed ch41-build   2>&1 | tee ch41-build-fixed.log
docker run --rm ch41-build:fixed

The failed vertex belongs to build input/context state. Restarting the Docker daemon would not create required.txt.

8. Optional daemon/system evidence — read only first

If a lab command instead fails because the daemon is unreachable, stop the scenario and capture daemon evidence appropriate to the platform. On systemd Linux, journalctl -xu docker.service is the documented starting point. Docker Desktop and Windows containers use different log locations; do not assume a Linux path.

# Native Linux/systemd example (read-only):
journalctl -xu docker.service --since '-15 min' --no-pager 2>/dev/null | tail -200 || true

9. Challenge — identify the owning layer before the command

Symptom First evidence Likely layer to test first
Context name not found context show/ls/inspect Client/context
RestartCount rises; exit 17 inspect State + logs + events Process/restart contract
Peer name NXDOMAIN on custom bridge network inspect + exact alias/name Docker DNS/network identity
Write says read-only filesystem inspect Mounts RW Mount/storage configuration
ENOSPC inside 4 MiB tmpfs mount options + df inside mount Bounded storage capacity
COPY source missing plain BuildKit vertex/context contents Build input/context

10. Cleanup only exact lab state

docker rm -f ch41-crash ch41-api 2>/dev/null || true
docker network rm ch41-net 2>/dev/null || true
docker image rm ch41-build:broken ch41-build:fixed 2>/dev/null || true
rm -rf ch41-host ch41-build
# Keep or explicitly remove the evidence files you created:
# ch41-*.log, ch41-*.json, ch41-*.events

No Docker-wide cleanup is required. Evidence files are left to you deliberately; incident preservation is a choice, not an accidental side effect.

Knowledge check

A crash-loop container has RestartCount 3 and exit code 17. Which layer should you investigate before daemon configuration?

Why does the bind lab use an explicit read-only mount?

What does the 4 MiB tmpfs incident teach that filling the host disk would not?

A BuildKit COPY fails because required.txt is absent. Why is restarting dockerd a poor response?

What proves the DNS fix rather than merely hiding the failure?

Next lesson

Next: Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents: Configuration, Design Choices, and Tradeoffs

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Version-sensitive explanations use Docker Engine/CLI 29.8.1 as the current Engine release baseline, while labs record the learner’s actual Engine, API, Compose, Buildx, BuildKit, containerd/runc exposure, kernel/cgroup mode, context, and image digest before interpreting evidence.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.