Chapter 41Lesson 01~205 minutes

Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents: Concepts, Architecture, and Mental Model

Build an evidence-first Docker incident model that localizes symptoms across context, daemon, build/image identity, runtime/process, storage, network, security, and external dependencies before changing state.

Troubleshooting modelEvidence firstContextsDaemonIncident timeline

Learning objectives

  • Turn an ambiguous Docker symptom into a timestamped, layer-by-layer evidence plan before changing state.
  • Prove host, Docker context, daemon/API, source/image identity, container/process, storage, network, permissions/security, and external dependency state independently.
  • Preserve first-failure evidence such as logs, inspect JSON, events, exit/health state, build progress, and external responses.
  • Distinguish a Docker object problem from an application problem and a local Docker problem from an external registry/service problem.
  • Choose the smallest reversible correction that matches the layer owning the failure.

1. Troubleshooting is localization, not command collection

When a developer reports “Docker is broken,” the symptom may have originated in the client, the selected context, the daemon, an image pull, the build graph, the container process, DNS, a mount, an LSM policy, a resource limit, or an external registry. A useful incident method narrows the ownership boundary before it changes anything. This prevents a daemon restart from erasing the timeline of an application crash, a rebuild from masking a wrong-context problem, or a broad cleanup from deleting the very object that proves the failure.

The core rule is simple: preserve, identify, localize, correct, verify. Commands are secondary to that sequence.

2. Mental model: symptom to smallest repair

Evidence-first incident path
flowchart TD
  A[Reported symptom + timestamp] --> B[Host / client / context]
  B --> C[Daemon / API / component state]
  C --> D[Source / build / image digest]
  D --> E[Container / runtime / process / health]
  E --> F[Mounts / storage / resources]
  F --> G[Network / DNS / published path]
  G --> H[Permissions / seccomp / LSM / identity]
  H --> I[Registry / provider / external service]
  I --> J[Smallest reversible repair]
  J --> K[Re-run same verification]
  K --> L[Runbook / prevention update]
            

The arrows represent evidence gates, not a mandatory delay through every layer. If the CLI says the selected context does not exist, you already have strong client/context evidence and should not start editing container DNS. If the container has a stable image digest but exits with code 17 and an application error in stdout, the daemon may be healthy while the process contract is not.

3. Build an incident evidence ledger

Layer Identity/evidence to preserve Typical question
Host/client/context UTC time, OS/kernel, Docker CLI, current context, endpoint Am I operating the host/daemon I think I am?
Daemon/API Server/API version, daemon logs, info/security/storage/network options Is dockerd reachable and internally healthy?
Build/image Source revision, Dockerfile/context, plain build log, image/index digest Did the expected bytes build/pull?
Container/process Container ID, image ID, state, exit code, restart count, health, stdout/stderr Did the process start, stay alive, and report healthy?
Storage/resources Mount source/target/RW, volume ID, writable size, limits/stats Can the process access the intended data and capacity?
Network/DNS Network IDs, aliases, IP/gateway/DNS, published ports, listener path Can the intended peer resolve and reach the endpoint?
Security/identity UID/GID, capabilities, seccomp/LSM, read-only rootfs, denials Is policy correctly denying or unexpectedly blocking the operation?
Registry/external Repository/digest, HTTP/status/error, credentials scope, upstream health Is the failure outside the local container runtime?

4. Preserve the first failure before “fixing” it

First-failure evidence is often the most informative state in the incident. A restart can reset process state. Recreating a container changes its ID, timestamps, writable layer, and possibly its image reference resolution. Deleting a failed build log removes the exact vertex that failed. Before any correction, capture the object identity and the failure in its original context.

date -u +%Y-%m-%dT%H:%M:%SZ
docker context show
docker version
docker info
docker ps -a --no-trunc
docker inspect <exact-container-id-or-name>
docker logs --timestamps --tail 200 <exact-container-id-or-name>

These are examples of read-only evidence collection. Use the smallest subset needed for the incident and redact tokens, private registry credentials, internal hostnames, or sensitive application output before sharing evidence.

5. Client, context, and host come before containers

The Docker CLI is a client. A command can be syntactically valid yet target the wrong daemon because of the active context, DOCKER_CONTEXT, --context, or DOCKER_HOST. Capture docker context show and docker context inspect before any mutating command during an incident. Also record the host/VM boundary: Docker Desktop containers run in a Linux VM, so “host filesystem,” daemon logs, network paths, and kernel evidence differ from native Linux.

docker context show
docker context inspect "$(docker context show)"
docker version
# On native Linux, host/kernel evidence may include:
uname -a 2>/dev/null || true

6. Daemon/API evidence: is Docker itself unhealthy?

A client connection failure is not automatically a daemon crash. The socket may be absent, permissions may be wrong, the context endpoint may be stale, or the service may have failed configuration validation. On systemd Linux, daemon logs are normally inspected with journalctl -xu docker.service. Depending on distribution, older systems may use syslog/messages. Docker Desktop stores dockerd/containerd VM logs in its managed log locations, while Windows containers use Windows Event Log.

If the daemon is unresponsive on Linux, Docker documents SIGUSR1 as a way to request goroutine stack traces without stopping the daemon. That is an advanced diagnostic action: capture service/host evidence first and use it only on an authorized host.

7. Build and image identity: separate “could not build” from “wrong artifact”

A build failure belongs to the build graph until evidence says otherwise. Record the exact Dockerfile, context, frontend, builder, build args that are not secrets, and plain progress. If the build succeeds, record the resulting image ID/digest. A container incident after a successful build is not automatically a build-cache problem.

docker buildx version
docker buildx inspect --bootstrap
docker buildx build --progress=plain --load -t example:incident .
docker image inspect example:incident --format 'Id={{.Id}} RepoDigests={{json .RepoDigests}}' 

For production investigations, prefer immutable registry digests. A mutable tag such as latest is a pointer, not stable incident identity.

8. Container, process, exit, restart, and health are different signals

A container can exist while its process is stopped; a process can run while the healthcheck is unhealthy; an Engine restart policy reacts to process exit, not generic health failure. Preserve .State, .RestartCount, configured healthcheck, and logs together. Exit codes need application/process context: code 137 may indicate SIGKILL/OOM or a forced kill path; code 126/127 often points to command/permission/not-found problems, but treat these as hypotheses until inspect/log evidence confirms them.

docker inspect <container> --format '{{json .State}}'
docker inspect <container> --format 'RestartCount={{.RestartCount}} Image={{.Image}}'
docker logs --timestamps --tail 200 <container>

9. Storage and permissions: prove the mount contract

Inside-container “permission denied,” “read-only file system,” and “no space left on device” are different failure classes. Inspect the mount type, source, destination, read/write flag, volume identity, container user, filesystem capacity, and resource limits before changing permissions. A bind mount source is resolved on the daemon host, not necessarily the CLI machine. A named volume persists independently of the container. The writable layer disappears with the container.

docker inspect <container> --format '{{json .Mounts}}'
docker inspect <container> --format 'User={{.Config.User}} ReadonlyRootfs={{.HostConfig.ReadonlyRootfs}}'
docker ps --size --filter id=<container-id-prefix>
docker system df -v

Do not “solve” an ownership problem by making the whole tree world-writable. Identify the numeric UID/GID and the intended writable path, then correct only that boundary.

10. Network, DNS, and published-path evidence

Ask separate questions: is the process listening? is the container attached to the expected network? can the peer resolve the intended name? can it route to the target? is the host port actually published and bound to the intended interface? On a user-defined bridge, Docker provides embedded DNS at 127.0.0.11 and resolves container names/aliases. The default bridge does not provide the same name-resolution behavior.

docker network ls
docker network inspect <network>
docker inspect <container> --format '{{json .NetworkSettings.Networks}}'
docker port <container>

Change only one layer at a time. Editing DNS and firewall rules simultaneously destroys causal evidence.

11. Registry and external-service failures are their own layer

A registry 401, 403, certificate error, timeout, or manifest/platform mismatch should not be mislabeled as a Dockerfile bug. Record repository, digest/reference, platform, response/error, current context, proxy path, and credential scope without printing secrets. Similarly, an application can be healthy inside its container while an upstream database or API is unavailable.

12. Smallest repair, then same verification

The repair should match the evidence owner. Wrong context: switch/qualify the context. Missing build input: correct the context/Dockerfile. DNS typo: correct the service name/alias. Read-only mount: make only the intended target writable. Crash loop: fix the process contract and preserve the old failure packet. After the change, repeat the exact check that failed; do not declare success merely because a new container started.

This is the DevOps connection: an incident is reproducible only when identities, inputs, environment, timestamps, side effects, and before/after evidence are attached to the same story.

Knowledge check

Why should docker context show be captured before a mutating incident command?

A container is running but its health status is unhealthy. Does Engine restart policy automatically restart it?

What evidence distinguishes a read-only mount from a Unix ownership problem?

Why is a mutable image tag weak incident identity?

What makes a repair “smallest safe”?

Next lesson

Next: Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents: Guided Hands-On Workflow and Core Operations

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Version-sensitive explanations use Docker Engine/CLI 29.8.1 as the current Engine release baseline, while labs record the learner’s actual Engine, API, Compose, Buildx, BuildKit, containerd/runc exposure, kernel/cgroup mode, context, and image digest before interpreting evidence.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.