Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents: Concepts, Architecture, and Mental Model
Build an evidence-first Docker incident model that localizes symptoms across context, daemon, build/image identity, runtime/process, storage, network, security, and external dependencies before changing state.
Learning objectives
- Turn an ambiguous Docker symptom into a timestamped, layer-by-layer evidence plan before changing state.
- Prove host, Docker context, daemon/API, source/image identity, container/process, storage, network, permissions/security, and external dependency state independently.
- Preserve first-failure evidence such as logs, inspect JSON, events, exit/health state, build progress, and external responses.
- Distinguish a Docker object problem from an application problem and a local Docker problem from an external registry/service problem.
- Choose the smallest reversible correction that matches the layer owning the failure.
1. Troubleshooting is localization, not command collection
When a developer reports “Docker is broken,” the symptom may have originated in the client, the selected context, the daemon, an image pull, the build graph, the container process, DNS, a mount, an LSM policy, a resource limit, or an external registry. A useful incident method narrows the ownership boundary before it changes anything. This prevents a daemon restart from erasing the timeline of an application crash, a rebuild from masking a wrong-context problem, or a broad cleanup from deleting the very object that proves the failure.
The core rule is simple: preserve, identify, localize, correct, verify. Commands are secondary to that sequence.
2. Mental model: symptom to smallest repair
flowchart TD
A[Reported symptom + timestamp] --> B[Host / client / context]
B --> C[Daemon / API / component state]
C --> D[Source / build / image digest]
D --> E[Container / runtime / process / health]
E --> F[Mounts / storage / resources]
F --> G[Network / DNS / published path]
G --> H[Permissions / seccomp / LSM / identity]
H --> I[Registry / provider / external service]
I --> J[Smallest reversible repair]
J --> K[Re-run same verification]
K --> L[Runbook / prevention update]
The arrows represent evidence gates, not a mandatory delay through every layer. If the CLI says the selected context does not exist, you already have strong client/context evidence and should not start editing container DNS. If the container has a stable image digest but exits with code 17 and an application error in stdout, the daemon may be healthy while the process contract is not.
3. Build an incident evidence ledger
| Layer | Identity/evidence to preserve | Typical question |
|---|---|---|
| Host/client/context | UTC time, OS/kernel, Docker CLI, current context, endpoint | Am I operating the host/daemon I think I am? |
| Daemon/API | Server/API version, daemon logs, info/security/storage/network options | Is dockerd reachable and internally healthy? |
| Build/image | Source revision, Dockerfile/context, plain build log, image/index digest | Did the expected bytes build/pull? |
| Container/process | Container ID, image ID, state, exit code, restart count, health, stdout/stderr | Did the process start, stay alive, and report healthy? |
| Storage/resources | Mount source/target/RW, volume ID, writable size, limits/stats | Can the process access the intended data and capacity? |
| Network/DNS | Network IDs, aliases, IP/gateway/DNS, published ports, listener path | Can the intended peer resolve and reach the endpoint? |
| Security/identity | UID/GID, capabilities, seccomp/LSM, read-only rootfs, denials | Is policy correctly denying or unexpectedly blocking the operation? |
| Registry/external | Repository/digest, HTTP/status/error, credentials scope, upstream health | Is the failure outside the local container runtime? |
4. Preserve the first failure before “fixing” it
First-failure evidence is often the most informative state in the incident. A restart can reset process state. Recreating a container changes its ID, timestamps, writable layer, and possibly its image reference resolution. Deleting a failed build log removes the exact vertex that failed. Before any correction, capture the object identity and the failure in its original context.
date -u +%Y-%m-%dT%H:%M:%SZ
docker context show
docker version
docker info
docker ps -a --no-trunc
docker inspect <exact-container-id-or-name>
docker logs --timestamps --tail 200 <exact-container-id-or-name>
These are examples of read-only evidence collection. Use the smallest subset needed for the incident and redact tokens, private registry credentials, internal hostnames, or sensitive application output before sharing evidence.
5. Client, context, and host come before containers
The Docker CLI is a client. A command can be syntactically valid yet
target the wrong daemon because of the active context,
DOCKER_CONTEXT, --context, or
DOCKER_HOST. Capture
docker context show and
docker context inspect before any mutating command
during an incident. Also record the host/VM boundary: Docker Desktop
containers run in a Linux VM, so “host filesystem,” daemon logs,
network paths, and kernel evidence differ from native Linux.
docker context show
docker context inspect "$(docker context show)"
docker version
# On native Linux, host/kernel evidence may include:
uname -a 2>/dev/null || true
6. Daemon/API evidence: is Docker itself unhealthy?
A client connection failure is not automatically a daemon crash. The
socket may be absent, permissions may be wrong, the context endpoint
may be stale, or the service may have failed configuration
validation. On systemd Linux, daemon logs are normally inspected
with journalctl -xu docker.service. Depending on
distribution, older systems may use syslog/messages. Docker Desktop
stores dockerd/containerd VM logs in its managed log locations,
while Windows containers use Windows Event Log.
If the daemon is unresponsive on Linux, Docker documents
SIGUSR1 as a way to request goroutine stack traces
without stopping the daemon. That is an advanced diagnostic action:
capture service/host evidence first and use it only on an authorized
host.
7. Build and image identity: separate “could not build” from “wrong artifact”
A build failure belongs to the build graph until evidence says otherwise. Record the exact Dockerfile, context, frontend, builder, build args that are not secrets, and plain progress. If the build succeeds, record the resulting image ID/digest. A container incident after a successful build is not automatically a build-cache problem.
docker buildx version
docker buildx inspect --bootstrap
docker buildx build --progress=plain --load -t example:incident .
docker image inspect example:incident --format 'Id={{.Id}} RepoDigests={{json .RepoDigests}}'
For production investigations, prefer immutable registry digests. A
mutable tag such as latest is a pointer, not stable
incident identity.
8. Container, process, exit, restart, and health are different signals
A container can exist while its process is stopped; a process can
run while the healthcheck is unhealthy; an Engine restart policy
reacts to process exit, not generic health failure. Preserve
.State, .RestartCount, configured
healthcheck, and logs together. Exit codes need application/process
context: code 137 may indicate SIGKILL/OOM or a forced kill path;
code 126/127 often points to command/permission/not-found problems,
but treat these as hypotheses until inspect/log evidence confirms
them.
docker inspect <container> --format '{{json .State}}'
docker inspect <container> --format 'RestartCount={{.RestartCount}} Image={{.Image}}'
docker logs --timestamps --tail 200 <container>
9. Storage and permissions: prove the mount contract
Inside-container “permission denied,” “read-only file system,” and “no space left on device” are different failure classes. Inspect the mount type, source, destination, read/write flag, volume identity, container user, filesystem capacity, and resource limits before changing permissions. A bind mount source is resolved on the daemon host, not necessarily the CLI machine. A named volume persists independently of the container. The writable layer disappears with the container.
docker inspect <container> --format '{{json .Mounts}}'
docker inspect <container> --format 'User={{.Config.User}} ReadonlyRootfs={{.HostConfig.ReadonlyRootfs}}'
docker ps --size --filter id=<container-id-prefix>
docker system df -v
Do not “solve” an ownership problem by making the whole tree world-writable. Identify the numeric UID/GID and the intended writable path, then correct only that boundary.
10. Network, DNS, and published-path evidence
Ask separate questions: is the process listening? is the container
attached to the expected network? can the peer resolve the intended
name? can it route to the target? is the host port actually
published and bound to the intended interface? On a user-defined
bridge, Docker provides embedded DNS at 127.0.0.11 and
resolves container names/aliases. The default bridge does not
provide the same name-resolution behavior.
docker network ls
docker network inspect <network>
docker inspect <container> --format '{{json .NetworkSettings.Networks}}'
docker port <container>
Change only one layer at a time. Editing DNS and firewall rules simultaneously destroys causal evidence.
11. Registry and external-service failures are their own layer
A registry 401, 403, certificate error,
timeout, or manifest/platform mismatch should not be mislabeled as a
Dockerfile bug. Record repository, digest/reference, platform,
response/error, current context, proxy path, and credential scope
without printing secrets. Similarly, an application can be healthy
inside its container while an upstream database or API is
unavailable.
12. Smallest repair, then same verification
The repair should match the evidence owner. Wrong context: switch/qualify the context. Missing build input: correct the context/Dockerfile. DNS typo: correct the service name/alias. Read-only mount: make only the intended target writable. Crash loop: fix the process contract and preserve the old failure packet. After the change, repeat the exact check that failed; do not declare success merely because a new container started.
This is the DevOps connection: an incident is reproducible only when identities, inputs, environment, timestamps, side effects, and before/after evidence are attached to the same story.
Knowledge check
Why should docker context show be captured before a mutating incident command?
Because the CLI can target different daemons. Without context identity, later evidence cannot prove which host/daemon was changed.
A container is running but its health status is unhealthy. Does Engine restart policy automatically restart it?
No. Standard Engine restart policies react to process exit; health is a separate signal unless another component acts on it.
What evidence distinguishes a read-only mount from a Unix ownership problem?
Inspect the mount RW flag and target first, then inspect runtime UID/GID and filesystem ownership. A read-only mount denies writes regardless of ordinary ownership.
Why is a mutable image tag weak incident identity?
The tag can move. Preserve the image ID and preferably the registry digest/platform so the investigated artifact is immutable.
What makes a repair “smallest safe”?
It changes only the state owned by the demonstrated failure layer, preserves unrelated objects/evidence, and can be verified with the same failed check.
Official references and version notes
2026-09-22. Version-sensitive explanations use Docker Engine/CLI 29.8.1 as the current Engine release baseline, while labs record the learner’s actual Engine, API, Compose, Buildx, BuildKit, containerd/runc exposure, kernel/cgroup mode, context, and image digest before interpreting evidence.
- Docker Docs — Troubleshoot the Docker daemon — daemon startup/configuration and networking/DNS troubleshooting.
- Docker Docs — Read the daemon logs — Linux journal/syslog locations, Docker Desktop VM logs, Windows Event Log, debug logging, and non-destructive SIGUSR1 stack traces.
- Docker Engine 29 release notes — current 29.8.1 fixes, security changes, regressions, and documented compatibility notes.
- docker context show and docker context inspect — prove the daemon endpoint before mutation.
- docker inspect — low-level Docker object state, IDs, mounts, network settings, health, and runtime configuration.
- docker logs — bounded container stdout/stderr evidence.
- docker events — daemon object event timeline and filters.
-
docker buildx build
—
--progress=plainand raw build progress for reproducible failure traces. - Docker Docs — Optimize build cache — input/cache boundaries relevant to failed and unexpectedly expensive builds.
- Docker Docs — Networking overview — container network namespaces, default versus user-defined networks, and embedded DNS at 127.0.0.11.
- Docker Docs — Bridge network driver — automatic name resolution on user-defined bridges and network isolation.
- Docker Docs — Storage — writable-layer versus persistent mount boundaries.
- Docker Docs — Volumes — Docker-managed persistent data and lifecycle evidence.
- Docker Docs — Bind mounts — daemon-host source paths, read/write authority, and host coupling.
- Docker Docs — Storage drivers — copy-on-write writable-layer behavior and performance implications.
- Docker Docs — Running containers — process state, exit behavior, health checks, resource and mount options.
- Docker Docs — Resource constraints — distinguish host pressure from configured CPU/memory/PID limits.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.