Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents: Configuration, Design Choices, and Tradeoffs
Choose troubleshooting actions deliberately: preserve versus restart, inspect versus recreate, targeted cleanup versus deletion, debug techniques, and daemon versus application rollback.
Learning objectives
- Choose between preserving evidence, restart, recreation, targeted cleanup, and rollback based on what each action destroys or changes.
-
Decide when
docker exec, a purpose-built helper/debug container, or host/daemon inspection is the appropriate boundary. - Separate application rollback from daemon/host rollback.
- Use current Engine 29 release notes and platform-specific daemon-log locations as compatibility evidence rather than folklore.
- Build a troubleshooting decision record with prerequisites, affected state, expected evidence, and rollback.
1. Every diagnostic action has an evidence cost
A restart, recreation, cleanup, or policy change is not “just troubleshooting.” It mutates the system being observed. Good incident work asks two questions before acting: what evidence will this action destroy? and what hypothesis will it test? If neither answer is clear, the action is probably premature.
2. Preserve versus restart
| Choice | Use when | Evidence/side effect | Guardrail |
|---|---|---|---|
| Preserve in place | Object still exists and evidence is readable | Keeps ID, timestamps, writable layer, state, logs | Default first move |
| Restart one container | Hypothesis is process initialization/transient runtime and evidence is saved | Changes PID/timestamps/process state | Record old state/logs first |
| Restart daemon | Daemon itself is proven unhealthy or config requires it | Affects every local workload; may reset daemon timeline | Maintenance plan + daemon logs/config captured |
| Host reboot | Kernel/host issue is evidenced and operationally approved | Widest local change | Last-resort planned action, not a generic Docker fix |
“Restart Docker first” is attractive because it sometimes changes symptoms. It is weak troubleshooting because it changes many variables at once.
3. Recreate versus inspect in place
Recreation is appropriate when immutable deployment is the intended correction—but only after evidence from the failed object has been preserved. Recreating from a mutable tag can also pull different bytes, turning a runtime test into an artifact change. Record the failed container ID, image ID/digest, mounts, networks, config, logs, and exit/health state first.
4. Targeted cleanup versus no deletion
Disk pressure does not justify deleting unknown Docker state. Start
with accounting: host bytes/inodes,
docker system df -v, docker ps --size,
docker buildx du, exact volume usage, and logging
configuration. If cleanup is required, select exact labeled/known
objects or a narrowly reviewed filter. During an incident,
“reclaimable” does not mean “operationally irrelevant.”
5. Exec debugging versus helper/debug container
docker exec runs a new process inside the existing
container namespaces/filesystem and is useful when the image
contains the needed tool and the container is running.
Minimal/distroless images may have no shell or diagnostic utilities;
installing tools into a live production container mutates evidence
and can violate immutability. A purpose-built helper can instead
join only the necessary network or mount boundary in an authorized
lab, preserving the target image. If Docker Debug is available in
your environment, treat it as an optional tooling path and still
preserve object identity/evidence first.
6. Daemon logs are platform-owned evidence
| Platform | Documented starting point | Important distinction |
|---|---|---|
| Linux with systemd | journalctl -xu docker.service | Host daemon/service logs |
| Older/non-systemd Linux | syslog/messages/daemon/docker logs depending on distro | Distribution-specific |
| Docker Desktop macOS | ~/Library/Containers/com.docker.docker/Data/log/vm/init.log | dockerd/containerd run inside Desktop VM |
| Docker Desktop Windows WSL2 | %LOCALAPPDATA%\Docker\log\vm\init.log | VM service multiplexed log |
| Windows containers | Windows Event Log | Windows daemon/runtime path |
Debug logging can be enabled, but that is a daemon configuration change. Capture existing logs/config and understand reload/restart behavior before changing log level during an incident.
7. Application rollback versus daemon rollback
If digest B crashes and digest A is known good, an application rollback should move only the application artifact/config boundary. If a daemon upgrade introduced a storage/network regression, rollback planning is host infrastructure work and may involve data-format compatibility. Do not downgrade Engine across storage/image-store migrations casually. Record Engine/containerd/storage state and consult release notes.
8. Engine 29 compatibility belongs in the incident record
As of the 2026-09-22 baseline, Engine 29.8.1 includes fixes around the containerd image store, managed-containerd startup warnings, OpenVZ user-namespace detection, and network filtering. Earlier Engine 29 releases also carried security fixes and regressions. A particularly instructive compatibility case is the seccomp hardening introduced around 29.4.2 for the AF_ALG/kernel crypto vulnerability: some 32-bit/Wine/SteamCMD workloads can break. Docker explicitly warns against solving that with an unconfined seccomp profile. This is why the exact Engine/kernel/workload version belongs in troubleshooting evidence.
9. Network/DNS correction versus host-policy mutation
If a user-defined bridge peer name is wrong, fix the name/alias. If
external DNS fails, inspect /etc/resolv.conf, the
embedded resolver, upstream server reachability, and daemon DNS
configuration. Do not change host DNS and firewall simultaneously.
Similarly, published-port reachability requires separate proof of
app listener, container endpoint, Docker publish state, host
bind/interface, firewall/routing, and client path.
10. Security denial: evidence, not an invitation to disable controls
A seccomp/AppArmor/SELinux/capability denial can be the correct enforcement of policy. Capture the requested operation, effective container security settings, host policy log, image/workload version, and intended privilege requirement. The correction should be the narrowest policy/capability change justified by the application—not blanket privilege.
11. Worked decision table
| Scenario | Preferred first action | Prerequisites/evidence | Why |
|---|---|---|---|
| Container exits repeatedly | Inspect state/logs/events in place | Exact container/image ID, restart policy | Preserves failure before recreation |
| Daemon API unavailable | Verify context/socket/service + daemon logs | Host/context identity | Separates wrong endpoint from daemon failure |
| Disk pressure | Inventory bytes/inodes + Docker usage | Host filesystem + object ownership | Prevents unrelated deletion |
| Distroless app has DNS issue | Use inspect + scoped helper or available debug tooling | Target network identity, no privilege escalation | Avoids mutating production image |
| Known-good app digest exists | Rollback app artifact only | Digest A/B + config compatibility | Keeps daemon/host constant |
| Possible Engine regression | Capture versions/release-note match before host change | Engine/kernel/storage/network evidence | Avoids blind downgrade/restart |
12. Troubleshooting decision record template
Symptom and first timestamp:
Affected object IDs/digests:
Context/daemon endpoint:
Host/kernel/Engine/API versions:
First-failure evidence preserved:
Owning layer hypothesis:
Alternative hypotheses rejected by:
Chosen action:
State changed by action:
Evidence that may be lost:
Rollback path:
Exact verification repeated:
Result:
Runbook/prevention update:
Knowledge check
When is container recreation appropriate during troubleshooting?
After failed-object evidence is preserved and recreation is itself the intended minimal correction or verification—not as a way to erase unexplained state.
Why can a daemon restart be riskier than restarting one container?
It changes state for all workloads on that daemon and can erase/alter daemon timeline evidence, so the scope is much wider.
What is the main advantage of a scoped helper/debug container for a minimal image?
It can provide tools without modifying the target image, if attached only to the necessary authorized boundary.
Why must Engine/kernel versions be recorded for a seccomp compatibility incident?
Default profiles and kernel vulnerabilities/behavior change over time; version-specific compatibility can explain a denial without disabling the sandbox.
What is wrong with treating reclaimable disk bytes as automatically safe to delete?
Reclaimable is a storage-accounting property, not an operational-value judgment. Incident or rollback artifacts may still matter.
Official references and version notes
2026-09-22. Version-sensitive explanations use Docker Engine/CLI 29.8.1 as the current Engine release baseline, while labs record the learner’s actual Engine, API, Compose, Buildx, BuildKit, containerd/runc exposure, kernel/cgroup mode, context, and image digest before interpreting evidence.
- Docker Docs — Troubleshoot the Docker daemon — daemon startup/configuration and networking/DNS troubleshooting.
- Docker Docs — Read the daemon logs — Linux journal/syslog locations, Docker Desktop VM logs, Windows Event Log, debug logging, and non-destructive SIGUSR1 stack traces.
- Docker Engine 29 release notes — current 29.8.1 fixes, security changes, regressions, and documented compatibility notes.
- docker context show and docker context inspect — prove the daemon endpoint before mutation.
- docker inspect — low-level Docker object state, IDs, mounts, network settings, health, and runtime configuration.
- docker logs — bounded container stdout/stderr evidence.
- docker events — daemon object event timeline and filters.
-
docker buildx build
—
--progress=plainand raw build progress for reproducible failure traces. - Docker Docs — Optimize build cache — input/cache boundaries relevant to failed and unexpectedly expensive builds.
- Docker Docs — Networking overview — container network namespaces, default versus user-defined networks, and embedded DNS at 127.0.0.11.
- Docker Docs — Bridge network driver — automatic name resolution on user-defined bridges and network isolation.
- Docker Docs — Storage — writable-layer versus persistent mount boundaries.
- Docker Docs — Volumes — Docker-managed persistent data and lifecycle evidence.
- Docker Docs — Bind mounts — daemon-host source paths, read/write authority, and host coupling.
- Docker Docs — Storage drivers — copy-on-write writable-layer behavior and performance implications.
- Docker Docs — Running containers — process state, exit behavior, health checks, resource and mount options.
- Docker Docs — Resource constraints — distinguish host pressure from configured CPU/memory/PID limits.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.