Chapter 41Lesson 03~190 minutes

Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents: Configuration, Design Choices, and Tradeoffs

Choose troubleshooting actions deliberately: preserve versus restart, inspect versus recreate, targeted cleanup versus deletion, debug techniques, and daemon versus application rollback.

TradeoffsPreserve evidenceDebuggingRollbackCleanup

Learning objectives

  • Choose between preserving evidence, restart, recreation, targeted cleanup, and rollback based on what each action destroys or changes.
  • Decide when docker exec, a purpose-built helper/debug container, or host/daemon inspection is the appropriate boundary.
  • Separate application rollback from daemon/host rollback.
  • Use current Engine 29 release notes and platform-specific daemon-log locations as compatibility evidence rather than folklore.
  • Build a troubleshooting decision record with prerequisites, affected state, expected evidence, and rollback.

1. Every diagnostic action has an evidence cost

A restart, recreation, cleanup, or policy change is not “just troubleshooting.” It mutates the system being observed. Good incident work asks two questions before acting: what evidence will this action destroy? and what hypothesis will it test? If neither answer is clear, the action is probably premature.

2. Preserve versus restart

Choice Use when Evidence/side effect Guardrail
Preserve in place Object still exists and evidence is readable Keeps ID, timestamps, writable layer, state, logs Default first move
Restart one container Hypothesis is process initialization/transient runtime and evidence is saved Changes PID/timestamps/process state Record old state/logs first
Restart daemon Daemon itself is proven unhealthy or config requires it Affects every local workload; may reset daemon timeline Maintenance plan + daemon logs/config captured
Host reboot Kernel/host issue is evidenced and operationally approved Widest local change Last-resort planned action, not a generic Docker fix

“Restart Docker first” is attractive because it sometimes changes symptoms. It is weak troubleshooting because it changes many variables at once.

3. Recreate versus inspect in place

Recreation is appropriate when immutable deployment is the intended correction—but only after evidence from the failed object has been preserved. Recreating from a mutable tag can also pull different bytes, turning a runtime test into an artifact change. Record the failed container ID, image ID/digest, mounts, networks, config, logs, and exit/health state first.

4. Targeted cleanup versus no deletion

Disk pressure does not justify deleting unknown Docker state. Start with accounting: host bytes/inodes, docker system df -v, docker ps --size, docker buildx du, exact volume usage, and logging configuration. If cleanup is required, select exact labeled/known objects or a narrowly reviewed filter. During an incident, “reclaimable” does not mean “operationally irrelevant.”

5. Exec debugging versus helper/debug container

docker exec runs a new process inside the existing container namespaces/filesystem and is useful when the image contains the needed tool and the container is running. Minimal/distroless images may have no shell or diagnostic utilities; installing tools into a live production container mutates evidence and can violate immutability. A purpose-built helper can instead join only the necessary network or mount boundary in an authorized lab, preserving the target image. If Docker Debug is available in your environment, treat it as an optional tooling path and still preserve object identity/evidence first.

Safety boundary. Do not turn “need more visibility” into host socket mounting, privileged mode, host PID/network namespaces, or disabled seccomp/LSM policy. Those change the trust boundary rather than merely improving observation.

6. Daemon logs are platform-owned evidence

Platform Documented starting point Important distinction
Linux with systemd journalctl -xu docker.service Host daemon/service logs
Older/non-systemd Linux syslog/messages/daemon/docker logs depending on distro Distribution-specific
Docker Desktop macOS ~/Library/Containers/com.docker.docker/Data/log/vm/init.log dockerd/containerd run inside Desktop VM
Docker Desktop Windows WSL2 %LOCALAPPDATA%\Docker\log\vm\init.log VM service multiplexed log
Windows containers Windows Event Log Windows daemon/runtime path

Debug logging can be enabled, but that is a daemon configuration change. Capture existing logs/config and understand reload/restart behavior before changing log level during an incident.

7. Application rollback versus daemon rollback

If digest B crashes and digest A is known good, an application rollback should move only the application artifact/config boundary. If a daemon upgrade introduced a storage/network regression, rollback planning is host infrastructure work and may involve data-format compatibility. Do not downgrade Engine across storage/image-store migrations casually. Record Engine/containerd/storage state and consult release notes.

8. Engine 29 compatibility belongs in the incident record

As of the 2026-09-22 baseline, Engine 29.8.1 includes fixes around the containerd image store, managed-containerd startup warnings, OpenVZ user-namespace detection, and network filtering. Earlier Engine 29 releases also carried security fixes and regressions. A particularly instructive compatibility case is the seccomp hardening introduced around 29.4.2 for the AF_ALG/kernel crypto vulnerability: some 32-bit/Wine/SteamCMD workloads can break. Docker explicitly warns against solving that with an unconfined seccomp profile. This is why the exact Engine/kernel/workload version belongs in troubleshooting evidence.

9. Network/DNS correction versus host-policy mutation

If a user-defined bridge peer name is wrong, fix the name/alias. If external DNS fails, inspect /etc/resolv.conf, the embedded resolver, upstream server reachability, and daemon DNS configuration. Do not change host DNS and firewall simultaneously. Similarly, published-port reachability requires separate proof of app listener, container endpoint, Docker publish state, host bind/interface, firewall/routing, and client path.

10. Security denial: evidence, not an invitation to disable controls

A seccomp/AppArmor/SELinux/capability denial can be the correct enforcement of policy. Capture the requested operation, effective container security settings, host policy log, image/workload version, and intended privilege requirement. The correction should be the narrowest policy/capability change justified by the application—not blanket privilege.

11. Worked decision table

Scenario Preferred first action Prerequisites/evidence Why
Container exits repeatedly Inspect state/logs/events in place Exact container/image ID, restart policy Preserves failure before recreation
Daemon API unavailable Verify context/socket/service + daemon logs Host/context identity Separates wrong endpoint from daemon failure
Disk pressure Inventory bytes/inodes + Docker usage Host filesystem + object ownership Prevents unrelated deletion
Distroless app has DNS issue Use inspect + scoped helper or available debug tooling Target network identity, no privilege escalation Avoids mutating production image
Known-good app digest exists Rollback app artifact only Digest A/B + config compatibility Keeps daemon/host constant
Possible Engine regression Capture versions/release-note match before host change Engine/kernel/storage/network evidence Avoids blind downgrade/restart

12. Troubleshooting decision record template

Symptom and first timestamp:
Affected object IDs/digests:
Context/daemon endpoint:
Host/kernel/Engine/API versions:
First-failure evidence preserved:
Owning layer hypothesis:
Alternative hypotheses rejected by:
Chosen action:
State changed by action:
Evidence that may be lost:
Rollback path:
Exact verification repeated:
Result:
Runbook/prevention update:

Knowledge check

When is container recreation appropriate during troubleshooting?

Why can a daemon restart be riskier than restarting one container?

What is the main advantage of a scoped helper/debug container for a minimal image?

Why must Engine/kernel versions be recorded for a seccomp compatibility incident?

What is wrong with treating reclaimable disk bytes as automatically safe to delete?

Next lesson

Next: Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents: Diagnostics, Failure Modes, Security, and Performance

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Version-sensitive explanations use Docker Engine/CLI 29.8.1 as the current Engine release baseline, while labs record the learner’s actual Engine, API, Compose, Buildx, BuildKit, containerd/runc exposure, kernel/cgroup mode, context, and image digest before interpreting evidence.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.