Checkpoint Lab — Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents
Run a multi-fault incident drill, build a timestamped evidence packet, localize each failure to its owning layer, apply minimal fixes, and convert findings into runbook prevention checks.
Learning objectives
- Create three simultaneous but independent disposable Docker failures and preserve their original object/evidence identities.
- Build a timestamped incident packet containing context, versions, IDs, inspect/log/event/network/mount evidence, and assumptions.
- Localize each fault to process, DNS/network identity, or mount authority without restarting Docker or changing host security policy.
- Apply one minimal correction per fault and repeat the exact failed check.
- Convert observed causes into concrete runbook/prevention controls before cleanup.
1. Scenario and safety boundary
You are the on-call engineer for a small local Docker workload. Three alerts arrive together: one container keeps restarting, a client cannot resolve its peer, and a writer cannot create a file. Your job is not to make everything green as quickly as possible; it is to identify three owners, preserve the first-failure packet, repair each with minimal scope, and prove that no unrelated Docker state was touched.
2. Preflight and immutable lab identity
set -eu
mkdir -p ch41-checkpoint/evidence ch41-checkpoint/hostdata
cd ch41-checkpoint
printf 'timestamp=%s
' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" | tee evidence/preflight.txt
printf 'context=%s
' "$(docker context show)" | tee -a evidence/preflight.txt
docker version | tee -a evidence/preflight.txt
docker info --format 'Server={{.ServerVersion}} Driver={{.Driver}} Cgroup={{.CgroupVersion}}' | tee -a evidence/preflight.txt
docker compose version 2>/dev/null | tee -a evidence/preflight.txt || true
docker buildx version | tee -a evidence/preflight.txt
docker pull alpine:3.22
BASE_REF="$(docker image inspect alpine:3.22 --format '{{index .RepoDigests 0}}')"
printf 'base=%s
' "$BASE_REF" | tee evidence/base.txt
If the context is shared/production, stop. The lab creates intentionally broken objects.
3. Predict before creating the incident
Prediction A: ch41cp-crash will exit 23 and restart only according to its bounded on-failure policy.
Prediction B: a probe on ch41cp-net will fail to resolve ch41cp-peer-typo but will resolve ch41cp-peer.
Prediction C: a process mounting hostdata read-only will fail to create /data/result.txt even though the host directory exists.
Invariant: daemon/context, base digest, host DNS/firewall/security policy, and unrelated Docker objects remain unchanged.
4. Create the three failures
docker rm -f ch41cp-crash ch41cp-peer 2>/dev/null || true
docker network rm ch41cp-net 2>/dev/null || true
docker network create --label devops.academy.lab=ch41cp ch41cp-net
docker run -d --name ch41cp-peer --network ch41cp-net --label devops.academy.lab=ch41cp "$BASE_REF" sleep 600
docker run -d \
--name ch41cp-crash \
--restart on-failure:2 \
--label devops.academy.lab=ch41cp "$BASE_REF" sh -c 'echo "checkpoint synthetic crash" >&2; exit 23'
# DNS failure: preserve output, but do not change host DNS.
docker run --rm --network ch41cp-net "$BASE_REF" nslookup ch41cp-peer-typo > evidence/dns-fail.txt 2>&1 || true
# Storage failure: explicit read-only bind.
docker run \
--rm \
--name ch41cp-ro \
--mount type=bind,source="$PWD/hostdata",target=/data,readonly "$BASE_REF" sh -c 'echo checkpoint > /data/result.txt' > evidence/storage-fail.txt 2>&1 || true
sleep 4
5. Freeze the first-failure packet
date -u +%Y-%m-%dT%H:%M:%SZ | tee evidence/first-failure-time.txt
docker context inspect "$(docker context show)" > evidence/context.json
docker ps -a --no-trunc --filter label=devops.academy.lab=ch41cp > evidence/ps.txt
docker inspect ch41cp-crash > evidence/crash-inspect.json
docker logs --timestamps ch41cp-crash > evidence/crash.log 2>&1 || true
END="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
docker events --since 10m --until "$END" --filter container=ch41cp-crash > evidence/crash.events || true
docker network inspect ch41cp-net > evidence/network.json
docker inspect ch41cp-peer --format '{{json .NetworkSettings.Networks}}' > evidence/peer-network.json
# Recreate only a stopped inspectable object for the mount contract; do not run it.
docker create --name ch41cp-ro-inspect --mount type=bind,source="$PWD/hostdata",target=/data,readonly "$BASE_REF" true >/dev/null
docker inspect ch41cp-ro-inspect --format '{{json .Mounts}}' > evidence/storage-mount.json
docker rm ch41cp-ro-inspect >/dev/null
At this point you should be able to explain all three failures without a daemon restart: process exit/restart, wrong DNS identity, and a read-only mount contract.
6. Localize fault A — process/restart layer
docker inspect ch41cp-crash --format 'Id={{.Id}} Status={{.State.Status}} Exit={{.State.ExitCode}} RestartCount={{.RestartCount}} Image={{.Image}}'
cat evidence/crash.log
Diagnosis: exit code 23 and synthetic stderr belong to the process contract. The daemon is reachable and has applied the bounded restart policy. Preserve the old ID, then replace only this lab container with a healthy process.
OLD_CRASH_ID="$(docker inspect ch41cp-crash --format '{{.Id}}')"
docker rm ch41cp-crash
docker run -d --name ch41cp-crash --label devops.academy.lab=ch41cp "$BASE_REF" sleep 600
NEW_CRASH_ID="$(docker inspect ch41cp-crash --format '{{.Id}}')"
printf 'old=%s
new=%s
' "$OLD_CRASH_ID" "$NEW_CRASH_ID" | tee evidence/crash-fix.txt
docker inspect ch41cp-crash --format 'Status={{.State.Status}} Exit={{.State.ExitCode}}' | tee -a evidence/crash-fix.txt
7. Localize fault B — Docker DNS identity layer
The network evidence shows ch41cp-peer attached to
ch41cp-net. The failing query asks for
ch41cp-peer-typo; no host DNS change is justified.
cat evidence/dns-fail.txt
docker run --rm --network ch41cp-net "$BASE_REF" nslookup ch41cp-peer | tee evidence/dns-fixed.txt
Diagnosis: identity/configuration typo on the user-defined network. The repair is the correct peer name or declared alias.
9. Verify the predicted state changes independently
# A: replacement process is running.
docker inspect ch41cp-crash --format 'Status={{.State.Status}} RestartCount={{.RestartCount}}'
# B: intended peer resolves on the same user-defined network.
docker run --rm --network ch41cp-net "$BASE_REF" nslookup ch41cp-peer
# C: intended host side effect exists only in the lab directory.
test -f hostdata/result.txt && printf 'storage_side_effect=present
'
# Invariant: current context did not change.
printf 'context_after=%s
' "$(docker context show)" | tee evidence/post.txt
10. Optional daemon-log snapshot
If you are on native Linux/systemd and are authorized to read service logs, add a bounded daemon log snapshot to the packet. On Docker Desktop or Windows containers, use the platform-specific documented log path instead; do not invent a Linux path.
journalctl -u docker.service --since '-20 min' --no-pager 2>/dev/null | tail -300 > evidence/daemon.log || true
11. Evidence packet checklist
| Evidence | File/command | What it proves |
|---|---|---|
| Timestamp/context/version | evidence/preflight.txt + context.json | Which client/daemon/platform was investigated |
| Base image identity | evidence/base.txt | Immutable lab image reference |
| Crash state/log/events | crash-inspect.json, crash.log, crash.events | Process/restart ownership |
| Network state + failure/success | network.json, dns-fail.txt, dns-fixed.txt | Name-resolution identity cause and repair |
| Mount contract + failure/success | storage-mount.json, storage-fail.txt, storage-fixed.txt | Read-only mount cause and narrow writable repair |
| Before/after object identity | crash-fix.txt | Failed container was preserved before replacement |
| Assumptions/limitations | incident-note below | What was not tested or generalized |
12. Write the incident and prevention note
Incident start UTC:
Context/endpoint:
Engine/API/base digest:
Fault A owner: process/restart
Fault A evidence:
Fault A minimal fix:
Fault B owner: Docker network/DNS identity
Fault B evidence:
Fault B minimal fix:
Fault C owner: mount read/write authority
Fault C evidence:
Fault C minimal fix:
Unchanged boundaries: daemon config, host DNS/firewall/security policy, unrelated Docker objects
Limitations: single-host disposable lab; no external registry/provider incident simulated
Prevention 1: capture context + exact IDs before mutation
Prevention 2: alert on restart count and preserve bounded logs/events
Prevention 3: integration-test expected service names on declared network
Prevention 4: contract-test mount RW/UID expectations
Runbook change: require same failed check to be repeated after every repair
13. Bounded cleanup and evidence retention
docker rm -f ch41cp-crash ch41cp-peer 2>/dev/null || true
docker network rm ch41cp-net 2>/dev/null || true
# Keep ch41-checkpoint/evidence if you want the incident packet.
# Remove only the disposable host side-effect directory when no longer needed:
rm -rf hostdata
No Docker-wide cleanup is necessary. If you remove the evidence directory, do so only after deciding that retention is no longer required.
14. What Chapter 41 adds to the operating model
You can now diagnose Docker incidents without starting from destruction: prove context and daemon identity, bind every claim to exact IDs/digests/timestamps, preserve logs/events/inspect/build evidence, localize process/network/storage/security/external layers, apply the smallest correction, and repeat the same failed verification. Chapter 42 turns these practices into the final production capstone: build, secure, publish, operate, observe, and recover a complete Dockerized application with an end-to-end evidence packet.
Knowledge check
Why does the checkpoint preserve the old crash container ID before replacement?
Replacement is the correction, but the old ID ties the original state/logs/events to the failure and prevents the new healthy container from rewriting incident history.
Which evidence proves the DNS incident is not a host-DNS outage?
The peer is present on the inspected user-defined network and the same scoped probe resolves the correct peer name while the typo fails.
Why is making only the checkpoint bind writable a valid minimal repair?
The writer’s contract requires a write and inspect evidence proves the failure is the read-only mount flag; unrelated permissions and host policy remain unchanged.
What important Docker layers did this checkpoint deliberately not modify?
Daemon configuration, host DNS/firewall/security policy, registry credentials, shared images/caches, and unrelated containers/networks.
What closes an incident more strongly than “the service looks okay now”?
Repeating the exact failed verification and preserving before/after evidence tied to the same context and intended object identity.
Official references and version notes
2026-09-22. Version-sensitive explanations use Docker Engine/CLI 29.8.1 as the current Engine release baseline, while labs record the learner’s actual Engine, API, Compose, Buildx, BuildKit, containerd/runc exposure, kernel/cgroup mode, context, and image digest before interpreting evidence.
- Docker Docs — Troubleshoot the Docker daemon — daemon startup/configuration and networking/DNS troubleshooting.
- Docker Docs — Read the daemon logs — Linux journal/syslog locations, Docker Desktop VM logs, Windows Event Log, debug logging, and non-destructive SIGUSR1 stack traces.
- Docker Engine 29 release notes — current 29.8.1 fixes, security changes, regressions, and documented compatibility notes.
- docker context show and docker context inspect — prove the daemon endpoint before mutation.
- docker inspect — low-level Docker object state, IDs, mounts, network settings, health, and runtime configuration.
- docker logs — bounded container stdout/stderr evidence.
- docker events — daemon object event timeline and filters.
-
docker buildx build
—
--progress=plainand raw build progress for reproducible failure traces. - Docker Docs — Optimize build cache — input/cache boundaries relevant to failed and unexpectedly expensive builds.
- Docker Docs — Networking overview — container network namespaces, default versus user-defined networks, and embedded DNS at 127.0.0.11.
- Docker Docs — Bridge network driver — automatic name resolution on user-defined bridges and network isolation.
- Docker Docs — Storage — writable-layer versus persistent mount boundaries.
- Docker Docs — Volumes — Docker-managed persistent data and lifecycle evidence.
- Docker Docs — Bind mounts — daemon-host source paths, read/write authority, and host coupling.
- Docker Docs — Storage drivers — copy-on-write writable-layer behavior and performance implications.
- Docker Docs — Running containers — process state, exit behavior, health checks, resource and mount options.
- Docker Docs — Resource constraints — distinguish host pressure from configured CPU/memory/PID limits.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.