Chapter 41Lesson 05~270 minutes

Checkpoint Lab — Troubleshooting Docker: Daemon Failures, Networking, DNS, Storage, Permissions, Builds, and Container Incidents

Run a multi-fault incident drill, build a timestamped evidence packet, localize each failure to its owning layer, apply minimal fixes, and convert findings into runbook prevention checks.

CheckpointMulti-fault drillEvidence packetMinimal repairRunbook

Learning objectives

  • Create three simultaneous but independent disposable Docker failures and preserve their original object/evidence identities.
  • Build a timestamped incident packet containing context, versions, IDs, inspect/log/event/network/mount evidence, and assumptions.
  • Localize each fault to process, DNS/network identity, or mount authority without restarting Docker or changing host security policy.
  • Apply one minimal correction per fault and repeat the exact failed check.
  • Convert observed causes into concrete runbook/prevention controls before cleanup.

1. Scenario and safety boundary

You are the on-call engineer for a small local Docker workload. Three alerts arrive together: one container keeps restarting, a client cannot resolve its peer, and a writer cannot create a file. Your job is not to make everything green as quickly as possible; it is to identify three owners, preserve the first-failure packet, repair each with minimal scope, and prove that no unrelated Docker state was touched.

Safety boundary. Use only an authorized local/disposable context. Do not restart Docker, change daemon configuration, alter host firewall/DNS, broaden host permissions, or run Docker-wide cleanup during this checkpoint.

2. Preflight and immutable lab identity

set -eu
mkdir -p ch41-checkpoint/evidence ch41-checkpoint/hostdata
cd ch41-checkpoint
printf 'timestamp=%s
' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" | tee evidence/preflight.txt
printf 'context=%s
' "$(docker context show)" | tee -a evidence/preflight.txt
docker version | tee -a evidence/preflight.txt
docker info --format 'Server={{.ServerVersion}} Driver={{.Driver}} Cgroup={{.CgroupVersion}}' | tee -a evidence/preflight.txt
docker compose version 2>/dev/null | tee -a evidence/preflight.txt || true
docker buildx version | tee -a evidence/preflight.txt

docker pull alpine:3.22
BASE_REF="$(docker image inspect alpine:3.22 --format '{{index .RepoDigests 0}}')"
printf 'base=%s
' "$BASE_REF" | tee evidence/base.txt

If the context is shared/production, stop. The lab creates intentionally broken objects.

3. Predict before creating the incident

Prediction A: ch41cp-crash will exit 23 and restart only according to its bounded on-failure policy.
Prediction B: a probe on ch41cp-net will fail to resolve ch41cp-peer-typo but will resolve ch41cp-peer.
Prediction C: a process mounting hostdata read-only will fail to create /data/result.txt even though the host directory exists.
Invariant: daemon/context, base digest, host DNS/firewall/security policy, and unrelated Docker objects remain unchanged.

4. Create the three failures

docker rm -f ch41cp-crash ch41cp-peer 2>/dev/null || true
docker network rm ch41cp-net 2>/dev/null || true
docker network create --label devops.academy.lab=ch41cp ch41cp-net

docker run -d --name ch41cp-peer --network ch41cp-net   --label devops.academy.lab=ch41cp "$BASE_REF" sleep 600

docker run -d \
  --name ch41cp-crash \
  --restart on-failure:2 \
  --label devops.academy.lab=ch41cp "$BASE_REF"   sh -c 'echo "checkpoint synthetic crash" >&2; exit 23'

# DNS failure: preserve output, but do not change host DNS.
docker run --rm --network ch41cp-net "$BASE_REF" nslookup ch41cp-peer-typo   > evidence/dns-fail.txt 2>&1 || true

# Storage failure: explicit read-only bind.
docker run \
  --rm \
  --name ch41cp-ro \
  --mount type=bind,source="$PWD/hostdata",target=/data,readonly   "$BASE_REF" sh -c 'echo checkpoint > /data/result.txt'   > evidence/storage-fail.txt 2>&1 || true
sleep 4

5. Freeze the first-failure packet

date -u +%Y-%m-%dT%H:%M:%SZ | tee evidence/first-failure-time.txt
docker context inspect "$(docker context show)" > evidence/context.json
docker ps -a --no-trunc --filter label=devops.academy.lab=ch41cp > evidence/ps.txt

docker inspect ch41cp-crash > evidence/crash-inspect.json
docker logs --timestamps ch41cp-crash > evidence/crash.log 2>&1 || true
END="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
docker events --since 10m --until "$END" --filter container=ch41cp-crash > evidence/crash.events || true

docker network inspect ch41cp-net > evidence/network.json
docker inspect ch41cp-peer --format '{{json .NetworkSettings.Networks}}' > evidence/peer-network.json

# Recreate only a stopped inspectable object for the mount contract; do not run it.
docker create --name ch41cp-ro-inspect   --mount type=bind,source="$PWD/hostdata",target=/data,readonly   "$BASE_REF" true >/dev/null
docker inspect ch41cp-ro-inspect --format '{{json .Mounts}}' > evidence/storage-mount.json
docker rm ch41cp-ro-inspect >/dev/null

At this point you should be able to explain all three failures without a daemon restart: process exit/restart, wrong DNS identity, and a read-only mount contract.

6. Localize fault A — process/restart layer

docker inspect ch41cp-crash --format 'Id={{.Id}} Status={{.State.Status}} Exit={{.State.ExitCode}} RestartCount={{.RestartCount}} Image={{.Image}}'
cat evidence/crash.log

Diagnosis: exit code 23 and synthetic stderr belong to the process contract. The daemon is reachable and has applied the bounded restart policy. Preserve the old ID, then replace only this lab container with a healthy process.

OLD_CRASH_ID="$(docker inspect ch41cp-crash --format '{{.Id}}')"
docker rm ch41cp-crash
docker run -d --name ch41cp-crash --label devops.academy.lab=ch41cp "$BASE_REF" sleep 600
NEW_CRASH_ID="$(docker inspect ch41cp-crash --format '{{.Id}}')"
printf 'old=%s
new=%s
' "$OLD_CRASH_ID" "$NEW_CRASH_ID" | tee evidence/crash-fix.txt
docker inspect ch41cp-crash --format 'Status={{.State.Status}} Exit={{.State.ExitCode}}' | tee -a evidence/crash-fix.txt

7. Localize fault B — Docker DNS identity layer

The network evidence shows ch41cp-peer attached to ch41cp-net. The failing query asks for ch41cp-peer-typo; no host DNS change is justified.

cat evidence/dns-fail.txt
docker run --rm --network ch41cp-net "$BASE_REF" nslookup ch41cp-peer   | tee evidence/dns-fixed.txt

Diagnosis: identity/configuration typo on the user-defined network. The repair is the correct peer name or declared alias.

8. Localize fault C — mount authority layer

cat evidence/storage-fail.txt
cat evidence/storage-mount.json

Diagnosis: the mount was explicitly read-only. The requirement says this synthetic writer must create one file, so the minimal repair is to make this exact lab mount writable—not change global permissions.

docker run \
  --rm \
  --name ch41cp-rw \
  --mount type=bind,source="$PWD/hostdata",target=/data   "$BASE_REF" sh -c 'echo checkpoint > /data/result.txt && cat /data/result.txt'   | tee evidence/storage-fixed.txt
cat hostdata/result.txt

9. Verify the predicted state changes independently

# A: replacement process is running.
docker inspect ch41cp-crash --format 'Status={{.State.Status}} RestartCount={{.RestartCount}}'

# B: intended peer resolves on the same user-defined network.
docker run --rm --network ch41cp-net "$BASE_REF" nslookup ch41cp-peer

# C: intended host side effect exists only in the lab directory.
test -f hostdata/result.txt && printf 'storage_side_effect=present
'

# Invariant: current context did not change.
printf 'context_after=%s
' "$(docker context show)" | tee evidence/post.txt

10. Optional daemon-log snapshot

If you are on native Linux/systemd and are authorized to read service logs, add a bounded daemon log snapshot to the packet. On Docker Desktop or Windows containers, use the platform-specific documented log path instead; do not invent a Linux path.

journalctl -u docker.service --since '-20 min' --no-pager 2>/dev/null | tail -300 > evidence/daemon.log || true

11. Evidence packet checklist

Evidence File/command What it proves
Timestamp/context/version evidence/preflight.txt + context.json Which client/daemon/platform was investigated
Base image identity evidence/base.txt Immutable lab image reference
Crash state/log/events crash-inspect.json, crash.log, crash.events Process/restart ownership
Network state + failure/success network.json, dns-fail.txt, dns-fixed.txt Name-resolution identity cause and repair
Mount contract + failure/success storage-mount.json, storage-fail.txt, storage-fixed.txt Read-only mount cause and narrow writable repair
Before/after object identity crash-fix.txt Failed container was preserved before replacement
Assumptions/limitations incident-note below What was not tested or generalized

12. Write the incident and prevention note

Incident start UTC:
Context/endpoint:
Engine/API/base digest:
Fault A owner: process/restart
Fault A evidence:
Fault A minimal fix:
Fault B owner: Docker network/DNS identity
Fault B evidence:
Fault B minimal fix:
Fault C owner: mount read/write authority
Fault C evidence:
Fault C minimal fix:
Unchanged boundaries: daemon config, host DNS/firewall/security policy, unrelated Docker objects
Limitations: single-host disposable lab; no external registry/provider incident simulated
Prevention 1: capture context + exact IDs before mutation
Prevention 2: alert on restart count and preserve bounded logs/events
Prevention 3: integration-test expected service names on declared network
Prevention 4: contract-test mount RW/UID expectations
Runbook change: require same failed check to be repeated after every repair

13. Bounded cleanup and evidence retention

docker rm -f ch41cp-crash ch41cp-peer 2>/dev/null || true
docker network rm ch41cp-net 2>/dev/null || true
# Keep ch41-checkpoint/evidence if you want the incident packet.
# Remove only the disposable host side-effect directory when no longer needed:
rm -rf hostdata

No Docker-wide cleanup is necessary. If you remove the evidence directory, do so only after deciding that retention is no longer required.

14. What Chapter 41 adds to the operating model

You can now diagnose Docker incidents without starting from destruction: prove context and daemon identity, bind every claim to exact IDs/digests/timestamps, preserve logs/events/inspect/build evidence, localize process/network/storage/security/external layers, apply the smallest correction, and repeat the same failed verification. Chapter 42 turns these practices into the final production capstone: build, secure, publish, operate, observe, and recover a complete Dockerized application with an end-to-end evidence packet.

Knowledge check

Why does the checkpoint preserve the old crash container ID before replacement?

Which evidence proves the DNS incident is not a host-DNS outage?

Why is making only the checkpoint bind writable a valid minimal repair?

What important Docker layers did this checkpoint deliberately not modify?

What closes an incident more strongly than “the service looks okay now”?

Next lesson

Next: Production Capstone: Build, Secure, Publish, Operate, Observe, and Recover a Complete Dockerized Application: Requirements, Constraints, and Target Architecture

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Version-sensitive explanations use Docker Engine/CLI 29.8.1 as the current Engine release baseline, while labs record the learner’s actual Engine, API, Compose, Buildx, BuildKit, containerd/runc exposure, kernel/cgroup mode, context, and image digest before interpreting evidence.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.