Chapter 24Lesson 04~120 minutes

Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Diagnostics, Failure Modes, Security, and Performance

Diagnose broken probes, shallow healthchecks, restart loops, dependency misuse, and aggressive probe load from preserved first-failure evidence.

Health checksRestart policiesGraceful shutdownDependency healthEvidence-first

Learning objectives

  • Diagnose missing probe tooling, shallow PID-only checks, restart loops, aggressive probe load, and dependency-gating misuse.
  • Preserve health history, logs, events, restart count, exit code, and first-failure evidence before corrective action.
  • Separate client/context, image, process, health, dependency, network, and external-availability failure domains.
  • Repair only the smallest faulty layer without disabling security or hiding the original cause.
Evidence-first rule. Before “fixing” health or restart behavior, preserve the image digest, container ID, configured probe/timing, recent health log, normal logs, Engine events, exit/OOM state, restart count, dependency state, and external reachability evidence. Automatic recovery can erase the clearest first-failure state.

1. Diagnostic sequence: start with the smallest causal boundary

  1. Preserve first-failure timestamps, logs, events, health history, and inspect state.
  2. Confirm Docker context, Engine/CLI/Compose versions, host/platform.
  3. Confirm exact image identity and whether health/STOPSIGNAL are inherited or overridden.
  4. Inspect process state, exit code, OOM flag, restart count, health status, and health log.
  5. Inspect dependency health and application retry behavior.
  6. Inspect networking/published path only after proving the local listener/probe layer.
  7. Check external proxy/TLS/DNS/load-balancer evidence last.
  8. Correct the smallest faulty layer and rerun only that scope.
docker inspect NAME --format '{{json .State}}' > state.json
docker inspect NAME --format '{{json .Config.Healthcheck}}' > health-config.json
docker inspect NAME --format '{{json .HostConfig.RestartPolicy}}' > restart-policy.json
docker logs --timestamps --since 10m NAME > app.log 2>&1
docker events --since 10m --filter container=NAME --format '{{json .}}' > events.jsonl

2. Failure mode: the healthcheck depends on a missing tool

A minimal image may not contain curl, wget, a shell, or even standard utilities. If the probe command cannot execute, health fails because the diagnostic dependency is missing—not because the service is unhealthy.

docker rm -f dkr24-broken-probe 2>/dev/null || true

docker run -d \
  --name dkr24-broken-probe \
  --health-cmd='curl -fsS http://127.0.0.1:8080/health' \
  --health-interval=2s \
  --health-retries=2   alpine:3.22 sh -c 'while :; do sleep 1; done'

sleep 6
docker inspect dkr24-broken-probe --format '{{json .State.Health}}'   | tee dkr24-broken-probe-health.json
Interpretation. The health log should show that curl is unavailable. The repair is to use a probe mechanism actually present (or deliberately add a small probe helper during image design), not to restart the container repeatedly or disable health checks blindly.

3. Failure mode: probing only PID existence

A PID-only check can remain green during a deadlock, stuck event loop, exhausted connection pool, or broken listener. If Docker already reports PID 1 running, a healthcheck that merely confirms the same fact adds little diagnostic value. Probe the smallest application behavior that represents useful service.

4. Failure mode: restart loops hide first-failure evidence

An aggressive policy can make a crash appear as “flapping.” Preserve the first exit code and earliest logs/events before changing restart behavior. Backoff reduces restart flood but does not fix the application.

Evidence Why it matters
.State.ExitCode / OOMKilled Separates process failure from memory kill.
.RestartCount + events Shows repeated Engine recovery attempts.
Earliest logs after each start Can expose deterministic startup failure before logs rotate.
Image digest + config Proves whether every restart used the same artifact/config.

5. Failure mode: treating health as an SLA

A healthy local probe does not guarantee that the service is reachable from a user network, meets latency/error-rate objectives, or can access all required dependencies. Health is one local signal. Availability and SLOs need external telemetry and request-level evidence.

6. Failure mode: using depends_on instead of resilient clients

Compose can gate initial startup, but once the stack is running a dependency can fail. If the application has no retry/reconnect behavior, it remains fragile. Repair the client resilience logic; do not keep adding more startup ordering.

7. Failure mode: healthchecks create their own load

A probe that performs an expensive database query every second across hundreds of replicas can become a workload. Measure probe latency/cost and choose a cheap endpoint. Failure detection time should be short enough for operations but not so aggressive that the check becomes the incident.

8. Repair the intentionally broken probe without hiding evidence

Keep the original inspect output, then replace the disposable container with a probe that uses only known tooling. Recreating a disposable lab container is appropriate because health configuration is part of container configuration; the important requirement is to preserve the failed object’s evidence first.

docker inspect dkr24-broken-probe > dkr24-broken-probe-inspect-before.json

docker rm -f dkr24-broken-probe

docker run -d \
  --name dkr24-fixed-probe \
  --health-cmd='test -f /tmp/ready' \
  --health-interval=2s \
  --health-timeout=1s \
  --health-retries=2   alpine:3.22 sh -c 'touch /tmp/ready; while :; do sleep 1; done'

sleep 5
docker inspect dkr24-fixed-probe --format '{{json .State.Health}}'   | tee dkr24-fixed-probe-health.json

9. Security and disruption boundaries

Health troubleshooting does not justify --privileged, Docker socket mounts, broad root access, disabled seccomp/AppArmor/SELinux, firewall disabling, or printing secrets in probe output. Probe output is inspectable and should be treated as potentially shareable diagnostics. Keep credentials out of health command arguments and output.

10. Exact cleanup

docker rm -f dkr24-broken-probe dkr24-fixed-probe 2>/dev/null || true

Knowledge check

A health log says “curl: not found.” Is the application proven unhealthy?

Why is a PID-existence healthcheck often weak?

What should you preserve before changing a restart policy during a crash loop?

Why can service_healthy still leave a fragile application?

What is a common performance mistake with probes?

Next lesson

Next: Checkpoint Lab — Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline used for compatibility notes. Current Dockerfile HEALTHCHECK supports interval, timeout, start_period, start_interval (Engine 25+), and retries. Engine 29 includes a fix for health checks delayed too long when start_interval exceeds start_period. The labs record the actual Engine/CLI/Compose versions before relying on these semantics.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.