Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Diagnostics, Failure Modes, Security, and Performance
Diagnose broken probes, shallow healthchecks, restart loops, dependency misuse, and aggressive probe load from preserved first-failure evidence.
Learning objectives
- Diagnose missing probe tooling, shallow PID-only checks, restart loops, aggressive probe load, and dependency-gating misuse.
- Preserve health history, logs, events, restart count, exit code, and first-failure evidence before corrective action.
- Separate client/context, image, process, health, dependency, network, and external-availability failure domains.
- Repair only the smallest faulty layer without disabling security or hiding the original cause.
1. Diagnostic sequence: start with the smallest causal boundary
- Preserve first-failure timestamps, logs, events, health history, and inspect state.
- Confirm Docker context, Engine/CLI/Compose versions, host/platform.
- Confirm exact image identity and whether health/STOPSIGNAL are inherited or overridden.
- Inspect process state, exit code, OOM flag, restart count, health status, and health log.
- Inspect dependency health and application retry behavior.
- Inspect networking/published path only after proving the local listener/probe layer.
- Check external proxy/TLS/DNS/load-balancer evidence last.
- Correct the smallest faulty layer and rerun only that scope.
docker inspect NAME --format '{{json .State}}' > state.json
docker inspect NAME --format '{{json .Config.Healthcheck}}' > health-config.json
docker inspect NAME --format '{{json .HostConfig.RestartPolicy}}' > restart-policy.json
docker logs --timestamps --since 10m NAME > app.log 2>&1
docker events --since 10m --filter container=NAME --format '{{json .}}' > events.jsonl
2. Failure mode: the healthcheck depends on a missing tool
A minimal image may not contain curl,
wget, a shell, or even standard utilities. If the probe
command cannot execute, health fails because the diagnostic
dependency is missing—not because the service is unhealthy.
docker rm -f dkr24-broken-probe 2>/dev/null || true
docker run -d \
--name dkr24-broken-probe \
--health-cmd='curl -fsS http://127.0.0.1:8080/health' \
--health-interval=2s \
--health-retries=2 alpine:3.22 sh -c 'while :; do sleep 1; done'
sleep 6
docker inspect dkr24-broken-probe --format '{{json .State.Health}}' | tee dkr24-broken-probe-health.json
curl is unavailable. The repair is to use a probe
mechanism actually present (or deliberately add a small probe helper
during image design), not to restart the container repeatedly or
disable health checks blindly.
3. Failure mode: probing only PID existence
A PID-only check can remain green during a deadlock, stuck event loop, exhausted connection pool, or broken listener. If Docker already reports PID 1 running, a healthcheck that merely confirms the same fact adds little diagnostic value. Probe the smallest application behavior that represents useful service.
4. Failure mode: restart loops hide first-failure evidence
An aggressive policy can make a crash appear as “flapping.” Preserve the first exit code and earliest logs/events before changing restart behavior. Backoff reduces restart flood but does not fix the application.
| Evidence | Why it matters |
|---|---|
.State.ExitCode / OOMKilled |
Separates process failure from memory kill. |
.RestartCount + events |
Shows repeated Engine recovery attempts. |
| Earliest logs after each start | Can expose deterministic startup failure before logs rotate. |
| Image digest + config | Proves whether every restart used the same artifact/config. |
5. Failure mode: treating health as an SLA
A healthy local probe does not guarantee that the service is reachable from a user network, meets latency/error-rate objectives, or can access all required dependencies. Health is one local signal. Availability and SLOs need external telemetry and request-level evidence.
6. Failure mode: using depends_on instead of resilient clients
Compose can gate initial startup, but once the stack is running a dependency can fail. If the application has no retry/reconnect behavior, it remains fragile. Repair the client resilience logic; do not keep adding more startup ordering.
7. Failure mode: healthchecks create their own load
A probe that performs an expensive database query every second across hundreds of replicas can become a workload. Measure probe latency/cost and choose a cheap endpoint. Failure detection time should be short enough for operations but not so aggressive that the check becomes the incident.
8. Repair the intentionally broken probe without hiding evidence
Keep the original inspect output, then replace the disposable container with a probe that uses only known tooling. Recreating a disposable lab container is appropriate because health configuration is part of container configuration; the important requirement is to preserve the failed object’s evidence first.
docker inspect dkr24-broken-probe > dkr24-broken-probe-inspect-before.json
docker rm -f dkr24-broken-probe
docker run -d \
--name dkr24-fixed-probe \
--health-cmd='test -f /tmp/ready' \
--health-interval=2s \
--health-timeout=1s \
--health-retries=2 alpine:3.22 sh -c 'touch /tmp/ready; while :; do sleep 1; done'
sleep 5
docker inspect dkr24-fixed-probe --format '{{json .State.Health}}' | tee dkr24-fixed-probe-health.json
9. Security and disruption boundaries
Health troubleshooting does not justify --privileged,
Docker socket mounts, broad root access, disabled
seccomp/AppArmor/SELinux, firewall disabling, or printing secrets in
probe output. Probe output is inspectable and should be treated as
potentially shareable diagnostics. Keep credentials out of health
command arguments and output.
10. Exact cleanup
docker rm -f dkr24-broken-probe dkr24-fixed-probe 2>/dev/null || true
Knowledge check
A health log says “curl: not found.” Is the application proven unhealthy?
No. The probe itself is invalid in that image. Diagnose the probe dependency first.
Why is a PID-existence healthcheck often weak?
Docker already knows whether PID 1 is running; the check may miss a live-but-unusable application.
What should you preserve before changing a restart policy during a crash loop?
First-failure logs, events, exit/OOM state, restart count, image/config identity, and timestamps.
Why can service_healthy still leave a fragile application?
It coordinates startup only; later dependency failures still require retry/reconnect behavior.
What is a common performance mistake with probes?
Running an expensive probe too frequently, creating avoidable load or even contributing to failure.
Official references and version notes
- Dockerfile reference — HEALTHCHECK — health states, timing, exit codes, output retention, and health-status events.
- Dockerfile reference — STOPSIGNAL — image-defined graceful-stop signal.
- Docker CLI — docker container stop — stop signal and timeout/escalation behavior.
-
Docker Docs — Restart policies
—
on-failure,always,unless-stopped, successful-start threshold, and manual-stop behavior. - Docker CLI — restart policies — backoff and retry semantics.
-
Docker Compose — startup/shutdown order
—
service_started,service_healthy, andservice_completed_successfully. - Compose services — healthcheck — Compose healthcheck fields and override behavior.
- Compose services — depends_on — dependency conditions and explicit Compose restart propagation.
- Docker Engine 29 release notes — current Engine baseline and healthcheck fixes.
Docker
Engine 29.8.1 is the current Engine baseline used for compatibility
notes. Current Dockerfile HEALTHCHECK supports
interval, timeout,
start_period, start_interval (Engine 25+),
and retries. Engine 29 includes a fix for health checks
delayed too long when start_interval exceeds
start_period. The labs record the actual
Engine/CLI/Compose versions before relying on these semantics.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.