Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Configuration, Design Choices, and Tradeoffs
Choose Docker healthcheck depth/timing, restart policy, application retries, dependency gating, and external checks using explicit tradeoffs.
Learning objectives
- Choose exec versus shell probe form and an appropriate probe depth.
- Tune interval, timeout, retries, start period, and start interval from expected startup/failure characteristics.
- Select a restart policy separately from health semantics and define when application retries remain necessary.
- Distinguish Docker health, Compose startup gating, and external load-balancer/readiness checks.
1. Exec versus shell healthcheck form
Exec form runs the probe directly and avoids shell parsing. Shell
form is appropriate when you deliberately need operators such as
&&, variable expansion, or multiple checks. The
shell form also introduces the shell itself as a dependency and
changes signal/quoting behavior.
| Choice | Prefer when | Tradeoff |
|---|---|---|
| Exec form | One executable can perform the whole probe. | Clear argument boundaries; no shell operators or implicit variable expansion. |
| Shell form / CMD-SHELL | Probe intentionally composes shell logic. | More expressive, but quoting and shell availability become part of the health contract. |
2. Choose probe depth to match the failure you need to detect
PID existence is extremely cheap but often too shallow: Docker already knows whether PID 1 is alive. A local protocol check proves more, such as “the HTTP handler returns a valid response.” A deep dependency check can prove end-to-end behavior but may mark many containers unhealthy because one shared dependency is down, causing synchronized recovery storms.
| Probe depth | Detects | Main risk |
|---|---|---|
| Process/PID-like | Basic process existence. | Duplicates runtime state and misses deadlocks/unusable service. |
| Local application endpoint | Listener + local handler path. | Requires tool/protocol support in image. |
| Critical internal dependency | Application + selected dependency path. | Can couple health to shared failure and amplify restarts. |
| External user path | Real availability through network/TLS/proxy. | Belongs in external monitoring/load balancing rather than only container health. |
3. Tune timing from measured startup and failure budgets
Probe timing should be derived from expected startup duration, request latency, transient-error tolerance, and acceptable failure-detection time. An aggressive one-second interval across hundreds of containers can create real load; a five-minute interval can make failure detection useless.
| Option | Question to ask |
|---|---|
start_period |
How long can normal initialization take before failures should count? |
start_interval |
How often should we look for early readiness during startup grace? |
interval |
How quickly must a steady-state failure be detected without creating excess load? |
timeout |
How long can a healthy probe legitimately take? |
retries |
How much transient failure should be tolerated before declaring unhealthy? |
4. Restart policy is a process-lifecycle choice
Use on-failure when non-zero exits represent
recoverable failures and a bounded retry count can prevent endless
loops. always can be appropriate for a service expected
to run whenever Docker runs. unless-stopped preserves
an operator’s intentional stop across daemon restarts. None of these
policies fixes a process that is alive but unhealthy.
5. Application retries versus Compose dependency gating
Compose service_healthy is useful for initial startup
ordering. Applications still need bounded retries/backoff,
connection re-establishment, and sensible timeouts because
dependencies can restart later. The two mechanisms are
complementary: gating prevents an avoidable cold-start race;
application resilience handles the long-running system.
6. Docker health versus external load-balancer health
A local probe should usually be fast, local, and cheap. An external health check can validate the route users actually take, including DNS, host networking, reverse proxy, TLS, and possibly a broader dependency chain. Use both only when each answers a different operational question.
7. Decision table
| Scenario | Recommended mechanism | Evidence before decision |
|---|---|---|
| App takes 20–30 seconds to initialize |
Healthcheck with measured
start_period/start_interval.
|
Startup traces + first-success health timestamp. |
| Process occasionally exits 42 due to recoverable transient failure |
Bounded on-failure:N plus root-cause logging.
|
Exit codes, restart count/backoff, incident rate. |
| Database may restart after app is already running | Client retry/reconnect logic; Compose health gating only for startup. | Connection failure timeline and dependency events. |
| Local HTTP handler works but users see 503 | External path check and proxy/network investigation. | Container health + proxy/TLS/DNS/firewall evidence. |
| Service must remain intentionally stopped after operator action |
unless-stopped often fits better than
always.
|
Operational runbook and daemon-restart expectation. |
8. Keep adjacent state domains distinct
- Host/client/context: which Docker daemon receives the command.
-
Image: inherited or declared healthcheck,
STOPSIGNAL, probe tooling. - Container runtime: current PID, exit code, restart count, stop timeout.
- Health: probe command, timing, status, failing streak, probe output.
- Dependency: Compose startup conditions and application retries.
- External availability: published path, DNS, firewall, proxy, TLS, upstreams.
Knowledge check
When is exec-form HEALTHCHECK preferable?
When a single executable can perform the probe and you want explicit arguments without shell parsing.
Why can a deep dependency healthcheck be dangerous?
A shared dependency failure can make many containers unhealthy at once and trigger correlated remediation or alert storms.
Which policy best preserves an intentional manual stop across daemon restart?
unless-stopped.
What evidence should determine start_period?
Measured normal startup duration and the time to first meaningful readiness, not a guessed round number.
Why can an external load-balancer check be valuable even when Docker health is present?
It verifies the path users actually traverse, including components outside the container namespace.
Official references and version notes
- Dockerfile reference — HEALTHCHECK — health states, timing, exit codes, output retention, and health-status events.
- Dockerfile reference — STOPSIGNAL — image-defined graceful-stop signal.
- Docker CLI — docker container stop — stop signal and timeout/escalation behavior.
-
Docker Docs — Restart policies
—
on-failure,always,unless-stopped, successful-start threshold, and manual-stop behavior. - Docker CLI — restart policies — backoff and retry semantics.
-
Docker Compose — startup/shutdown order
—
service_started,service_healthy, andservice_completed_successfully. - Compose services — healthcheck — Compose healthcheck fields and override behavior.
- Compose services — depends_on — dependency conditions and explicit Compose restart propagation.
- Docker Engine 29 release notes — current Engine baseline and healthcheck fixes.
Docker
Engine 29.8.1 is the current Engine baseline used for compatibility
notes. Current Dockerfile HEALTHCHECK supports
interval, timeout,
start_period, start_interval (Engine 25+),
and retries. Engine 29 includes a fix for health checks
delayed too long when start_interval exceeds
start_period. The labs record the actual
Engine/CLI/Compose versions before relying on these semantics.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.