Chapter 24Lesson 03~120 minutes

Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Configuration, Design Choices, and Tradeoffs

Choose Docker healthcheck depth/timing, restart policy, application retries, dependency gating, and external checks using explicit tradeoffs.

Health checksRestart policiesGraceful shutdownDependency healthEvidence-first

Learning objectives

  • Choose exec versus shell probe form and an appropriate probe depth.
  • Tune interval, timeout, retries, start period, and start interval from expected startup/failure characteristics.
  • Select a restart policy separately from health semantics and define when application retries remain necessary.
  • Distinguish Docker health, Compose startup gating, and external load-balancer/readiness checks.
Design rule. Recovery mechanisms should be intentionally non-overlapping. A healthcheck says “this container’s probe result”; a restart policy says “what to do after process exit”; application retries say “what to do when a dependency fails”; an external load balancer says “whether to route traffic.” Do not make one mechanism impersonate all four.

1. Exec versus shell healthcheck form

Exec form runs the probe directly and avoids shell parsing. Shell form is appropriate when you deliberately need operators such as &&, variable expansion, or multiple checks. The shell form also introduces the shell itself as a dependency and changes signal/quoting behavior.

Choice Prefer when Tradeoff
Exec form One executable can perform the whole probe. Clear argument boundaries; no shell operators or implicit variable expansion.
Shell form / CMD-SHELL Probe intentionally composes shell logic. More expressive, but quoting and shell availability become part of the health contract.

2. Choose probe depth to match the failure you need to detect

PID existence is extremely cheap but often too shallow: Docker already knows whether PID 1 is alive. A local protocol check proves more, such as “the HTTP handler returns a valid response.” A deep dependency check can prove end-to-end behavior but may mark many containers unhealthy because one shared dependency is down, causing synchronized recovery storms.

Probe depth Detects Main risk
Process/PID-like Basic process existence. Duplicates runtime state and misses deadlocks/unusable service.
Local application endpoint Listener + local handler path. Requires tool/protocol support in image.
Critical internal dependency Application + selected dependency path. Can couple health to shared failure and amplify restarts.
External user path Real availability through network/TLS/proxy. Belongs in external monitoring/load balancing rather than only container health.

3. Tune timing from measured startup and failure budgets

Probe timing should be derived from expected startup duration, request latency, transient-error tolerance, and acceptable failure-detection time. An aggressive one-second interval across hundreds of containers can create real load; a five-minute interval can make failure detection useless.

Option Question to ask
start_period How long can normal initialization take before failures should count?
start_interval How often should we look for early readiness during startup grace?
interval How quickly must a steady-state failure be detected without creating excess load?
timeout How long can a healthy probe legitimately take?
retries How much transient failure should be tolerated before declaring unhealthy?

4. Restart policy is a process-lifecycle choice

Use on-failure when non-zero exits represent recoverable failures and a bounded retry count can prevent endless loops. always can be appropriate for a service expected to run whenever Docker runs. unless-stopped preserves an operator’s intentional stop across daemon restarts. None of these policies fixes a process that is alive but unhealthy.

5. Application retries versus Compose dependency gating

Compose service_healthy is useful for initial startup ordering. Applications still need bounded retries/backoff, connection re-establishment, and sensible timeouts because dependencies can restart later. The two mechanisms are complementary: gating prevents an avoidable cold-start race; application resilience handles the long-running system.

6. Docker health versus external load-balancer health

A local probe should usually be fast, local, and cheap. An external health check can validate the route users actually take, including DNS, host networking, reverse proxy, TLS, and possibly a broader dependency chain. Use both only when each answers a different operational question.

7. Decision table

Scenario Recommended mechanism Evidence before decision
App takes 20–30 seconds to initialize Healthcheck with measured start_period/start_interval. Startup traces + first-success health timestamp.
Process occasionally exits 42 due to recoverable transient failure Bounded on-failure:N plus root-cause logging. Exit codes, restart count/backoff, incident rate.
Database may restart after app is already running Client retry/reconnect logic; Compose health gating only for startup. Connection failure timeline and dependency events.
Local HTTP handler works but users see 503 External path check and proxy/network investigation. Container health + proxy/TLS/DNS/firewall evidence.
Service must remain intentionally stopped after operator action unless-stopped often fits better than always. Operational runbook and daemon-restart expectation.

8. Keep adjacent state domains distinct

  • Host/client/context: which Docker daemon receives the command.
  • Image: inherited or declared healthcheck, STOPSIGNAL, probe tooling.
  • Container runtime: current PID, exit code, restart count, stop timeout.
  • Health: probe command, timing, status, failing streak, probe output.
  • Dependency: Compose startup conditions and application retries.
  • External availability: published path, DNS, firewall, proxy, TLS, upstreams.

Knowledge check

When is exec-form HEALTHCHECK preferable?

Why can a deep dependency healthcheck be dangerous?

Which policy best preserves an intentional manual stop across daemon restart?

What evidence should determine start_period?

Why can an external load-balancer check be valuable even when Docker health is present?

Next lesson

Next: Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Diagnostics, Failure Modes, Security, and Performance

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline used for compatibility notes. Current Dockerfile HEALTHCHECK supports interval, timeout, start_period, start_interval (Engine 25+), and retries. Engine 29 includes a fix for health checks delayed too long when start_interval exceeds start_period. The labs record the actual Engine/CLI/Compose versions before relying on these semantics.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.