Chapter 24Lesson 02~120 minutes

Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Guided Hands-On Workflow and Core Operations

Run safe Docker health, restart-policy, signal, and Compose dependency-health labs while preserving inspect, logs, and restart evidence.

Health checksRestart policiesGraceful shutdownDependency healthEvidence-first

Learning objectives

  • Create a bounded healthcheck whose dependency is already present in the image.
  • Drive a container from healthy to unhealthy and prove that unhealthy state alone does not restart it.
  • Exercise a synthetic non-zero exit under on-failure and inspect restart evidence.
  • Demonstrate graceful SIGTERM handling and a Compose service_healthy dependency without production endpoints or privileged settings.
Lab safety. The workflow uses only disposable local containers, an isolated Compose project, Alpine shell built-ins, fake state files, and exact cleanup names. No daemon configuration, privileged mode, Docker socket mount, production endpoint, broad prune, or real credential is used.

1. Preflight and evidence directory

Record the execution context before changing state. Health semantics are version-sensitive enough that the actual Engine/Compose versions belong in the evidence packet.

mkdir -p dkr24-evidence

docker context show | tee dkr24-evidence/context.txt
docker version | tee dkr24-evidence/docker-version.txt
docker info | tee dkr24-evidence/docker-info.txt
docker compose version | tee dkr24-evidence/compose-version.txt 2>&1 || true
docker buildx version | tee dkr24-evidence/buildx-version.txt 2>&1 || true

docker pull alpine:3.22
docker image inspect alpine:3.22 --format '{{json .RepoDigests}}'   | tee dkr24-evidence/alpine-digests.txt

2. Create a bounded healthcheck and inspect its history

The probe uses test -f, which already exists in Alpine’s shell environment. This avoids the classic mistake of defining a healthcheck that depends on curl or another utility absent from a minimal image.

docker rm -f dkr24-health 2>/dev/null || true

docker run -d \
  --name dkr24-health \
  --label academy=docker-ch24 \
  --health-cmd='test -f /tmp/healthy' \
  --health-interval=2s \
  --health-timeout=1s \
  --health-retries=2 \
  --health-start-period=3s \
  --health-start-interval=1s   alpine:3.22 sh -c 'touch /tmp/healthy; while :; do sleep 1; done'

sleep 5
docker inspect dkr24-health \
  --format 'process={{.State.Status}} health={{.State.Health.Status}} streak={{.State.Health.FailingStreak}} restarts={{.RestartCount}}'
docker inspect dkr24-health --format '{{json .State.Health.Log}}'   | tee dkr24-evidence/health-history-initial.json

3. Make the application “unhealthy” without stopping PID 1

Remove the synthetic readiness file. The main loop keeps PID 1 alive, so the container should transition to unhealthy while Running=true. Record the restart count before and after to prove that health failure alone does not activate the Engine restart policy.

BEFORE=$(docker inspect dkr24-health --format '{{.RestartCount}}')
docker exec dkr24-health rm -f /tmp/healthy
sleep 6

docker inspect dkr24-health \
  --format 'running={{.State.Running}} status={{.State.Status}} health={{.State.Health.Status}} streak={{.State.Health.FailingStreak}} restarts={{.RestartCount}}'   | tee dkr24-evidence/unhealthy-state.txt
AFTER=$(docker inspect dkr24-health --format '{{.RestartCount}}')
printf 'restart_count_before=%s after=%s
' "$BEFORE" "$AFTER"   | tee dkr24-evidence/unhealthy-restart-proof.txt

docker inspect dkr24-health --format '{{json .State.Health.Log}}'   | tee dkr24-evidence/health-history-unhealthy.json
Expected observation. The container remains running, health becomes unhealthy, and the restart count remains unchanged. That is the central distinction of this chapter.

4. Recover health without replacing the container

Restore only the condition the probe checks. The process never exited, so the container identity and restart count should remain unchanged.

docker exec dkr24-health touch /tmp/healthy
sleep 4
docker inspect dkr24-health --format 'id={{.Id}} health={{.State.Health.Status}} restarts={{.RestartCount}}'   | tee dkr24-evidence/recovered-health.txt

5. Demonstrate process-exit recovery with on-failure

This container intentionally stays up for more than ten seconds before each non-zero exit so restart-policy monitoring is clearly active. The retry count is bounded at two. It is a process-exit test—not a healthcheck test.

docker rm -f dkr24-restart 2>/dev/null || true

docker run -d --name dkr24-restart   --label academy=docker-ch24   --restart on-failure:2   alpine:3.22 sh -c '
    echo "start $(date -u +%FT%TZ)";
    sleep 11;
    echo "synthetic crash exit=42" >&2;
    exit 42
  '

# Allow enough time for the initial run and bounded restart attempts.
sleep 36
docker inspect dkr24-restart \
  --format 'status={{.State.Status}} exit={{.State.ExitCode}} restart_count={{.RestartCount}} policy={{.HostConfig.RestartPolicy.Name}} max={{.HostConfig.RestartPolicy.MaximumRetryCount}}'   | tee dkr24-evidence/restart-policy-state.txt
docker logs --timestamps dkr24-restart 2>&1   | tee dkr24-evidence/restart-policy-logs.txt

6. Prove graceful SIGTERM handling

The shell is PID 1 and installs a TERM trap. docker stop sends the configured signal and waits up to five seconds. Successful evidence is the TERM log followed by a clean exit rather than timeout escalation.

docker rm -f dkr24-graceful 2>/dev/null || true

docker run -d --name dkr24-graceful   --label academy=docker-ch24   --stop-signal SIGTERM   --stop-timeout 5   alpine:3.22 sh -c '
    trap "echo graceful-term-received; exit 0" TERM INT;
    echo ready;
    while :; do sleep 1; done
  '

sleep 2
docker stop dkr24-graceful | tee dkr24-evidence/graceful-stop-command.txt
docker inspect dkr24-graceful \
  --format 'exit={{.State.ExitCode}} finished={{.State.FinishedAt}} stop_signal={{.Config.StopSignal}} timeout={{.HostConfig.StopTimeout}}'   | tee dkr24-evidence/graceful-stop-state.txt
docker logs --timestamps dkr24-graceful   | tee dkr24-evidence/graceful-stop-logs.txt

7. Gate a dependent service on health with Compose

The dependency becomes ready only after creating a local marker. Compose waits for service_healthy before creating the one-shot worker. This demonstrates startup gating without claiming permanent resilience.

mkdir -p dkr24-compose
cat > dkr24-compose/compose.yaml <<'YAML'
name: dkr24-health-gate
services:
  dependency:
    image: alpine:3.22
    command: ["sh", "-c", "sleep 4; touch /tmp/ready; echo dependency-ready; sleep 30"]
    healthcheck:
      test: ["CMD-SHELL", "test -f /tmp/ready"]
      interval: 1s
      timeout: 1s
      retries: 5
      start_period: 2s
      start_interval: 1s
  worker:
    image: alpine:3.22
    depends_on:
      dependency:
        condition: service_healthy
    command: ["sh", "-c", "echo worker-started-after-dependency-health"]
YAML

cd dkr24-compose
docker compose config | tee ../dkr24-evidence/compose-normalized.yaml
docker compose up --abort-on-container-exit --exit-code-from worker   | tee ../dkr24-evidence/compose-up.txt
cd ..

8. Challenge: choose the failing layer

Suppose the dependency reports healthy, the worker starts, and five minutes later the dependency process stays alive but stops accepting requests. Which layer must change?

  • If the local probe is too shallow, improve the healthcheck.
  • If the worker assumes the dependency can never fail after startup, add application-level retry/reconnect logic.
  • If users still cannot reach a healthy service, investigate external DNS/network/TLS/proxy state.
  • Do not reach first for a restart policy; it only sees process exit.

9. Exact cleanup

Capture evidence first, then remove only named lab resources.

docker rm -f dkr24-health dkr24-restart dkr24-graceful 2>/dev/null || true
(cd dkr24-compose && docker compose down --remove-orphans)
rm -rf dkr24-compose

Knowledge check

Why did dkr24-health become unhealthy without restarting?

Why does the restart-policy demo sleep more than ten seconds before exiting?

What proves graceful shutdown in the signal lab?

What does Compose service_healthy prove?

Why avoid curl in the first healthcheck lab?

Next lesson

Next: Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Configuration, Design Choices, and Tradeoffs

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline used for compatibility notes. Current Dockerfile HEALTHCHECK supports interval, timeout, start_period, start_interval (Engine 25+), and retries. Engine 29 includes a fix for health checks delayed too long when start_interval exceeds start_period. The labs record the actual Engine/CLI/Compose versions before relying on these semantics.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.