Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Guided Hands-On Workflow and Core Operations
Run safe Docker health, restart-policy, signal, and Compose dependency-health labs while preserving inspect, logs, and restart evidence.
Learning objectives
- Create a bounded healthcheck whose dependency is already present in the image.
- Drive a container from healthy to unhealthy and prove that unhealthy state alone does not restart it.
-
Exercise a synthetic non-zero exit under
on-failureand inspect restart evidence. -
Demonstrate graceful SIGTERM handling and a Compose
service_healthydependency without production endpoints or privileged settings.
1. Preflight and evidence directory
Record the execution context before changing state. Health semantics are version-sensitive enough that the actual Engine/Compose versions belong in the evidence packet.
mkdir -p dkr24-evidence
docker context show | tee dkr24-evidence/context.txt
docker version | tee dkr24-evidence/docker-version.txt
docker info | tee dkr24-evidence/docker-info.txt
docker compose version | tee dkr24-evidence/compose-version.txt 2>&1 || true
docker buildx version | tee dkr24-evidence/buildx-version.txt 2>&1 || true
docker pull alpine:3.22
docker image inspect alpine:3.22 --format '{{json .RepoDigests}}' | tee dkr24-evidence/alpine-digests.txt
2. Create a bounded healthcheck and inspect its history
The probe uses test -f, which already exists in
Alpine’s shell environment. This avoids the classic mistake of
defining a healthcheck that depends on curl or another
utility absent from a minimal image.
docker rm -f dkr24-health 2>/dev/null || true
docker run -d \
--name dkr24-health \
--label academy=docker-ch24 \
--health-cmd='test -f /tmp/healthy' \
--health-interval=2s \
--health-timeout=1s \
--health-retries=2 \
--health-start-period=3s \
--health-start-interval=1s alpine:3.22 sh -c 'touch /tmp/healthy; while :; do sleep 1; done'
sleep 5
docker inspect dkr24-health \
--format 'process={{.State.Status}} health={{.State.Health.Status}} streak={{.State.Health.FailingStreak}} restarts={{.RestartCount}}'
docker inspect dkr24-health --format '{{json .State.Health.Log}}' | tee dkr24-evidence/health-history-initial.json
3. Make the application “unhealthy” without stopping PID 1
Remove the synthetic readiness file. The main loop keeps PID 1
alive, so the container should transition to
unhealthy while Running=true. Record the
restart count before and after to prove that health failure alone
does not activate the Engine restart policy.
BEFORE=$(docker inspect dkr24-health --format '{{.RestartCount}}')
docker exec dkr24-health rm -f /tmp/healthy
sleep 6
docker inspect dkr24-health \
--format 'running={{.State.Running}} status={{.State.Status}} health={{.State.Health.Status}} streak={{.State.Health.FailingStreak}} restarts={{.RestartCount}}' | tee dkr24-evidence/unhealthy-state.txt
AFTER=$(docker inspect dkr24-health --format '{{.RestartCount}}')
printf 'restart_count_before=%s after=%s
' "$BEFORE" "$AFTER" | tee dkr24-evidence/unhealthy-restart-proof.txt
docker inspect dkr24-health --format '{{json .State.Health.Log}}' | tee dkr24-evidence/health-history-unhealthy.json
4. Recover health without replacing the container
Restore only the condition the probe checks. The process never exited, so the container identity and restart count should remain unchanged.
docker exec dkr24-health touch /tmp/healthy
sleep 4
docker inspect dkr24-health --format 'id={{.Id}} health={{.State.Health.Status}} restarts={{.RestartCount}}' | tee dkr24-evidence/recovered-health.txt
5. Demonstrate process-exit recovery with on-failure
This container intentionally stays up for more than ten seconds before each non-zero exit so restart-policy monitoring is clearly active. The retry count is bounded at two. It is a process-exit test—not a healthcheck test.
docker rm -f dkr24-restart 2>/dev/null || true
docker run -d --name dkr24-restart --label academy=docker-ch24 --restart on-failure:2 alpine:3.22 sh -c '
echo "start $(date -u +%FT%TZ)";
sleep 11;
echo "synthetic crash exit=42" >&2;
exit 42
'
# Allow enough time for the initial run and bounded restart attempts.
sleep 36
docker inspect dkr24-restart \
--format 'status={{.State.Status}} exit={{.State.ExitCode}} restart_count={{.RestartCount}} policy={{.HostConfig.RestartPolicy.Name}} max={{.HostConfig.RestartPolicy.MaximumRetryCount}}' | tee dkr24-evidence/restart-policy-state.txt
docker logs --timestamps dkr24-restart 2>&1 | tee dkr24-evidence/restart-policy-logs.txt
6. Prove graceful SIGTERM handling
The shell is PID 1 and installs a TERM trap.
docker stop sends the configured signal and waits up to
five seconds. Successful evidence is the TERM log followed by a
clean exit rather than timeout escalation.
docker rm -f dkr24-graceful 2>/dev/null || true
docker run -d --name dkr24-graceful --label academy=docker-ch24 --stop-signal SIGTERM --stop-timeout 5 alpine:3.22 sh -c '
trap "echo graceful-term-received; exit 0" TERM INT;
echo ready;
while :; do sleep 1; done
'
sleep 2
docker stop dkr24-graceful | tee dkr24-evidence/graceful-stop-command.txt
docker inspect dkr24-graceful \
--format 'exit={{.State.ExitCode}} finished={{.State.FinishedAt}} stop_signal={{.Config.StopSignal}} timeout={{.HostConfig.StopTimeout}}' | tee dkr24-evidence/graceful-stop-state.txt
docker logs --timestamps dkr24-graceful | tee dkr24-evidence/graceful-stop-logs.txt
7. Gate a dependent service on health with Compose
The dependency becomes ready only after creating a local marker.
Compose waits for service_healthy before creating the
one-shot worker. This demonstrates startup gating without claiming
permanent resilience.
mkdir -p dkr24-compose
cat > dkr24-compose/compose.yaml <<'YAML'
name: dkr24-health-gate
services:
dependency:
image: alpine:3.22
command: ["sh", "-c", "sleep 4; touch /tmp/ready; echo dependency-ready; sleep 30"]
healthcheck:
test: ["CMD-SHELL", "test -f /tmp/ready"]
interval: 1s
timeout: 1s
retries: 5
start_period: 2s
start_interval: 1s
worker:
image: alpine:3.22
depends_on:
dependency:
condition: service_healthy
command: ["sh", "-c", "echo worker-started-after-dependency-health"]
YAML
cd dkr24-compose
docker compose config | tee ../dkr24-evidence/compose-normalized.yaml
docker compose up --abort-on-container-exit --exit-code-from worker | tee ../dkr24-evidence/compose-up.txt
cd ..
8. Challenge: choose the failing layer
Suppose the dependency reports healthy, the worker starts, and five minutes later the dependency process stays alive but stops accepting requests. Which layer must change?
- If the local probe is too shallow, improve the healthcheck.
- If the worker assumes the dependency can never fail after startup, add application-level retry/reconnect logic.
- If users still cannot reach a healthy service, investigate external DNS/network/TLS/proxy state.
- Do not reach first for a restart policy; it only sees process exit.
9. Exact cleanup
Capture evidence first, then remove only named lab resources.
docker rm -f dkr24-health dkr24-restart dkr24-graceful 2>/dev/null || true
(cd dkr24-compose && docker compose down --remove-orphans)
rm -rf dkr24-compose
Knowledge check
Why did dkr24-health become unhealthy without restarting?
Its main process stayed alive. Engine restart policies react to process exit, not the health state transition.
Why does the restart-policy demo sleep more than ten seconds before exiting?
Docker restart-policy monitoring considers a container successfully started after it has remained up long enough; the lab makes that prerequisite explicit before the synthetic crash.
What proves graceful shutdown in the signal lab?
The application log records TERM receipt and inspect records a clean exit before the configured timeout, rather than a forced kill.
What does Compose service_healthy prove?
It proves the dependency satisfied its configured healthcheck before the dependent was created. It does not prove future availability.
Why avoid curl in the first healthcheck lab?
A probe must rely on tooling actually present in the image. Using a missing utility creates a probe failure rather than an application-health signal.
Official references and version notes
- Dockerfile reference — HEALTHCHECK — health states, timing, exit codes, output retention, and health-status events.
- Dockerfile reference — STOPSIGNAL — image-defined graceful-stop signal.
- Docker CLI — docker container stop — stop signal and timeout/escalation behavior.
-
Docker Docs — Restart policies
—
on-failure,always,unless-stopped, successful-start threshold, and manual-stop behavior. - Docker CLI — restart policies — backoff and retry semantics.
-
Docker Compose — startup/shutdown order
—
service_started,service_healthy, andservice_completed_successfully. - Compose services — healthcheck — Compose healthcheck fields and override behavior.
- Compose services — depends_on — dependency conditions and explicit Compose restart propagation.
- Docker Engine 29 release notes — current Engine baseline and healthcheck fixes.
Docker
Engine 29.8.1 is the current Engine baseline used for compatibility
notes. Current Dockerfile HEALTHCHECK supports
interval, timeout,
start_period, start_interval (Engine 25+),
and retries. Engine 29 includes a fix for health checks
delayed too long when start_interval exceeds
start_period. The labs record the actual
Engine/CLI/Compose versions before relying on these semantics.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.