Chapter 23Lesson 05~120 minutes

Checkpoint Lab — Logs, Logging Drivers, Rotation, stdout/stderr Contracts, Events, stats, and Container Observability

Generate a controlled container failure and assemble a correlated evidence packet from application logs, Engine events, stats, inspect state, and daemon evidence.

Container observabilityLogging driversEvents & statsRotation & retentionEvidence-first

Learning objectives

  • Generate a controlled failure and assemble a time-correlated incident packet from logs, events, stats, inspect state, and daemon evidence.
  • Define a bounded rotation/retention policy and state where data can be lost.
  • Verify image/container identity and logging configuration before interpreting telemetry.
  • Document sensitivity, sharing, and platform assumptions so evidence is useful without becoming a credential leak.
Checkpoint outcome. Produce a compact incident dossier, not a pile of commands. The dossier must answer: what artifact ran, what container/process state changed, what the application emitted, what Engine events occurred, what resources looked like, what logging policy applied, and which evidence is safe to share.

1. Scenario: controlled failure and correlated timeline

A synthetic service emits structured heartbeats, briefly creates CPU load, then exits with code 42 after writing a final stderr record. You will capture a stats sample while it runs, query Engine events after it exits, preserve logs and inspect state, and define a bounded retention policy. This is a local simulation; no cloud logging or admin-only service is required.

2. Preflight and assumptions packet

Record Engine/CLI/Compose/Buildx versions even though this lab does not build an image. Record the image repository digest, daemon logging default, and context. On Docker Desktop, note that daemon logs belong to the Desktop VM.

mkdir -p dkr23-checkpoint

docker context show | tee dkr23-checkpoint/context.txt
docker version | tee dkr23-checkpoint/docker-version.txt
docker info | tee dkr23-checkpoint/docker-info.txt
docker compose version | tee dkr23-checkpoint/compose-version.txt 2>&1 || true
docker buildx version | tee dkr23-checkpoint/buildx-version.txt 2>&1 || true
docker info --format 'default_logging_driver={{.LoggingDriver}}' | tee dkr23-checkpoint/default-driver.txt

docker pull alpine:3.22
docker image inspect alpine:3.22 --format '{{json .RepoDigests}}' | tee dkr23-checkpoint/image-digests.txt

3. Write predictions before execution

Prediction Independent verification
The app will run with the per-container local driver regardless of daemon default. docker inspect ... HostConfig.LogConfig.
The stats sample will exist only while the container is running. Capture docker stats --no-stream before failure; stopped containers do not return live stats data.
The final process exit will be code 42, not an OOM kill. .State.ExitCode, .State.OOMKilled, final stderr record.
Engine events will contain lifecycle evidence for this labeled container. Bounded docker events query by label/time after the run.
Rotation settings cap local Docker retention but do not prove remote retention. Inspect per-container options and state this limitation explicitly.

4. Run the observed workload

The workload is bounded to roughly 10 seconds. It writes normal heartbeats to stdout, one warning to stderr, performs a short CPU burst, and exits deliberately. The failure is application-controlled, not a daemon crash or resource attack.

docker rm -f dkr23-check 2>/dev/null || true
START_UTC=$(date -u +%Y-%m-%dT%H:%M:%SZ)

docker run -d --name dkr23-check   --label academy=docker-ch23   --label scenario=checkpoint   --log-driver local   --log-opt max-size=1m   --log-opt max-file=3   alpine:3.22 sh -c '
    i=1
    while [ "$i" -le 6 ]; do
      echo "{\"level\":\"info\",\"seq\":$i,\"msg\":\"heartbeat\"}"
      if [ "$i" -eq 4 ]; then
        echo "{\"level\":\"warn\",\"seq\":4,\"msg\":\"synthetic degradation\"}" >&2
        end=$(( $(date +%s) + 2 )); while [ $(date +%s) -lt $end ]; do :; done
      fi
      sleep 1
      i=$((i+1))
    done
    echo "{\"level\":\"error\",\"msg\":\"controlled exit\",\"exit\":42}" >&2
    exit 42
  '

sleep 3
docker stats --no-stream --format '{{json .}}' dkr23-check   | tee dkr23-checkpoint/stats-during-run.json || true

docker wait dkr23-check | tee dkr23-checkpoint/wait-exit.txt
END_UTC=$(date -u +%Y-%m-%dT%H:%M:%SZ)

5. Capture the evidence packet after failure

Do not remove the failed container yet. Preserve state, logs, and events first. The failed object itself is evidence.

docker inspect dkr23-check > dkr23-checkpoint/container-inspect.json
docker inspect dkr23-check \
  --format 'id={{.Id}} image={{.Image}} status={{.State.Status}} exit={{.State.ExitCode}} oom={{.State.OOMKilled}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} log={{json .HostConfig.LogConfig}}'   | tee dkr23-checkpoint/container-summary.txt

docker logs --timestamps --since "$START_UTC" dkr23-check 2>&1   | tee dkr23-checkpoint/application-logs.txt

docker events \
  --since "$START_UTC" \
  --until "$END_UTC" \
  --filter type=container \
  --filter label=academy=docker-ch23 \
  --filter label=scenario=checkpoint \
  --format '{{json .}}'   | tee dkr23-checkpoint/engine-events.jsonl

6. Add daemon evidence without broad privilege changes

Capture a bounded daemon-log window only if your platform and permissions already allow it. Do not enable debug mode just to satisfy the lab. If unavailable, record “not accessible in this environment” as a limitation.

# Optional on Linux/systemd.
journalctl -u docker.service \
  --since "$START_UTC" \
  --until "$END_UTC" \
  --no-pager   > dkr23-checkpoint/daemon-window.txt 2>&1 ||   printf 'daemon log window not accessible through journalctl in this environment
'   > dkr23-checkpoint/daemon-window.txt

7. Build the incident timeline

Read the packet in time order. A strong timeline connects the application’s final stderr record, the Engine’s die event, and inspect’s exit code 42. The stats sample shows runtime behavior before failure but is not a post-mortem metric store.

Evidence Question answered
image-digests.txt + inspect image ID What artifact/runtime identity did this container use?
application-logs.txt What did the application emit, and when?
engine-events.jsonl What lifecycle operations/events did the Engine observe?
stats-during-run.json What resource snapshot existed before the controlled failure?
container-summary.txt What final exit/OOM/logging configuration state was preserved?
daemon-window.txt Did the Engine/driver report platform-side errors during the same window?

8. Define a bounded retention policy

For this lab, local with three files of at most 1 MB each demonstrates a finite per-container policy. Production numbers must be chosen from measured output rate, incident-detection latency, host disk budget, and external retention requirements.

A sample policy statement: “Retain enough local Docker logs for first-response troubleshooting, capped by driver rotation; stream selected production logs to a protected external system for longer retention; alert on external delivery failure; never log credentials; treat non-blocking mode as lossy and use it only when application availability is more important than complete log capture.”

9. Sensitivity and sharing notes

  • Application logs can contain customer data, tokens, URLs, query strings, stack traces, or internal topology.
  • docker inspect can include environment variables, labels, mount paths, network addresses, and command arguments.
  • docker info and daemon logs can reveal host/runtime metadata.
  • Review and redact evidence according to policy before sharing. Redaction must preserve timestamps, IDs/digests, and causal fields needed for diagnosis.

10. Verification checklist

  • The exact Docker context, Engine/CLI versions, daemon default driver, image digest, container ID, and per-container driver/options are recorded.
  • The process exit is shown as code 42 with OOMKilled=false unless the environment produced an unexpected result that is documented.
  • Application logs, Engine events, stats, and inspect state are kept as separate evidence sources.
  • No real secret, production endpoint, daemon mutation, Docker socket mount, privileged mode, broad prune, or direct manipulation of Docker-managed log files is used.
  • The retention statement includes disk budget, expected log rate, external retention, delivery/backpressure behavior, and sensitivity.

11. Cleanup and archive

Remove only the exact lab container after the packet is complete. Keep or archive the evidence directory only after reviewing it for sensitive metadata.

docker rm dkr23-check 2>/dev/null || true
docker ps -a --filter label=academy=docker-ch23

tar -czf dkr23-checkpoint.tgz dkr23-checkpoint/

12. What Chapter 23 adds to the production model

You now have an observability contract that separates producer output, driver behavior, Engine lifecycle, runtime counters, daemon state, and external delivery. That makes failures diagnosable without conflating “no log line” with “no event,” “no event” with “no metric,” or “local cache” with “remote delivery.” Chapter 24 builds on this evidence model with health checks, restart policies, dependency health, graceful shutdown, signals, and self-healing patterns.

Knowledge check

The final log says controlled exit 42 and inspect reports exit 42, OOMKilled=false. What is the primary cause?

Why capture stats before docker wait completes?

Why preserve Engine events if application logs already show the failure?

Does a 3 MB local retention cap mean the external system retains 3 MB?

What must be reviewed before sharing the evidence archive?

Next lesson

Next: Health Checks, Restart Policies, Dependency Health, Graceful Shutdown, Signals, and Self-Healing Patterns: Concepts, Architecture, and Mental Model

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline used for compatibility notes. The daemon default logging driver remains json-file; Docker recommends local for general use because it rotates by default. The mandatory labs configure logging per container so they do not require editing daemon.json or restarting Docker. Always record docker version, docker info, context, the actual per-container HostConfig.LogConfig, and platform-specific daemon-log location instead of assuming defaults.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.