Checkpoint Lab — Logs, Logging Drivers, Rotation, stdout/stderr Contracts, Events, stats, and Container Observability
Generate a controlled container failure and assemble a correlated evidence packet from application logs, Engine events, stats, inspect state, and daemon evidence.
Learning objectives
- Generate a controlled failure and assemble a time-correlated incident packet from logs, events, stats, inspect state, and daemon evidence.
- Define a bounded rotation/retention policy and state where data can be lost.
- Verify image/container identity and logging configuration before interpreting telemetry.
- Document sensitivity, sharing, and platform assumptions so evidence is useful without becoming a credential leak.
1. Scenario: controlled failure and correlated timeline
A synthetic service emits structured heartbeats, briefly creates CPU load, then exits with code 42 after writing a final stderr record. You will capture a stats sample while it runs, query Engine events after it exits, preserve logs and inspect state, and define a bounded retention policy. This is a local simulation; no cloud logging or admin-only service is required.
2. Preflight and assumptions packet
Record Engine/CLI/Compose/Buildx versions even though this lab does not build an image. Record the image repository digest, daemon logging default, and context. On Docker Desktop, note that daemon logs belong to the Desktop VM.
mkdir -p dkr23-checkpoint
docker context show | tee dkr23-checkpoint/context.txt
docker version | tee dkr23-checkpoint/docker-version.txt
docker info | tee dkr23-checkpoint/docker-info.txt
docker compose version | tee dkr23-checkpoint/compose-version.txt 2>&1 || true
docker buildx version | tee dkr23-checkpoint/buildx-version.txt 2>&1 || true
docker info --format 'default_logging_driver={{.LoggingDriver}}' | tee dkr23-checkpoint/default-driver.txt
docker pull alpine:3.22
docker image inspect alpine:3.22 --format '{{json .RepoDigests}}' | tee dkr23-checkpoint/image-digests.txt
3. Write predictions before execution
| Prediction | Independent verification |
|---|---|
The app will run with the per-container
local driver regardless of daemon default.
|
docker inspect ... HostConfig.LogConfig. |
| The stats sample will exist only while the container is running. |
Capture docker stats --no-stream before
failure; stopped containers do not return live stats data.
|
| The final process exit will be code 42, not an OOM kill. |
.State.ExitCode, .State.OOMKilled,
final stderr record.
|
| Engine events will contain lifecycle evidence for this labeled container. |
Bounded docker events query by label/time after
the run.
|
| Rotation settings cap local Docker retention but do not prove remote retention. | Inspect per-container options and state this limitation explicitly. |
4. Run the observed workload
The workload is bounded to roughly 10 seconds. It writes normal heartbeats to stdout, one warning to stderr, performs a short CPU burst, and exits deliberately. The failure is application-controlled, not a daemon crash or resource attack.
docker rm -f dkr23-check 2>/dev/null || true
START_UTC=$(date -u +%Y-%m-%dT%H:%M:%SZ)
docker run -d --name dkr23-check --label academy=docker-ch23 --label scenario=checkpoint --log-driver local --log-opt max-size=1m --log-opt max-file=3 alpine:3.22 sh -c '
i=1
while [ "$i" -le 6 ]; do
echo "{\"level\":\"info\",\"seq\":$i,\"msg\":\"heartbeat\"}"
if [ "$i" -eq 4 ]; then
echo "{\"level\":\"warn\",\"seq\":4,\"msg\":\"synthetic degradation\"}" >&2
end=$(( $(date +%s) + 2 )); while [ $(date +%s) -lt $end ]; do :; done
fi
sleep 1
i=$((i+1))
done
echo "{\"level\":\"error\",\"msg\":\"controlled exit\",\"exit\":42}" >&2
exit 42
'
sleep 3
docker stats --no-stream --format '{{json .}}' dkr23-check | tee dkr23-checkpoint/stats-during-run.json || true
docker wait dkr23-check | tee dkr23-checkpoint/wait-exit.txt
END_UTC=$(date -u +%Y-%m-%dT%H:%M:%SZ)
5. Capture the evidence packet after failure
Do not remove the failed container yet. Preserve state, logs, and events first. The failed object itself is evidence.
docker inspect dkr23-check > dkr23-checkpoint/container-inspect.json
docker inspect dkr23-check \
--format 'id={{.Id}} image={{.Image}} status={{.State.Status}} exit={{.State.ExitCode}} oom={{.State.OOMKilled}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} log={{json .HostConfig.LogConfig}}' | tee dkr23-checkpoint/container-summary.txt
docker logs --timestamps --since "$START_UTC" dkr23-check 2>&1 | tee dkr23-checkpoint/application-logs.txt
docker events \
--since "$START_UTC" \
--until "$END_UTC" \
--filter type=container \
--filter label=academy=docker-ch23 \
--filter label=scenario=checkpoint \
--format '{{json .}}' | tee dkr23-checkpoint/engine-events.jsonl
6. Add daemon evidence without broad privilege changes
Capture a bounded daemon-log window only if your platform and permissions already allow it. Do not enable debug mode just to satisfy the lab. If unavailable, record “not accessible in this environment” as a limitation.
# Optional on Linux/systemd.
journalctl -u docker.service \
--since "$START_UTC" \
--until "$END_UTC" \
--no-pager > dkr23-checkpoint/daemon-window.txt 2>&1 || printf 'daemon log window not accessible through journalctl in this environment
' > dkr23-checkpoint/daemon-window.txt
7. Build the incident timeline
Read the packet in time order. A strong timeline connects the
application’s final stderr record, the Engine’s
die event, and inspect’s exit code 42. The stats sample
shows runtime behavior before failure but is not a post-mortem
metric store.
| Evidence | Question answered |
|---|---|
image-digests.txt + inspect image ID |
What artifact/runtime identity did this container use? |
application-logs.txt |
What did the application emit, and when? |
engine-events.jsonl |
What lifecycle operations/events did the Engine observe? |
stats-during-run.json |
What resource snapshot existed before the controlled failure? |
container-summary.txt |
What final exit/OOM/logging configuration state was preserved? |
daemon-window.txt |
Did the Engine/driver report platform-side errors during the same window? |
8. Define a bounded retention policy
For this lab, local with three files of at most 1 MB
each demonstrates a finite per-container policy. Production numbers
must be chosen from measured output rate, incident-detection
latency, host disk budget, and external retention requirements.
A sample policy statement: “Retain enough local Docker logs for first-response troubleshooting, capped by driver rotation; stream selected production logs to a protected external system for longer retention; alert on external delivery failure; never log credentials; treat non-blocking mode as lossy and use it only when application availability is more important than complete log capture.”
9. Sensitivity and sharing notes
- Application logs can contain customer data, tokens, URLs, query strings, stack traces, or internal topology.
-
docker inspectcan include environment variables, labels, mount paths, network addresses, and command arguments. -
docker infoand daemon logs can reveal host/runtime metadata. - Review and redact evidence according to policy before sharing. Redaction must preserve timestamps, IDs/digests, and causal fields needed for diagnosis.
10. Verification checklist
- The exact Docker context, Engine/CLI versions, daemon default driver, image digest, container ID, and per-container driver/options are recorded.
-
The process exit is shown as code 42 with
OOMKilled=falseunless the environment produced an unexpected result that is documented. - Application logs, Engine events, stats, and inspect state are kept as separate evidence sources.
- No real secret, production endpoint, daemon mutation, Docker socket mount, privileged mode, broad prune, or direct manipulation of Docker-managed log files is used.
- The retention statement includes disk budget, expected log rate, external retention, delivery/backpressure behavior, and sensitivity.
11. Cleanup and archive
Remove only the exact lab container after the packet is complete. Keep or archive the evidence directory only after reviewing it for sensitive metadata.
docker rm dkr23-check 2>/dev/null || true
docker ps -a --filter label=academy=docker-ch23
tar -czf dkr23-checkpoint.tgz dkr23-checkpoint/
12. What Chapter 23 adds to the production model
You now have an observability contract that separates producer output, driver behavior, Engine lifecycle, runtime counters, daemon state, and external delivery. That makes failures diagnosable without conflating “no log line” with “no event,” “no event” with “no metric,” or “local cache” with “remote delivery.” Chapter 24 builds on this evidence model with health checks, restart policies, dependency health, graceful shutdown, signals, and self-healing patterns.
Knowledge check
The final log says controlled exit 42 and inspect reports exit 42, OOMKilled=false. What is the primary cause?
The application deliberately exited with code 42. The stats sample can describe pre-exit resource use but does not override the preserved exit evidence.
Why capture stats before docker wait completes?
docker stats provides live data for running containers; once the process is stopped, a stopped container does not provide a live usage sample.
Why preserve Engine events if application logs already show the failure?
Events independently establish Engine-observed lifecycle actions and can expose restart, OOM, health, kill, or operator actions not represented in application output.
Does a 3 MB local retention cap mean the external system retains 3 MB?
No. Local driver retention and external sink retention are separate policies and must be verified independently.
What must be reviewed before sharing the evidence archive?
Logs, inspect output, docker info, daemon logs, paths, addresses, labels/environment metadata, and any customer/security-sensitive data; preserve causal IDs/timestamps while redacting according to policy.
Official references and version notes
- Docker Docs — Configure logging drivers — daemon/container driver selection, delivery modes, labels/tags, and current default-driver behavior.
- Docker Docs — Local file logging driver — default rotation/compression behavior and supported options.
- Docker Docs — JSON file logging driver — JSON framing and explicit rotation options.
-
Docker Docs — Dual logging
— how
docker logscan remain available with remote drivers and when it does not. -
Docker CLI — docker container logs
— timestamps,
--since,--until, follow, and tail behavior. - Docker CLI — docker system events — event scope, filters, JSON Lines formatting, and bounded event history.
- Docker CLI — docker container stats — CPU, memory, network, block I/O, PIDs, and Linux cache-reporting notes.
- Docker Docs — Read daemon logs — platform-specific locations and systemd/desktop guidance.
- Docker Engine 29 release notes — current Engine baseline and logging-related fixes/features.
Docker
Engine 29.8.1 is the current Engine baseline used for compatibility
notes. The daemon default logging driver remains
json-file; Docker recommends local for
general use because it rotates by default. The mandatory labs
configure logging per container so they do not require editing
daemon.json or restarting Docker. Always record
docker version, docker info, context, the
actual per-container HostConfig.LogConfig, and
platform-specific daemon-log location instead of assuming defaults.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.