Chapter 33Lesson 04~185 minutes

Daemon Configuration, daemon.json, systemd, Proxies, Registry Mirrors, Live Restore, and Host Integration: Diagnostics, Failure Modes, Security, and Performance

Diagnose configuration conflicts, invalid JSON, proxy and registry trust failures, unsafe reload assumptions, credential exposure, and restart incidents without destroying evidence or weakening security.

DiagnosticsInvalid JSONConflictsCredentialsRecovery

Learning objectives

  • Preserve daemon/service/config evidence before attempting recovery from configuration failure.
  • Diagnose duplicate flag/JSON ownership, malformed or unsupported JSON, proxy mistakes, registry trust failures, and ineffective reloads causally.
  • Keep credentials, TLS verification, LSMs, firewalls, and daemon API protections intact while troubleshooting.
  • Distinguish configuration parse/validation failure from daemon runtime, registry/proxy, container, and external-service failures.
  • Apply the least destructive correction and verify only the smallest affected scope before wider rollout.

1. Evidence-first daemon diagnostic sequence

Daemon incidents are dangerous because the most obvious “fix” is often a restart—which can destroy the evidence needed to explain why startup, pull, reload, or continuity failed. Use this order:

  1. Preserve first-failure output and daemon logs.
  2. Confirm host/platform, context, and Engine/CLI versions.
  3. Record service-manager ownership and exact dockerd command line.
  4. Preserve/redact daemon.json and relevant drop-ins.
  5. Validate the proposed/current file offline when possible.
  6. Confirm container identities and external health before daemon lifecycle actions.
  7. Separate proxy/registry reachability from daemon parsing/config ownership.
  8. Apply one minimal correction.
  9. Revalidate, then reload or restart only if the corrected setting requires it.
  10. Compare daemon PID/start time, container state, logs, and external request evidence.

2. Failure: same option in startup flags and daemon.json

Docker explicitly rejects options defined both as daemon flags and in the configuration file. A common Linux example is attempting to set hosts in daemon.json while a packaged systemd unit already starts Docker with -H. The fix is not to keep editing JSON until startup works; first decide which layer owns the option.

# Read-only ownership evidence
systemctl show docker --property=FragmentPath,DropInPaths
ps -eo pid,args | grep '[d]ockerd'

# Inspect JSON privately, then validate a corrected copy before rollout:
sudo dockerd --validate --config-file=/tmp/corrected-daemon.json

3. Intentionally broken example: unsupported directive

cat > /tmp/dca33-broken.json <<'EOF'
{
  "labels": ["devops-academy.chapter=33-diagnostic"],
  "imaginary-setting": true
}
EOF

set +e
sudo dockerd --validate --config-file=/tmp/dca33-broken.json   > /tmp/dca33-broken.out 2>&1
RC=$?
set -e
cat /tmp/dca33-broken.out
printf 'exit=%s
' "$RC"

Interpretation: validation should fail before any daemon change because the directive is not recognized. Preserve the output. Repair by removing/replacing the unsupported key, then rerun validation. Do not restart the current daemon just to discover a parse/config error.

4. Failure: malformed JSON

A missing comma, duplicated structural fragment, or invalid JSON syntax can prevent startup. Use a JSON parser plus Docker’s own validator because generic JSON validity is necessary but not sufficient.

python -m json.tool /tmp/candidate-daemon.json >/dev/null
sudo dockerd --validate --config-file=/tmp/candidate-daemon.json

If Python is unavailable, use another trusted local JSON parser; Docker validation remains the authoritative Engine-level check.

5. Failure: proxy configured at the wrong layer

Observed symptom Likely layer Evidence Minimal correction
docker pull cannot reach registry; app proxy works daemon proxy daemon config/service env + daemon log + network test from host configure daemon proxy, not application env
app cannot reach internet; daemon pulls work container/application proxy container inspect/env + app logs runtime/app proxy policy
build dependency fetch fails; runtime works build/BuildKit network/proxy plain build log + build config build-specific proxy/secret/network policy

6. Failure: proxy credentials exposed in unit files or evidence

If a systemd drop-in includes credentials in a proxy URL, systemctl cat or configuration backups can expose them to users or logs with access. Do not “fix” this by printing values during troubleshooting. Revoke/rotate leaked credentials, move secret handling to the organization’s approved protected mechanism, and retain only redacted endpoint evidence.

7. Failure: broad insecure-registry entry as a TLS shortcut

An insecure-registry setting weakens transport assumptions for the named endpoint. It is not a cure for an unknown CA or expired certificate. For production, repair the registry certificate/CA trust and confirm hostname matching. Keep any development exception narrow, documented, time-bounded, and isolated.

Do not troubleshoot by disabling TLS verification globally. Preserve the original certificate error and fix the trust chain.

8. Failure: “I reloaded Docker, but the change did nothing”

First ask whether Docker documents the key as reloadable. If not, no amount of SIGHUP repetition will transform a startup-only setting into a hot change. Capture the current daemon PID and effective state, classify the change as restart-required, and move it into an approved maintenance/migration path.

If the key is reloadable, inspect daemon logs around the reload. Docker can reject a reload conflict without terminating the daemon, so “Docker is still running” is not proof that the new config took effect.

9. Failure: restart before preserving daemon logs

Daemon startup/reload failures often emit the clearest explanation once. On systemd Linux, capture the relevant journal window before retrying:

date -u +%Y-%m-%dT%H:%M:%SZ
sudo journalctl -u docker --since '-15 min' --no-pager > /tmp/dca33-docker-journal.txt
systemctl show docker --property=MainPID,ActiveEnterTimestamp > /tmp/dca33-service-state.txt

Redact secrets before sharing. On Docker Desktop or non-systemd platforms, use the documented platform-specific daemon logs instead of inventing a systemd path.

10. Failure: live restore is enabled but continuity still breaks

Check Why it matters
Container is standalone Linux container Live restore scope differs for Windows containers and Swarm services
Daemon options stayed compatible Bridge/storage/network option changes can prevent reattachment
Upgrade is within supported patch scope Docker documents live restore across patch releases, not arbitrary major upgrades
Application logging rate during outage Full FIFO can block logging process activity
External health, not only container state A surviving process can still be unhealthy/unreachable

11. Performance diagnosis: mirror/proxy can be the bottleneck

Adding a mirror or corporate proxy inserts another network/cache/TLS component. Measure DNS, TCP/TLS establishment, cache hit/miss behavior, upstream latency, and registry response codes before raising Docker concurrency limits. Tuning max-concurrent-downloads without endpoint evidence can amplify load and make throttling worse.

12. Minimal-correction runbook

1. Preserve the first error and daemon log window.
2. Record context, versions, systemd owner, dockerd command line, and redacted config.
3. Identify the failing layer: parse/ownership, daemon runtime, proxy, registry/TLS, container, or external service.
4. Correct one cause in a candidate file.
5. Run JSON parsing + dockerd --validate.
6. Decide reload versus restart from current docs.
7. Capture daemon/container identities before the transition.
8. Apply only in authorized scope.
9. Verify effective config, daemon logs, container continuity, and external request.
10. Roll back if predicted state does not match evidence.

13. Security red lines for daemon troubleshooting

  • Do not expose an unauthenticated Docker TCP API.
  • Do not mount the Docker socket into an untrusted helper container.
  • Do not disable TLS, firewalls, seccomp, AppArmor, or SELinux to “see if it works.”
  • Do not use broad root filesystem permissions or chmod 777.
  • Do not mark all registries insecure.
  • Do not print secrets into terminal logs or CI output.
  • Do not restart a production daemon before capturing first-failure evidence and continuity expectations.

Knowledge check

Docker stays running after a SIGHUP, but your registry mirror did not change. What evidence do you need next?

Why is generic JSON validation not enough for daemon.json?

What is the correct response to an unknown registry CA in production?

Why can live restore still produce an application incident?

A proxy password appears in a captured systemd drop-in. What two actions are required?

Next lesson

Next: Checkpoint Lab — Daemon Configuration, daemon.json, systemd, Proxies, Registry Mirrors, Live Restore, and Host Integration

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Diagnostic baseline date: 2026-09-22. The runbook intentionally keeps parse/validation, daemon runtime, service-manager, proxy/registry, and container health evidence separate so a restart cannot hide the original causal layer.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.