Chapter 02Lesson 04~105 minutes

Installing Docker Engine and Desktop, CLI Basics, Daemon Connectivity, Contexts, and Environment Setup: Diagnostics, Failure Modes, Security, and Performance

Troubleshoot Docker installation and connectivity by preserving first failure and localizing it to host/platform, client/environment, context/endpoint, authorization, daemon service/configuration, storage, proxy/network, Desktop backend, or version/plugin layers before making changes.

DiagnosticsContextsPermissionsDaemon configDesktop/WSL

Learning objectives

  • Apply an evidence-first diagnostic sequence before restart, reinstall, configuration mutation, or cleanup.
  • Diagnose wrong-context success, socket permission denial, daemon startup failure, registry/proxy failure, Desktop/WSL confusion, and version-skew misconceptions.
  • Use dockerd --validate with a temporary intentionally broken configuration instead of disrupting the real daemon.
  • Reject insecure troubleshooting shortcuts such as world-writable sockets, plaintext remote APIs, data-root deletion, disabled TLS/firewalls, and blind reinstall.
  • Separate installation correctness from performance symptoms and measure the actual bottleneck boundary before tuning.
Chapter 02 technical baseline — verified 2026-09-20. Docker Engine 29.8.1 and Docker Desktop 4.91.0 are current releases at this verification date. Current upstream Buildx, BuildKit, and Compose releases are 0.37.1, 0.33.0, and 5.5.1 respectively. Do not infer those exact bundled versions from the product name: every lab records the actual client/server/API/context/storage/component baseline first.

1. Diagnose the first broken layer, not the loudest symptom

Docker installation failures are often misdiagnosed because multiple layers produce similar messages. “Cannot connect” might mean the daemon is stopped, the active context targets another endpoint, socket permissions deny access, an environment variable overrides the context, or Docker Desktop's VM is not running. A package install can be correct while client targeting is wrong.

Use a fixed sequence: preserve the first error → identify host/platform → inspect client and context → inspect endpoint/socket → inspect daemon service/Desktop backend → inspect daemon configuration/logs → inspect storage/proxy/network dependencies → run the smallest verification workload. Do not reinstall before locating the failing layer.

2. First-response triage

date -u +%Y-%m-%dT%H:%M:%SZ
uname -a 2>/dev/null || true
docker --version 2>/dev/null || true
docker context show 2>/dev/null || true
docker context ls 2>/dev/null || true
printf 'DOCKER_CONTEXT=%s
' "${DOCKER_CONTEXT:-}"
printf 'DOCKER_HOST=%s
' "${DOCKER_HOST:-}"
printf 'DOCKER_API_VERSION=%s
' "${DOCKER_API_VERSION:-}"
docker version 2>&1 || true
docker info 2>&1 || true

Preserve this output before changing contexts, restarting services, editing configuration, or re-running installers. A second failure after mutation is not the same evidence as the original one.

3. Failure mode: the command succeeded against the wrong daemon

This is more dangerous than a connection failure because it can create or delete valid resources somewhere else. A stale DOCKER_CONTEXT environment variable overrides the persistent context, and a command-line --context override can target yet another endpoint.

docker context show
docker context inspect "$(docker context show)"
env | grep '^DOCKER_' || true

# Read-only comparison; do not switch persistent defaults just to inspect.
docker --context default version 2>/dev/null || true
docker --context desktop-linux version 2>/dev/null || true

Corrective action is to remove the unintended override or choose the intended context explicitly. Do not delete “unexpected” containers until you know which daemon owns them.

4. Failure mode: permission denied on the Docker socket

On a normal rootful Linux installation, the default Unix socket is owned by root and commonly accessible to the docker group. Permission denial is not evidence that the daemon is broken. It can simply mean your user lacks authorization to that privileged socket.

ls -l /var/run/docker.sock 2>/dev/null || true
id
getent group docker || true
sudo docker version

If sudo docker version works while non-sudo access fails, the causal layer is local authorization. The unsafe fix is chmod 666 /var/run/docker.sock. The design choices are explicit sudo, trusted docker-group membership with root-equivalent implications, or a planned rootless installation.

5. Failure mode: daemon does not start

Collect service status and journal evidence before editing anything. A daemon can fail because of invalid JSON, conflicting flags/config keys, unsupported storage/network settings, missing directories, permissions, or environmental dependencies.

sudo systemctl status docker --no-pager || true
sudo journalctl -u docker --since '-15 min' --no-pager || true
sudo systemctl cat docker || true
sudo test -f /etc/docker/daemon.json && sudo cat /etc/docker/daemon.json || true

Do not delete /var/lib/docker because the daemon will not start. Data-root destruction is not a diagnostic method.

6. Intentionally broken example — validate daemon configuration offline

This lab creates a temporary invalid configuration file and asks dockerd to validate it without starting another daemon or touching the real configuration. This is a safe way to learn the error shape.

cat >/tmp/da-ch02-invalid-daemon.json <<'EOF'
{
  "log-level": "info",
  "this-option-does-not-exist": true
}
EOF

set +e
sudo dockerd --validate --config-file=/tmp/da-ch02-invalid-daemon.json
STATUS=$?
set -e
printf 'validation exit=%s
' "$STATUS"
rm -f /tmp/da-ch02-invalid-daemon.json

The expected result is a non-zero exit with an unknown-directive error. The repair is to correct the configuration before any daemon restart. dockerd --validate is therefore a change-risk control, not just a troubleshooting convenience.

7. Failure mode: daemon is healthy but pulls fail

Registry access belongs to the daemon/build network path, not merely the shell that runs the CLI. A corporate proxy configured only in the user's terminal may not configure the system daemon. Preserve the registry error, DNS resolution, proxy configuration, and daemon logs before changing certificate or firewall settings.

Never solve a corporate TLS interception or registry trust issue by globally disabling TLS verification or declaring arbitrary registries insecure. Identify the organization's approved CA/proxy/registry path and configure that specific trust relationship.

8. Failure mode: Windows/WSL/Desktop boundary confusion

On Windows, three different facts can be mistaken for each other: WSL is installed, Docker Desktop is running, and the Docker CLI is targeting the Desktop backend. Verify each independently.

wsl --version
docker --version
docker context ls
docker context show
docker version

If Docker Desktop does not start, inspect Desktop diagnostics and host virtualization/WSL prerequisites rather than installing another Docker Engine inside the same WSL distribution. Docker's WSL guidance warns against running a separate Engine/CLI stack inside WSL alongside Desktop because it can conflict with the intended integration.

9. Failure mode: version skew misunderstood as corruption

A CLI and daemon can have different versions. First read docker version and the negotiated API. If one command is unknown or a field is missing, check whether it is a client feature, server API feature, CLI plugin feature, or Desktop-only feature. Reinstalling both sides may hide the question without answering it.

docker version
docker buildx version 2>/dev/null || true
docker compose version 2>/dev/null || true

Buildx and Compose are independently released CLI plugins. Their versions do not need to numerically match Engine.

10. Failure mode: an unreviewed convenience script became production state

The risk is not that the official script is inherently malicious; the risk is lifecycle opacity. The production host may have received whichever stable packages were current at bootstrap time, with no reviewed version pin or upgrade plan. The correct recovery is to inventory the resulting repositories/packages, document ownership, and transition to normal package-manager upgrades—not repeatedly re-run the bootstrap script.

apt-cache policy docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin 2>/dev/null || true
dpkg-query -W -f='${Package} ${Version}
'   docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin 2>/dev/null || true

11. Performance during setup: measure before “tuning Docker”

Slow pulls, high disk use, or poor Desktop build performance can come from registry latency, proxy behavior, VM disk placement, filesystem sharing, storage drivers, image size, or host resource pressure. More CPU/memory is not a universal fix. Record download timing, disk free space, data-root location, storage driver/image store, and Desktop VM allocation before changing settings.

docker info --format 'DockerRootDir={{.DockerRootDir}} Driver={{.Driver}}'
df -h 2>/dev/null | sed -n '1,12p'
docker system df 2>/dev/null || true

Chapter 40 performs full Docker performance engineering. Chapter 02 only establishes the discipline: measure the installation boundary before tuning it.

12. Evidence-first incident runbook

  1. Capture the exact error and timestamp.
  2. Record host OS/architecture and Docker client version.
  3. Record DOCKER_* overrides, context, and endpoint.
  4. Confirm server reachability and API/component baseline.
  5. For native Engine, inspect service status/journal and validate daemon configuration before restart.
  6. Check permissions, disk/data-root, DNS/proxy/registry only if evidence points there.
  7. Apply one correction and repeat the smallest failed check.
  8. Do not prune/delete/reinstall until resource ownership and recovery are proven.

Knowledge check

Why is a command that succeeds against the wrong context more dangerous than “Cannot connect to the daemon”?

If sudo docker version works but docker version returns permission denied, which layer is failing?

Why is dockerd --validate preferable to “edit daemon.json and restart until it works”?

Why can a shell HTTP_PROXY setting fail to fix image pulls?

Why should you not install a second Docker Engine inside WSL merely because Docker Desktop cannot connect?

Summary

Installation troubleshooting is a layer-localization exercise: host/platform, client/environment, context/endpoint, permissions, daemon service/configuration, storage, proxy/network, and Desktop backend are separate. Preserve first failure, validate before restart, and never use privilege escalation, broad socket permissions, data-root deletion, or blind reinstall as routine fixes. The checkpoint lesson now asks you to prove an environment through a controlled restart and evidence packet.

Next lesson

Next: Checkpoint Lab — Installing Docker Engine and Desktop, CLI Basics, Daemon Connectivity, Contexts, and Environment Setup

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version-sensitive statements in this chapter were rechecked against primary documentation on 2026-09-20. Installation support, package names, Desktop requirements, component versions, context behavior, rootless prerequisites, and security guidance evolve. Always record the versions and platform facts actually reported by the environment you are operating.

Current baseline, not a frozen requirement

At this chapter's verification date, Docker Engine 29.8.1 and Docker Desktop 4.91.0 are current releases, while Buildx 0.37.1, BuildKit 0.33.0, and Compose 5.5.1 are current upstream releases. Bundled component versions can differ by Engine/Desktop/package source. Use docker version, docker info, docker buildx version, and docker compose version as execution evidence rather than assuming those upstream versions are installed.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.