Chapter 22Lesson 04~120 minutes

CPU, Memory, PIDs, Devices, ulimits, cgroups v2, Reservations, and Container Resource Governance: Diagnostics, Failure Modes, Security, and Performance

Diagnose Docker throttling, OOM, PID exhaustion, ulimit exhaustion, and device failures from preserved first-failure evidence.

Resource governancecgroups v2CPU & memoryPIDs & ulimitsEvidence-first

Learning objectives

  • Diagnose resource incidents from preserved state instead of changing limits blindly.
  • Differentiate CPU throttling, host contention, memory OOM, PID exhaustion, ulimit exhaustion, and device-denial failures.
  • Interpret misleading signals such as high CPU percentages or application crashes after OOM pressure.
  • Repair the smallest safe scope without privileged mode, blanket daemon changes, or destructive cleanup.
Incident rule. Preserve the failed container, its limits, stats snapshot, logs, exit/OOM state, and relevant cgroup files before changing anything. “Increase the limit and retry” is not diagnosis.

1. Evidence-first diagnostic sequence

  1. Preserve first-failure evidence: container ID, image digest, timestamps, logs, exit code, OOMKilled, and current HostConfig.
  2. Confirm host/platform and docker version/info/context; on Desktop, record VM resource settings.
  3. Confirm cgroup mode and daemon/runtime state.
  4. Confirm exact image/build identity; do not assume a new application release uses the same resources.
  5. Inspect process/health and runtime metrics.
  6. Inspect cgroup files and rlimits for the suspected controller.
  7. Only then inspect unrelated storage/network/external dependencies if the evidence points there.
  8. Apply the least destructive correction and rerun the smallest safe scope.

2. Failure mode: “the app crashed” but the memory cgroup killed it

A memory failure can surface as an abrupt exit, commonly code 137 when a process receives SIGKILL. Do not infer OOM from 137 alone; inspect .State.OOMKilled, logs, memory.events, the configured memory/swap values, and observed working set. A deployment wrapper or explicit kill -9 can also produce a SIGKILL-style exit.

docker inspect APP --format 'image={{.Image}} exit={{.State.ExitCode}} OOMKilled={{.State.OOMKilled}} error={{json .State.Error}}'
docker inspect APP --format 'memory={{.HostConfig.Memory}} reservation={{.HostConfig.MemoryReservation}} memorySwap={{.HostConfig.MemorySwap}}'
docker stats --no-stream APP
docker logs --timestamps APP > first-failure.log 2>&1
# If still running and cgroup v2:
docker exec APP sh -c 'test -f /sys/fs/cgroup/memory.events && cat /sys/fs/cgroup/memory.events || true'

3. Failure mode: CPU throttling misread as an application regression

First distinguish high demand from quota throttling and host contention. Inspect cpu.max and, where available, cpu.stat counters such as throttled periods. A low CPU percentage may be perfectly consistent with a very low quota. Conversely, high CPU percentage is not automatically a problem if no latency/error SLO is violated.

docker inspect APP \
  --format 'NanoCpus={{.HostConfig.NanoCpus}} quota={{.HostConfig.CpuQuota}} period={{.HostConfig.CpuPeriod}} cpuset={{.HostConfig.CpusetCpus}} shares={{.HostConfig.CpuShares}}'
docker exec APP sh -c 'test -f /sys/fs/cgroup/cpu.max && { cat /sys/fs/cgroup/cpu.max; cat /sys/fs/cgroup/cpu.stat; } || true'

4. Intentionally broken example: safe PID exhaustion

This disposable example proves that process-creation failure belongs to the PIDs controller, not to DNS or image corruption. It attempts a fixed number of short-lived children under a low limit, preserves the logs, then removes only that container.

docker rm -f dkr22-broken-pids 2>/dev/null || true
docker run --name dkr22-broken-pids --label academy=docker-ch22   --pids-limit 8 alpine:3.22 sh -c '
    i=1
    while [ "$i" -le 14 ]; do
      sleep 5 &
      rc=$?
      echo "spawn=$i rc=$rc"
      i=$((i+1))
    done
    wait
  ' || true

docker inspect dkr22-broken-pids --format 'exit={{.State.ExitCode}} pids={{.HostConfig.PidsLimit}}'
docker logs dkr22-broken-pids 2>&1 | tee dkr22-pids-first-failure.log
docker rm dkr22-broken-pids
Repair. Do not blindly set --pids-limit=-1. Measure the expected process/thread count, identify leaks or unbounded spawning, then choose a bounded value with operational headroom.

5. Failure mode: open-file exhaustion

Applications can report “too many open files” even when memory and PIDs look healthy. Preserve the container's Ulimits configuration and read the process limit. Engine 29's current runtime behavior makes explicit nofile verification important because defaults can follow systemd rather than older Docker-era assumptions.

docker inspect APP --format '{{json .HostConfig.Ulimits}}'
docker exec APP sh -c 'ulimit -n; cat /proc/1/limits | grep -i "open files" || true'

6. Failure mode: one device is missing, so someone suggests privileged mode

Reject the shortcut. Inspect the device requirement, daemon-host device presence, HostConfig device mapping, runtime/plugin configuration, ownership, and LSM policy. A device denial is a security/identity boundary; turning off the boundary hides the cause and grants unrelated authority.

7. Common causal mistakes

Symptom Wrong shortcut Evidence-first interpretation
High CPU % “CPU leak—restart it.” Compare quota, host CPUs, cpu.stat, workload throughput, and latency before judging.
Exit 137 “Definitely OOM.” Check OOMKilled, memory events, and whether an operator/orchestrator sent SIGKILL.
Cannot fork “Container runtime is broken.” Inspect PidsLimit, pids.current/pids.max, and thread/process profile.
Too many open files “Raise memory.” Inspect rlimits and file-descriptor usage.
Device permission error “Use privileged.” Inspect exact device mapping, runtime integration, UID/GID/LSM policy.

8. Performance: limits should reveal contention, not create hidden queues

Too-tight CPU quota can stretch latency; too-tight memory can turn normal cache behavior into reclaim/OOM churn; too-low PIDs or nofile limits can create burst failures. Treat the limit and its workload metrics as one experiment. A successful benchmark on an idle developer host does not prove production headroom under contention.

9. Least-destructive correction patterns

  • Application leak: fix/revert the release; keep the protective memory ceiling while validating.
  • Legitimate larger working set: raise the envelope only after measured evidence and host capacity review.
  • CPU contention: tune quota/weight or placement based on latency and throughput evidence.
  • PID/thread growth: fix unbounded concurrency, then set a ceiling with headroom.
  • Open files: set an explicit tested nofile limit, not a host-wide unlimited value.
  • Device requirement: use the narrow runtime/device mechanism; never blanket privileged mode.
Next lesson

Next: assemble a full evidence packet, profile a synthetic workload, apply bounded controls, and justify a resource envelope.

Knowledge check

Does exit code 137 prove an OOM kill?

A container cannot create a child process. Which evidence is most relevant first?

Why preserve cpu.stat before raising a CPU quota?

What is the safe response to a single missing device?

Why can a nofile incident appear after an Engine/runtime upgrade?

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline used here. Earlier Engine 29 packaging moved to containerd 2.x; one important consequence documented in Engine 29 release notes is that the default container nofile limit can follow systemd's default instead of older very-high Docker defaults. Docker Desktop runs Linux containers inside a VM, so host capacity, cgroup paths, devices, and swap evidence belong to that VM boundary. Record docker version, docker info, and the actual cgroup mode on your host rather than assuming these values.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.