CPU, Memory, PIDs, Devices, ulimits, cgroups v2, Reservations, and Container Resource Governance: Diagnostics, Failure Modes, Security, and Performance
Diagnose Docker throttling, OOM, PID exhaustion, ulimit exhaustion, and device failures from preserved first-failure evidence.
Learning objectives
- Diagnose resource incidents from preserved state instead of changing limits blindly.
- Differentiate CPU throttling, host contention, memory OOM, PID exhaustion, ulimit exhaustion, and device-denial failures.
- Interpret misleading signals such as high CPU percentages or application crashes after OOM pressure.
- Repair the smallest safe scope without privileged mode, blanket daemon changes, or destructive cleanup.
1. Evidence-first diagnostic sequence
-
Preserve first-failure evidence: container ID, image digest,
timestamps, logs, exit code,
OOMKilled, and current HostConfig. -
Confirm host/platform and
docker version/info/context; on Desktop, record VM resource settings. - Confirm cgroup mode and daemon/runtime state.
- Confirm exact image/build identity; do not assume a new application release uses the same resources.
- Inspect process/health and runtime metrics.
- Inspect cgroup files and rlimits for the suspected controller.
- Only then inspect unrelated storage/network/external dependencies if the evidence points there.
- Apply the least destructive correction and rerun the smallest safe scope.
2. Failure mode: “the app crashed” but the memory cgroup killed it
A memory failure can surface as an abrupt exit, commonly code 137
when a process receives SIGKILL. Do not infer OOM from 137 alone;
inspect .State.OOMKilled, logs,
memory.events, the configured memory/swap values, and
observed working set. A deployment wrapper or explicit
kill -9 can also produce a SIGKILL-style exit.
docker inspect APP --format 'image={{.Image}} exit={{.State.ExitCode}} OOMKilled={{.State.OOMKilled}} error={{json .State.Error}}'
docker inspect APP --format 'memory={{.HostConfig.Memory}} reservation={{.HostConfig.MemoryReservation}} memorySwap={{.HostConfig.MemorySwap}}'
docker stats --no-stream APP
docker logs --timestamps APP > first-failure.log 2>&1
# If still running and cgroup v2:
docker exec APP sh -c 'test -f /sys/fs/cgroup/memory.events && cat /sys/fs/cgroup/memory.events || true'
3. Failure mode: CPU throttling misread as an application regression
First distinguish high demand from quota throttling and host
contention. Inspect cpu.max and, where available,
cpu.stat counters such as throttled periods. A low CPU
percentage may be perfectly consistent with a very low quota.
Conversely, high CPU percentage is not automatically a problem if no
latency/error SLO is violated.
docker inspect APP \
--format 'NanoCpus={{.HostConfig.NanoCpus}} quota={{.HostConfig.CpuQuota}} period={{.HostConfig.CpuPeriod}} cpuset={{.HostConfig.CpusetCpus}} shares={{.HostConfig.CpuShares}}'
docker exec APP sh -c 'test -f /sys/fs/cgroup/cpu.max && { cat /sys/fs/cgroup/cpu.max; cat /sys/fs/cgroup/cpu.stat; } || true'
4. Intentionally broken example: safe PID exhaustion
This disposable example proves that process-creation failure belongs to the PIDs controller, not to DNS or image corruption. It attempts a fixed number of short-lived children under a low limit, preserves the logs, then removes only that container.
docker rm -f dkr22-broken-pids 2>/dev/null || true
docker run --name dkr22-broken-pids --label academy=docker-ch22 --pids-limit 8 alpine:3.22 sh -c '
i=1
while [ "$i" -le 14 ]; do
sleep 5 &
rc=$?
echo "spawn=$i rc=$rc"
i=$((i+1))
done
wait
' || true
docker inspect dkr22-broken-pids --format 'exit={{.State.ExitCode}} pids={{.HostConfig.PidsLimit}}'
docker logs dkr22-broken-pids 2>&1 | tee dkr22-pids-first-failure.log
docker rm dkr22-broken-pids
--pids-limit=-1. Measure the expected process/thread
count, identify leaks or unbounded spawning, then choose a bounded
value with operational headroom.
5. Failure mode: open-file exhaustion
Applications can report “too many open files” even when memory and
PIDs look healthy. Preserve the container's
Ulimits configuration and read the process limit.
Engine 29's current runtime behavior makes explicit
nofile verification important because defaults can
follow systemd rather than older Docker-era assumptions.
docker inspect APP --format '{{json .HostConfig.Ulimits}}'
docker exec APP sh -c 'ulimit -n; cat /proc/1/limits | grep -i "open files" || true'
6. Failure mode: one device is missing, so someone suggests privileged mode
Reject the shortcut. Inspect the device requirement, daemon-host device presence, HostConfig device mapping, runtime/plugin configuration, ownership, and LSM policy. A device denial is a security/identity boundary; turning off the boundary hides the cause and grants unrelated authority.
7. Common causal mistakes
| Symptom | Wrong shortcut | Evidence-first interpretation |
|---|---|---|
| High CPU % | “CPU leak—restart it.” |
Compare quota, host CPUs, cpu.stat, workload
throughput, and latency before judging.
|
| Exit 137 | “Definitely OOM.” |
Check OOMKilled, memory events, and whether an
operator/orchestrator sent SIGKILL.
|
| Cannot fork | “Container runtime is broken.” |
Inspect PidsLimit,
pids.current/pids.max, and
thread/process profile.
|
| Too many open files | “Raise memory.” | Inspect rlimits and file-descriptor usage. |
| Device permission error | “Use privileged.” | Inspect exact device mapping, runtime integration, UID/GID/LSM policy. |
8. Performance: limits should reveal contention, not create hidden queues
Too-tight CPU quota can stretch latency; too-tight memory can turn normal cache behavior into reclaim/OOM churn; too-low PIDs or nofile limits can create burst failures. Treat the limit and its workload metrics as one experiment. A successful benchmark on an idle developer host does not prove production headroom under contention.
9. Least-destructive correction patterns
- Application leak: fix/revert the release; keep the protective memory ceiling while validating.
- Legitimate larger working set: raise the envelope only after measured evidence and host capacity review.
- CPU contention: tune quota/weight or placement based on latency and throughput evidence.
- PID/thread growth: fix unbounded concurrency, then set a ceiling with headroom.
- Open files: set an explicit tested nofile limit, not a host-wide unlimited value.
- Device requirement: use the narrow runtime/device mechanism; never blanket privileged mode.
Knowledge check
Does exit code 137 prove an OOM kill?
No. It is consistent with SIGKILL, but you should verify
.State.OOMKilled, memory events, logs, and external
kill/restart actions.
A container cannot create a child process. Which evidence is most relevant first?
Its PIDs policy and cgroup state:
HostConfig.PidsLimit,
pids.current/pids.max, plus
application logs.
Why preserve cpu.stat before raising a CPU
quota?
It can show throttling evidence and helps distinguish a real quota bottleneck from another latency cause.
What is the safe response to a single missing device?
Inspect and grant only the required device/runtime capability. Do not use privileged mode or disable LSM controls as a shortcut.
Why can a nofile incident appear after an Engine/runtime upgrade?
Runtime/default ulimit behavior can change; Engine 29 release
notes document a shift toward systemd default
LimitNOFILE in containerd 2.x packaging.
Official references and version notes
- Docker Docs — Resource constraints — CPU and memory hard/soft controls, swap semantics, and OOM guidance.
-
Docker Docs — Runtime metrics
—
docker stats, cgroup discovery, and cgroups v2 behavior. - Docker CLI — docker container run — CPU, memory, PID, ulimit, device, and cgroup options.
- Docker CLI — docker container stats — CPU, memory, PIDs, I/O, and one-shot JSON/template output.
- Compose Deploy Specification — resources — limits, reservations, PIDs, and device reservations.
- Docker Desktop — Resources — CPU/memory/swap capacity of the Docker Desktop Linux VM.
- Docker Engine 29 release notes — current Engine baseline and recent runtime/ulimit changes.
-
Linux kernel — Control Group v2
— controller semantics for
cpu.*,memory.*, andpids.*.
Docker
Engine 29.8.1 is the current Engine baseline used here. Earlier
Engine 29 packaging moved to containerd 2.x; one important
consequence documented in Engine 29 release notes is that the
default container nofile limit can follow systemd's
default instead of older very-high Docker defaults. Docker Desktop
runs Linux containers inside a VM, so host capacity, cgroup paths,
devices, and swap evidence belong to that VM boundary. Record
docker version, docker info, and the
actual cgroup mode on your host rather than assuming these values.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.