Chapter 22Lesson 01~120 minutes

CPU, Memory, PIDs, Devices, ulimits, cgroups v2, Reservations, and Container Resource Governance: Concepts, Architecture, and Mental Model

Translate Docker CPU, memory, PID, ulimit, and device declarations into cgroup v2 and kernel enforcement evidence.

Resource governancecgroups v2CPU & memoryPIDs & ulimitsEvidence-first

Learning objectives

  • Explain how Docker resource flags become kernel scheduling, cgroup, rlimit, and device policy.
  • Distinguish hard ceilings, relative weights, soft memory pressure controls, and host capacity.
  • Read cgroup v2 evidence such as cpu.max, memory.max, memory.events, and pids.max.
  • Interpret docker stats, exit state, and OOMKilled without guessing from application symptoms.
Chapter 22 principle. A container is not automatically resource-limited. Docker configuration is only the declaration; the kernel scheduler, cgroup controllers, rlimits, device policy, and the capacity of the daemon host are the enforcement boundary. Measure first, then set a limit with a specific failure mode in mind.

1. The practical problem: “containerized” does not mean “bounded”

Earlier chapters established image identity, process lifecycle, mounts, networking, and persistent data. None of those concepts automatically prevents a single container from consuming most available CPU, memory, or process slots. Docker's current resource-constraint documentation is explicit: by default a container can use as much of a resource as the host kernel scheduler allows.

This chapter therefore separates capacity from policy. Host or Docker Desktop VM capacity answers “what exists?”; a container limit answers “what is this workload allowed to consume?”; observability answers “what did it actually consume?”; and application behavior answers “what happened when pressure appeared?” Those are four different facts.

2. Mental model: declaration → kernel control → pressure → evidence

A Docker resource flag changes container configuration in the daemon. On Linux, the runtime turns much of that configuration into cgroup controller state or process rlimits. The kernel then schedules CPU, accounts and constrains memory, limits process creation, and enforces device access. Application symptoms appear after those controls act, so diagnosis must trace backward from evidence rather than infer the cause from a crash.

Resource-control enforcement path
  flowchart TD
    A[Docker CLI / Compose resource declaration] --> B[Engine container configuration]
    B --> C[cgroups v2: cpu / memory / pids]
    B --> D[process rlimits]
    B --> E[device allow / mapping policy]
    C --> F[Linux scheduler + memory controller]
    D --> G[syscall resource ceilings]
    E --> H[device access decision]
    F --> I[throttling / reclaim / OOM / PID denial]
    G --> I
    H --> I
    I --> J[docker stats + inspect + cgroup files + app evidence]
    J --> K[measured resource envelope]
            

3. State to identify before changing a limit

State Evidence Why it matters
Daemon host and context docker context show, docker info Limits apply where the daemon executes. Docker Desktop means the Linux VM, not directly the Windows/macOS host.
Cgroup version / namespace docker info --format {{.CgroupVersion}}; /sys/fs/cgroup/cgroup.controllers File names and some semantics differ between cgroups v1 and v2.
CPU capacity docker info --format {{.NCPU}}; host/VM CPU allocation A --cpus ceiling must be interpreted against available cores and contention.
Memory capacity docker info; Desktop Resources settings A 1 GiB container ceiling means little if the VM itself has only 2 GiB available.
Container policy docker inspect .HostConfig Records quota, memory, PIDs, ulimits, cpuset, and device configuration.
Runtime state docker stats --no-stream, .State.OOMKilled, exit code Shows observed use and how the process terminated.

4. cgroups v2: the enforcement files are evidence, not implementation trivia

Docker has supported cgroups v2 since 20.10. Current Docker documentation notes that the default cgroup driver on v2 is systemd, the default cgroup namespace mode is private, and --oom-kill-disable is discarded on v2. A simple portable probe is the existence of /sys/fs/cgroup/cgroup.controllers.

Inside a Linux container on cgroups v2, files such as cpu.max, cpu.weight, memory.max, memory.high, memory.swap.max, pids.max, and memory.events provide direct evidence. Do not hard-code a host path such as /sys/fs/cgroup/system.slice/docker-<id>.scope unless you first prove the cgroup driver and host topology.

docker info --format 'cgroup={{.CgroupVersion}} driver={{.CgroupDriver}} cpus={{.NCPU}} memory={{.MemTotal}}'
docker run --rm alpine:3.22 sh -c '
  if [ -f /sys/fs/cgroup/cgroup.controllers ]; then
    echo "cgroup=v2"
    cat /proc/self/cgroup
    printf "cpu.max="; cat /sys/fs/cgroup/cpu.max
    printf "memory.max="; cat /sys/fs/cgroup/memory.max
    printf "pids.max="; cat /sys/fs/cgroup/pids.max
  else
    echo "cgroup=v1 or non-v2 layout"
    cat /proc/self/cgroup
  fi
'

5. CPU: ceiling, placement, and relative weight are different controls

--cpus is a convenient quota-style ceiling; Docker documents --cpus=1.5 as equivalent to a 100000 μs period with a 150000 μs quota. --cpuset-cpus answers a different question: which logical CPUs may schedule the workload? --cpu-shares is a relative weight that matters during contention and does not reserve a guaranteed fraction when CPUs are idle.

Control Question it answers Common mistake
--cpus How much CPU time may this container consume? Treating it as a reservation or assuming it selects a specific core.
--cpuset-cpus Which CPUs may run the workload? Using affinity to solve a quota or noisy-neighbor problem without evidence.
--cpu-shares Who gets relatively more CPU during contention? Expecting a hard maximum when the host is otherwise idle.

6. Memory: hard maximum, soft pressure signal, and swap are separate

--memory sets a hard maximum. --memory-reservation is a soft control that becomes relevant under pressure and must be lower than the hard limit to take precedence. --memory-swap represents the combined memory-plus-swap allowance when used with --memory. If both values are equal, Docker documents that the container has no swap allowance. Do not infer a container's usable swap from free; Docker warns that such tools can report host swap rather than the container-specific allowance.

On cgroups v2, read memory.events as part of the failure record. An application can fail allocation itself, or the memory cgroup can kill a process; those are not identical incidents.

7. PIDs, ulimits, and devices constrain different kernel surfaces

--pids-limit limits process/thread tasks in the container's pids cgroup. --ulimit sets per-process resource limits such as open files. Engine 29 release notes document an important current change: with containerd 2.x/systemd defaults, the default nofile limit may be much lower than old Docker defaults, so applications that assumed 1,048,576 file descriptors need explicit evidence.

--device is not “just another resource limit”: it grants a container access to a host device node. Use the narrowest device, permissions, and runtime integration that works. Do not replace a single-device requirement with --privileged.

8. docker stats is a view, not a diagnosis by itself

docker stats is excellent first evidence for CPU, memory, PIDs, block I/O, and network I/O. But every number needs a denominator and a policy context. For example, CPU percentage must be interpreted alongside host/VM CPU count and the container quota; the PIDS column can include threads; and memory usage needs the configured limit plus application semantics.

docker stats --no-stream --format 'table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}\t{{.PIDs}}'

9. DevOps connection: the resource envelope is part of the release contract

A reproducible deployment records image digest, command, mounts, networks, and security posture—but also the resource envelope and the capacity assumptions behind it. A limit copied from another host is not evidence. A useful operating model records baseline load, observed peaks, throttling/OOM/PID behavior, host or VM capacity, and the reason each hard or soft control exists.

Next lesson

Next: apply one bounded control at a time, capture Docker and cgroup evidence, and deliberately trigger a safe memory-limit failure without destabilizing the host.

Knowledge check

Does Docker impose CPU and memory limits automatically?

What is the difference between --cpus and --cpu-shares?

How can you identify cgroups v2 from inside a Linux environment?

Why is OOMKilled=true useful evidence?

Why is --privileged the wrong default for device access?

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline used here. Earlier Engine 29 packaging moved to containerd 2.x; one important consequence documented in Engine 29 release notes is that the default container nofile limit can follow systemd's default instead of older very-high Docker defaults. Docker Desktop runs Linux containers inside a VM, so host capacity, cgroup paths, devices, and swap evidence belong to that VM boundary. Record docker version, docker info, and the actual cgroup mode on your host rather than assuming these values.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.