Chapter 22Lesson 03~120 minutes

CPU, Memory, PIDs, Devices, ulimits, cgroups v2, Reservations, and Container Resource Governance: Configuration, Design Choices, and Tradeoffs

Choose resource limits, soft controls, affinity, swap policy, host capacity, and device access using explicit tradeoffs and observable prerequisites.

Resource governancecgroups v2CPU & memoryPIDs & ulimitsEvidence-first

Learning objectives

  • Choose between quota, cpuset, relative CPU weight, memory limit/reservation, and host-level capacity controls.
  • Explain memory/swap combinations and reservations without treating soft controls as guarantees.
  • Compare per-container device access with mediated services and platform-managed device reservation.
  • Document platform prerequisites and observable evidence for each resource-governance decision.
Design rule. Resource governance is a workload-and-platform contract. Hard limits, soft controls, placement, and reservations answer different questions; none should be chosen because a copied example “looks reasonable.”

1. Start with the question, not the flag

Before selecting a Docker option, state the failure you are trying to prevent or the service objective you are trying to protect. “Use less CPU” is vague. “Prevent this batch job from consuming more than one core while latency-sensitive services share the host” leads naturally to a CPU quota. “Keep a licensed single-thread workload on CPUs 2–3” points toward affinity. “Prefer this service under contention” points toward relative weight.

2. Hard limits versus reservations and soft controls

Mechanism Semantics Use when Do not claim
--memory Hard memory ceiling A runaway process must not consume unbounded RAM That it guarantees application performance.
--memory-reservation Soft memory target under pressure You want reclaim/pressure behavior below a hard maximum That the workload can never exceed it.
--cpus CPU time ceiling You need a predictable maximum CPU share That CPU is pre-reserved when idle or contended.
--cpu-shares Relative weight under contention Several workloads compete for CPU and priority matters That it is a hard cap.
Compose deploy reservations Platform-level reservation model The target platform honors the declared reservation semantics That standalone Engine creates physical capacity that is not present.

3. CPU quota versus cpuset

A quota changes how much CPU time is available. A cpuset changes where the process may run. They can be combined, but each narrows a different dimension. Affinity can improve isolation or cache locality in specialized workloads, yet it can also strand capacity and reduce scheduler flexibility. Prefer quota first unless measurements justify topology-aware placement.

Relative CPU weight is meaningful only when workloads contend. A container with low shares can still use idle CPU. This distinction is essential when interpreting a benchmark on an otherwise quiet developer machine.

4. Memory and swap policy: choose a failure mode deliberately

Swap can absorb bursts but adds latency. No swap makes the memory ceiling more deterministic but increases the chance that the cgroup kills a process during spikes. Docker documents that --memory-swap is total memory plus swap; setting it equal to --memory disables additional swap, while -1 allows unlimited swap up to host availability.

For latency-sensitive services, “no swap” may be defensible. For batch work, some swap may be acceptable. The correct policy depends on SLOs, working-set behavior, host storage, and failure recovery—not ideology.

5. Per-container governance versus host/VM governance

Container limits partition available capacity; they do not create capacity. On Docker Desktop, the Linux VM itself has configurable CPU, RAM, swap, and disk limits. A four-CPU container cannot consume four host CPUs if the Docker VM has only two. Similarly, several containers can each have 2 GiB hard limits while the VM has less aggregate RAM, making simultaneous peaks impossible.

Capacity planning therefore has two levels: host/VM envelope and workload envelopes. Keep both in the evidence packet.

6. Device access versus a mediated service

Direct device mapping couples a workload to one daemon host and expands authority. Prefer a mediated service, network API, or runtime-specific device plugin/driver when it provides the function without direct device-node access. If direct mapping is required, pin the exact device and access mode; do not mount all of /dev and do not use privileged mode as a convenience.

7. Compose resource declarations: portable model, platform-specific enforcement

The Compose Deploy Specification models limits and reservations for CPU, memory, PIDs, and devices. That declarative model is valuable, but actual enforcement depends on the platform executing it. Validate the normalized Compose model and then inspect the resulting container HostConfig/cgroups on the target platform.

# Example model; verify support on the target platform before relying on reservations.
services:
  api:
    image: example/api@sha256:RECORD_THE_REAL_DIGEST
    deploy:
      resources:
        limits:
          cpus: "0.50"
          memory: 256M
          pids: 128
        reservations:
          memory: 128M

8. Decision table: choose an approach with prerequisites and evidence

Scenario Choice Prerequisites / trust boundary Verify
CPU-heavy batch job on shared host --cpus quota Linux scheduler/cgroup support; known host CPU capacity HostConfig, cpu.max, stats, completion time.
Latency service with brief memory spikes Hard memory ceiling + measured soft reservation Measured working set; tested pressure behavior memory.max, memory.high where applicable, memory.events, app latency.
Worker creates many threads PIDs limit with headroom Thread/process profile known pids.max, stats PIDS, app error path.
Needs one GPU/device Runtime-specific narrow device reservation/mapping Authorized device and host driver/runtime HostConfig/device reservation plus successful minimal capability test.
Developer laptop Docker Desktop Tune Desktop VM first, then containers Known VM CPU/RAM/swap configuration Desktop resources + docker info + per-container policy.

9. Reproducibility and rollback

Record the image digest and resource policy together. If a release changes memory behavior, compare the same workload and policy before deciding the limit is wrong. Rollback can mean restoring the previous image, restoring the previous envelope, or both. Changing limits without preserving the prior configuration destroys the experiment.

Next lesson

Next: diagnose throttling, OOM, PID exhaustion, ulimit exhaustion, and device failures while preserving first-failure evidence.

Knowledge check

Is --memory-reservation a hard ceiling?

When should you prefer --cpuset-cpus over --cpus?

What does setting --memory-swap=-1 mean?

Can per-container limits compensate for an undersized Docker Desktop VM?

Why can Compose reservations require platform verification?

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline used here. Earlier Engine 29 packaging moved to containerd 2.x; one important consequence documented in Engine 29 release notes is that the default container nofile limit can follow systemd's default instead of older very-high Docker defaults. Docker Desktop runs Linux containers inside a VM, so host capacity, cgroup paths, devices, and swap evidence belong to that VM boundary. Record docker version, docker info, and the actual cgroup mode on your host rather than assuming these values.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.