Chapter 25Lesson 03~120 minutes

Container Security Model: Namespaces, Capabilities, Seccomp, AppArmor, SELinux, Devices, and Privilege Boundaries: Configuration, Design Choices, and Tradeoffs

Choose default or custom seccomp, capability allowlists, LSM policy, read-only filesystems, rootless/user namespaces, and stronger isolation using explicit tradeoffs.

Container securityCapabilitiesSeccompAppArmor / SELinuxLeast privilege

Learning objectives

  • Choose between default and custom seccomp policy, a capability allowlist, LSM policy, read-only filesystem design, rootless/user namespaces, and stronger VM/sandbox isolation.
  • Connect each control to portability, least privilege, diagnosability, failure isolation, host cost, and multi-tenant threat assumptions.
  • State explicit prerequisites and observable evidence for a security design rather than treating a flag as proof of safety.
  • Recognize when hostile multi-tenancy requires a stronger boundary than a shared kernel container.
Design rule. Start with workload and threat model, then choose controls. Security flags are implementation mechanisms; they are not a substitute for deciding which principals, kernel operations, host resources, and external systems the workload must be allowed to reach.

1. Default seccomp versus a custom profile

The built-in profile tracks Docker’s runtime expectations and current hardening. A custom profile can solve a legitimate compatibility requirement or enforce a narrower application policy, but it creates maintenance responsibility: architecture differences, new application versions, kernel behavior, and Docker security updates must all be reviewed. Prefer the default unless there is evidence for a deliberate change.

2. Capability allowlist versus default capability set

A practical hardening path is to test with cap-drop=ALL and then add back the smallest capability set required. This creates an explicit allowlist and makes authority changes visible in code review. The cost is compatibility work for software that silently assumes traditional root privileges. Treat capability additions as documented requirements with a reproducer.

3. LSM policy: host/distribution context matters

AppArmor and SELinux solve related mandatory-access-control problems through different policy models and are integrated differently by distributions. Do not ship one host’s policy assumption as if it were portable everywhere. Instead, document the required host control, provide a safe baseline on unsupported platforms, and test enforcement where the production environment actually runs.

4. Read-only filesystem versus writable runtime state

A read-only root filesystem makes unauthorized or accidental mutation harder and exposes software that improperly writes into its installation tree. It works best when writable state is classified: ephemeral state goes to tmpfs, persistent state goes to an explicit managed volume, and configuration comes from declared runtime inputs. This also improves incident reasoning because writable surfaces are enumerable.

5. Rootless Docker and user namespaces are additional layers

User namespaces remap container identities; rootless Docker runs the daemon and containers without a rootful daemon. Both can reduce host impact from certain compromise paths, but each changes networking, storage, device, and privileged-operation assumptions. They do not make broad mounts, weak application authorization, or exposed credentials safe. Chapter 26 treats these tradeoffs in depth.

6. Shared kernel versus stronger VM/sandbox boundary

For mutually hostile or strongly isolated tenants, the shared host kernel can be the wrong trust boundary even when container controls are hardened. A microVM, VM, sandboxed runtime, or dedicated host may provide a more appropriate isolation layer. This costs density, startup time, operational complexity, or hardware resources, but it changes the failure domain rather than only tightening one container policy.

7. Decision table: choose the control from the failure boundary

Need / threat Preferred control Prerequisites/evidence Tradeoff
Ordinary app does not need kernel privileges Non-root + drop capabilities; add back narrowly Application tests; inspect user/caps May expose hidden privilege assumptions.
Reduce syscall surface Built-in seccomp first; custom reviewed profile only when justified Kernel seccomp support; denied-syscall evidence Custom policy maintenance burden.
Constrain filesystem/network/object access beyond Unix permissions AppArmor or SELinux policy Host LSM active; audit evidence Distribution-specific policy work.
Prevent rootfs mutation Read-only rootfs + explicit tmpfs/volumes Writable-path inventory Requires application path cleanup.
Reduce rootful daemon/container host impact Rootless Docker or user namespaces Supported kernel/userspace; mapped IDs; feature review Networking/storage/device limitations.
Run hostile multi-tenant workloads VM/microVM/sandbox/dedicated host boundary Threat model and platform support Higher resource/operational cost.

8. Worked scenario: image-processing API

Assume an API reads an uploaded image, writes temporary conversion output, and returns the result. It needs no host devices, no host namespace, no persistent rootfs writes, and no Docker daemon access. A justified baseline is a non-root UID, dropped capabilities, default seccomp/LSM policy, read-only rootfs, and a bounded tmpfs for temporary conversion. If the conversion library later needs one blocked syscall, preserve the failing syscall evidence and review the narrowest policy change instead of expanding unrelated permissions.

Prediction Observable proof
Root filesystem mutations fail Attempted write to a non-tmpfs root path fails and inspect shows ReadonlyRootfs=true.
Temporary conversion succeeds Write/read within the dedicated tmpfs and inspect shows that tmpfs target.
No Linux capabilities retained Inspect requested cap-drop and runtime capability evidence where available.
Default security layers remain enabled Docker info/host policy state and absence of a container override that disables them.

9. Portability and Docker Desktop

Docker Desktop runs Linux containers inside a managed Linux VM. The effective kernel, cgroup, namespace, AppArmor/SELinux availability, device surface, and host filesystem path semantics therefore differ from a native Linux Engine. Record the daemon context and platform before claiming that a host security control is active.

Knowledge check

When should you prefer Docker’s built-in seccomp profile?

What is the value of cap-drop=ALL plus narrow add-back?

Why is read-only rootfs easier to operate when writable paths are classified?

Does rootless Docker replace application-level authorization?

When is a VM/sandbox boundary preferable to further container hardening?

Next lesson

Next: Container Security Model: Namespaces, Capabilities, Seccomp, AppArmor, SELinux, Devices, and Privilege Boundaries: Diagnostics, Failure Modes, Security, and Performance

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Version/security baseline, verified 2026-09-21.

Docker Engine 29.8.1 is the current Engine baseline. The built-in seccomp policy remains a moderately protective allowlist-style default and currently blocks roughly 44 syscalls out of 300+. Docker’s AppArmor integration uses a generated docker-default profile where AppArmor is active; Engine 29.8 adds support for generating that profile from a custom daemon template. The 2026 CVE-2026-31431 hardening changed default seccomp/LSM handling for AF_ALG/socketcall; Docker explicitly warns against disabling seccomp as a workaround. Actual host LSM, rootless/userns, kernel, daemon, and Desktop/VM state must therefore be inspected rather than assumed.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.