Chapter 38Lesson 04~205 minutes

Production Container Patterns, Immutable Delivery, Configuration, Statelessness, Sidecars, and Operational Contracts: Diagnostics, Failure Modes, Security, and Performance

Diagnose mutable-container drift, writable-layer data loss, secret leakage, unbounded logs/resources, misleading health, brittle companion coupling, and replacement/recovery failures without erasing first-failure evidence.

DiagnosticsDriftData durabilityHealthRecovery

Learning objectives

  • Diagnose production-container failures by preserving immutable identity, runtime state, storage, security, network, health, resource, log, and external evidence before repair.
  • Recognize runtime mutation and writable-layer state as hidden drift rather than durable fixes.
  • Differentiate a healthcheck failure from an externally unreachable or dependency-broken application.
  • Diagnose resource/log pressure and companion-service coupling without disabling security controls or using broad cleanup.
  • Repair an intentionally broken read-only filesystem example by declaring the correct writable state boundary.

1. Evidence-first incident sequence

Before replacing or restarting anything, record the context/host, exact container ID, configured image reference and image ID, health history, restart count, exit/OOM state, mounts, security settings, resources, logs/events, network publication, volume identity, and the external request result. Then change the smallest layer that explains the failure.

2. Failure: SSH/exec patching becomes the “fix”

An operator edits a Python file inside the running container and service returns to normal. The immediate symptom is gone, but the release artifact is unchanged and the patch will disappear on replacement. Preserve the diff/evidence, reproduce the correction in source, rebuild a new digest, test it, and replace the container. Use the running mutation only as incident evidence, not as release state.

3. Failure: database or business state in writable layer

Container recreation appears to “randomly” delete data. Inspect .Mounts and application paths. If the data directory is absent from declared persistent storage, the behavior is deterministic: it lived in the writable layer. Fix the data architecture and restore from valid backup; do not copy random Engine storage directories by hand.

4. Failure: secret baked into image or environment

Finding a credential in image history, image config, environment inspection, logs, or source is a credential incident. Preserve evidence, revoke/rotate the secret, remove it from future build/runtime paths, and verify the replacement digest/config. Merely deleting a file in a later image layer does not erase it from prior layers.

5. Failure: unbounded logs or resources

Disk exhaustion and host OOM can masquerade as application instability. Inspect logging driver/options, Docker disk use, container resource limits, docker stats, OOM state, and host filesystem capacity. Apply bounded per-service logging and measured resource limits; do not start with global prune or arbitrary restarts.

6. Failure: healthcheck treated as SLA proof

A container is healthy, yet users cannot connect. Preserve the health history and then inspect the external path: bound interface, published host port, firewall/load balancer, DNS/TLS, and upstream dependencies. Healthy proves only the configured probe path succeeded.

7. Failure: restart loop hides the root cause

A restart policy can make a failed process repeatedly disappear and reappear. Capture the first exit code, logs, events, restart count, and OOM status before tuning the policy. If the application exits because configuration is invalid, more restarts add noise rather than resilience.

8. Failure: companion lifecycle assumption

A companion expects the primary app to exist continuously and exits permanently during an ordinary app replacement. The main app returns healthy but the observer/forwarder remains absent. A robust companion should retry the app endpoint and have its own restart/health contract, or the function should move to infrastructure whose lifecycle matches the requirement.

9. Intentionally broken example: undeclared write on read-only rootfs

# BROKEN LAB EXAMPLE — safe local demonstration
services:
  app:
    image: python:3.13-alpine
    user: "10001:10001"
    read_only: true
    command:
      - python
      - -c
      - |
        from pathlib import Path
        Path('/app/state.txt').write_text('business-state')
Expected first failure (exact wording varies):
OSError: [Errno 30] Read-only file system: '/app/state.txt'

The correct diagnosis is not “read-only is broken.” The application attempted a persistent write to an undeclared immutable path. Decide whether the data is ephemeral (tmpfs), durable (named volume/external database), or should not be written at runtime. Keep the root filesystem read-only and declare only the required writable path.

10. Safe repair of the broken example

services:
  app:
    image: python:3.13-alpine
    user: "10001:10001"
    read_only: true
    tmpfs:
      - /tmp:rw,noexec,nosuid,size=8m
    volumes:
      - state:/data
    command:
      - python
      - -c
      - |
        from pathlib import Path
        Path('/data/state.txt').write_text('synthetic-state')
volumes:
  state: {}

In a complete application, ensure the mounted data path has correct numeric ownership for the runtime user and add backup/retention rules. Do not weaken the whole filesystem to make one path writable.

11. Layer-by-layer diagnostic matrix

Symptom Likely layer Evidence before correction
different code than release image/runtime drift configured image ref, image ID, file checksum
data missing after recreate storage lifecycle mounts, volume ID, backup record
permission denied on intended data identity/storage/security UID/GID, mount owner, LSM/security opts
healthy but unreachable network/external path health history, published port, request/DNS/TLS evidence
exit 137 / OOM resource/kernel OOMKilled, memory limit, stats, host pressure
disk fills quickly logging/storage log driver/options, Docker df, host free space/inodes
companion absent after app change service lifecycle container IDs, restart policy, logs/events

12. Smallest-safe-correction rule

Fix the layer that violated the contract: new digest for code, config revision for configuration, ownership/mount for data, resource envelope for pressure, logging policy for retention, network rule for reachability, or companion retry/lifecycle for coupling. Do not change daemon security, host filesystem permissions, or unrelated resources to compensate for an application-level defect.

13. Preserve recovery evidence

After the correction, retain both before and after identities: old/new digest, old/new container ID, volume identity, health/restart state, external request, state checksum, logs/events, and the specific change that explains recovery. Recovery is not “it works now”; it is a causal, reproducible state transition.

Knowledge check

A container is healthy but users cannot connect. What should you do next?

Why is an exec-based production patch not an acceptable final correction?

A read-only rootfs causes /app/state.txt writes to fail. What is the right question?

What evidence distinguishes resource failure from application exception?

Why capture the old container ID and digest before replacement?

Next lesson

Next: Checkpoint Lab — Production Container Patterns, Immutable Delivery, Configuration, Statelessness, Sidecars, and Operational Contracts

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Diagnostic baseline: 2026-09-22. The lesson deliberately avoids broad cleanup, disabling security controls, socket mounting, privileged mode, or daemon-wide mutations as troubleshooting shortcuts.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.