Chapter 20Lesson 04~135 minutes

Volumes, Named Volumes, Volume Drivers, Backup/Restore, Sharing, and Persistent Data Patterns: Diagnostics, Failure Modes, Security, and Performance

Diagnose volume data loss, permission failures, unsafe backup assumptions, driver mismatches, and destructive cleanup by preserving storage evidence before correction.

DiagnosticsPermissionsData lossPrune safetyFailure evidence

Learning objectives

  • Preserve volume, container, ownership, backup, and driver evidence before attempting repair.
  • Diagnose accidental data deletion risk from Compose teardown and broad cleanup commands.
  • Reproduce and repair a numeric UID/GID permission mismatch without hiding the original cause.
  • Explain why live database file copying may be inconsistent and choose an application-supported backup boundary.
  • Separate Docker local-driver behavior from plugin/remote-driver behavior and avoid unsupported assumptions.
Diagnostic rule. Do not “fix storage” by deleting it. First preserve volume inspect output, consuming-container mount records, numeric ownership/mode, application logs, backup checksum/metadata, driver identity, and the exact command or lifecycle event that produced the failure.

1. Evidence-first storage diagnostic sequence

  1. Record host/platform and docker version, docker info, active context, Compose version.
  2. Capture the exact image/container IDs and application logs before restart or replacement.
  3. Capture docker volume inspect, driver/options, labels, and consuming-container mounts.
  4. Inspect numeric UID/GID and modes from inside a safe helper/container.
  5. Identify whether the failed path is the writable layer, named volume, bind mount, tmpfs, or driver/plugin storage.
  6. Preserve backup artifact checksum, creation time, source data version, and consistency method.
  7. For remote/plugin storage, capture provider/driver health without exposing credentials.
  8. Apply the smallest correction and rerun only the affected verification.

2. Failure mode: confusing Compose teardown with data cleanup

Ordinary docker compose down removes project containers and networks but preserves named volumes. The --volumes option changes the data lifecycle and removes project-declared named volumes plus attached anonymous volumes. That is a destructive storage decision, not a normal application stop.

Before any teardown in a stateful project, render the normalized model, list project-labeled volumes, record whether each volume is project-managed or external, and decide retention explicitly.

docker compose config
# Read-only inventory example for a disposable project:
docker volume ls --filter label=com.docker.compose.project=myproject

# Safe default teardown preserves named volumes:
docker compose down

# Do not add the volume-deletion option unless data destruction is the explicit, verified intent.

3. Failure mode: blind volume pruning

An “unused” volume means Docker sees no current container reference. It does not mean the data has no business owner, retention requirement, backup dependency, or planned future reuse. Broad prune commands are therefore excluded from this academy's routine troubleshooting.

Prefer label-based inventory, exact dependency checks, and exact docker volume rm NAME only after backup/retention decisions are proven. Docker's refusal to remove an in-use volume is a helpful guard, but it cannot protect an unreferenced volume whose data still matters.

4. Intentionally broken example: restored data has the wrong owner

Create a disposable volume containing a file that only root may read. Then run a consumer as UID/GID 1001. Preserve the error and ownership evidence before fixing it.

docker volume create   --label devops-academy.lab=ch20-diag   da20-perm

docker run --rm --mount type=volume,src=da20-perm,dst=/data   busybox:1.36.1 sh -c '
    echo protected-state > /data/state.txt
    chmod 600 /data/state.txt
    chown 0:0 /data/state.txt
    ls -ln /data/state.txt
  '

# Preserve the first failure:
docker run --rm --user 1001:1001   --mount type=volume,src=da20-perm,dst=/data,readonly   busybox:1.36.1 cat /data/state.txt || true

# Evidence, still without changing the file:
docker run --rm --mount type=volume,src=da20-perm,dst=/data,readonly   busybox:1.36.1 ls -ln /data/state.txt

docker volume inspect da20-perm

The failure is not a network or image-pull problem: the process UID does not match the file authorization. In this synthetic lab, correct the intended owner explicitly, then rerun the same consumer.

docker run --rm --mount type=volume,src=da20-perm,dst=/data   busybox:1.36.1 chown 1001:1001 /data/state.txt

docker run --rm --user 1001:1001   --mount type=volume,src=da20-perm,dst=/data,readonly   busybox:1.36.1 cat /data/state.txt

docker rm -f da20-perm-consumer 2>/dev/null || true
docker volume rm da20-perm

In production, the correct ownership may instead be established at image startup, provisioning, restore tooling, or application initialization. Do not apply recursive ownership changes blindly to unknown data.

5. Failure mode: copying a live database directory

A tar command can read files while an application is modifying them. The resulting archive may contain a combination of moments that the database cannot recover as a coherent transaction state. A successful tar exit code only proves the archive command completed.

Use the database's documented dump/snapshot/checkpoint/backup mechanism, or deliberately quiesce/stop writes when that is the supported consistency method. Record the application checkpoint/LSN/transaction marker or backup job ID alongside the volume snapshot/archive.

6. Failure mode: assuming every driver behaves like local storage

Remote/plugin drivers may have their own mount lifecycle, snapshot capability, access modes, network requirements, credentials, capacity constraints, and failure recovery. A local-driver Mountpoint path is not a universal concept. Always inspect Driver/Options first and consult the selected driver's authoritative documentation.

7. Storage performance without superstition

Measure the application workload, filesystem/backend latency, host resource pressure, and driver path before blaming Docker. A remote driver adds network/storage latency; Docker Desktop introduces a managed VM/storage path; synchronous database durability can be intentionally slower than buffered writes. Performance tuning that weakens durability is a data-loss tradeoff and must be explicit.

Capture baseline operation latency, IOPS/throughput where available, host disk space, application fsync/checkpoint behavior, and backup duration. Optimize the measured bottleneck without disabling durability guarantees casually.

8. Failure matrix

Symptom Likely first layer Preserve Least-destructive correction
Replacement sees empty data Wrong volume name/project/mount volume inspect + container Mounts + Compose config Attach intended exact volume; do not recreate/delete first.
Permission denied Filesystem ownership/mode or user mapping numeric UID/GID/mode + process UID/GID Correct declared ownership/access in disposable scope.
Backup restores but app rejects it Consistency/schema/application state backup metadata + app logs/version Use supported app-consistent backup/restore path.
Volume missing after teardown Destructive lifecycle option/cleanup shell/CI logs + Compose model + backup evidence Restore from verified backup; fix teardown policy.
Remote volume intermittently unavailable Driver/provider/network boundary driver name/options + plugin/provider health Fix external dependency; avoid local-driver assumptions.
Disk pressure from old volumes Retention/inventory problem labels, age, ownership, backups, consumers Delete exact reviewed objects; do not blindly prune.

9. Diagnostic cleanup rule

Remove only disposable diagnostic objects with exact names after verification. Never delete unknown Docker storage directories manually, never replace ownership analysis with indiscriminate world-writable permissions, and never force a volume removal merely to clear an error.

Next lesson

Next: Checkpoint Lab — Volumes, Named Volumes, Volume Drivers, Backup/Restore, Sharing, and Persistent Data Patterns

Run the full persistence-and-recovery checkpoint: replacement, archive, checksum, restore, independent verification, and an explicit RPO/consistency statement.

Knowledge check

Why is “unused volume” not equivalent to “safe to delete”?

A restored file is owned by UID 0 but the app runs as UID 1001. What evidence matters first?

Why can a tar archive of a running database be unusable even if tar exits successfully?

What should you inspect before assuming a volume behaves like local storage?

What makes down --volumes qualitatively different from ordinary down?

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is current; Docker Compose 5.5.1, Buildx 0.37.1, and BuildKit 0.33.0 are the current upstream baselines used for compatibility discussion. The mandatory labs use the built-in local volume driver and BusyBox, require no paid service, and record the learner's actual installed versions rather than assuming they match upstream.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.