Chapter 20Lesson 01~125 minutes

Volumes, Named Volumes, Volume Drivers, Backup/Restore, Sharing, and Persistent Data Patterns: Concepts, Architecture, and Mental Model

Model Docker volumes as persistent storage objects with an independent lifecycle, explicit driver/ownership semantics, and evidence-backed backup, restore, retention, and deletion decisions.

VolumesPersistenceDriversLifecycleData identity

Learning objectives

  • Explain why a Docker volume is an independently managed storage object rather than part of a container writable layer.
  • Distinguish named and anonymous volumes, driver/options, mount target, consuming containers, ownership/permissions, and retention intent.
  • Inspect volume metadata and container mount state without changing data.
  • Explain default volume population, volume-nocopy, read-only mounts, and why driver mountpoints are implementation evidence rather than portable application paths.
  • Connect persistent-data identity, backup artifacts, checksums, and restore verification to reproducible DevOps operations.
Chapter 20 principle. A container is replaceable compute; persistent data is a separate lifecycle. Removing, rebuilding, or upgrading a container must never be used as a proxy for deciding what happens to its data.

1. The practical problem: container lifecycle is not data lifecycle

Earlier chapters deliberately treated containers as disposable objects. That model becomes dangerous if application state is left in the writable layer: deleting the container deletes that state. A volume moves the data lifecycle into a Docker-managed storage object that can remain after one container exits and be attached to a replacement container later.

This does not make the data magically safe. You still need an owner, a driver, permission rules, consistency assumptions, backup artifacts, restore tests, and an explicit retention/deletion decision. Persistence means the bytes can outlive a container; recoverability means you can prove a usable copy exists and restore it.

2. Mental model: object → storage → mount → writes → retained data

Start with the named volume object. Docker associates it with a driver and driver-specific options. When a container starts, the driver makes storage available and Docker mounts it at the declared container target. Application writes then land in that mounted storage rather than in the container writable layer. Stopping or replacing the container does not, by itself, delete the volume. Backup, restore, migration, or deletion are separate operations.

Persistent-data causality
flowchart TD
  V[Volume object name + driver + options] --> S[Driver-provided storage local or remote]
  S --> M[Container mount target /data]
  M --> A[Application reads/writes]
  A --> R[Retained data independent of container]
  R --> B[Backup artifact checksum + timestamp]
  B --> X[Restore target new volume]
  X --> Q[Verification content + ownership + app semantics]
  R --> D[Intentional deletion after dependency check]
            

3. State you must identify before changing anything

State Evidence Why it matters
Volume name and labels docker volume ls, docker volume inspect Human identity and ownership boundary.
Driver and options .Driver, .Options Determines storage semantics; plugin/remote behavior is not interchangeable.
Mountpoint .Mountpoint where meaningful Host/driver implementation evidence; not a portable application contract.
Consumers container .Mounts, label-filtered docker ps -a Prevents deleting storage still depended on by a container.
Permissions numeric UID/GID/mode from a helper container Container names do not solve filesystem authorization.
Data version application marker/schema/version file Tells you which logical state a backup or restore represents.
Backup identity filename, timestamp, checksum, source volume Makes a backup independently auditable.
Restore target new volume name + verification result Separates recovery proof from the original storage object.
Retention decision documented owner/RPO/expiry Prevents “unused” from being confused with “safe to delete.”

4. Named versus anonymous volumes

A named volume has an explicit stable name such as da20-data. An anonymous volume is created without an explicit name, receives a generated identifier, and is harder to reconnect intentionally after container replacement. Both are separate volume objects. For state that must survive application updates, a named volume makes ownership and recovery intent far easier to express.

Container removal flags can clean anonymous volumes associated with a container, while named volumes survive ordinary container deletion. That asymmetry is one reason to avoid treating docker rm -v or Compose teardown as a casual storage-cleanup mechanism.

5. Mount behavior and the copy-up trap

Mounting a volume over a non-empty image directory obscures the underlying image files while the mount is active. When an empty volume is mounted at a path that already contains files in the container image, Docker normally copies that existing content into the volume. The volume-nocopy option disables that population behavior. Neither behavior should be discovered accidentally in production.

Read-only volume mounts are useful when one container should consume but not mutate shared data. The mount access mode is separate from the filesystem ownership/mode stored inside the volume.

6. Read-only inspection first

docker version
docker info
docker context show
docker compose version || true
docker buildx version || true

docker volume ls
docker ps -a --no-trunc

docker pull busybox:1.36.1
docker image inspect busybox:1.36.1 --format 'ID={{.Id}} RepoDigests={{json .RepoDigests}}'
# Example if a named volume already exists:
docker volume inspect da20-data 2>/dev/null || true

docker ps -a --filter volume=da20-data --no-trunc

docker ps -a --filter volume=da20-data -q | while read id; do
  [ -n "$id" ] && docker inspect "$id" --format 'Container={{.Name}} Mounts={{json .Mounts}}'
done

Do not enter Docker's internal storage directories and edit files there. The daemon and selected driver own that implementation. Use containers and documented volume APIs as the data-access boundary.

7. A backup is an artifact, not a belief

A useful backup record answers: which volume, which logical data version, when, under what consistency method, which archive checksum, and was restore actually tested? A tar file copied somewhere is not sufficient evidence if the source application was mutating files while the archive was created or if restored ownership makes the application unable to read them.

For ordinary files, a stopped/quiesced writer plus an archive and checksum can be enough. Databases often need their own snapshot, dump, checkpoint, replication, or backup API to achieve application consistency. Docker volume mechanics do not replace database-specific consistency guarantees.

8. DevOps connection: persistence is an operational contract

In a reproducible operating model, image digest, container configuration, volume identity, data version, ownership, backup checksum, restore target, and external storage assumptions are independently verifiable. This separation enables safe rollbacks: you can replace compute while retaining data, or restore data into a new volume without mutating the original evidence.

Next lesson

Next: Volumes, Named Volumes, Volume Drivers, Backup/Restore, Sharing, and Persistent Data Patterns: Guided Hands-On Workflow and Core Operations

Create a disposable volume, persist data across container replacement, produce a checksum-backed archive, restore it into a fresh volume, and delete only exact lab resources.

Knowledge check

If a container using a named volume is removed, is the named volume automatically deleted?

Why is a volume mountpoint not a portable application path?

What does volume-nocopy change?

Does persistence prove recoverability?

Why should volume ownership be captured as numeric UID/GID?

Official references and version notes

Version baseline, verified 2026-09-21.

Docker Engine 29.8.1 is current; Docker Compose 5.5.1, Buildx 0.37.1, and BuildKit 0.33.0 are the current upstream baselines used for compatibility discussion. The mandatory labs use the built-in local volume driver and BusyBox, require no paid service, and record the learner's actual installed versions rather than assuming they match upstream.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.