Checkpoint Lab — Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support
Run an evidence-driven Swarm checkpoint: deploy a stack, bind image digests, attach a secret and overlay, force a controlled failed update, roll back, and document the estate decision.
Learning objectives
- Deploy a disposable Swarm stack whose service, secret, overlay network, and image digest are independently verifiable.
- Record source/tool/component assumptions before mutation and predict state transitions before executing them.
- Roll from digest A to B with a controlled failure, preserve failed-task evidence, and roll back to A.
- Produce an evidence packet that links swarm, service, task, image, secret, network, update, and cleanup identities.
- Write a legacy-support/migration note that distinguishes current Swarm support from platform-fit requirements.
1. Assumptions and preflight packet
date -u +%Y-%m-%dT%H:%M:%SZ
docker version
docker compose version
docker buildx version
docker context show
docker info --format 'Swarm={{.Swarm.LocalNodeState}} Driver={{.Driver}} Cgroup={{.CgroupVersion}}'
docker ps --format 'table {{.ID}} {{.Names}} {{.Image}}'
containerd --version 2>/dev/null || true
runc --version 2>/dev/null | head -1 || true
Record output rather than assuming package versions. The mutation precondition is an authorized disposable context with Swarm inactive.
2. Predict before execution
| Prediction | State layer | How you will verify |
|---|---|---|
swarm init creates one manager and cluster ID
|
daemon/control plane | docker info + node ls |
| stack deploy creates one service and overlay network | service/network | stack services + network inspect |
| secret value is granted to task but not placed in environment | security/secret | service inspect grant + task success; never print value |
| failed B update creates failed task generation while preserving service history | scheduler/update | service ps --no-trunc + UpdateStatus |
| rollback returns service image to digest A | desired state/image | service inspect image + converged task |
3. Initialize the disposable swarm and capture identity
docker swarm init
SWARM_ID=$(docker info --format '{{.Swarm.Cluster.ID}}')
NODE_ID=$(docker info --format '{{.Swarm.NodeID}}')
printf 'swarm=%s node=%s
' "$SWARM_ID" "$NODE_ID"
docker node ls
Do not retrieve or record join tokens; this single-node checkpoint does not need them.
4. Resolve digest A and B
docker pull alpine:3.21
docker pull alpine:3.22
IMAGE_A=$(docker image inspect alpine:3.21 --format '{{index .RepoDigests 0}}')
IMAGE_B=$(docker image inspect alpine:3.22 --format '{{index .RepoDigests 0}}')
printf 'IMAGE_A=%s
IMAGE_B=%s
' "$IMAGE_A" "$IMAGE_B"
If those minor tags age out, substitute two available current Alpine minor tags and record their exact digests. The checkpoint requirement is two immutable subjects, not those human tag strings.
5. Create the fake secret without retaining plaintext
printf 'chapter39-training-secret
' > ch39-secret.txt
docker secret create --label academy.chapter=39 ch39_checkpoint_secret ch39-secret.txt
rm -f ch39-secret.txt
docker secret inspect ch39_checkpoint_secret --format 'ID={{.ID}} Name={{.Spec.Name}} Labels={{json .Spec.Labels}}'
6. Create the stack artifact
cat > ch39-checkpoint.yml <<EOF
version: "3.8"
services:
app:
image: ${IMAGE_A}
command: ["sh", "-c", "test -s /run/secrets/ch39_checkpoint_secret && while true; do sleep 15; done"]
secrets: [ch39_checkpoint_secret]
networks: [private]
deploy:
replicas: 1
labels:
academy.chapter: "39"
restart_policy:
condition: on-failure
update_config:
parallelism: 1
delay: 2s
failure_action: pause
rollback_config:
parallelism: 1
delay: 2s
secrets:
ch39_checkpoint_secret:
external: true
networks:
private:
driver: overlay
EOF
sed -n '1,220p' ch39-checkpoint.yml
The file contains no real secret. It references the secret object by name. The service image is digest A.
7. Deploy and verify the A generation
docker stack deploy -c ch39-checkpoint.yml ch39cp
docker stack services ch39cp
docker stack ps --no-trunc ch39cp
SERVICE=ch39cp_app
docker service inspect "$SERVICE" --format 'ID={{.ID}} Version={{.Version.Index}} Image={{.Spec.TaskTemplate.ContainerSpec.Image}}'
docker service inspect "$SERVICE" --format '{{json .Spec.TaskTemplate.ContainerSpec.Secrets}}'
docker network inspect ch39cp_private --format 'ID={{.Id}} Driver={{.Driver}} Scope={{.Scope}}'
8. Controlled B failure: change image and command together
docker service update \
--image "$IMAGE_B" \
--args sh -c 'exit 42' \
--update-parallelism 1 \
--update-delay 2s \
--update-failure-action pause "$SERVICE" || true
# Preserve first-failure evidence before rollback
docker service inspect "$SERVICE" --format '{{json .UpdateStatus}}'
docker service ps --no-trunc "$SERVICE"
docker service inspect "$SERVICE" --format 'CurrentImage={{.Spec.TaskTemplate.ContainerSpec.Image}} Version={{.Version.Index}}'
The expected failure is synthetic: digest B starts with a command that exits 42. The point is to preserve the failed task generation and update status before any correction.
9. Roll back and verify restoration to A
docker service rollback "$SERVICE"
docker service ps --no-trunc "$SERVICE"
docker service inspect "$SERVICE" --format '{{json .UpdateStatus}}'
docker service inspect "$SERVICE" --format 'RestoredImage={{.Spec.TaskTemplate.ContainerSpec.Image}}'
Verify that the restored service image matches
$IMAGE_A and that a task reaches Running. Rollback
command success alone is not the acceptance criterion.
10. Evidence packet
printf 'swarm=%s
node=%s
A=%s
B=%s
' "$SWARM_ID" "$NODE_ID" "$IMAGE_A" "$IMAGE_B"
docker node ls
docker service inspect "$SERVICE" --pretty
docker service ps --no-trunc "$SERVICE"
docker secret inspect ch39_checkpoint_secret --format 'ID={{.ID}} Name={{.Spec.Name}}'
docker network inspect ch39cp_private --format 'ID={{.Id}} Driver={{.Driver}} Scope={{.Scope}}'
docker events --since 15m --until 0s --filter type=service --filter type=secret --filter type=network 2>/dev/null || true
Keep: timestamp, Docker/context versions, swarm/node IDs, service ID/spec version, task history, A/B digests, secret object ID without value, overlay network ID, UpdateStatus/rollback result, and the explicit limitation that a one-node swarm does not test HA or multi-host overlay behavior.
11. Legacy-support / migration decision note
Current status: Swarm mode is present in current Docker Engine and documented for production use.
Estate being evaluated: <single-node checkpoint / real estate inventory>
Requirements met today: <list>
Requirements not met / externalized: <list>
Quorum design: <N managers, zones, tolerance>
Release identity: <digest-bound yes/no>
Stack compatibility risks: <fields/extensions>
Secrets/config rotation: <procedure>
Operational evidence quality: <task history, backups, alerts>
Reason to keep: <requirements/evidence>
Reason to migrate: <requirements/evidence>
Target platform prerequisites: <if applicable>
Migration rehearsal / rollback path: <if applicable>
Decision owner and review date: <name/date>
12. Verification checklist before cleanup
- the stack service initially ran digest A;
- the secret grant and overlay network were independently inspected;
- digest B produced a controlled failed task generation;
- first-failure task/update evidence was preserved before rollback;
- rollback restored digest A and a running task;
- the migration/support note distinguishes current support from platform fit;
- no join token or secret plaintext entered the evidence packet.
13. Bounded cleanup and rollback
docker stack rm ch39cp
# Wait until the service is gone before removing the external secret
docker service ls --filter label=com.docker.stack.namespace=ch39cp
docker secret rm ch39_checkpoint_secret
rm -f ch39-checkpoint.yml
# Confirm this is still the one-node disposable lab before destroying membership
docker node ls
docker swarm leave --force
docker info --format '{{.Swarm.LocalNodeState}}'
--force is
guarded by the checkpoint’s explicit single-manager disposable
precondition. In a real estate, leaving/removing a manager is a
quorum change and requires a maintenance/recovery plan.
14. What Chapter 39 adds to the production operating model
You can now distinguish Docker Swarm’s desired service state from node-local containers, preserve Raft/quorum evidence, bind distributed updates to immutable digests, operate overlay/secrets/stack objects safely, and support an existing Swarm without pretending its platform decision is automatic. Chapter 40 turns to Docker performance engineering: build speed, image size, runtime overhead, network, storage, and resource profiling.
Knowledge check
Why did the checkpoint intentionally fail digest B before rollback?
To create a controlled task-generation failure so the learner can preserve UpdateStatus and task history, then prove rollback restores the previous desired service specification.
What evidence proves the secret existed without exposing it?
Secret object ID/name/labels and the service grant. The plaintext value is neither required nor appropriate evidence.
What does the one-node checkpoint not prove?
Manager high availability, cross-node scheduling during failure, or multi-host overlay/routing behavior.
If rollback returns zero but the service has no Running task, did the checkpoint pass?
No. The acceptance criterion includes observed service convergence and restored digest A, not just CLI exit status.
What should a migration note say about Swarm support?
State the current fact: Swarm mode remains built into current Docker Engine and documented for production use; then separately evaluate whether it fits the estate’s requirements.
Official references and version notes
2026-09-22. Course baseline: Docker Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, BuildKit 0.33.0. The executable labs record the learner’s actual installed component versions. Swarm mode remains built into current Docker Engine and current Docker documentation explicitly describes it as a production runtime option; this chapter uses “legacy estate” to mean an existing platform that must be operated or evaluated deliberately, not that Swarm mode is removed.
- Docker Docs — Swarm mode — current support position and core feature set.
- Docker Docs — How services work — desired state, services, tasks, replicas, constraints, and update behavior.
- Docker Docs — How nodes work — managers, workers, scheduling, and manager-count guidance.
- Docker Docs — Administer and maintain a swarm — quorum, manager distribution, backup, and disaster recovery.
- Docker Docs — Raft consensus — replicated manager state and majority requirements.
- Docker Docs — Manage swarm service networks — overlay, ingress, routing mesh, and control/data-plane traffic.
- Docker Docs — Manage sensitive data with Docker secrets — encrypted Raft storage and in-memory task mounts.
- Docker Docs — Docker configs — immutable configs and service/stack lifecycle.
- Docker Docs — Apply rolling updates to a service — update delay, parallelism, and task replacement.
- Docker Docs — Rolling update tutorial — observing task transitions during updates.
- Docker Docs — Deploy a stack to a swarm — manager-only stack deployment and Compose-file compatibility warning.
- Docker CLI — docker stack deploy — current flags including image digest resolution and registry auth propagation.
- Docker CLI — docker service update — rolling-update, rollback, image, secret, config, and publish controls.
- Docker Engine 29 release notes — current Engine-era compatibility baseline.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.