Chapter 39Lesson 05~240 minutes

Checkpoint Lab — Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support

Run an evidence-driven Swarm checkpoint: deploy a stack, bind image digests, attach a secret and overlay, force a controlled failed update, roll back, and document the estate decision.

CheckpointDigest A→BRollbackSecretMigration note

Learning objectives

  • Deploy a disposable Swarm stack whose service, secret, overlay network, and image digest are independently verifiable.
  • Record source/tool/component assumptions before mutation and predict state transitions before executing them.
  • Roll from digest A to B with a controlled failure, preserve failed-task evidence, and roll back to A.
  • Produce an evidence packet that links swarm, service, task, image, secret, network, update, and cleanup identities.
  • Write a legacy-support/migration note that distinguishes current Swarm support from platform-fit requirements.
Safety boundary. This checkpoint mutates Swarm membership and service desired state. Use only a disposable Engine/VM with no production workloads. If the current context already belongs to any swarm you did not create specifically for this checkpoint, do not proceed with the mutation path.

1. Assumptions and preflight packet

date -u +%Y-%m-%dT%H:%M:%SZ
docker version
docker compose version
docker buildx version
docker context show
docker info --format 'Swarm={{.Swarm.LocalNodeState}} Driver={{.Driver}} Cgroup={{.CgroupVersion}}'
docker ps --format 'table {{.ID}}	{{.Names}}	{{.Image}}'
containerd --version 2>/dev/null || true
runc --version 2>/dev/null | head -1 || true

Record output rather than assuming package versions. The mutation precondition is an authorized disposable context with Swarm inactive.

2. Predict before execution

Prediction State layer How you will verify
swarm init creates one manager and cluster ID daemon/control plane docker info + node ls
stack deploy creates one service and overlay network service/network stack services + network inspect
secret value is granted to task but not placed in environment security/secret service inspect grant + task success; never print value
failed B update creates failed task generation while preserving service history scheduler/update service ps --no-trunc + UpdateStatus
rollback returns service image to digest A desired state/image service inspect image + converged task

3. Initialize the disposable swarm and capture identity

docker swarm init
SWARM_ID=$(docker info --format '{{.Swarm.Cluster.ID}}')
NODE_ID=$(docker info --format '{{.Swarm.NodeID}}')
printf 'swarm=%s node=%s
' "$SWARM_ID" "$NODE_ID"
docker node ls

Do not retrieve or record join tokens; this single-node checkpoint does not need them.

4. Resolve digest A and B

docker pull alpine:3.21
docker pull alpine:3.22
IMAGE_A=$(docker image inspect alpine:3.21 --format '{{index .RepoDigests 0}}')
IMAGE_B=$(docker image inspect alpine:3.22 --format '{{index .RepoDigests 0}}')
printf 'IMAGE_A=%s
IMAGE_B=%s
' "$IMAGE_A" "$IMAGE_B"

If those minor tags age out, substitute two available current Alpine minor tags and record their exact digests. The checkpoint requirement is two immutable subjects, not those human tag strings.

5. Create the fake secret without retaining plaintext

printf 'chapter39-training-secret
' > ch39-secret.txt
docker secret create --label academy.chapter=39 ch39_checkpoint_secret ch39-secret.txt
rm -f ch39-secret.txt
docker secret inspect ch39_checkpoint_secret --format 'ID={{.ID}} Name={{.Spec.Name}} Labels={{json .Spec.Labels}}' 

6. Create the stack artifact

cat > ch39-checkpoint.yml <<EOF
version: "3.8"
services:
  app:
    image: ${IMAGE_A}
    command: ["sh", "-c", "test -s /run/secrets/ch39_checkpoint_secret && while true; do sleep 15; done"]
    secrets: [ch39_checkpoint_secret]
    networks: [private]
    deploy:
      replicas: 1
      labels:
        academy.chapter: "39"
      restart_policy:
        condition: on-failure
      update_config:
        parallelism: 1
        delay: 2s
        failure_action: pause
      rollback_config:
        parallelism: 1
        delay: 2s
secrets:
  ch39_checkpoint_secret:
    external: true
networks:
  private:
    driver: overlay
EOF

sed -n '1,220p' ch39-checkpoint.yml

The file contains no real secret. It references the secret object by name. The service image is digest A.

7. Deploy and verify the A generation

docker stack deploy -c ch39-checkpoint.yml ch39cp
docker stack services ch39cp
docker stack ps --no-trunc ch39cp
SERVICE=ch39cp_app
docker service inspect "$SERVICE" --format 'ID={{.ID}} Version={{.Version.Index}} Image={{.Spec.TaskTemplate.ContainerSpec.Image}}'
docker service inspect "$SERVICE" --format '{{json .Spec.TaskTemplate.ContainerSpec.Secrets}}'
docker network inspect ch39cp_private --format 'ID={{.Id}} Driver={{.Driver}} Scope={{.Scope}}' 

8. Controlled B failure: change image and command together

docker service update \
  --image "$IMAGE_B" \
  --args sh -c 'exit 42' \
  --update-parallelism 1 \
  --update-delay 2s \
  --update-failure-action pause   "$SERVICE" || true

# Preserve first-failure evidence before rollback
docker service inspect "$SERVICE" --format '{{json .UpdateStatus}}'
docker service ps --no-trunc "$SERVICE"
docker service inspect "$SERVICE" --format 'CurrentImage={{.Spec.TaskTemplate.ContainerSpec.Image}} Version={{.Version.Index}}' 

The expected failure is synthetic: digest B starts with a command that exits 42. The point is to preserve the failed task generation and update status before any correction.

9. Roll back and verify restoration to A

docker service rollback "$SERVICE"
docker service ps --no-trunc "$SERVICE"
docker service inspect "$SERVICE" --format '{{json .UpdateStatus}}'
docker service inspect "$SERVICE" --format 'RestoredImage={{.Spec.TaskTemplate.ContainerSpec.Image}}' 

Verify that the restored service image matches $IMAGE_A and that a task reaches Running. Rollback command success alone is not the acceptance criterion.

10. Evidence packet

printf 'swarm=%s
node=%s
A=%s
B=%s
' "$SWARM_ID" "$NODE_ID" "$IMAGE_A" "$IMAGE_B"
docker node ls
docker service inspect "$SERVICE" --pretty
docker service ps --no-trunc "$SERVICE"
docker secret inspect ch39_checkpoint_secret --format 'ID={{.ID}} Name={{.Spec.Name}}'
docker network inspect ch39cp_private --format 'ID={{.Id}} Driver={{.Driver}} Scope={{.Scope}}'
docker events --since 15m --until 0s --filter type=service --filter type=secret --filter type=network 2>/dev/null || true

Keep: timestamp, Docker/context versions, swarm/node IDs, service ID/spec version, task history, A/B digests, secret object ID without value, overlay network ID, UpdateStatus/rollback result, and the explicit limitation that a one-node swarm does not test HA or multi-host overlay behavior.

11. Legacy-support / migration decision note

Current status: Swarm mode is present in current Docker Engine and documented for production use.
Estate being evaluated: <single-node checkpoint / real estate inventory>
Requirements met today: <list>
Requirements not met / externalized: <list>
Quorum design: <N managers, zones, tolerance>
Release identity: <digest-bound yes/no>
Stack compatibility risks: <fields/extensions>
Secrets/config rotation: <procedure>
Operational evidence quality: <task history, backups, alerts>
Reason to keep: <requirements/evidence>
Reason to migrate: <requirements/evidence>
Target platform prerequisites: <if applicable>
Migration rehearsal / rollback path: <if applicable>
Decision owner and review date: <name/date>

12. Verification checklist before cleanup

Checkpoint passes only if:
  • the stack service initially ran digest A;
  • the secret grant and overlay network were independently inspected;
  • digest B produced a controlled failed task generation;
  • first-failure task/update evidence was preserved before rollback;
  • rollback restored digest A and a running task;
  • the migration/support note distinguishes current support from platform fit;
  • no join token or secret plaintext entered the evidence packet.

13. Bounded cleanup and rollback

docker stack rm ch39cp
# Wait until the service is gone before removing the external secret
docker service ls --filter label=com.docker.stack.namespace=ch39cp
docker secret rm ch39_checkpoint_secret
rm -f ch39-checkpoint.yml

# Confirm this is still the one-node disposable lab before destroying membership
docker node ls
docker swarm leave --force
docker info --format '{{.Swarm.LocalNodeState}}' 
Safety boundary. The final --force is guarded by the checkpoint’s explicit single-manager disposable precondition. In a real estate, leaving/removing a manager is a quorum change and requires a maintenance/recovery plan.

14. What Chapter 39 adds to the production operating model

You can now distinguish Docker Swarm’s desired service state from node-local containers, preserve Raft/quorum evidence, bind distributed updates to immutable digests, operate overlay/secrets/stack objects safely, and support an existing Swarm without pretending its platform decision is automatic. Chapter 40 turns to Docker performance engineering: build speed, image size, runtime overhead, network, storage, and resource profiling.

Knowledge check

Why did the checkpoint intentionally fail digest B before rollback?

What evidence proves the secret existed without exposing it?

What does the one-node checkpoint not prove?

If rollback returns zero but the service has no Running task, did the checkpoint pass?

What should a migration note say about Swarm support?

Next lesson

Next: Docker Performance Engineering: Build Speed, Image Size, Runtime Overhead, Network, Storage, and Resource Profiling: Concepts, Architecture, and Mental Model

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Course baseline: Docker Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, BuildKit 0.33.0. The executable labs record the learner’s actual installed component versions. Swarm mode remains built into current Docker Engine and current Docker documentation explicitly describes it as a production runtime option; this chapter uses “legacy estate” to mean an existing platform that must be operated or evaluated deliberately, not that Swarm mode is removed.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.