Chapter 39Lesson 01~185 minutes

Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support: Concepts, Architecture, and Mental Model

Understand Docker Swarm desired state, manager quorum, service/task identity, overlay networking, secrets, routing, rolling updates, and legacy-estate operating boundaries.

Swarm modeServices & tasksRaft quorumOverlay networksSecrets

Learning objectives

  • Explain the manager desired-state loop from service specification to tasks and containers.
  • Distinguish manager Raft state, node role, service identity, task identity, and standalone container state.
  • Explain overlay networks, ingress/routing mesh, secrets/configs, image digests, and update/rollback state as separate evidence domains.
  • Reason about quorum without treating an odd manager count as a substitute for failure-domain design.
  • Inspect an existing Swarm read-only before any membership, service, network, or secret change.

1. Why Swarm operations need a different mental model

A standalone container is primarily a daemon-local object. A Swarm service is a cluster desired-state object. The manager records what should exist, the scheduler creates tasks to satisfy that specification, and worker Engines materialize those tasks as containers. Troubleshooting therefore starts with the service and task graph—not with whichever container happens to be visible on one node.

Current Docker still ships Swarm mode inside Docker Engine. Docker’s own guidance says to use Swarm mode when Swarm is the intended production runtime; otherwise use Compose for non-Swarm deployments. “Legacy estate” in this chapter means an existing Swarm environment that deserves competent support and an explicit platform decision, not an unsupported or removed feature.

2. Mental model: desired state becomes tasks

Swarm control loop and runtime dataflow
flowchart TD
  A[Manager Raft state] --> B[Service spec + version]
  B --> C[Scheduler]
  C --> D[Task on node]
  D --> E[Container process]
  F[Overlay network] --> D
  G[Secret / config grants] --> D
  H[Image digest] --> D
  E --> I[Observed task state]
  I --> J[Manager reconciliation]
  J --> C
  B --> K[Rolling update / rollback policy]
  K --> C
            

The manager’s service specification is authoritative desired state. A task is a one-way scheduling unit tied to one service slot; when a task fails or is replaced, the orchestrator creates a new task rather than “repairing” the old task in place. The container is the node-local runtime realization of that task.

3. Evidence domains you must not collapse

Domain Identity or state Read-only evidence
Cluster Swarm ID, manager set, quorum docker info, docker node ls
Node Node ID, role, availability, manager reachability docker node inspect
Service Service ID, Spec.Version, desired replicas, update policy docker service inspect
Task Task ID, slot, node, desired/current state, error docker service ps --no-trunc
Image Resolved digest and platform service ContainerSpec.Image, registry metadata
Network Overlay/ingress identity and attachment docker network inspect
Secret/config Object ID and service grant, never plaintext evidence docker secret ls, service inspect
Update UpdateStatus, rollback config, task generations service inspect + task history

4. Managers, workers, and Raft quorum

Managers maintain replicated cluster state through Raft and also expose the Swarm management API. Workers execute tasks but do not participate in Raft. A majority of managers must remain available to change cluster state. With three managers the majority is two; with five it is three. Docker recommends an odd number for fault tolerance and generally a maximum of seven managers.

If quorum is lost, already-running tasks can continue, but the swarm cannot perform management operations or reschedule work. That distinction is operationally critical: “containers still answer” does not mean the control plane is healthy.

5. Services, tasks, and containers

A service is the declarative workload contract: image, command, replicas or global mode, networks, secrets, placement, resources, restart/update policy, and publications. A task is a scheduled attempt to realize one part of that service. Containers are created from tasks on worker Engines. Because failed tasks remain in task history, the task list is often more informative than docker ps.

6. Replicated and global services

Mode Desired state Typical fit Failure question
Replicated N interchangeable task slots API/web workers Can another eligible node run a replacement?
Global One task on every eligible node node agent / collector Which nodes match constraints and availability?

7. Overlay networking and ingress

User-defined overlay networks provide service-to-service connectivity across participating Docker daemons. Swarm also creates an ingress overlay used by the routing mesh for published services. Control-plane traffic is encrypted; application data-plane encryption is a separate design choice. A single-node lab can demonstrate object semantics, but it cannot prove cross-host dataplane behavior or availability.

8. Routing mesh versus host-mode publishing

With routing-mesh publication, a published port can be accepted on swarm nodes and routed to an active service task. Host-mode publication bypasses that routing behavior and exposes only where a task is running. Host mode can be useful for deterministic node-local integrations, but it shifts load-balancing and port-collision responsibilities outward.

9. Secrets and configs are Swarm objects

Swarm secrets are stored in the encrypted Raft log and are delivered only to tasks granted access, mounted into an in-memory filesystem by default. Configs are also cluster objects but are not secrets. Both should be versioned as identities in operational evidence; a secret value should never be printed to prove its existence.

10. Rolling update is a controlled desired-state transition

Update parallelism, delay, order, monitor interval, maximum failure ratio, and failure action decide how quickly task generations change and what happens when new tasks fail. Rollback is another desired-state transition. Capture the service spec version and task history before updating so you can prove what changed.

11. Image tags are especially dangerous during distributed updates

A mutable tag can resolve differently across time. For a controlled rollout, record the digest that the service specification uses. docker stack deploy can resolve image references through the registry; after deployment, inspect the service’s ContainerSpec.Image rather than assuming the tag still represents the deployed bytes.

13. Read-only first: existing-estate inspection

docker version
docker context show
docker info --format '{{json .Swarm}}'
docker node ls
docker service ls
docker stack ls
docker network ls --filter driver=overlay
docker secret ls
docker config ls

These commands reveal cluster identity and objects without changing membership or desired state. On a worker node, manager-only commands fail; that failure is itself useful evidence about the context you are targeting.

14. Current support boundary

Swarm mode remains a current Docker Engine feature with maintained documentation. Platform selection is nevertheless a separate decision: a team may keep an existing Swarm because its operational requirements fit, or migrate because ecosystem, multi-cluster, policy, autoscaling, or organizational requirements point elsewhere. This course does not rank platforms; it teaches you to bind that decision to requirements and evidence.

Knowledge check

A service has three replicas. One task exits. Does Swarm normally restart that exact task?

A five-manager swarm loses three managers but worker containers are still serving traffic. Is the control plane healthy?

Is a Swarm secret equivalent to an environment variable?

Why inspect the service image field after deployment?

Does a successful local docker compose up prove the same file is fully supported by docker stack deploy?

Next lesson

Next: Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support: Guided Hands-On Workflow and Core Operations

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Course baseline: Docker Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, BuildKit 0.33.0. The executable labs record the learner’s actual installed component versions. Swarm mode remains built into current Docker Engine and current Docker documentation explicitly describes it as a production runtime option; this chapter uses “legacy estate” to mean an existing platform that must be operated or evaluated deliberately, not that Swarm mode is removed.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.