Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support: Configuration, Design Choices, and Tradeoffs
Choose Swarm service, networking, manager-quorum, publication, and update strategies by explicit prerequisites, failure domains, and observable evidence.
Learning objectives
- Choose between standalone Compose, Swarm service/stack, and an external orchestrator from explicit requirements.
- Design replicated/global services, manager quorum, publication mode, and update policy with failure-domain evidence.
- Explain stack/Compose compatibility and secret/config lifecycle constraints.
- Separate current Swarm support from organizational migration decisions.
- Write a decision record with prerequisites, predicted state, and observable verification.
1. Design starts with the control-plane requirement
Do not choose Swarm because a Compose file already exists, and do not migrate away solely because the estate is old. Ask what the workload requires: one host or many, desired-state rescheduling, rolling updates, secrets/config distribution, placement constraints, service discovery, routing, policy ecosystem, cross-cluster operations, and team skills.
2. Standalone Compose versus Swarm versus Kubernetes boundary
| Choice | Control plane | Best-fit evidence | Important boundary |
|---|---|---|---|
| Compose | single Docker Engine project | single-host lifecycle and dependency model | no Raft scheduler or cross-node rescheduling |
| Swarm | Docker Engine integrated cluster | services/tasks, Raft, overlays, rolling updates | smaller ecosystem; stack deploy compatibility differs from current Compose |
| Kubernetes or another orchestrator | separate cluster control plane | broader ecosystem/policy/multi-workload platform needs | different API and operating model; migration is not a file-format toggle |
3. Replicated versus global
Replicated mode expresses a target count. Global mode expresses one task on every eligible node. Use replicated mode when capacity is a count you want the scheduler to place; use global mode when node membership itself defines desired instances, such as a host-level agent. Constraints and node availability still determine eligibility.
4. Manager count and quorum design
| Managers | Quorum | Simultaneous manager failures tolerated | Operational note |
|---|---|---|---|
| 1 | 1 | 0 | lab or non-HA only |
| 3 | 2 | 1 | common minimum HA layout |
| 5 | 3 | 2 | more failure tolerance, more Raft write coordination |
| 7 | 4 | 3 | Docker-recommended upper range for manager count |
Place managers across independent failure domains. An odd count helps fault tolerance but cannot protect against correlated infrastructure failure or bad operational actions.
5. Routing mesh versus host-mode publish
| Publish mode | What accepts traffic | Benefit | Cost / prerequisite |
|---|---|---|---|
| ingress/routing mesh | swarm nodes according to ingress behavior | built-in service publication and load distribution | extra routing layer; understand firewall/LB path |
| host | node where a task binds the port | direct node-local path | external LB/discovery and port-placement discipline become your responsibility |
6. Update policy is a failure-containment policy
parallelism=1 limits blast radius but lengthens
rollout. start-first can reduce interruption but
temporarily increases resource and port pressure.
stop-first minimizes overlap but may reduce
availability. pause preserves the failed rollout state
for diagnosis; automatic rollback can reduce exposure but may erase
the moment when humans should capture first-failure evidence. Choose
deliberately.
8. Secret lifecycle and rotation
Secrets are immutable objects. Rotation is typically “create new secret → update service grant → verify new tasks → remove old secret after no service references it.” This is different from editing a file in place. The secret value is not evidence; object ID/name, grant, task generation, and application success are.
9. Stack compatibility is a deployment contract
Current Docker documentation still warns that
docker stack deploy uses the legacy Compose V3-era
format and is not compatible with the entire latest Compose
Specification. Treat the stack file as its own validated artifact.
Provider-specific Compose extensions, local bind assumptions,
profiles, and development/watch features require explicit review.
10. Platform support versus platform fit
Swarm mode is currently supported in Docker Engine and has maintained documentation. That factual status does not decide whether a new system should use it. A decision record should list application requirements, operational capabilities, ecosystem dependencies, migration cost, skill availability, failure modes, and support obligations—then document the chosen platform without rewriting history.
11. Worked scenario: existing three-manager estate
Suppose an estate has three managers across three zones, eight workers, ten services, and a stable operational team. A proposed migration promises a richer policy ecosystem but introduces a new control plane, observability stack, deployment API, and training burden. The correct engineering move is not “Swarm is old, migrate” or “it works, never change.” First capture current service SLOs, manager/quorum health, release/rollback performance, security gaps, missing platform capabilities, and migration test evidence.
12. Decision matrix
| Question | Swarm-oriented answer | Alternative-platform trigger | Evidence |
|---|---|---|---|
| Need multi-host desired state? | Swarm services provide it | No multi-host need → Compose may suffice | service/task topology |
| Need advanced admission/policy ecosystem? | limited native policy surface | strong requirement may favor another orchestrator | control inventory + compliance requirement |
| Need simple integrated Docker operations? | Swarm can minimize platform layers | existing org standard may outweigh simplicity | operator skills + incident MTTR |
| Need cross-cluster/fleet governance? | requires external operational design | mature multi-cluster control may be decisive | fleet inventory + automation |
| Migration justified? | only with measured requirement/gap | gap + tested target + rollback plan | migration ADR and rehearsal evidence |
13. Evidence-first design record template
Decision: keep / migrate / reduce scope
Current swarm: <ID>, managers=<N>, workers=<N>
Requirements: <availability, policy, scale, compliance, ecosystem>
Image policy: digest-bound? <yes/no>
Quorum design: <manager count + failure domains>
Network publication: <ingress/host/external LB>
Secret/config lifecycle: <rotation procedure>
Update policy: <parallelism/order/failure action>
Known gaps: <measured>
Target prerequisites: <versions, skills, integrations>
Validation evidence: <tests/incident drills>
Rollback or coexistence path: <documented>
Knowledge check
Does an odd number of managers guarantee quorum during any failure?
No. It improves tolerance to manager loss, but correlated failures or loss beyond the majority threshold still remove quorum.
When can host-mode publication be preferable to the routing mesh?
When deterministic node-local binding is required and an external load balancer/discovery system explicitly owns distribution and failover.
Is automatic rollback always safer than update pause?
Not universally. Rollback can reduce exposure, while pause can preserve failure evidence. The choice depends on service risk, observability, and operational policy.
What factual statement can you make about Swarm support today?
Swarm mode remains built into current Docker Engine and is documented by Docker as a production runtime option.
Why should a stack file be validated separately from a local Compose file?
Because docker stack deploy has a documented
Compose compatibility boundary and does not implement every
modern Compose Specification feature.
Official references and version notes
2026-09-22. Course baseline: Docker Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, BuildKit 0.33.0. The executable labs record the learner’s actual installed component versions. Swarm mode remains built into current Docker Engine and current Docker documentation explicitly describes it as a production runtime option; this chapter uses “legacy estate” to mean an existing platform that must be operated or evaluated deliberately, not that Swarm mode is removed.
- Docker Docs — Swarm mode — current support position and core feature set.
- Docker Docs — How services work — desired state, services, tasks, replicas, constraints, and update behavior.
- Docker Docs — How nodes work — managers, workers, scheduling, and manager-count guidance.
- Docker Docs — Administer and maintain a swarm — quorum, manager distribution, backup, and disaster recovery.
- Docker Docs — Raft consensus — replicated manager state and majority requirements.
- Docker Docs — Manage swarm service networks — overlay, ingress, routing mesh, and control/data-plane traffic.
- Docker Docs — Manage sensitive data with Docker secrets — encrypted Raft storage and in-memory task mounts.
- Docker Docs — Docker configs — immutable configs and service/stack lifecycle.
- Docker Docs — Apply rolling updates to a service — update delay, parallelism, and task replacement.
- Docker Docs — Rolling update tutorial — observing task transitions during updates.
- Docker Docs — Deploy a stack to a swarm — manager-only stack deployment and Compose-file compatibility warning.
- Docker CLI — docker stack deploy — current flags including image digest resolution and registry auth propagation.
- Docker CLI — docker service update — rolling-update, rollback, image, secret, config, and publish controls.
- Docker Engine 29 release notes — current Engine-era compatibility baseline.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.