Chapter 39Lesson 03~185 minutes

Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support: Configuration, Design Choices, and Tradeoffs

Choose Swarm service, networking, manager-quorum, publication, and update strategies by explicit prerequisites, failure domains, and observable evidence.

TradeoffsRouting meshQuorumCompose vs stackLegacy support

Learning objectives

  • Choose between standalone Compose, Swarm service/stack, and an external orchestrator from explicit requirements.
  • Design replicated/global services, manager quorum, publication mode, and update policy with failure-domain evidence.
  • Explain stack/Compose compatibility and secret/config lifecycle constraints.
  • Separate current Swarm support from organizational migration decisions.
  • Write a decision record with prerequisites, predicted state, and observable verification.

1. Design starts with the control-plane requirement

Do not choose Swarm because a Compose file already exists, and do not migrate away solely because the estate is old. Ask what the workload requires: one host or many, desired-state rescheduling, rolling updates, secrets/config distribution, placement constraints, service discovery, routing, policy ecosystem, cross-cluster operations, and team skills.

2. Standalone Compose versus Swarm versus Kubernetes boundary

Choice Control plane Best-fit evidence Important boundary
Compose single Docker Engine project single-host lifecycle and dependency model no Raft scheduler or cross-node rescheduling
Swarm Docker Engine integrated cluster services/tasks, Raft, overlays, rolling updates smaller ecosystem; stack deploy compatibility differs from current Compose
Kubernetes or another orchestrator separate cluster control plane broader ecosystem/policy/multi-workload platform needs different API and operating model; migration is not a file-format toggle

3. Replicated versus global

Replicated mode expresses a target count. Global mode expresses one task on every eligible node. Use replicated mode when capacity is a count you want the scheduler to place; use global mode when node membership itself defines desired instances, such as a host-level agent. Constraints and node availability still determine eligibility.

4. Manager count and quorum design

Managers Quorum Simultaneous manager failures tolerated Operational note
1 1 0 lab or non-HA only
3 2 1 common minimum HA layout
5 3 2 more failure tolerance, more Raft write coordination
7 4 3 Docker-recommended upper range for manager count

Place managers across independent failure domains. An odd count helps fault tolerance but cannot protect against correlated infrastructure failure or bad operational actions.

5. Routing mesh versus host-mode publish

Publish mode What accepts traffic Benefit Cost / prerequisite
ingress/routing mesh swarm nodes according to ingress behavior built-in service publication and load distribution extra routing layer; understand firewall/LB path
host node where a task binds the port direct node-local path external LB/discovery and port-placement discipline become your responsibility

6. Update policy is a failure-containment policy

parallelism=1 limits blast radius but lengthens rollout. start-first can reduce interruption but temporarily increases resource and port pressure. stop-first minimizes overlap but may reduce availability. pause preserves the failed rollout state for diagnosis; automatic rollback can reduce exposure but may erase the moment when humans should capture first-failure evidence. Choose deliberately.

7. Mutable tags versus digest-bound service specs

Human release tags can remain part of the workflow, but the service should be traceable to the exact digest selected during deployment. During incident review, record both the release alias and the stored service image reference. Rebuilding a tag and calling it “the same release” breaks evidence continuity.

8. Secret lifecycle and rotation

Secrets are immutable objects. Rotation is typically “create new secret → update service grant → verify new tasks → remove old secret after no service references it.” This is different from editing a file in place. The secret value is not evidence; object ID/name, grant, task generation, and application success are.

9. Stack compatibility is a deployment contract

Current Docker documentation still warns that docker stack deploy uses the legacy Compose V3-era format and is not compatible with the entire latest Compose Specification. Treat the stack file as its own validated artifact. Provider-specific Compose extensions, local bind assumptions, profiles, and development/watch features require explicit review.

10. Platform support versus platform fit

Swarm mode is currently supported in Docker Engine and has maintained documentation. That factual status does not decide whether a new system should use it. A decision record should list application requirements, operational capabilities, ecosystem dependencies, migration cost, skill availability, failure modes, and support obligations—then document the chosen platform without rewriting history.

11. Worked scenario: existing three-manager estate

Suppose an estate has three managers across three zones, eight workers, ten services, and a stable operational team. A proposed migration promises a richer policy ecosystem but introduces a new control plane, observability stack, deployment API, and training burden. The correct engineering move is not “Swarm is old, migrate” or “it works, never change.” First capture current service SLOs, manager/quorum health, release/rollback performance, security gaps, missing platform capabilities, and migration test evidence.

12. Decision matrix

Question Swarm-oriented answer Alternative-platform trigger Evidence
Need multi-host desired state? Swarm services provide it No multi-host need → Compose may suffice service/task topology
Need advanced admission/policy ecosystem? limited native policy surface strong requirement may favor another orchestrator control inventory + compliance requirement
Need simple integrated Docker operations? Swarm can minimize platform layers existing org standard may outweigh simplicity operator skills + incident MTTR
Need cross-cluster/fleet governance? requires external operational design mature multi-cluster control may be decisive fleet inventory + automation
Migration justified? only with measured requirement/gap gap + tested target + rollback plan migration ADR and rehearsal evidence

13. Evidence-first design record template

Decision: keep / migrate / reduce scope
Current swarm: <ID>, managers=<N>, workers=<N>
Requirements: <availability, policy, scale, compliance, ecosystem>
Image policy: digest-bound? <yes/no>
Quorum design: <manager count + failure domains>
Network publication: <ingress/host/external LB>
Secret/config lifecycle: <rotation procedure>
Update policy: <parallelism/order/failure action>
Known gaps: <measured>
Target prerequisites: <versions, skills, integrations>
Validation evidence: <tests/incident drills>
Rollback or coexistence path: <documented>

Knowledge check

Does an odd number of managers guarantee quorum during any failure?

When can host-mode publication be preferable to the routing mesh?

Is automatic rollback always safer than update pause?

What factual statement can you make about Swarm support today?

Why should a stack file be validated separately from a local Compose file?

Next lesson

Next: Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support: Diagnostics, Failure Modes, Security, and Performance

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Course baseline: Docker Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, BuildKit 0.33.0. The executable labs record the learner’s actual installed component versions. Swarm mode remains built into current Docker Engine and current Docker documentation explicitly describes it as a production runtime option; this chapter uses “legacy estate” to mean an existing platform that must be operated or evaluated deliberately, not that Swarm mode is removed.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.