Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support: Diagnostics, Failure Modes, Security, and Performance
Diagnose Swarm incidents without destroying evidence: quorum loss, mutable tags, exposed tokens, update failures, stack/Compose mismatches, and unsafe cleanup.
Learning objectives
- Preserve Swarm control-plane, service, task, image, network, secret/config, and event evidence before repair.
- Diagnose quorum loss without casually forcing a new cluster or removing managers.
- Recognize exposed join tokens, mutable-tag drift, update failures, and stack/Compose mismatches.
- Repair the smallest responsible layer instead of pruning or recreating the cluster.
- Explain why running application containers can coexist with an unhealthy Swarm control plane.
1. Incident rule: preserve desired state and task history first
Swarm’s most valuable failure evidence often lives in manager state: service spec/version, UpdateStatus, task history, node reachability, and scheduler errors. If you immediately delete services, force-leave managers, or reinitialize the cluster, you destroy the relationships needed to explain the failure.
2. Evidence-first diagnostic sequence
docker version
docker context show
docker info --format '{{json .Swarm}}'
docker node ls
docker service ls
docker service inspect <service> --pretty
docker service ps --no-trunc <service>
docker service inspect <service> --format '{{json .UpdateStatus}}'
docker network ls --filter driver=overlay
docker secret ls
docker config ls
docker events --since 20m --filter type=service --filter type=node
Only after preserving these layers should you inspect node-local container logs, mounts, health, resources, registry access, or external load-balancer state.
3. Failure mode: lost manager quorum
Symptoms include timeouts or management operations that cannot be
committed, while existing tasks may continue to run. Confirm the
manager count and reachability first. The preferred recovery is to
restore failed managers to communication.
--force-new-cluster is disaster recovery, not a
first-line “make it work” command.
4. Failure mode: join token exposure
A join token authorizes a node to join as worker or manager. Never paste it into tickets, lesson screenshots, shell history exports, or logs. If exposure is suspected, rotate the relevant token from a healthy manager and update the secure distribution path. Do not “prove” a token by printing it into an evidence packet.
5. Failure mode: mutable image tag during update
A service update using a mutable tag can make incident reasoning
ambiguous: which digest did the manager resolve, and which tasks
pulled which content? Preserve ContainerSpec.Image,
task image errors, registry availability, and deployment timestamp.
Repair by selecting an explicit intended digest, not by repeatedly
forcing updates against the same tag.
6. Failure mode: update pauses or repeatedly fails
Start with docker service ps --no-trunc and
UpdateStatus. A failed task may point to image pull,
command exit, health, mount, secret/config, resource, placement, or
network errors. The service-level failure action controls
orchestration behavior; it does not explain root cause.
7. Intentionally broken example: impossible placement
# Disposable lab service only
docker service create --name ch39_broken --constraint 'node.labels.chapter39==yes' alpine:3.22 sleep 1d
docker service ps --no-trunc ch39_broken
Expected evidence is a task remaining pending with a placement error because no node satisfies the constraint. The repair is not to restart Docker or recreate the swarm; either add the intended label to an authorized disposable node or correct/remove the mistaken constraint:
docker service update --constraint-rm 'node.labels.chapter39==yes' ch39_broken
docker service ps --no-trunc ch39_broken
docker service rm ch39_broken
8. Failure mode: docker compose up mistaken for stack
semantics
Compose can start local containers successfully while Swarm stack deployment ignores or rejects unsupported fields. Capture the exact deployment command and warnings such as “Ignoring unsupported options.” Fix the deployment artifact or tool choice; do not silently assume the local Compose result proves cluster behavior.
9. Failure mode: secrets placed in environment variables
Swarm secrets exist specifically to provide a narrower delivery path. Environment variables can surface through process inspection, diagnostics, crash dumps, and application logs. Convert the application to read from the mounted secret file where possible, and rotate the credential if it was exposed.
10. Failure mode: pruning overlay resources during an incident
Broad prune commands can delete unused-looking objects that are part of rollback, recovery, or another team’s workload. During a Swarm incident, inventory by exact service, stack namespace, labels, and network attachments. Remove only objects with a confirmed owner and dependency graph.
11. Node-local container state is downstream evidence
A container may be healthy while its node is unavailable to the manager, or a container may have exited while the service already created a replacement elsewhere. Always correlate container ID with task ID and service slot. Without that link, node logs can be correct but operationally irrelevant.
12. Minimal repair ladder
| Observed evidence | Smallest next action | Avoid |
|---|---|---|
| manager unreachable but quorum exists | restore manager connectivity or replace through planned membership change | force-new-cluster |
| pending task: no suitable node | fix constraint/capacity/availability | daemon restart |
| task: secret/config missing | repair service grant/object reference | printing secret |
| image pull digest error | verify registry/digest/auth path | rebuild unrelated image |
| update paused after failure | diagnose task error, then resume/rollback explicitly | force-update loops |
| stack warning unsupported field | adapt stack artifact or deployment model | assuming Compose parity |
13. Security and performance caveats
Managers are sensitive to resource starvation because they run Raft and scheduling. Separate manager duties from heavy application workload in critical estates. Overlay encryption, ingress routing, and service discovery add operational layers; measure latency and throughput before tuning. Do not disable security controls or bypass the orchestrator just to make a benchmark faster.
Knowledge check
The site is serving traffic but
docker service update times out after three of five
managers fail. What layer is broken?
The Swarm control plane has lost manager quorum. Existing tasks can keep serving, but management changes cannot be committed.
A task is Pending with “no suitable node.” Should you restart Docker first?
No. Inspect constraints, node labels/availability, and resource requirements. The scheduler is telling you a placement predicate cannot be satisfied.
Why is --force-new-cluster dangerous as a routine
quorum fix?
It is disaster-recovery behavior that rewrites the manager set around one node. Use it only when normal quorum recovery is impossible and backups/evidence are understood.
What is the first evidence to capture for a paused rolling update?
Service spec/version, UpdateStatus, and full task history/errors. Those identify which generation failed and why.
Why should join tokens be absent from incident bundles?
They are credentials that authorize new nodes to join the swarm. Evidence should record token rotation/status, not token values.
Official references and version notes
2026-09-22. Course baseline: Docker Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, BuildKit 0.33.0. The executable labs record the learner’s actual installed component versions. Swarm mode remains built into current Docker Engine and current Docker documentation explicitly describes it as a production runtime option; this chapter uses “legacy estate” to mean an existing platform that must be operated or evaluated deliberately, not that Swarm mode is removed.
- Docker Docs — Swarm mode — current support position and core feature set.
- Docker Docs — How services work — desired state, services, tasks, replicas, constraints, and update behavior.
- Docker Docs — How nodes work — managers, workers, scheduling, and manager-count guidance.
- Docker Docs — Administer and maintain a swarm — quorum, manager distribution, backup, and disaster recovery.
- Docker Docs — Raft consensus — replicated manager state and majority requirements.
- Docker Docs — Manage swarm service networks — overlay, ingress, routing mesh, and control/data-plane traffic.
- Docker Docs — Manage sensitive data with Docker secrets — encrypted Raft storage and in-memory task mounts.
- Docker Docs — Docker configs — immutable configs and service/stack lifecycle.
- Docker Docs — Apply rolling updates to a service — update delay, parallelism, and task replacement.
- Docker Docs — Rolling update tutorial — observing task transitions during updates.
- Docker Docs — Deploy a stack to a swarm — manager-only stack deployment and Compose-file compatibility warning.
- Docker CLI — docker stack deploy — current flags including image digest resolution and registry auth propagation.
- Docker CLI — docker service update — rolling-update, rollback, image, secret, config, and publish controls.
- Docker Engine 29 release notes — current Engine-era compatibility baseline.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.