Chapter 39Lesson 04~195 minutes

Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support: Diagnostics, Failure Modes, Security, and Performance

Diagnose Swarm incidents without destroying evidence: quorum loss, mutable tags, exposed tokens, update failures, stack/Compose mismatches, and unsafe cleanup.

DiagnosticsSecurityQuorum lossUpdate failureIncident evidence

Learning objectives

  • Preserve Swarm control-plane, service, task, image, network, secret/config, and event evidence before repair.
  • Diagnose quorum loss without casually forcing a new cluster or removing managers.
  • Recognize exposed join tokens, mutable-tag drift, update failures, and stack/Compose mismatches.
  • Repair the smallest responsible layer instead of pruning or recreating the cluster.
  • Explain why running application containers can coexist with an unhealthy Swarm control plane.

1. Incident rule: preserve desired state and task history first

Swarm’s most valuable failure evidence often lives in manager state: service spec/version, UpdateStatus, task history, node reachability, and scheduler errors. If you immediately delete services, force-leave managers, or reinitialize the cluster, you destroy the relationships needed to explain the failure.

2. Evidence-first diagnostic sequence

docker version
docker context show
docker info --format '{{json .Swarm}}'
docker node ls
docker service ls
docker service inspect <service> --pretty
docker service ps --no-trunc <service>
docker service inspect <service> --format '{{json .UpdateStatus}}'
docker network ls --filter driver=overlay
docker secret ls
docker config ls
docker events --since 20m --filter type=service --filter type=node

Only after preserving these layers should you inspect node-local container logs, mounts, health, resources, registry access, or external load-balancer state.

3. Failure mode: lost manager quorum

Symptoms include timeouts or management operations that cannot be committed, while existing tasks may continue to run. Confirm the manager count and reachability first. The preferred recovery is to restore failed managers to communication. --force-new-cluster is disaster recovery, not a first-line “make it work” command.

Safety boundary. Do not copy another manager’s Raft directory, casually demote/remove managers, or force a new cluster before preserving state and confirming the quorum calculation. Docker documents node Raft state as unique to the node identity.

4. Failure mode: join token exposure

A join token authorizes a node to join as worker or manager. Never paste it into tickets, lesson screenshots, shell history exports, or logs. If exposure is suspected, rotate the relevant token from a healthy manager and update the secure distribution path. Do not “prove” a token by printing it into an evidence packet.

5. Failure mode: mutable image tag during update

A service update using a mutable tag can make incident reasoning ambiguous: which digest did the manager resolve, and which tasks pulled which content? Preserve ContainerSpec.Image, task image errors, registry availability, and deployment timestamp. Repair by selecting an explicit intended digest, not by repeatedly forcing updates against the same tag.

6. Failure mode: update pauses or repeatedly fails

Start with docker service ps --no-trunc and UpdateStatus. A failed task may point to image pull, command exit, health, mount, secret/config, resource, placement, or network errors. The service-level failure action controls orchestration behavior; it does not explain root cause.

7. Intentionally broken example: impossible placement

# Disposable lab service only
docker service create --name ch39_broken --constraint 'node.labels.chapter39==yes' alpine:3.22 sleep 1d
docker service ps --no-trunc ch39_broken

Expected evidence is a task remaining pending with a placement error because no node satisfies the constraint. The repair is not to restart Docker or recreate the swarm; either add the intended label to an authorized disposable node or correct/remove the mistaken constraint:

docker service update --constraint-rm 'node.labels.chapter39==yes' ch39_broken
docker service ps --no-trunc ch39_broken
docker service rm ch39_broken

8. Failure mode: docker compose up mistaken for stack semantics

Compose can start local containers successfully while Swarm stack deployment ignores or rejects unsupported fields. Capture the exact deployment command and warnings such as “Ignoring unsupported options.” Fix the deployment artifact or tool choice; do not silently assume the local Compose result proves cluster behavior.

9. Failure mode: secrets placed in environment variables

Swarm secrets exist specifically to provide a narrower delivery path. Environment variables can surface through process inspection, diagnostics, crash dumps, and application logs. Convert the application to read from the mounted secret file where possible, and rotate the credential if it was exposed.

10. Failure mode: pruning overlay resources during an incident

Broad prune commands can delete unused-looking objects that are part of rollback, recovery, or another team’s workload. During a Swarm incident, inventory by exact service, stack namespace, labels, and network attachments. Remove only objects with a confirmed owner and dependency graph.

11. Node-local container state is downstream evidence

A container may be healthy while its node is unavailable to the manager, or a container may have exited while the service already created a replacement elsewhere. Always correlate container ID with task ID and service slot. Without that link, node logs can be correct but operationally irrelevant.

12. Minimal repair ladder

Observed evidence Smallest next action Avoid
manager unreachable but quorum exists restore manager connectivity or replace through planned membership change force-new-cluster
pending task: no suitable node fix constraint/capacity/availability daemon restart
task: secret/config missing repair service grant/object reference printing secret
image pull digest error verify registry/digest/auth path rebuild unrelated image
update paused after failure diagnose task error, then resume/rollback explicitly force-update loops
stack warning unsupported field adapt stack artifact or deployment model assuming Compose parity

13. Security and performance caveats

Managers are sensitive to resource starvation because they run Raft and scheduling. Separate manager duties from heavy application workload in critical estates. Overlay encryption, ingress routing, and service discovery add operational layers; measure latency and throughput before tuning. Do not disable security controls or bypass the orchestrator just to make a benchmark faster.

Knowledge check

The site is serving traffic but docker service update times out after three of five managers fail. What layer is broken?

A task is Pending with “no suitable node.” Should you restart Docker first?

Why is --force-new-cluster dangerous as a routine quorum fix?

What is the first evidence to capture for a paused rolling update?

Why should join tokens be absent from incident bundles?

Next lesson

Next: Checkpoint Lab — Docker Swarm Fundamentals, Services, Stacks, Secrets, Overlay Networks, Rolling Updates, and Legacy Estate Support

Continue with the next lesson in the course sequence and carry forward the evidence-first Docker operating model.

Official references and version notes

Baseline checked:

2026-09-22. Course baseline: Docker Engine/CLI 29.8.1, Compose 5.5.1, Buildx 0.37.1, BuildKit 0.33.0. The executable labs record the learner’s actual installed component versions. Swarm mode remains built into current Docker Engine and current Docker documentation explicitly describes it as a production runtime option; this chapter uses “legacy estate” to mean an existing platform that must be operated or evaluated deliberately, not that Swarm mode is removed.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.