Docker DNS, Service Discovery, Aliases, IPv4/IPv6, Custom Subnets, and Multi-Network Applications: Diagnostics, Failure Modes, Security, and Performance
Diagnose Docker DNS, alias scope, subnet overlap, multi-network, and IPv6 failures with first-failure evidence and least-destructive corrections.
Learning objectives
- Preserve network, resolver, container, and application evidence before changing a failing Docker topology.
- Diagnose hard-coded IPs, alias-scope mistakes, subnet overlap, resolver-boundary confusion, IPv6 capability gaps, and DNS-cache symptoms causally.
- Distinguish lookup failure from connection failure, application failure, host routing/firewall failure, and external DNS/provider failure.
- Apply the least destructive correction to the smallest failing layer rather than flattening networks or adding static hosts entries.
- Recognize performance symptoms caused by DNS lookup behavior or emulated external dependencies without weakening security controls.
/etc/resolv.conf, lookup output, application test, and
relevant logs. A fix without first-failure evidence teaches nothing.
1. Diagnostic sequence
- Preserve the first failing lookup/connection and timestamp.
-
Confirm
docker context show, Engine/CLI/Compose version, host/platform. - Confirm daemon/API reachability and exact container/image IDs.
- Inspect container health/process/listener before blaming DNS.
- Inspect network memberships, aliases, endpoint addresses, subnet/gateway, and resolver configuration.
- Run a name lookup from the failing caller.
- If lookup succeeds, test the exact application port/path.
- Only then inspect host routing/firewall/published-port or external DNS/provider state when relevant.
- Apply the smallest correction and rerun only the failed layer plus one end-to-end verification.
2. Failure: a hard-coded container IP survives in configuration
Symptom: after recreating api, the
client still tries the old endpoint address. The new container is
healthy by service name.
docker inspect api --format '{{json .NetworkSettings.Networks}}'
docker exec client nslookup api
docker exec client wget -T 2 -qO- http://api:8080/
# Compare the application's configured destination; do not change it yet.
docker exec client env | grep -E 'API|HOST|URL' || true
Cause: application configuration used an incidental endpoint allocation as identity. Correction: point the client to the stable service name/alias, redeploy the smallest disposable scope, and retain old/new config plus lookup evidence.
3. Intentionally broken example: alias exists—but on the other network
Create an isolated two-network topology where
api-internal exists only on back. Query it
from a front-only client and preserve the failure.
docker network create --label devops-academy.lab=ch18diag da18d-front
docker network create --label devops-academy.lab=ch18diag da18d-back
docker run -d --name da18d-api --network da18d-front \
--label devops-academy.lab=ch18diag busybox:1.36.1 sleep 3600
docker network connect --alias api-internal da18d-back da18d-api
docker run -d --name da18d-client --network da18d-front \
--label devops-academy.lab=ch18diag busybox:1.36.1 sleep 3600
# Preserve expected failure and topology evidence.
docker exec da18d-client nslookup api-internal || true
docker inspect da18d-api --format '{{json .NetworkSettings.Networks}}'
docker inspect da18d-client --format '{{json .NetworkSettings.Networks}}'
docker network inspect da18d-front da18d-back
The cause is not “Docker DNS is broken.” The alias is valid only on
da18d-back, while the caller has no endpoint there. The
correction depends on intent: use the API's front-network
name/alias, or deliberately attach the client to back only if policy
allows it. Do not add a global hosts entry.
4. Failure: custom subnet overlaps another route or Docker network
Overlap can appear as network-creation rejection, ambiguous routing, or traffic unexpectedly following a host/VPN path. Preserve IPAM and host-route evidence.
docker network ls -q | xargs -r docker network inspect \
--format '{{.Name}} {{json .IPAM.Config}}'
ip route 2>/dev/null || true
# If a proposed subnet conflicts, choose a non-overlapping lab subnet.
# Do not delete unrelated networks to make the range available.
Changing DNS cannot repair an IP routing conflict. The owning layer is address planning/IPAM.
5. Failure: confusing host DNS with Docker embedded DNS
A name that resolves on the host may fail in a container because the host uses a VPN split-DNS resolver, search suffix, local hosts file, or other policy not visible the same way inside the Docker environment. Conversely, a Docker service name can resolve inside a user-defined network but not on the host.
# Compare evidence, not assumptions.
getent hosts some-name 2>/dev/null || true
docker exec <container> cat /etc/resolv.conf
docker exec <container> nslookup some-name || true
docker network inspect <network> --format '{{json .Containers}}'
Decide whether some-name belongs to Docker discovery or
external/host DNS. Then fix the authoritative layer. Do not point
every container at an arbitrary public resolver; that can break
internal names and bypass organizational DNS policy.
6. Failure: Compose says IPv6, but the platform/path cannot deliver it
There are several distinct states: Compose accepts
enable_ipv6; the daemon creates an IPv6-enabled
network; the container receives an IPv6 endpoint; local peer traffic
works; the host forwards IPv6; upstream routing exists; an external
DNS AAAA record exists. Diagnose each separately.
docker version
docker info
docker network inspect <network> --format '{{json .IPAM.Config}}'
docker inspect <container> --format '{{json .NetworkSettings.Networks}}'
docker exec <container> cat /etc/resolv.conf
# Only test external IPv6 after local endpoint/routing evidence succeeds.
If the daemon/platform does not support the requested IPv6 path, record the limitation. Do not disable firewall policy or perform broad daemon mutations just to force a lab result.
7. Failure: DNS cache looks like a Docker network outage
Applications and language runtimes may cache DNS responses beyond what you expect. If Docker's current lookup returns the new endpoint but the long-running application still connects to the old one, the cache may be inside the application/runtime rather than Docker DNS.
Preserve: current nslookup result from the same
container, application logs showing the destination, process uptime,
and restart behavior. Prefer a bounded application-level
reload/reconnect strategy; do not restart the Docker daemon as a
blind cache flush.
8. Unsafe “fixes” to reject
- Do not connect every container to every network merely to make names resolve.
-
Do not hard-code endpoint IPs or copy them into
/etc/hostsas a production discovery mechanism. - Do not disable host firewalls/TLS/security profiles to debug DNS.
- Do not mount the Docker socket or use privileged containers for ordinary network inspection.
- Do not delete unrelated networks or run unbounded prune commands to resolve subnet conflicts.
- Do not print credentials/tokens while debugging external DNS/registry/provider integrations.
9. Performance: identify whether DNS is actually the latency source
Slow requests are often attributed to “Docker DNS” without timing
the lookup separately from connect/TLS/application work. Use
application timestamps and bounded lookup/connect tests. A fast
nslookup plus a slow TCP connection points away from
DNS. Repeated lookups may be affected by application resolver
behavior and external forwarders.
Do not optimize by bypassing names with static IPs unless measurements prove discovery is the bottleneck and the address contract is intentionally managed.
10. Clean the diagnostic lab only
docker rm -f da18d-client da18d-api 2>/dev/null || true
docker network rm da18d-front da18d-back 2>/dev/null || true
docker ps -a --filter label=devops-academy.lab=ch18diag
docker network ls --filter label=devops-academy.lab=ch18diag
Knowledge check
A Docker name resolves to the current IP, but the application still connects to an old IP. Which layer is now suspect?
The application/runtime DNS cache or its own configuration, not Docker’s current embedded DNS answer.
Why is adding a hosts entry a poor fix for a network-scoped alias failure?
It bypasses the intended membership/discovery model, hard-codes an endpoint, and can accidentally grant a name outside its intended scope.
What evidence distinguishes subnet overlap from a DNS problem?
Docker network IPAM ranges plus host/VPN routes show address overlap; DNS changes cannot repair routing ambiguity.
Which IPv6 state should be tested before public IPv6 reachability?
First verify daemon capability, network IPAM, container IPv6 endpoint, local routes, and local peer connectivity. External routing is a later boundary.
Why avoid restarting Docker during DNS troubleshooting?
It is broad and destructive, can erase transient first-failure evidence, and rarely targets the actual application/resolver/network-scope cause.
Official references and version notes
-
Docker Docs — Networking overview
— resolver behavior, custom-network embedded DNS, and
127.0.0.11. - Docker Docs — Bridge network driver — user-defined bridge isolation and automatic name/alias resolution.
- Docker Docs — docker network connect — endpoint attachment and network-scoped aliases.
- Docker Docs — Networking in Compose — project networks and service-name discovery.
- Compose Specification — service network aliases — aliases are scoped to the network on which they are declared.
-
Compose Specification — networks
— IPAM,
enable_ipv4,enable_ipv6, internal/external networks, and custom names. - Docker Docs — IPv6 networking — Linux-daemon support, IPv6 network creation, ULA allocation, and Compose examples.
- Docker Engine 29 release notes — current Engine behavior and networking fixes.
Docker Engine 29.8.1 (released 2026-09-15) is the current Engine 29 release in the primary release notes, and Docker Compose v5.5.1 is the current upstream Compose release. The runnable labs deliberately record your Engine/CLI/Compose versions and context because DNS, IPv6, Desktop networking, firewall integration, and helper-tool availability vary by platform. The mandatory path uses only local disposable networks and containers; IPv6 is an optional capability-gated extension.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.