Docker and OCI Repositories, Connectors, Registry Paths, Authentication, Layers, and Image Metadata: Diagnostics, Failure Modes, Security, and Performance
Diagnose Docker/OCI failures from evidence: wrong registry path, TLS/insecure-registry mismatch, missing bearer-token realm, tag drift, 401/403/404, upstream failure, hostname/reverse-proxy mismatch, and cache/layer assumptions.
Learning objectives
- Use a repeatable evidence-first diagnostic sequence for Docker/OCI requests.
- Interpret 401/403/404, TLS, manifest, and upstream failures without hiding the original cause.
- Distinguish registry routing errors from authorization, client cache, Nexus proxy cache, database/blob, and upstream failures.
- Diagnose tag drift by comparing resolved manifest digests rather than relying on tag names.
- Apply least-destructive corrections and avoid direct Nexus database/blob manipulation.
Incident rule. Never “fix Docker” by globally disabling TLS verification, broadening Nexus privileges to administrator, deleting blob files, wiping a production proxy cache, or changing reverse-proxy headers without evidence. Reproduce against a disposable repository/client first.
1. Diagnostic sequence
Preserve evidence in a fixed order so each layer can be falsified instead of guessed:
- Record the exact image reference, timestamp, client version, and error text.
- Confirm Nexus version/edition and the intended routing mode.
- Inspect registry host/path/port and TLS trust.
- Confirm bearer-token realm and client login state.
- Confirm repository format/type, group order, and privileges.
- Resolve tag → manifest digest; inspect manifest/config/layer state.
- For proxies, inspect remote URL, cache state, upstream reachability/rate limits.
- Inspect database/blob/disk health and request/application logs.
- Apply the smallest correction and retry one controlled request.
2. Broken example: using the ordinary Nexus repository path with a registry client
Path-based Docker/OCI routing deliberately differs from the ordinary
/repository/<name> URL used by Maven/npm/raw HTTP
clients. This intentionally broken request adds
/repository/ to the image name:
# Intentionally wrong for path-based OCI routing:
podman pull --tls-verify=false 127.0.0.1:8081/repository/academy-ch08-group/library/alpine:3.20
Depending on client/version, the symptom may be 404, name unknown,
manifest unknown, or an authentication scope that clearly contains
the wrong path. Preserve that output. The repair is not to create a
repository named repository; use the documented
path-based reference:
podman pull --tls-verify=false 127.0.0.1:8081/academy-ch08-group/library/alpine:3.20
3. TLS / insecure-registry mismatch
Docker Engine expects secure registry transport unless a host is
explicitly configured as an insecure registry. Podman can disable
TLS verification per command, which is why the disposable HTTP lab
uses it. In production, the correction is a valid HTTPS endpoint and
trusted CA—not a blanket --tls-verify=false policy.
| Symptom | Likely layer | Evidence to collect | Correction |
|---|---|---|---|
| HTTPS client talks to plain HTTP endpoint | Transport/TLS | Client error, listener scheme/port, Nexus/reverse-proxy config | Use correct HTTPS endpoint or disposable client-specific plain-HTTP option. |
| Certificate name mismatch | TLS/DNS/reverse proxy | Certificate SAN, requested hostname, Host/X-Forwarded-* headers | Issue/use certificate for the actual registry host and preserve host headers correctly. |
| Unknown private CA | Client trust store | Certificate chain and client trust config | Install only the lab/organizational CA into the appropriate trust store; do not disable verification globally. |
4. Bearer-token realm not enabled
If the registry endpoint is reachable but login/pull repeatedly fails during the token challenge, inspect the active realms. For self-hosted Docker repositories the Docker Bearer Token Realm is required; native OCI uses the OCI Bearer Token Realm. A wrong or missing realm can produce authentication failures even when the username/password works in the Nexus browser UI.
Do not respond by granting * privileges or sharing the
admin account. Enable the correct realm on the disposable instance,
then test with a least-privilege user and one image reference.
5. 401 versus 403: identity versus authorization
A 401 challenge can be part of normal token flow. A final 401 often
means login/token acquisition failed. A 403 usually means Nexus
recognized the identity but the requested repository action is not
permitted. Capture the requested repository scope and compare it
with
nx-repository-view-<format>-<repo>-read/add/edit
privileges as appropriate. Group read privileges do not
automatically grant direct member access outside a group request.
6. Tag points to a new digest unexpectedly
Suppose deployment A recorded app:prod yesterday and
deployment B uses the same tag today, but their resolved manifest
digests differ. That is tag drift, not “Docker
cache corruption.” First verify the registry/tag metadata and digest
from a fresh client or registry inspection. Then determine whether
mutable tags are intentional.
# Example evidence shape; substitute your lab refs.
podman image inspect 127.0.0.1:8081/academy-ch08-group/learner-example/ch08-app:prod --format '{{json .RepoDigests}}'
# Better release record: the manifest digest captured at push/pull time.
cat manifest-digest.txt
If immutability is required, enforce Disable redeploy on the hosted repository and use new version tags. If channel tags must move, record digest-at-promotion and deploy by digest.
7. Proxy upstream outage or rate limit
When Docker Hub or another upstream is unavailable/rate-limited, already-cached manifests/layers may continue to work while uncached content fails. Determine whether the request was a cache hit, metadata refresh, or new blob fetch. Do not globally disable routing rules, TLS, or authentication to test the upstream.
| Question | Evidence |
|---|---|
| Was this exact manifest cached? | Nexus Browse/Search/proxy repository content and request logs. |
| Did a fresh client still fail? | Use isolated client state to remove local cache ambiguity. |
| Is upstream reachable from Nexus host? | Controlled HEAD/registry request from the Nexus network; respect upstream terms/rate limits. |
| Is failure only one tag? | Resolve by digest or inspect tag metadata; a cached digest can exist while a moving tag refresh fails. |
8. Stale client state versus Nexus proxy state
Container engines have local image/layer stores. Nexus proxy repositories have their own cached registry content. A client can report “image is up to date” without contacting Nexus for every blob; conversely, a fresh client can still receive stale/failed data from Nexus proxy state. Isolate one layer at a time and preserve timestamps/request logs.
9. Cleanup and referenced-layer assumptions
Operators sometimes assume a “dangling” or apparently unreferenced blob may be deleted from the filesystem. Never do this. Registry manifests, tags, configs, layers, soft-delete/reclamation, and repository tasks have internal consistency requirements. Use supported Nexus delete/cleanup/task mechanisms for the exact pinned version. A successful repository deletion may not immediately equal filesystem space recovery because blob reclamation can be a separate stage.
10. Hostname mismatch behind a reverse proxy
The bearer challenge, redirect/location headers, TLS certificate,
and forwarded host/scheme must agree on the externally visible
registry address. A reverse proxy that rewrites to Nexus correctly
but sends the wrong Host or
X-Forwarded-Proto can create authentication loops or
clients that request tokens from an unreachable internal URL.
Compare the client-visible WWW-Authenticate header with
the expected public hostname.
11. Performance: identify the slow layer before tuning
| Layer | Typical bottleneck | Evidence |
|---|---|---|
| Client | Local storage decompression, concurrent pull behavior | Client debug/timing and local disk metrics. |
| Network | Latency/bandwidth/TLS proxy | Request timings, path MTU, proxy/load-balancer metrics. |
| Nexus JVM | CPU/heap/direct-memory pressure | Nexus metrics/logs; do not blindly increase heap. |
| Database | Metadata/query latency | PostgreSQL/H2 evidence appropriate to deployment; no direct edits. |
| Blob storage | Layer read/write latency and throughput | Storage latency/IOPS/free-space metrics. |
| Proxy upstream | Registry latency/rate limiting/outage | Proxy request logs, upstream status, cache-hit versus miss evidence. |
12. Mini incident drill
A CI job reports manifest unknown for
127.0.0.1:8081/academy-ch08-group/learner-example/ch08-app:1.0.0. Another developer can pull the digest directly from hosted. Work
the sequence:
- Capture the failing exact group reference and request time.
- Confirm group member order and that hosted is a member.
- Inspect whether the tag exists in hosted and what digest it resolves to.
- Check the CI identity has group read privilege; do not infer this from hosted direct access.
- Use a fresh local client state to rule out stale local naming/cache behavior.
- Check request logs for repository resolution and HTTP status.
- Correct only the missing group membership/order/privilege/tag issue identified.
- Pull once by tag and once by the expected digest; preserve both results.
Knowledge check
A path-based pull includes /repository/ before the repository name and returns 404. What should you check first?
The registry routing syntax. Path-based Docker/OCI references omit the ordinary Nexus /repository/ prefix.
Why is a 401 not sufficient evidence of a wrong password?
The registry bearer-token handshake intentionally begins with a 401 challenge containing WWW-Authenticate. Inspect whether token acquisition/retry succeeds.
A tag resolves to a different digest than yesterday. Is clearing the local layer cache the primary fix?
No. First determine whether the tag was intentionally or accidentally moved. The registry tag-to-digest mapping is the key evidence.
Why should you never delete an apparently unused layer file directly from the blob store?
Nexus maintains repository/database/blob consistency and reclamation semantics. Manual filesystem deletion can corrupt referenced content or internal state.
What evidence helps distinguish an upstream outage from a Nexus cache hit?
Whether the exact manifest/layers already exist in the proxy, request logs showing proxy behavior, and a controlled fresh-client request.
Official references and version notes
- Nexus Repository Download and 3.95.0 release notes — current downloadable self-hosted baseline for this chapter.
- Sonatype: Docker Registry — Docker registry paths, routing methods, API support, and self-hosted connector guidance.
- Sonatype: Docker Authentication — Docker Bearer Token Realm, login behavior, and anonymous-access prerequisites.
- Sonatype: OCI Repositories, Create an OCI Repository, and Configure OCI Repository.
- Sonatype: OCI CLI Usage — Docker/Podman/OCI-compatible client examples.
- Sonatype: Proxy Repository for Docker and reverse-proxy strategies.
- Docker: docker image pull — tag versus digest pulls and layer reuse.
- OCI Image Manifest Specification and OCI Image Configuration.
Version-sensitive statements were rechecked against Sonatype,
Docker, and OCI primary documentation on 2026-08-26. The mandatory
lab assumes self-hosted Nexus Repository Community Edition 3.95.0,
Java 21 on the Nexus side, native OCI repositories introduced in
Nexus 3.94.0, and a current Docker-compatible client. Record
docker version or podman version locally;
client behavior evolves independently of Nexus. Re-check the live
release and format documentation before executing the lab.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.