Chapter 08Lesson 04165–215 min

Docker and OCI Repositories, Connectors, Registry Paths, Authentication, Layers, and Image Metadata: Diagnostics, Failure Modes, Security, and Performance

Diagnose Docker/OCI failures from evidence: wrong registry path, TLS/insecure-registry mismatch, missing bearer-token realm, tag drift, 401/403/404, upstream failure, hostname/reverse-proxy mismatch, and cache/layer assumptions.

401 / 403 / 404TLSBearer realmTag driftDiagnostics

Learning objectives

  • Use a repeatable evidence-first diagnostic sequence for Docker/OCI requests.
  • Interpret 401/403/404, TLS, manifest, and upstream failures without hiding the original cause.
  • Distinguish registry routing errors from authorization, client cache, Nexus proxy cache, database/blob, and upstream failures.
  • Diagnose tag drift by comparing resolved manifest digests rather than relying on tag names.
  • Apply least-destructive corrections and avoid direct Nexus database/blob manipulation.

Incident rule. Never “fix Docker” by globally disabling TLS verification, broadening Nexus privileges to administrator, deleting blob files, wiping a production proxy cache, or changing reverse-proxy headers without evidence. Reproduce against a disposable repository/client first.

1. Diagnostic sequence

Preserve evidence in a fixed order so each layer can be falsified instead of guessed:

  1. Record the exact image reference, timestamp, client version, and error text.
  2. Confirm Nexus version/edition and the intended routing mode.
  3. Inspect registry host/path/port and TLS trust.
  4. Confirm bearer-token realm and client login state.
  5. Confirm repository format/type, group order, and privileges.
  6. Resolve tag → manifest digest; inspect manifest/config/layer state.
  7. For proxies, inspect remote URL, cache state, upstream reachability/rate limits.
  8. Inspect database/blob/disk health and request/application logs.
  9. Apply the smallest correction and retry one controlled request.

2. Broken example: using the ordinary Nexus repository path with a registry client

Path-based Docker/OCI routing deliberately differs from the ordinary /repository/<name> URL used by Maven/npm/raw HTTP clients. This intentionally broken request adds /repository/ to the image name:

# Intentionally wrong for path-based OCI routing:
podman pull --tls-verify=false   127.0.0.1:8081/repository/academy-ch08-group/library/alpine:3.20

Depending on client/version, the symptom may be 404, name unknown, manifest unknown, or an authentication scope that clearly contains the wrong path. Preserve that output. The repair is not to create a repository named repository; use the documented path-based reference:

podman pull --tls-verify=false   127.0.0.1:8081/academy-ch08-group/library/alpine:3.20

3. TLS / insecure-registry mismatch

Docker Engine expects secure registry transport unless a host is explicitly configured as an insecure registry. Podman can disable TLS verification per command, which is why the disposable HTTP lab uses it. In production, the correction is a valid HTTPS endpoint and trusted CA—not a blanket --tls-verify=false policy.

Symptom Likely layer Evidence to collect Correction
HTTPS client talks to plain HTTP endpoint Transport/TLS Client error, listener scheme/port, Nexus/reverse-proxy config Use correct HTTPS endpoint or disposable client-specific plain-HTTP option.
Certificate name mismatch TLS/DNS/reverse proxy Certificate SAN, requested hostname, Host/X-Forwarded-* headers Issue/use certificate for the actual registry host and preserve host headers correctly.
Unknown private CA Client trust store Certificate chain and client trust config Install only the lab/organizational CA into the appropriate trust store; do not disable verification globally.

4. Bearer-token realm not enabled

If the registry endpoint is reachable but login/pull repeatedly fails during the token challenge, inspect the active realms. For self-hosted Docker repositories the Docker Bearer Token Realm is required; native OCI uses the OCI Bearer Token Realm. A wrong or missing realm can produce authentication failures even when the username/password works in the Nexus browser UI.

Do not respond by granting * privileges or sharing the admin account. Enable the correct realm on the disposable instance, then test with a least-privilege user and one image reference.

5. 401 versus 403: identity versus authorization

A 401 challenge can be part of normal token flow. A final 401 often means login/token acquisition failed. A 403 usually means Nexus recognized the identity but the requested repository action is not permitted. Capture the requested repository scope and compare it with nx-repository-view-<format>-<repo>-read/add/edit privileges as appropriate. Group read privileges do not automatically grant direct member access outside a group request.

6. Tag points to a new digest unexpectedly

Suppose deployment A recorded app:prod yesterday and deployment B uses the same tag today, but their resolved manifest digests differ. That is tag drift, not “Docker cache corruption.” First verify the registry/tag metadata and digest from a fresh client or registry inspection. Then determine whether mutable tags are intentional.

# Example evidence shape; substitute your lab refs.
podman image inspect   127.0.0.1:8081/academy-ch08-group/learner-example/ch08-app:prod   --format '{{json .RepoDigests}}'

# Better release record: the manifest digest captured at push/pull time.
cat manifest-digest.txt

If immutability is required, enforce Disable redeploy on the hosted repository and use new version tags. If channel tags must move, record digest-at-promotion and deploy by digest.

7. Proxy upstream outage or rate limit

When Docker Hub or another upstream is unavailable/rate-limited, already-cached manifests/layers may continue to work while uncached content fails. Determine whether the request was a cache hit, metadata refresh, or new blob fetch. Do not globally disable routing rules, TLS, or authentication to test the upstream.

Question Evidence
Was this exact manifest cached? Nexus Browse/Search/proxy repository content and request logs.
Did a fresh client still fail? Use isolated client state to remove local cache ambiguity.
Is upstream reachable from Nexus host? Controlled HEAD/registry request from the Nexus network; respect upstream terms/rate limits.
Is failure only one tag? Resolve by digest or inspect tag metadata; a cached digest can exist while a moving tag refresh fails.

8. Stale client state versus Nexus proxy state

Container engines have local image/layer stores. Nexus proxy repositories have their own cached registry content. A client can report “image is up to date” without contacting Nexus for every blob; conversely, a fresh client can still receive stale/failed data from Nexus proxy state. Isolate one layer at a time and preserve timestamps/request logs.

9. Cleanup and referenced-layer assumptions

Operators sometimes assume a “dangling” or apparently unreferenced blob may be deleted from the filesystem. Never do this. Registry manifests, tags, configs, layers, soft-delete/reclamation, and repository tasks have internal consistency requirements. Use supported Nexus delete/cleanup/task mechanisms for the exact pinned version. A successful repository deletion may not immediately equal filesystem space recovery because blob reclamation can be a separate stage.

10. Hostname mismatch behind a reverse proxy

The bearer challenge, redirect/location headers, TLS certificate, and forwarded host/scheme must agree on the externally visible registry address. A reverse proxy that rewrites to Nexus correctly but sends the wrong Host or X-Forwarded-Proto can create authentication loops or clients that request tokens from an unreachable internal URL. Compare the client-visible WWW-Authenticate header with the expected public hostname.

11. Performance: identify the slow layer before tuning

Layer Typical bottleneck Evidence
Client Local storage decompression, concurrent pull behavior Client debug/timing and local disk metrics.
Network Latency/bandwidth/TLS proxy Request timings, path MTU, proxy/load-balancer metrics.
Nexus JVM CPU/heap/direct-memory pressure Nexus metrics/logs; do not blindly increase heap.
Database Metadata/query latency PostgreSQL/H2 evidence appropriate to deployment; no direct edits.
Blob storage Layer read/write latency and throughput Storage latency/IOPS/free-space metrics.
Proxy upstream Registry latency/rate limiting/outage Proxy request logs, upstream status, cache-hit versus miss evidence.

12. Mini incident drill

A CI job reports manifest unknown for 127.0.0.1:8081/academy-ch08-group/learner-example/ch08-app:1.0.0. Another developer can pull the digest directly from hosted. Work the sequence:

  1. Capture the failing exact group reference and request time.
  2. Confirm group member order and that hosted is a member.
  3. Inspect whether the tag exists in hosted and what digest it resolves to.
  4. Check the CI identity has group read privilege; do not infer this from hosted direct access.
  5. Use a fresh local client state to rule out stale local naming/cache behavior.
  6. Check request logs for repository resolution and HTTP status.
  7. Correct only the missing group membership/order/privilege/tag issue identified.
  8. Pull once by tag and once by the expected digest; preserve both results.

Knowledge check

A path-based pull includes /repository/ before the repository name and returns 404. What should you check first?

Why is a 401 not sufficient evidence of a wrong password?

A tag resolves to a different digest than yesterday. Is clearing the local layer cache the primary fix?

Why should you never delete an apparently unused layer file directly from the blob store?

What evidence helps distinguish an upstream outage from a Nexus cache hit?

Next lesson

Integrated checkpoint

Build a fresh topology, predict state changes, push once, retag without rebuilding, prove digest identity/layer reuse, test authentication failure, and clean up through supported controls.

Official references and version notes

Version-sensitive statements were rechecked against Sonatype, Docker, and OCI primary documentation on 2026-08-26. The mandatory lab assumes self-hosted Nexus Repository Community Edition 3.95.0, Java 21 on the Nexus side, native OCI repositories introduced in Nexus 3.94.0, and a current Docker-compatible client. Record docker version or podman version locally; client behavior evolves independently of Nexus. Re-check the live release and format documentation before executing the lab.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.