Chapter 29Lesson 01220–300 min

Logging, Metrics, Support ZIPs, Request Auditing, Performance Tuning, and Capacity Diagnostics: Concepts, Architecture, and Mental Model

Observability is the ability to explain what the repository service is doing from evidence rather than intuition. In Nexus Repository, that evidence is distributed across HTTP request records, application/audit logs, metrics, JVM memory, database behavior, blob-store state, task history, client timing, and support bundles. This lesson builds the causal model before any performance tuning is attempted.

ObservabilityLogsMetricsJVM memoryDB/blob evidence

Learning objectives

  • Distinguish application, request, outbound, JVM, audit, database, blob, task, and client evidence.
  • Explain the difference between a request symptom and a root cause.
  • Use current Prometheus, Service Metrics Data, and Status API boundaries correctly.
  • Distinguish JVM heap from direct memory and host headroom.
  • Build a capacity profile that preserves version, workload, cache, DB, blob, and network context.
Dated baseline (27 August 2026). Lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. Re-check the live release notes, logging/metrics pages, system requirements, and instance-specific configuration before applying production tuning.
Measure before tuning. A slow package request is an end-to-end symptom, not proof of a JVM problem. Preserve the request, Nexus version/edition/runtime, cache state, database latency/pool state, blob latency/capacity, task activity, network/upstream timing, and client-cache state before changing a knob.
Privacy boundary. Logs and support bundles can contain usernames, client IPs, internal repository names/URLs, request paths, configuration, security information, dependency coordinates, and other operationally sensitive evidence. Sonatype removes password-related information from generated support files, but operators must still review/redact and share on a need-to-know basis.

1. The practical problem: “Nexus is slow” is not a diagnosis

A developer reports that mvn dependency:resolve, npm install, or a container pull is slow. That sentence tells you only that an end-to-end transaction took longer than expected. The same symptom can be caused by a cold proxy fetch, a slow public upstream, an overloaded reverse proxy, exhausted PostgreSQL connections, high blob-storage latency, JVM memory pressure, a cleanup task competing for IO, client-side DNS, or the client’s own cache behavior.

Production troubleshooting therefore starts with a correlation key: which request, at what time, to which repository endpoint, from which client, with what response status and duration? Once that event is anchored, other signals can be aligned around it.

2. The evidence map

One request, many evidence planes
flowchart LR
C[Package \nclient / CI] --> RP[Reverse \nproxy / network]
RP --> NX[Nexus request \npath]
NX --> AUTH[Authorization + repository routing]
NX --> DB[(H2 / PostgreSQL)]
NX --> BS[(Blob store)]
NX --> UP[Remote upstream for proxy]
NX --> T[Scheduled tasks]
NX --> JVM[JVM heap + direct memory]
NX --> LOG[request / nexus / audit logs]
NX --> MET[Prometheus + service metrics]
DB --> DBL[PostgreSQL logs / pool evidence]
BS --> CAP[Capacity / latency / quota evidence]
LOG --> CORR[Correlated \nincident \ntimeline]
MET --> CORR
DBL --> CORR
CAP --> CORR

The diagram is intentionally not a “Nexus internals” diagram. It is a diagnostic map. The same HTTP response is influenced by routing/authentication, metadata lookups, blob IO, upstream behavior, tasks, JVM resources, and network paths. A useful investigation preserves evidence from all relevant planes before changing state.

3. Current self-hosted log families

Current Sonatype documentation places self-hosted logs under $data-dir/log/. The important files have different questions:

Evidence Primary question Typical signal
nexus.log What was the application doing? Startup/shutdown, warnings, exceptions, subsystem activity.
request.log What HTTP request reached Nexus? Timestamp, client IP/user, method/path, status, bytes, response duration.
outbound-request.log What did Nexus ask an upstream? Remote URL/method/status/bytes/timing for outbound calls.
jvm.log What did the JVM emit outside ordinary app logging? stdout/stderr and explicitly triggered thread dumps.
log/audit/audit.log Who or what changed configuration/content? JSON audit records for configuration and component/asset modifications.
Task logs/history Was background maintenance competing or failing? Task start/end/status/duration and task-specific errors.
PostgreSQL logs Was the external DB slow, blocked, or connection-constrained? Slow statements, locks, connections/disconnections, checkpoints/autovacuum.

Do not replace one evidence source with another. request.log can prove a 12-second response without telling you whether the delay was database, blob, upstream, or queueing. nexus.log can contain an exception without proving that exception affected the user’s request.

4. Audit log is change evidence, not request tracing

The Audit capability is created and enabled by default in current Nexus Repository. It records configuration changes and component/asset additions or removals as JSON events under $data-dir/log/audit/, with daily rotation and a current maximum retention of 90 days. That makes it valuable for questions such as “who changed this repository?” or “when was this component deleted?”

It does not replace request.log. A read-only package download may be operationally important while producing no configuration-change audit event. Conversely, a configuration edit can be visible in audit evidence even if the package client never issued a repository request.

5. Metrics: health trends versus individual requests

Metrics answer aggregate questions: Is request rate rising? Is component count growing? Is memory pressure persistent? Are errors increasing? Current 3.81+ Nexus exposes Prometheus-format metrics at /service/rest/metrics/prometheus with the nx-metrics-all privilege. The Service Metrics Data API uses /service/rest/metrics/data in 3.81+ and requires nexus:metrics:read; Sonatype explicitly notes that this API is not shown in the instance Swagger documentation.

# Read-only examples against a disposable/local instance.
# Inject credentials through your shell/secret store; do not paste real tokens into history.
curl -fsS -u "$NEXUS_USER:$NEXUS_PASS" \
  http://127.0.0.1:8081/service/rest/metrics/prometheus \
  -o metrics.prom

curl -fsS -u "$NEXUS_USER:$NEXUS_PASS" \
  http://127.0.0.1:8081/service/rest/metrics/data \
  -o service-metrics.json

Prometheus is best for time series; a one-time scrape is not a baseline. The Service Metrics Data API includes usage-oriented gauges such as component totals and request counts. Neither substitutes for client/request timing when diagnosing one slow transaction.

6. Status endpoints are deliberately narrow

/service/rest/v1/status indicates whether the application can serve reads, and /service/rest/v1/status/writable checks read/write availability. Sonatype explicitly warns that these are not hardware checks and do not validate an external database or disk. Therefore:

  • 200 status: useful application signal.
  • Not proven: healthy object-store latency, adequate disk headroom, PostgreSQL failover readiness, or upstream registry health.

This distinction prevents a common operational mistake: treating one green endpoint as proof that every dependency is healthy.

7. Heap, direct memory, and host memory are separate budgets

Heap (-Xmx) stores ordinary JVM objects. Direct memory is outside heap and is used for off-heap buffers and other native-memory work. The operating system also needs memory for filesystem/network buffers and other processes. Increasing heap can therefore make an off-heap or host-pressure problem worse.

Documentation nuance. Sonatype’s current memory page gives both a high-level rule that maximum heap plus maximum direct memory should stay within roughly two-thirds of physical RAM and example/percentage guidance that must be interpreted in context. Do not turn those numbers into a universal formula. Use the current profile/reference architecture, exact Nexus/database topology, measured memory behavior, and verified headroom for your deployment.

For a diagnostic, preserve at minimum: host RAM, -Xms, -Xmx, -XX:MaxDirectMemorySize, process RSS, swap/pressure evidence, GC behavior if available, and workload rate. “Heap is 90% used” without collection behavior or direct-memory/host context is insufficient.

8. PostgreSQL connection pool and query latency

External PostgreSQL is its own service with its own capacity limits. Nexus currently uses a default pool size of 100 connections; increasing it can be appropriate under measured heavy load, but the database server must be sized to support the aggregate pools for all Nexus nodes plus administrative/tooling headroom. A pool increase without database capacity can merely move the bottleneck.

Sonatype’s PostgreSQL guidance recommends useful logging such as connection/disconnection, lock waits, and statements exceeding a chosen duration. Correlate the request timestamp with DB evidence before changing pool size or query-related settings.

9. Blob IO and disk capacity

Blob stores contain artifact bytes; their latency can dominate uploads/downloads even when CPU and heap are healthy. The Blob Stores UI exposes state, blob count, and approximate used size. Soft quotas can create warnings but do not block reads/writes, and running out of storage while saving content can damage writes. Capacity monitoring is therefore a reliability control, not housekeeping.

For object stores, network locality matters. Current Sonatype storage guidance warns that cross-region storage produces unacceptable latency and that object storage can be materially slower than file-based blob storage. Record backend type and region/locality whenever reporting timings.

10. A capacity profile is part of every benchmark

{
  "nexusVersion": "3.95.2-01",
  "java": "21",
  "edition": "Community-or-Pro-record-exactly",
  "database": "H2-or-PostgreSQL",
  "blobBackend": "file-or-supported-object-store",
  "nodes": 1,
  "requestRate": "measured, not guessed",
  "cacheState": "cold-or-warm",
  "taskLoad": "none-or-listed",
  "client": "name-and-version",
  "testArtifact": "synthetic coordinate/digest",
  "networkPath": "loopback/LAN/proxy/upstream",
  "window": "timestamp + duration"
}

Without this context, comparing “300 ms yesterday” to “800 ms today” may compare different workloads, cache states, clients, or storage paths rather than a true regression.

11. Read-only preflight

  1. Record Nexus version, edition, Java, database, blob backend, node count, and current tasks.
  2. Check application status endpoints.
  3. Confirm request.log is receiving new entries using a harmless request.
  4. Confirm Prometheus/service-metrics access with a scoped read-only identity.
  5. Inspect blob state/used size and OS free space.
  6. For PostgreSQL, capture connection/pool and slow-query evidence from the DB side.
  7. Record current log level/retention controls before enabling extra verbosity.

12. DevOps operating principle

An observable artifact platform makes incidents reviewable and capacity decisions defensible. The desired loop is request → evidence → hypothesis → one controlled change → repeated measurement. That is the same discipline used for build reproducibility and artifact promotion: preserve identity and evidence before changing state.

Knowledge check

Why is “Nexus is slow” not a root cause?

What is the current Prometheus endpoint for Nexus 3.81+?

Can a 200 from /service/rest/v1/status prove PostgreSQL and disk are healthy?

Why can increasing heap make performance worse?

What should every timing result include?

Summary and next step

You now have a state model for request, application, audit, JVM, database, blob, task, and usage evidence and know why no single metric proves a root cause.

Lesson 2 turns this model into a disposable workflow that produces controlled traffic, captures evidence, inspects a support ZIP safely, and compares cold versus warm request behavior.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.