Chapter 29Lesson 03220–300 min

Logging, Metrics, Support ZIPs, Request Auditing, Performance Tuning, and Capacity Diagnostics: Configuration, Design Choices, and Tradeoffs

Performance configuration is an engineering tradeoff, not a list of “faster” settings. Verbose logs improve evidence but consume IO and may expose more data; larger memory allocations can starve the OS; larger PostgreSQL pools can overload the database; aggressive caching can improve latency while changing freshness behavior. This lesson turns evidence into design choices.

TradeoffsMemoryDatabase poolLogging retentionCapacity planning

Learning objectives

  • Choose logging verbosity and retention according to evidence value, overhead, and privacy.
  • Reason about heap, direct memory, and OS headroom as separate budgets.
  • Explain why PostgreSQL pool size must be coordinated with server max_connections and node count.
  • Distinguish vertical scaling, HA, cache tuning, and storage/network improvements.
  • Define capacity thresholds from trends and recovery headroom rather than disk-full emergencies.
Dated baseline (27 August 2026). Lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. Re-check the live release notes, logging/metrics pages, system requirements, and instance-specific configuration before applying production tuning.
Measure before tuning. A slow package request is an end-to-end symptom, not proof of a JVM problem. Preserve the request, Nexus version/edition/runtime, cache state, database latency/pool state, blob latency/capacity, task activity, network/upstream timing, and client-cache state before changing a knob.
Privacy boundary. Logs and support bundles can contain usernames, client IPs, internal repository names/URLs, request paths, configuration, security information, dependency coordinates, and other operationally sensitive evidence. Sonatype removes password-related information from generated support files, but operators must still review/redact and share on a need-to-know basis.

1. Tuning starts with a falsifiable hypothesis

A good tuning statement has four parts: observed evidence, suspected bottleneck, one change, and success criterion. Example: “p95 request latency rises only when blob read latency rises above 80 ms; move the disposable blob fixture to low-latency local storage; success is p95 below 250 ms under the same request mix.”

“Increase heap to make Nexus faster” has no causal evidence and no success criterion.

2. Logging verbosity versus overhead and privacy

Changing a logger to DEBUG/TRACE can reveal a failure path, but high-volume logging creates disk writes, larger support bundles, and more sensitive material. The production pattern is:

  1. Capture current logger configuration.
  2. Enable additional verbosity only for the narrow subsystem and incident window.
  3. Reproduce once or for a bounded period.
  4. Collect/redact evidence.
  5. Restore the previous level and verify normal log volume.

Never leave “everything DEBUG” as a permanent observability strategy.

3. Audit detail: preserve accountability without uncontrolled growth

Audit logging is valuable for configuration/content changes, but format metadata can generate very large attribute.changes entries. Current 3.95 runtime guidance provides nexus.audit.attribute.changes.enabled=false to suppress that field for all formats when excessive growth is a problem. That changes evidence content, so treat it as a governance decision: document why it was disabled, what alternative evidence remains, and how incident investigations will compensate.

4. Memory design: heap is not total process memory

Budget Pressure symptom Wrong shortcut
JVM heap Frequent/long GC, allocation pressure, OOM in heap Raise -Xmx without host headroom evidence
Direct memory Off-heap allocation failures/native pressure Assume low heap usage means memory is healthy
OS/file/network buffers Swap, reclaim pressure, poor IO caching Allocate nearly all RAM to JVM

The current Sonatype memory page includes both high-level host-headroom guidance and specific example settings. Use the exact current profile/reference architecture plus measured process/host evidence. Do not infer that “more memory is always faster.”

5. PostgreSQL pool versus database capacity

Nexus currently defaults to a 100-connection pool. A three-node HA deployment can therefore demand roughly 300 Nexus connections before DBA tooling and operational buffer. Sonatype’s max-connections example uses 350 for three nodes at the default pool size. Increasing maximumPoolSize changes concurrency pressure on PostgreSQL and must be paired with server CPU/memory/connection capacity evidence.

# Example only: use current supported configuration and measured need.
# Do not copy this value into production merely because it is larger.
maximumPoolSize=200

Pool exhaustion and slow queries are different problems. If all connections are busy because queries are slow, increasing the pool can increase DB contention rather than improve throughput.

6. CPU versus IO bottleneck

A CPU-bound workload shows sustained compute saturation correlated with latency while DB/blob/network waits remain reasonable. An IO-bound workload shows time spent waiting on database, blob, network, or disk. The correction must follow the dominant wait:

  • High CPU + low IO wait: profile workload, scale compute, or reduce expensive work.
  • Low/moderate CPU + high blob latency: fix storage locality/performance.
  • Moderate Nexus CPU + slow PostgreSQL: fix database/query/connection bottleneck.
  • Proxy-only slowness + high remote time: investigate upstream/network/cache freshness.

7. Vertical scaling versus HA

More CPU/RAM on one node can increase single-node capacity but does not create application-node availability. HA adds active nodes and failure tolerance but also multiplies database connections, local logs, operational complexity, and cost. Choose based on measured throughput and availability objectives, not because “cluster” sounds more scalable.

8. Cache tuning versus upstream behavior

Proxy caches reduce repeated upstream work, but metadata freshness, negative cache, auto-blocking, routing rules, and package-client caches all influence request patterns. If a public upstream is slow, extending a cache/freshness value may reduce remote calls but also changes how quickly upstream updates are observed. Record the artifact/metadata type and freshness requirement before tuning.

9. Blob storage: latency, capacity, and recovery are linked

Fast storage with no headroom is not healthy. Capacity planning must reserve space for current content, growth, soft-deleted/reclaimable bytes, temporary task/upgrade work, logs, backups/staging where relevant, and operational margin. Monitor both used and remaining space, and forecast time-to-threshold from a trend rather than waiting for an alert at the edge.

10. Retained logs versus disk budget

Audit logs have a documented 90-day maximum retention; other logs have their own rotation/retention behavior. Your evidence-retention policy must account for local disk, centralized log shipping (if used), privacy/legal requirements, and incident lookback. Do not delete logs manually during an incident merely to free disk without first preserving the evidence and fixing capacity pressure.

11. Decision table

Observed pattern Likely next design action Do not do first
Heap/GC healthy; direct-memory alert Inspect direct-memory/native allocation and host headroom Raise heap
DB pool wait + DB max connections near limit Measure query duration, DB resources, pool/server sizing together Raise pool alone
Cold proxy slow, warm proxy fast Investigate upstream/network/freshness and expected cache behavior Scale CPU blindly
Blob latency rises with p95 request latency Improve storage locality/performance or capacity Increase client retries indefinitely
Disk growth dominated by debug/audit logs Bound verbosity/retention and preserve required audit evidence Delete random log files during incident

12. Worked capacity forecast

Suppose a disposable blob store grows from 120 GB to 150 GB in 30 days: roughly 1 GB/day net growth. A 200 GB operational threshold is about 50 days away if growth remains linear. But cleanup cadence, release bursts, proxy cache behavior, and build volume can break linearity. The right output is not “50 days exactly”; it is “current trend reaches threshold in ~50 days, reforecast weekly, reserve maintenance/recovery headroom, and investigate dominant growth categories.”

used_start_gb = 120
used_end_gb = 150
days = 30
threshold_gb = 200
rate = (used_end_gb-used_start_gb)/days
eta_days = (threshold_gb-used_end_gb)/rate
print(rate, eta_days)
assert rate == 1.0
assert eta_days == 50.0

13. Compatibility and edition checklist

  • Metrics endpoint changed in 3.81+: record Nexus version before reusing monitoring configs.
  • HA changes how many node-local logs/support bundles and DB connections exist.
  • H2 has current workload limits; if capacity exceeds those limits, “tuning H2 harder” is not a supported scaling plan.
  • Object-store behavior depends on backend, region/locality, SDK/support matrix, and network.
  • Pro features can add operational capabilities, but mandatory observability does not require a paid product.

Knowledge check

Why is a larger PostgreSQL pool not automatically faster?

What does disabling audit attribute.changes trade away?

What evidence distinguishes CPU-bound from blob-IO-bound latency?

Does HA replace vertical scaling?

Why forecast time-to-capacity threshold instead of waiting for disk-full?

Summary and next step

Tuning choices now connect explicit evidence to bounded logging, memory, database, storage, cache, scaling, and retention decisions rather than one-size-fits-all settings.

Lesson 4 deliberately breaks the most common assumptions—heap, disk, support-bundle privacy, storage, request evidence, and database pools—and diagnoses each from preserved evidence.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.