Logging, Metrics, Support ZIPs, Request Auditing, Performance Tuning, and Capacity Diagnostics: Configuration, Design Choices, and Tradeoffs
Performance configuration is an engineering tradeoff, not a list of “faster” settings. Verbose logs improve evidence but consume IO and may expose more data; larger memory allocations can starve the OS; larger PostgreSQL pools can overload the database; aggressive caching can improve latency while changing freshness behavior. This lesson turns evidence into design choices.
Learning objectives
- Choose logging verbosity and retention according to evidence value, overhead, and privacy.
- Reason about heap, direct memory, and OS headroom as separate budgets.
- Explain why PostgreSQL pool size must be coordinated with server max_connections and node count.
- Distinguish vertical scaling, HA, cache tuning, and storage/network improvements.
- Define capacity thresholds from trends and recovery headroom rather than disk-full emergencies.
1. Tuning starts with a falsifiable hypothesis
A good tuning statement has four parts: observed evidence, suspected bottleneck, one change, and success criterion. Example: “p95 request latency rises only when blob read latency rises above 80 ms; move the disposable blob fixture to low-latency local storage; success is p95 below 250 ms under the same request mix.”
“Increase heap to make Nexus faster” has no causal evidence and no success criterion.
2. Logging verbosity versus overhead and privacy
Changing a logger to DEBUG/TRACE can reveal a failure path, but high-volume logging creates disk writes, larger support bundles, and more sensitive material. The production pattern is:
- Capture current logger configuration.
- Enable additional verbosity only for the narrow subsystem and incident window.
- Reproduce once or for a bounded period.
- Collect/redact evidence.
- Restore the previous level and verify normal log volume.
Never leave “everything DEBUG” as a permanent observability strategy.
3. Audit detail: preserve accountability without uncontrolled growth
Audit logging is valuable for configuration/content changes, but
format metadata can generate very large
attribute.changes entries. Current 3.95 runtime
guidance provides
nexus.audit.attribute.changes.enabled=false to suppress
that field for all formats when excessive growth is a problem. That
changes evidence content, so treat it as a governance decision:
document why it was disabled, what alternative evidence remains, and
how incident investigations will compensate.
4. Memory design: heap is not total process memory
| Budget | Pressure symptom | Wrong shortcut |
|---|---|---|
| JVM heap | Frequent/long GC, allocation pressure, OOM in heap | Raise -Xmx without host headroom evidence |
| Direct memory | Off-heap allocation failures/native pressure | Assume low heap usage means memory is healthy |
| OS/file/network buffers | Swap, reclaim pressure, poor IO caching | Allocate nearly all RAM to JVM |
The current Sonatype memory page includes both high-level host-headroom guidance and specific example settings. Use the exact current profile/reference architecture plus measured process/host evidence. Do not infer that “more memory is always faster.”
5. PostgreSQL pool versus database capacity
Nexus currently defaults to a 100-connection pool. A three-node HA
deployment can therefore demand roughly 300 Nexus connections before
DBA tooling and operational buffer. Sonatype’s max-connections
example uses 350 for three nodes at the default pool size.
Increasing maximumPoolSize changes concurrency pressure
on PostgreSQL and must be paired with server CPU/memory/connection
capacity evidence.
# Example only: use current supported configuration and measured need.
# Do not copy this value into production merely because it is larger.
maximumPoolSize=200
Pool exhaustion and slow queries are different problems. If all connections are busy because queries are slow, increasing the pool can increase DB contention rather than improve throughput.
6. CPU versus IO bottleneck
A CPU-bound workload shows sustained compute saturation correlated with latency while DB/blob/network waits remain reasonable. An IO-bound workload shows time spent waiting on database, blob, network, or disk. The correction must follow the dominant wait:
- High CPU + low IO wait: profile workload, scale compute, or reduce expensive work.
- Low/moderate CPU + high blob latency: fix storage locality/performance.
- Moderate Nexus CPU + slow PostgreSQL: fix database/query/connection bottleneck.
- Proxy-only slowness + high remote time: investigate upstream/network/cache freshness.
7. Vertical scaling versus HA
More CPU/RAM on one node can increase single-node capacity but does not create application-node availability. HA adds active nodes and failure tolerance but also multiplies database connections, local logs, operational complexity, and cost. Choose based on measured throughput and availability objectives, not because “cluster” sounds more scalable.
8. Cache tuning versus upstream behavior
Proxy caches reduce repeated upstream work, but metadata freshness, negative cache, auto-blocking, routing rules, and package-client caches all influence request patterns. If a public upstream is slow, extending a cache/freshness value may reduce remote calls but also changes how quickly upstream updates are observed. Record the artifact/metadata type and freshness requirement before tuning.
9. Blob storage: latency, capacity, and recovery are linked
Fast storage with no headroom is not healthy. Capacity planning must reserve space for current content, growth, soft-deleted/reclaimable bytes, temporary task/upgrade work, logs, backups/staging where relevant, and operational margin. Monitor both used and remaining space, and forecast time-to-threshold from a trend rather than waiting for an alert at the edge.
10. Retained logs versus disk budget
Audit logs have a documented 90-day maximum retention; other logs have their own rotation/retention behavior. Your evidence-retention policy must account for local disk, centralized log shipping (if used), privacy/legal requirements, and incident lookback. Do not delete logs manually during an incident merely to free disk without first preserving the evidence and fixing capacity pressure.
11. Decision table
| Observed pattern | Likely next design action | Do not do first |
|---|---|---|
| Heap/GC healthy; direct-memory alert | Inspect direct-memory/native allocation and host headroom | Raise heap |
| DB pool wait + DB max connections near limit | Measure query duration, DB resources, pool/server sizing together | Raise pool alone |
| Cold proxy slow, warm proxy fast | Investigate upstream/network/freshness and expected cache behavior | Scale CPU blindly |
| Blob latency rises with p95 request latency | Improve storage locality/performance or capacity | Increase client retries indefinitely |
| Disk growth dominated by debug/audit logs | Bound verbosity/retention and preserve required audit evidence | Delete random log files during incident |
12. Worked capacity forecast
Suppose a disposable blob store grows from 120 GB to 150 GB in 30 days: roughly 1 GB/day net growth. A 200 GB operational threshold is about 50 days away if growth remains linear. But cleanup cadence, release bursts, proxy cache behavior, and build volume can break linearity. The right output is not “50 days exactly”; it is “current trend reaches threshold in ~50 days, reforecast weekly, reserve maintenance/recovery headroom, and investigate dominant growth categories.”
used_start_gb = 120
used_end_gb = 150
days = 30
threshold_gb = 200
rate = (used_end_gb-used_start_gb)/days
eta_days = (threshold_gb-used_end_gb)/rate
print(rate, eta_days)
assert rate == 1.0
assert eta_days == 50.0
13. Compatibility and edition checklist
- Metrics endpoint changed in 3.81+: record Nexus version before reusing monitoring configs.
- HA changes how many node-local logs/support bundles and DB connections exist.
- H2 has current workload limits; if capacity exceeds those limits, “tuning H2 harder” is not a supported scaling plan.
- Object-store behavior depends on backend, region/locality, SDK/support matrix, and network.
- Pro features can add operational capabilities, but mandatory observability does not require a paid product.
Knowledge check
Why is a larger PostgreSQL pool not automatically faster?
It increases concurrent DB demand; if queries or DB resources are the bottleneck, a larger pool can increase contention and exhaust server connections.
What does disabling audit attribute.changes trade away?
It reduces potentially excessive audit-log growth but removes detail from audit evidence, so the governance/compensating-evidence decision must be documented.
What evidence distinguishes CPU-bound from blob-IO-bound latency?
Correlated request latency with sustained CPU saturation versus correlated blob/storage wait latency while CPU remains moderate.
Does HA replace vertical scaling?
No. HA primarily addresses service continuity/failure tolerance; it also changes capacity and operational topology. Choose both from measured throughput and availability goals.
Why forecast time-to-capacity threshold instead of waiting for disk-full?
Artifact writes and maintenance need headroom; disk exhaustion can cause failed/corrupted writes and leaves too little room for recovery/upgrade work.
Summary and next step
Tuning choices now connect explicit evidence to bounded logging, memory, database, storage, cache, scaling, and retention decisions rather than one-size-fits-all settings.
Lesson 4 deliberately breaks the most common assumptions—heap, disk, support-bundle privacy, storage, request evidence, and database pools—and diagnoses each from preserved evidence.
Official references and version notes
- Sonatype: Logging — current self-hosted log files, rotation concepts, Log Viewer behavior, request/outbound/audit evidence, and PostgreSQL logging boundary.
- Sonatype: Auditing — audit capability, JSON event records, daily rotation, and current 90-day maximum retention.
- Sonatype: Prometheus — current 3.81+ Prometheus endpoint and required metrics privilege.
- Sonatype: Service Metrics Data API — current component/request usage metrics endpoint, privilege requirement, and the fact that this API is not listed in instance Swagger.
- Sonatype: Status API — read/writable application state and explicit limits: these checks do not validate external DB or disk health.
- Sonatype: Support Features — Support ZIP creation, storage location, and HA per-node bundle behavior.
- Sonatype: Support API — support ZIP API payload categories including system information, thread dump, metrics, configuration, security, logs, audit logs, and JMX.
- Sonatype: Nexus Repository Memory Overview — heap, direct memory, host headroom, and memory-related JVM arguments.
- Sonatype: System Requirements — current workload profiles, H2 limits, PostgreSQL guidance, CPU/RAM/storage baselines, and low-latency DB requirements.
- Sonatype: PostgreSQL Installation/Configuration — slow-query logging guidance and connection-pool configuration.
- Sonatype: PostgreSQL Max Connections — sizing the server connection limit against Nexus node pool sizes and troubleshooting exhausted connections.
- Sonatype: Blob Stores — blob count/used-size status, soft quotas, and storage-health evidence.
- Sonatype: Storage Planning — locality/latency implications and object-store performance tradeoffs.
- Sonatype: Usage Metrics — current usage-center request/component measurements and edition differences.
- Sonatype: Self-Hosted Usage Guide — request-log fallback, per-node aggregation, and request-metric interpretation.
- Sonatype: Configuring the Runtime Environment — current runtime properties, including the 3.95 audit-log attribute-change control.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.