Logging, Metrics, Support ZIPs, Request Auditing, Performance Tuning, and Capacity Diagnostics: Diagnostics, Failure Modes, Security, and Performance
Performance incidents become dangerous when operators change multiple resource settings before preserving evidence. This lesson uses deliberately broken cases to practice the chapter diagnostic sequence: preserve the request and topology, identify the dominant wait or exhaustion, make the least-destructive correction, and repeat the same measurement.
Learning objectives
- Diagnose blind heap tuning and direct-memory pressure.
- Detect support-bundle privacy risk before external sharing.
- Distinguish storage latency from CPU saturation.
- Recover observability when request/audit evidence is missing or incomplete.
- Diagnose disk thresholds and PostgreSQL pool/connection exhaustion without unsafe shortcuts.
1. The diagnostic sequence
Use the same sequence from the master contract, narrowed to observability:
- Preserve the exact request/client timestamp/status/duration.
- Record Nexus version, edition, Java, DB, blob backend, node count, and cache state.
- Confirm repository URL/auth/routing before treating a 401/404 as “performance.”
- Inspect request/application/outbound/audit/task evidence.
- Inspect JVM heap/direct memory and host headroom.
- Inspect DB pool/slow-query/lock/connection evidence.
- Inspect blob latency/capacity/soft quota and network/upstream timing.
- Change one smallest variable.
- Repeat the same controlled request and compare.
2. Broken case A: blindly increasing heap
Observed:
request p95: 1800 ms
CPU: 32%
heap after GC: 48%
direct memory: 96% of configured max
blob latency p95: 22 ms
DB query p95: 18 ms
Proposed change: Xmx 8G -> 16G
The proposal is not supported. Heap is not the pressure signal;
direct memory is. Raising heap can reduce host headroom available to
direct/native memory. First inspect direct-memory use,
buffers/workload, current MaxDirectMemorySize, host
RSS/swap pressure, and exact Sonatype memory guidance for the
profile.
3. Broken case B: support ZIP shared without review
An operator creates a support ZIP and uploads it to a public issue because “passwords are removed.” The archive contains internal repository URLs, usernames, node IDs, package coordinates, and security/configuration context.
Correction: remove the public copy, treat the disclosure according to incident policy, rotate any potentially exposed secrets even if not expected, and re-create a minimally scoped/redacted evidence packet through the approved support channel. Password sanitization is not comprehensive data classification.
4. Broken case C: slow storage blamed on CPU
{
"request_ms": 2400,
"cpu_pct": 41,
"heap_pct": 55,
"db_ms": 23,
"blob_read_ms": 2180,
"upstream_ms": 0,
"task_overlap": false
}
The blob read explains most of the request. Investigate backend health, same-region/locality, disk/object-store latency, throughput, network, and capacity. Scaling CPU cannot remove a 2.18-second storage wait.
5. Broken case D: request evidence disappears during the incident
A customized request-log configuration stopped working after an upgrade, or someone disabled it to reduce disk writes during the incident. Now client latency cannot be correlated to Nexus response timing.
Correction: preserve remaining evidence, verify current request-logging configuration against the deployed Jetty/Nexus version, restore bounded request logging, and compensate with reverse-proxy/client logs for the missing window. Do not claim exact Nexus request latency for a window that has no Nexus request evidence.
6. Broken case E: disk threshold reached
Blob storage or the data/log filesystem is almost full. The unsafe reaction is to delete files directly from blob-store directories or purge random database/log files. Current supported response:
- Stop or throttle nonessential writes if needed through supported operational controls.
- Preserve incident/support evidence.
- Identify which filesystem/backend is full: application data/logs, blob store, DB storage, temp/backup area.
- Expand supported capacity or perform supported cleanup/reclamation only after preview/backup context.
- Validate repository content after recovery.
Direct blob/database deletion can turn capacity pressure into content corruption.
7. Broken case F: PostgreSQL pool exhausted
Nexus nodes: 3
maximumPoolSize per node: 100
PostgreSQL max_connections: 250
Observed DB error: FATAL: sorry, too many clients already
Even before DBA/tooling connections, the Nexus theoretical pool
demand can exceed the DB limit. Fix the system: measure
actual pool utilization/query duration, set PostgreSQL
max_connections with adequate server resources and
headroom, and keep node pool sizing coherent. Merely reducing client
traffic or blindly increasing Nexus pools does not correct the
configuration mismatch.
8. Broken case G: audit log grows unexpectedly
PyPI/NuGet metadata updates generate extremely large
attribute.changes entries and the audit filesystem
grows rapidly. First confirm that this field is the dominant source.
Current 3.95 guidance provides a property to suppress it globally.
If used, document the lost detail and retain
configuration/content-change accountability through the remaining
audit fields and complementary evidence.
9. Broken case H: status is green, users still fail
/service/rest/v1/status returns 200, but artifact
writes fail because the blob backend is out of capacity. This is not
contradictory: Sonatype states the status API does not validate
external DB/disk health. Add dedicated storage/database monitoring
rather than redefining 200 as “everything is healthy.”
10. A deterministic diagnosis fixture
cases = {
"heap_folklore": {"heap":48,"direct":96,"cpu":32,"db":18,"blob":22,"remote":0},
"blob_wait": {"heap":55,"direct":48,"cpu":41,"db":23,"blob":2180,"remote":0},
"remote_wait": {"heap":51,"direct":44,"cpu":33,"db":20,"blob":30,"remote":1700},
}
def dominant(m):
waits={"database":m["db"],"blob":m["blob"],"upstream":m["remote"]}
if m["direct"] >= 90: return "direct-memory-pressure"
name,value=max(waits.items(), key=lambda kv:kv[1])
return name if value > 500 else "no-dominant-wait-yet"
for name,m in cases.items():
print(name, "=>", dominant(m))
assert dominant(cases["heap_folklore"]) == "direct-memory-pressure"
assert dominant(cases["blob_wait"]) == "blob"
assert dominant(cases["remote_wait"]) == "upstream"
The fixture is deliberately simple. Its purpose is to enforce the reasoning order: choose the next evidence/control from the dominant signal, not from a favorite tuning knob.
11. Security-sensitive observability actions
- Never turn on logging that writes Authorization headers, passwords, private keys, or tokens for convenience.
- Do not expose Prometheus or support endpoints anonymously to the public internet.
-
Use scoped monitoring identities
(
nx-metrics-all/nexus:metrics:readas applicable), not shared administrator credentials. - Thread dumps, audit logs, request logs, and configuration exports are access-controlled operational data.
- When sharing evidence, keep an immutable internal original and a redacted external copy with a manifest of redactions.
12. Performance acceptance after a correction
A correction is not complete when the alert disappears. Repeat the same synthetic request mix and compare:
- p50/p95/p99 or another stated latency statistic;
- error rate/status distribution;
- CPU/heap/direct-memory/host pressure;
- DB pool/query/lock indicators;
- blob/upstream latency;
- task overlap and log volume;
- artifact digest/content correctness.
If several variables changed, you cannot confidently attribute improvement to one cause.
Knowledge check
Heap is 48% but direct memory is 96%. Should you increase Xmx first?
No. The evidence points to off-heap/direct-memory pressure; inspect direct/native memory and host headroom first.
Why can a green Status API coexist with failed artifact writes?
The Status API reports Nexus application state and does not validate external DB/disk/blob health; capacity/storage can fail independently.
Three nodes each have pool size 100 but PostgreSQL max_connections is 250. What is wrong?
The possible Nexus pool demand already exceeds the server limit before DBA/tooling headroom; pool and DB max_connections/resources must be designed together.
Why not delete blob files when disk is nearly full?
Direct blob deletion is unsupported and can create metadata/content inconsistency. Use supported capacity expansion and cleanup/reclamation with backup/verification.
How do you prove a performance change helped?
Repeat the same controlled workload with the same cache/topology conditions and compare stated metrics while changing only one intended variable.
Summary and next step
You diagnosed eight common observability/performance failures by preserving evidence, respecting security boundaries, identifying the dominant signal, and applying a least-destructive correction.
Lesson 5 turns the entire chapter into a baseline → bottleneck → one-change → remeasure checkpoint and produces the observability/capacity handoff needed for the production capstone.
Official references and version notes
- Sonatype: Logging — current self-hosted log files, rotation concepts, Log Viewer behavior, request/outbound/audit evidence, and PostgreSQL logging boundary.
- Sonatype: Auditing — audit capability, JSON event records, daily rotation, and current 90-day maximum retention.
- Sonatype: Prometheus — current 3.81+ Prometheus endpoint and required metrics privilege.
- Sonatype: Service Metrics Data API — current component/request usage metrics endpoint, privilege requirement, and the fact that this API is not listed in instance Swagger.
- Sonatype: Status API — read/writable application state and explicit limits: these checks do not validate external DB or disk health.
- Sonatype: Support Features — Support ZIP creation, storage location, and HA per-node bundle behavior.
- Sonatype: Support API — support ZIP API payload categories including system information, thread dump, metrics, configuration, security, logs, audit logs, and JMX.
- Sonatype: Nexus Repository Memory Overview — heap, direct memory, host headroom, and memory-related JVM arguments.
- Sonatype: System Requirements — current workload profiles, H2 limits, PostgreSQL guidance, CPU/RAM/storage baselines, and low-latency DB requirements.
- Sonatype: PostgreSQL Installation/Configuration — slow-query logging guidance and connection-pool configuration.
- Sonatype: PostgreSQL Max Connections — sizing the server connection limit against Nexus node pool sizes and troubleshooting exhausted connections.
- Sonatype: Blob Stores — blob count/used-size status, soft quotas, and storage-health evidence.
- Sonatype: Storage Planning — locality/latency implications and object-store performance tradeoffs.
- Sonatype: Usage Metrics — current usage-center request/component measurements and edition differences.
- Sonatype: Self-Hosted Usage Guide — request-log fallback, per-node aggregation, and request-metric interpretation.
- Sonatype: Configuring the Runtime Environment — current runtime properties, including the 3.95 audit-log attribute-change control.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.