Chapter 29Lesson 04250–330 min

Logging, Metrics, Support ZIPs, Request Auditing, Performance Tuning, and Capacity Diagnostics: Diagnostics, Failure Modes, Security, and Performance

Performance incidents become dangerous when operators change multiple resource settings before preserving evidence. This lesson uses deliberately broken cases to practice the chapter diagnostic sequence: preserve the request and topology, identify the dominant wait or exhaustion, make the least-destructive correction, and repeat the same measurement.

DiagnosticsFailure analysisPool exhaustionDisk pressureLeast-destructive fix

Learning objectives

  • Diagnose blind heap tuning and direct-memory pressure.
  • Detect support-bundle privacy risk before external sharing.
  • Distinguish storage latency from CPU saturation.
  • Recover observability when request/audit evidence is missing or incomplete.
  • Diagnose disk thresholds and PostgreSQL pool/connection exhaustion without unsafe shortcuts.
Dated baseline (27 August 2026). Lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. Re-check the live release notes, logging/metrics pages, system requirements, and instance-specific configuration before applying production tuning.
Measure before tuning. A slow package request is an end-to-end symptom, not proof of a JVM problem. Preserve the request, Nexus version/edition/runtime, cache state, database latency/pool state, blob latency/capacity, task activity, network/upstream timing, and client-cache state before changing a knob.
Privacy boundary. Logs and support bundles can contain usernames, client IPs, internal repository names/URLs, request paths, configuration, security information, dependency coordinates, and other operationally sensitive evidence. Sonatype removes password-related information from generated support files, but operators must still review/redact and share on a need-to-know basis.

1. The diagnostic sequence

Use the same sequence from the master contract, narrowed to observability:

  1. Preserve the exact request/client timestamp/status/duration.
  2. Record Nexus version, edition, Java, DB, blob backend, node count, and cache state.
  3. Confirm repository URL/auth/routing before treating a 401/404 as “performance.”
  4. Inspect request/application/outbound/audit/task evidence.
  5. Inspect JVM heap/direct memory and host headroom.
  6. Inspect DB pool/slow-query/lock/connection evidence.
  7. Inspect blob latency/capacity/soft quota and network/upstream timing.
  8. Change one smallest variable.
  9. Repeat the same controlled request and compare.

2. Broken case A: blindly increasing heap

Observed:
  request p95: 1800 ms
  CPU: 32%
  heap after GC: 48%
  direct memory: 96% of configured max
  blob latency p95: 22 ms
  DB query p95: 18 ms
Proposed change: Xmx 8G -> 16G

The proposal is not supported. Heap is not the pressure signal; direct memory is. Raising heap can reduce host headroom available to direct/native memory. First inspect direct-memory use, buffers/workload, current MaxDirectMemorySize, host RSS/swap pressure, and exact Sonatype memory guidance for the profile.

3. Broken case B: support ZIP shared without review

An operator creates a support ZIP and uploads it to a public issue because “passwords are removed.” The archive contains internal repository URLs, usernames, node IDs, package coordinates, and security/configuration context.

Correction: remove the public copy, treat the disclosure according to incident policy, rotate any potentially exposed secrets even if not expected, and re-create a minimally scoped/redacted evidence packet through the approved support channel. Password sanitization is not comprehensive data classification.

4. Broken case C: slow storage blamed on CPU

{
  "request_ms": 2400,
  "cpu_pct": 41,
  "heap_pct": 55,
  "db_ms": 23,
  "blob_read_ms": 2180,
  "upstream_ms": 0,
  "task_overlap": false
}

The blob read explains most of the request. Investigate backend health, same-region/locality, disk/object-store latency, throughput, network, and capacity. Scaling CPU cannot remove a 2.18-second storage wait.

5. Broken case D: request evidence disappears during the incident

A customized request-log configuration stopped working after an upgrade, or someone disabled it to reduce disk writes during the incident. Now client latency cannot be correlated to Nexus response timing.

Correction: preserve remaining evidence, verify current request-logging configuration against the deployed Jetty/Nexus version, restore bounded request logging, and compensate with reverse-proxy/client logs for the missing window. Do not claim exact Nexus request latency for a window that has no Nexus request evidence.

6. Broken case E: disk threshold reached

Blob storage or the data/log filesystem is almost full. The unsafe reaction is to delete files directly from blob-store directories or purge random database/log files. Current supported response:

  1. Stop or throttle nonessential writes if needed through supported operational controls.
  2. Preserve incident/support evidence.
  3. Identify which filesystem/backend is full: application data/logs, blob store, DB storage, temp/backup area.
  4. Expand supported capacity or perform supported cleanup/reclamation only after preview/backup context.
  5. Validate repository content after recovery.

Direct blob/database deletion can turn capacity pressure into content corruption.

7. Broken case F: PostgreSQL pool exhausted

Nexus nodes: 3
maximumPoolSize per node: 100
PostgreSQL max_connections: 250
Observed DB error: FATAL: sorry, too many clients already

Even before DBA/tooling connections, the Nexus theoretical pool demand can exceed the DB limit. Fix the system: measure actual pool utilization/query duration, set PostgreSQL max_connections with adequate server resources and headroom, and keep node pool sizing coherent. Merely reducing client traffic or blindly increasing Nexus pools does not correct the configuration mismatch.

8. Broken case G: audit log grows unexpectedly

PyPI/NuGet metadata updates generate extremely large attribute.changes entries and the audit filesystem grows rapidly. First confirm that this field is the dominant source. Current 3.95 guidance provides a property to suppress it globally. If used, document the lost detail and retain configuration/content-change accountability through the remaining audit fields and complementary evidence.

9. Broken case H: status is green, users still fail

/service/rest/v1/status returns 200, but artifact writes fail because the blob backend is out of capacity. This is not contradictory: Sonatype states the status API does not validate external DB/disk health. Add dedicated storage/database monitoring rather than redefining 200 as “everything is healthy.”

10. A deterministic diagnosis fixture

cases = {
 "heap_folklore": {"heap":48,"direct":96,"cpu":32,"db":18,"blob":22,"remote":0},
 "blob_wait": {"heap":55,"direct":48,"cpu":41,"db":23,"blob":2180,"remote":0},
 "remote_wait": {"heap":51,"direct":44,"cpu":33,"db":20,"blob":30,"remote":1700},
}

def dominant(m):
    waits={"database":m["db"],"blob":m["blob"],"upstream":m["remote"]}
    if m["direct"] >= 90: return "direct-memory-pressure"
    name,value=max(waits.items(), key=lambda kv:kv[1])
    return name if value > 500 else "no-dominant-wait-yet"

for name,m in cases.items():
    print(name, "=>", dominant(m))
assert dominant(cases["heap_folklore"]) == "direct-memory-pressure"
assert dominant(cases["blob_wait"]) == "blob"
assert dominant(cases["remote_wait"]) == "upstream"

The fixture is deliberately simple. Its purpose is to enforce the reasoning order: choose the next evidence/control from the dominant signal, not from a favorite tuning knob.

11. Security-sensitive observability actions

  • Never turn on logging that writes Authorization headers, passwords, private keys, or tokens for convenience.
  • Do not expose Prometheus or support endpoints anonymously to the public internet.
  • Use scoped monitoring identities (nx-metrics-all/nexus:metrics:read as applicable), not shared administrator credentials.
  • Thread dumps, audit logs, request logs, and configuration exports are access-controlled operational data.
  • When sharing evidence, keep an immutable internal original and a redacted external copy with a manifest of redactions.

12. Performance acceptance after a correction

A correction is not complete when the alert disappears. Repeat the same synthetic request mix and compare:

  • p50/p95/p99 or another stated latency statistic;
  • error rate/status distribution;
  • CPU/heap/direct-memory/host pressure;
  • DB pool/query/lock indicators;
  • blob/upstream latency;
  • task overlap and log volume;
  • artifact digest/content correctness.

If several variables changed, you cannot confidently attribute improvement to one cause.

Knowledge check

Heap is 48% but direct memory is 96%. Should you increase Xmx first?

Why can a green Status API coexist with failed artifact writes?

Three nodes each have pool size 100 but PostgreSQL max_connections is 250. What is wrong?

Why not delete blob files when disk is nearly full?

How do you prove a performance change helped?

Summary and next step

You diagnosed eight common observability/performance failures by preserving evidence, respecting security boundaries, identifying the dominant signal, and applying a least-destructive correction.

Lesson 5 turns the entire chapter into a baseline → bottleneck → one-change → remeasure checkpoint and produces the observability/capacity handoff needed for the production capstone.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.