Checkpoint Lab — Logging, Metrics, Support ZIPs, Request Auditing, Performance Tuning, and Capacity Diagnostics
This checkpoint is a miniature production incident and capacity review. You will establish a clean baseline, inject one controlled bottleneck in synthetic telemetry, diagnose it from correlated request/DB/blob/JVM evidence, apply exactly one correction, prove improvement, and package the result as a privacy-reviewed operating handoff.
Learning objectives
- Build a complete baseline containing request, JVM, DB, blob, task, and capacity signals.
- Predict which signals change before injecting a controlled bottleneck.
- Diagnose the bottleneck without changing multiple variables.
- Prove the correction with repeated comparable measurements and artifact integrity checks.
- Produce a redacted evidence packet and capacity/alerting runbook for Chapter 30.
1. Scenario
You operate a synthetic repository service called
learner-platform. It serves a stable 10 MiB artifact
and a small metadata asset. The mandatory exercise uses generated
telemetry so no public service or real Nexus instance is stressed.
Optional live validation may use a disposable local Nexus with the
same evidence checklist.
Your goal is not to obtain the smallest possible latency. Your goal is to prove that you can explain a regression and improve it without causing a second unknown change.
2. Preflight
- Nexus reference: 3.95.2-01, Java 21; record actual version if live.
- Edition: mandatory simulation is Community-compatible.
- Database: synthetic PostgreSQL signals; a live H2 extension must respect H2 workload limits.
- Blob backend: synthetic file/object latency profile; live path must be disposable.
- Credentials: none in mandatory simulation; live metrics use scoped secrets injected at runtime.
- Safety: no production load test, no direct DB/blob edits, no public-registry stress, no disabling TLS/auth checks.
3. Define acceptance before the test
Use explicit synthetic objectives:
| Metric | Baseline objective | Why it matters |
|---|---|---|
| Request p95 | < 250 ms | User/CI experience under the stated workload |
| Error rate | 0% synthetic errors | Do not trade speed for correctness |
| Heap after GC | < 70% | Headroom signal, not sole performance target |
| Direct memory | < 80% | Off-heap headroom |
| DB p95 | < 50 ms | Metadata/configuration dependency |
| Blob p95 | < 60 ms | Artifact byte path |
| Free capacity | > 25% | Operational/recovery headroom |
These are exercise objectives, not Sonatype guarantees.
4. Predict two changes
Before injection, write predictions. Example bottleneck: blob latency.
- Prediction A: request p95 will rise roughly with blob p95 while CPU/heap/DB remain near baseline.
- Prediction B: artifact SHA-256 will remain unchanged because latency changes delivery time, not artifact identity.
5. Generate deterministic baseline and incident telemetry
import hashlib, statistics
artifact=b"learner-platform-release-1.0.0\n" * 400000
digest=hashlib.sha256(artifact).hexdigest()
baseline=[
{"req":95,"cpu":38,"heap":52,"direct":44,"db":18,"blob":31,"err":0},
{"req":110,"cpu":41,"heap":54,"direct":45,"db":20,"blob":36,"err":0},
{"req":125,"cpu":39,"heap":53,"direct":46,"db":19,"blob":42,"err":0},
{"req":105,"cpu":40,"heap":52,"direct":45,"db":17,"blob":34,"err":0},
]
incident=[
{"req":820,"cpu":42,"heap":54,"direct":47,"db":21,"blob":745,"err":0},
{"req":910,"cpu":43,"heap":55,"direct":48,"db":22,"blob":830,"err":0},
{"req":790,"cpu":41,"heap":54,"direct":47,"db":20,"blob":720,"err":0},
{"req":860,"cpu":42,"heap":55,"direct":48,"db":21,"blob":780,"err":0},
]
def p95(xs):
xs=sorted(xs)
return xs[max(0, int(len(xs)*0.95)-1)]
def summarize(rows):
return {k:p95([r[k] for r in rows]) for k in ["req","cpu","heap","direct","db","blob","err"]}
print("sha256", digest)
print("baseline", summarize(baseline))
print("incident", summarize(incident))
assert summarize(incident)["blob"] > 700
assert summarize(incident)["db"] < 50
assert summarize(incident)["cpu"] < 60
The small sample makes the percentile arithmetic easy to inspect; it is a teaching fixture, not a statistical production benchmark.
6. Diagnose before correcting
The incident request latency rises by hundreds of milliseconds. CPU moves only slightly, heap/direct memory remain stable, DB remains near baseline, errors remain zero, and blob latency rises by roughly the same amount as request latency. The dominant hypothesis is blob/storage latency.
Rejected changes: increasing heap, increasing PostgreSQL pool size, disabling request logs, or changing cleanup policy. None addresses the observed wait.
7. Apply exactly one correction
In the synthetic exercise, the one change is “move the blob service to the same low-latency locality as Nexus,” represented by corrected blob timing. In a real deployment this could mean fixing storage/network placement according to the supported architecture; do not literally migrate a production blob store as an ad-hoc performance experiment.
corrected=[
{"req":120,"cpu":42,"heap":54,"direct":47,"db":21,"blob":39,"err":0},
{"req":135,"cpu":43,"heap":55,"direct":48,"db":22,"blob":45,"err":0},
{"req":115,"cpu":41,"heap":54,"direct":47,"db":20,"blob":35,"err":0},
{"req":128,"cpu":42,"heap":55,"direct":48,"db":21,"blob":42,"err":0},
]
print("corrected", summarize(corrected))
assert summarize(corrected)["req"] < 250
assert summarize(corrected)["blob"] < 60
assert hashlib.sha256(artifact).hexdigest() == digest
8. Capacity trend exercise
history_gb=[118,121,124,128,132,136,141,146,151]
days_between=5
rates=[(history_gb[i]-history_gb[i-1])/days_between for i in range(1,len(history_gb))]
rate=sum(rates)/len(rates)
threshold=190
eta=(threshold-history_gb[-1])/rate
print("avg_growth_gb_per_day", round(rate,2))
print("days_to_threshold", round(eta,1))
assert rate > 0
Record the estimate as a forecast with assumptions, not a guarantee. Add actions at earlier warning thresholds so cleanup/capacity expansion can be planned with backup/recovery headroom.
9. Build the evidence packet
Create a small directory containing only:
chapter29-evidence/
00-context.txt # version, Java, edition, DB, blob, cache, client, time window
01-baseline.json # synthetic summary
02-incident.json # synthetic summary
03-diagnosis.md # evidence -> hypothesis; rejected alternatives
04-change.md # exactly one controlled correction
05-validation.json # corrected summary + SHA-256 proof
06-capacity.csv # dated usage trend + threshold assumptions
07-redaction-manifest.md# what was removed/pseudonymized and why
Do not include credentials, Authorization headers, private keys, real production URLs, or a complete unreviewed support ZIP.
10. Optional live support-bundle extension
On a disposable local Nexus, generate a support ZIP using the UI or
current API, review its selected categories, and produce only a
redacted manifest. Do not upload it anywhere for the course. Verify
that the live instance’s metrics/request logs agree with the context
captured in 00-context.txt.
11. Validation matrix
| Claim | Evidence | Pass |
|---|---|---|
| Bottleneck was blob latency | Request rise tracks blob rise while CPU/heap/DB remain stable | Yes |
| One change corrected it | Corrected run changes only synthetic blob locality/timing | Yes |
| Artifact identity preserved | SHA-256 identical before/after | Yes |
| Capacity is managed proactively | Trend and threshold ETA recorded | Yes |
| Evidence can be shared safely | Redaction manifest + no secrets/production endpoints | Yes |
12. Rollback and cleanup
The mandatory simulation changes no Nexus state, so cleanup is deleting the local fixture/evidence directory after you have retained the checkpoint artifact if desired. In a live extension, restore only temporary logging levels/monitoring changes you explicitly made; do not delete repository content, logs, database state, or blob files as part of chapter cleanup.
13. Production observability runbook
- For incidents: preserve request + topology + cache state first.
- Correlate request/application/outbound/audit/task evidence.
- Correlate JVM, DB, blob, network/upstream signals.
- Classify data before sharing support evidence.
- Change one causal variable.
- Repeat the same workload and verify artifact correctness.
- Record capacity trend and next threshold/action date.
14. Bridge to the production capstone
Chapter 30 will require an artifact platform that is not merely designed and secured but operable. The observability packet from this chapter becomes the evidence layer for capstone security validation, migration/upgrade gates, failure injection, recovery drills, capacity planning, and final operational handoff.
Knowledge check
What made blob latency the supported root-cause hypothesis in the checkpoint?
Request latency rose with blob latency while CPU, heap/direct memory, DB timing, and errors remained near baseline.
Why was increasing heap rejected?
No heap/GC pressure explained the regression; changing heap would add an uncontrolled variable and could reduce host/direct-memory headroom.
What proves the performance correction did not alter artifact identity?
The same artifact bytes produce the same recorded SHA-256 before and after the correction.
Why is the capacity ETA a forecast rather than a guarantee?
Growth can change with releases, cleanup, proxy/cache behavior, users, and retention; the estimate depends on stated assumptions and must be refreshed.
What does Chapter 29 contribute to Chapter 30?
A repeatable evidence and capacity operating model so the final platform can be validated, diagnosed, tuned, and handed off rather than merely configured.
Summary and next step
The checkpoint proved a complete evidence-driven operating loop: baseline, prediction, controlled bottleneck, diagnosis, one correction, comparable remeasurement, integrity verification, capacity forecasting, and privacy-safe handoff.
Chapter 30 integrates every prior layer—topology, formats, security, automation, supply-chain governance, recovery, migration, upgrades, HA, and this observability discipline—into the final production artifact-platform capstone.
Official references and version notes
- Sonatype: Logging — current self-hosted log files, rotation concepts, Log Viewer behavior, request/outbound/audit evidence, and PostgreSQL logging boundary.
- Sonatype: Auditing — audit capability, JSON event records, daily rotation, and current 90-day maximum retention.
- Sonatype: Prometheus — current 3.81+ Prometheus endpoint and required metrics privilege.
- Sonatype: Service Metrics Data API — current component/request usage metrics endpoint, privilege requirement, and the fact that this API is not listed in instance Swagger.
- Sonatype: Status API — read/writable application state and explicit limits: these checks do not validate external DB or disk health.
- Sonatype: Support Features — Support ZIP creation, storage location, and HA per-node bundle behavior.
- Sonatype: Support API — support ZIP API payload categories including system information, thread dump, metrics, configuration, security, logs, audit logs, and JMX.
- Sonatype: Nexus Repository Memory Overview — heap, direct memory, host headroom, and memory-related JVM arguments.
- Sonatype: System Requirements — current workload profiles, H2 limits, PostgreSQL guidance, CPU/RAM/storage baselines, and low-latency DB requirements.
- Sonatype: PostgreSQL Installation/Configuration — slow-query logging guidance and connection-pool configuration.
- Sonatype: PostgreSQL Max Connections — sizing the server connection limit against Nexus node pool sizes and troubleshooting exhausted connections.
- Sonatype: Blob Stores — blob count/used-size status, soft quotas, and storage-health evidence.
- Sonatype: Storage Planning — locality/latency implications and object-store performance tradeoffs.
- Sonatype: Usage Metrics — current usage-center request/component measurements and edition differences.
- Sonatype: Self-Hosted Usage Guide — request-log fallback, per-node aggregation, and request-metric interpretation.
- Sonatype: Configuring the Runtime Environment — current runtime properties, including the 3.95 audit-log attribute-change control.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.