Chapter 29Lesson 05300–420 min

Checkpoint Lab — Logging, Metrics, Support ZIPs, Request Auditing, Performance Tuning, and Capacity Diagnostics

This checkpoint is a miniature production incident and capacity review. You will establish a clean baseline, inject one controlled bottleneck in synthetic telemetry, diagnose it from correlated request/DB/blob/JVM evidence, apply exactly one correction, prove improvement, and package the result as a privacy-reviewed operating handoff.

CheckpointBaselineBottleneck injectionRemeasureCapacity handoff

Learning objectives

  • Build a complete baseline containing request, JVM, DB, blob, task, and capacity signals.
  • Predict which signals change before injecting a controlled bottleneck.
  • Diagnose the bottleneck without changing multiple variables.
  • Prove the correction with repeated comparable measurements and artifact integrity checks.
  • Produce a redacted evidence packet and capacity/alerting runbook for Chapter 30.
Dated baseline (27 August 2026). Lessons use Nexus Repository 3.95.2-01 and Java 21 as the reference line. Re-check the live release notes, logging/metrics pages, system requirements, and instance-specific configuration before applying production tuning.
Measure before tuning. A slow package request is an end-to-end symptom, not proof of a JVM problem. Preserve the request, Nexus version/edition/runtime, cache state, database latency/pool state, blob latency/capacity, task activity, network/upstream timing, and client-cache state before changing a knob.
Privacy boundary. Logs and support bundles can contain usernames, client IPs, internal repository names/URLs, request paths, configuration, security information, dependency coordinates, and other operationally sensitive evidence. Sonatype removes password-related information from generated support files, but operators must still review/redact and share on a need-to-know basis.

1. Scenario

You operate a synthetic repository service called learner-platform. It serves a stable 10 MiB artifact and a small metadata asset. The mandatory exercise uses generated telemetry so no public service or real Nexus instance is stressed. Optional live validation may use a disposable local Nexus with the same evidence checklist.

Your goal is not to obtain the smallest possible latency. Your goal is to prove that you can explain a regression and improve it without causing a second unknown change.

2. Preflight

  • Nexus reference: 3.95.2-01, Java 21; record actual version if live.
  • Edition: mandatory simulation is Community-compatible.
  • Database: synthetic PostgreSQL signals; a live H2 extension must respect H2 workload limits.
  • Blob backend: synthetic file/object latency profile; live path must be disposable.
  • Credentials: none in mandatory simulation; live metrics use scoped secrets injected at runtime.
  • Safety: no production load test, no direct DB/blob edits, no public-registry stress, no disabling TLS/auth checks.

3. Define acceptance before the test

Use explicit synthetic objectives:

Metric Baseline objective Why it matters
Request p95 < 250 ms User/CI experience under the stated workload
Error rate 0% synthetic errors Do not trade speed for correctness
Heap after GC < 70% Headroom signal, not sole performance target
Direct memory < 80% Off-heap headroom
DB p95 < 50 ms Metadata/configuration dependency
Blob p95 < 60 ms Artifact byte path
Free capacity > 25% Operational/recovery headroom

These are exercise objectives, not Sonatype guarantees.

4. Predict two changes

Before injection, write predictions. Example bottleneck: blob latency.

  1. Prediction A: request p95 will rise roughly with blob p95 while CPU/heap/DB remain near baseline.
  2. Prediction B: artifact SHA-256 will remain unchanged because latency changes delivery time, not artifact identity.

5. Generate deterministic baseline and incident telemetry

import hashlib, statistics

artifact=b"learner-platform-release-1.0.0\n" * 400000
digest=hashlib.sha256(artifact).hexdigest()

baseline=[
 {"req":95,"cpu":38,"heap":52,"direct":44,"db":18,"blob":31,"err":0},
 {"req":110,"cpu":41,"heap":54,"direct":45,"db":20,"blob":36,"err":0},
 {"req":125,"cpu":39,"heap":53,"direct":46,"db":19,"blob":42,"err":0},
 {"req":105,"cpu":40,"heap":52,"direct":45,"db":17,"blob":34,"err":0},
]
incident=[
 {"req":820,"cpu":42,"heap":54,"direct":47,"db":21,"blob":745,"err":0},
 {"req":910,"cpu":43,"heap":55,"direct":48,"db":22,"blob":830,"err":0},
 {"req":790,"cpu":41,"heap":54,"direct":47,"db":20,"blob":720,"err":0},
 {"req":860,"cpu":42,"heap":55,"direct":48,"db":21,"blob":780,"err":0},
]

def p95(xs):
    xs=sorted(xs)
    return xs[max(0, int(len(xs)*0.95)-1)]

def summarize(rows):
    return {k:p95([r[k] for r in rows]) for k in ["req","cpu","heap","direct","db","blob","err"]}

print("sha256", digest)
print("baseline", summarize(baseline))
print("incident", summarize(incident))
assert summarize(incident)["blob"] > 700
assert summarize(incident)["db"] < 50
assert summarize(incident)["cpu"] < 60

The small sample makes the percentile arithmetic easy to inspect; it is a teaching fixture, not a statistical production benchmark.

6. Diagnose before correcting

The incident request latency rises by hundreds of milliseconds. CPU moves only slightly, heap/direct memory remain stable, DB remains near baseline, errors remain zero, and blob latency rises by roughly the same amount as request latency. The dominant hypothesis is blob/storage latency.

Rejected changes: increasing heap, increasing PostgreSQL pool size, disabling request logs, or changing cleanup policy. None addresses the observed wait.

7. Apply exactly one correction

In the synthetic exercise, the one change is “move the blob service to the same low-latency locality as Nexus,” represented by corrected blob timing. In a real deployment this could mean fixing storage/network placement according to the supported architecture; do not literally migrate a production blob store as an ad-hoc performance experiment.

corrected=[
 {"req":120,"cpu":42,"heap":54,"direct":47,"db":21,"blob":39,"err":0},
 {"req":135,"cpu":43,"heap":55,"direct":48,"db":22,"blob":45,"err":0},
 {"req":115,"cpu":41,"heap":54,"direct":47,"db":20,"blob":35,"err":0},
 {"req":128,"cpu":42,"heap":55,"direct":48,"db":21,"blob":42,"err":0},
]
print("corrected", summarize(corrected))
assert summarize(corrected)["req"] < 250
assert summarize(corrected)["blob"] < 60
assert hashlib.sha256(artifact).hexdigest() == digest

8. Capacity trend exercise

history_gb=[118,121,124,128,132,136,141,146,151]
days_between=5
rates=[(history_gb[i]-history_gb[i-1])/days_between for i in range(1,len(history_gb))]
rate=sum(rates)/len(rates)
threshold=190
eta=(threshold-history_gb[-1])/rate
print("avg_growth_gb_per_day", round(rate,2))
print("days_to_threshold", round(eta,1))
assert rate > 0

Record the estimate as a forecast with assumptions, not a guarantee. Add actions at earlier warning thresholds so cleanup/capacity expansion can be planned with backup/recovery headroom.

9. Build the evidence packet

Create a small directory containing only:

chapter29-evidence/
  00-context.txt          # version, Java, edition, DB, blob, cache, client, time window
  01-baseline.json        # synthetic summary
  02-incident.json        # synthetic summary
  03-diagnosis.md         # evidence -> hypothesis; rejected alternatives
  04-change.md            # exactly one controlled correction
  05-validation.json      # corrected summary + SHA-256 proof
  06-capacity.csv          # dated usage trend + threshold assumptions
  07-redaction-manifest.md# what was removed/pseudonymized and why

Do not include credentials, Authorization headers, private keys, real production URLs, or a complete unreviewed support ZIP.

10. Optional live support-bundle extension

On a disposable local Nexus, generate a support ZIP using the UI or current API, review its selected categories, and produce only a redacted manifest. Do not upload it anywhere for the course. Verify that the live instance’s metrics/request logs agree with the context captured in 00-context.txt.

11. Validation matrix

Claim Evidence Pass
Bottleneck was blob latency Request rise tracks blob rise while CPU/heap/DB remain stable Yes
One change corrected it Corrected run changes only synthetic blob locality/timing Yes
Artifact identity preserved SHA-256 identical before/after Yes
Capacity is managed proactively Trend and threshold ETA recorded Yes
Evidence can be shared safely Redaction manifest + no secrets/production endpoints Yes

12. Rollback and cleanup

The mandatory simulation changes no Nexus state, so cleanup is deleting the local fixture/evidence directory after you have retained the checkpoint artifact if desired. In a live extension, restore only temporary logging levels/monitoring changes you explicitly made; do not delete repository content, logs, database state, or blob files as part of chapter cleanup.

13. Production observability runbook

  1. For incidents: preserve request + topology + cache state first.
  2. Correlate request/application/outbound/audit/task evidence.
  3. Correlate JVM, DB, blob, network/upstream signals.
  4. Classify data before sharing support evidence.
  5. Change one causal variable.
  6. Repeat the same workload and verify artifact correctness.
  7. Record capacity trend and next threshold/action date.

14. Bridge to the production capstone

Chapter 30 will require an artifact platform that is not merely designed and secured but operable. The observability packet from this chapter becomes the evidence layer for capstone security validation, migration/upgrade gates, failure injection, recovery drills, capacity planning, and final operational handoff.

Knowledge check

What made blob latency the supported root-cause hypothesis in the checkpoint?

Why was increasing heap rejected?

What proves the performance correction did not alter artifact identity?

Why is the capacity ETA a forecast rather than a guarantee?

What does Chapter 29 contribute to Chapter 30?

Summary and next step

The checkpoint proved a complete evidence-driven operating loop: baseline, prediction, controlled bottleneck, diagnosis, one correction, comparable remeasurement, integrity verification, capacity forecasting, and privacy-safe handoff.

Chapter 30 integrates every prior layer—topology, formats, security, automation, supply-chain governance, recovery, migration, upgrades, HA, and this observability discipline—into the final production artifact-platform capstone.

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.