Production Capstone: Design, Secure, Automate, Migrate, and Recover an Enterprise Artifact Platform: Failure Injection, Troubleshooting, and Recovery Drill
Production confidence comes from controlled failure, not from a green dashboard. This lesson runs a disposable incident drill through the same diagnostic sequence used throughout the course: preserve concise evidence, confirm version/topology, isolate routing and authorization, inspect content/cache/upstream, inspect database/blob/disk, inspect logs/tasks/metrics, apply the least destructive correction, and verify with a controlled client request.
Learning objectives
- Diagnose credential, namespace, upstream, storage, cleanup, migration/upgrade, database/blob, network, policy, and observability failures.
- Preserve evidence before mutation and reject fixes that do not match the observed layer.
- Restore coherent database/blob state in an isolated recovery target.
- Distinguish service failover from data recovery and prove artifact identity after correction.
- Produce an incident timeline and recovery evidence bundle suitable for operational handoff.
capstone-lab-*, and every change has
preflight, evidence, verification, and rollback.
1. Failure-injection rules
-
Use only synthetic
capstone-lab-*resources and fixture files. -
Never expose a real credential. “Leak” only
token-FAKE_DO_NOT_USE. - Never attack or overload public registries; simulate upstream failure locally.
- Never corrupt a real Nexus database/blob store; use copied fixtures.
- Record a baseline fingerprint before every destructive simulation.
- Apply one correction per hypothesis and repeat the same verification request.
2. Incident matrix
| Incident | Primary symptom | First evidence | Wrong shortcut |
|---|---|---|---|
| Credential leak | secret appears in synthetic log | scope/identity + access history | hide line but keep credential valid |
| Namespace collision | unexpected public artifact selected | client endpoint + group/order/routing | delete random public cache files |
| Upstream outage | proxy miss fails; warm cache may work | outbound request/cache state | disable TLS/routing controls |
| Storage exhaustion | writes/tasks fail or quota alerts | blob/disk usage + task activity | delete blob files manually |
| Cleanup misconfiguration | needed disposable release candidate deleted | policy criteria/preview/task history | run compaction again |
| Upgrade failure | startup/search/client regression | version/runtime/release-note gate | blind downgrade of upgraded data |
| DB/blob mismatch | metadata exists but asset bytes missing | component/asset + blob/hash evidence | edit DB rows/files |
| Node/network failure | service path unavailable | LB/node/DB/blob dependency health | call independent servers “HA” |
| Policy bypass | direct upstream succeeds outside Nexus | client/network route | tighten Nexus policy only |
| Observability gap | cannot correlate request | logging/metrics configuration | change multiple settings before preserving evidence |
3. Diagnostic sequence as an executable decision discipline
PRESERVE EVIDENCE
→ VERSION / EDITION / JAVA / DB / BLOB / NODE TOPOLOGY
→ CLIENT URL + AUTH
→ REPOSITORY TYPE / GROUP / ROUTING
→ AUTHORIZATION
→ COMPONENT / ASSET / METADATA / DIGEST
→ PROXY CACHE / UPSTREAM
→ DATABASE / BLOB / DISK
→ LOGS / TASKS / METRICS
→ LEAST-DESTRUCTIVE CORRECTION
→ SAME CONTROLLED REQUEST + HASH/STATE VERIFICATION
4. Synthetic incident fixture
baseline={
"version":"3.95.2-01","java":21,"client_url":"nexus-group",
"auth":"reader-ok","route":"internal-first","db":"ok","blob":"ok",
"upstream":"ok","policy_path":"nexus-only","artifact_sha":"abc123"
}
incidents={
"credential_leak": {**baseline,"auth":"token-exposed"},
"namespace_bypass": {**baseline,"client_url":"public-direct","policy_path":"bypassed"},
"upstream_outage": {**baseline,"upstream":"down"},
"blob_loss": {**baseline,"blob":"missing","artifact_sha":None},
"upgrade_gate": {**baseline,"java":17},
}
def classify(s):
if s["auth"] == "token-exposed": return "credential"
if s["client_url"] == "public-direct": return "routing-bypass"
if s["java"] != 21: return "runtime-upgrade-gate"
if s["blob"] != "ok": return "storage-consistency"
if s["upstream"] != "ok": return "proxy-upstream"
return "unknown"
for name,state in incidents.items():
print(name, classify(state))
assert classify(incidents["blob_loss"]) == "storage-consistency"
5. Incident A: synthetic credential leak
Evidence contains token-FAKE_DO_NOT_USE. Correct
response: identify the owning service identity, revoke/rotate the
disposable credential, search the synthetic evidence/history for
use, replace it through the secret injection path, and rerun the
least-privilege test. Deleting the log line without invalidating the
credential does not remediate exposure.
6. Incident B: internal namespace bypass
The client config was edited to point directly to a public registry. Nexus logs show no request. That absence is itself evidence: changing Nexus roles or routing rules cannot control a request that bypasses Nexus. Correct the client/egress path, then repeat the request and confirm it traverses the approved group/proxy boundary.
7. Incident C: upstream outage with cache distinction
A known cached artifact remains available while a never-requested version fails during a simulated upstream outage. Diagnose proxy cache state, negative cache, auto-block/remote health, upstream/network evidence, and package-client cache separately. Do not globally disable TLS validation to “fix” remote access.
8. Incident D: storage exhaustion
When a disposable blob fixture reaches its operational threshold, stop new lab writes, preserve task/storage evidence, identify reclaimable versus live content, evaluate cleanup policy on synthetic data, then run supported reclamation only after logical deletion where appropriate. Direct filesystem deletion is never the recovery mechanism.
9. Incident E: cleanup deleted a needed candidate
This is a recovery/governance failure, not a reason to run more cleanup. Preserve the policy and task history, determine whether the artifact exists in a coherent backup/promotion target, restore or republish only from the authoritative build evidence, narrow the policy, preview again, and add a retention test to change review.
10. Incident F: failed migration/upgrade
If a rehearsal fails a Java/version/database/plugin/truststore gate, stop. Do not “see if production works anyway.” Restore the rehearsal clone or untouched source, correct the prerequisite, and rerun. If newer binaries already changed persistent state, rollback means restoring the verified pre-change checkpoint according to the supported procedure—not starting old binaries blindly.
11. Incident G: database/blob inconsistency
Fixture database metadata lists
capstone-lib-1.0.0.jar but the corresponding blob is
missing. A database-only restore cannot fix the content. Recover a
coherent database/blob pair from the same recovery point and
validate through Nexus/client requests plus SHA-256.
from hashlib import sha256
backup={"db":{"asset":"capstone-lib-1.0.0.jar","sha":""},"blob":b"capstone-release-bytes\n"}
backup["db"]["sha"] = sha256(backup["blob"]).hexdigest()
restored_blob = backup["blob"]
assert sha256(restored_blob).hexdigest() == backup["db"]["sha"]
print("coherent-restore: PASS")
12. Incident H: node/network failure versus data failure
In the optional HA simulation, a node failure can be masked by healthy peers. A database writer failure requires database-layer failover. A shared blob outage can affect every node. Logical corruption can be served successfully from every node. Diagnose failure domain before declaring “HA worked” or “HA failed.”
13. Incident I: insufficient observability
If you cannot correlate a controlled request, add the smallest bounded evidence source required for the next hypothesis—such as request logging, a narrow logger, or DB/blob timing—then reproduce once. Do not enable all DEBUG logging indefinitely. Review any support ZIP before sharing because it can contain sensitive operational context even after password sanitization.
14. Timed recovery drill
events=[
(0,"declare incident"),
(4,"freeze writes and preserve evidence"),
(12,"select verified recovery point"),
(24,"restore database and blob fixture"),
(38,"start isolated target"),
(46,"repository/client/hash validation complete"),
]
rto_minutes=events[-1][0]
backup_age_minutes=22
print("RTO",rto_minutes,"minutes")
print("observed data loss window",backup_age_minutes,"minutes")
assert rto_minutes <= 60
assert backup_age_minutes <= 30
This proves the synthetic drill meets the capstone's RTO ≤ 60 min and RPO ≤ 30 min assumptions. Real production targets require measured production-grade recovery rehearsals.
15. Incident evidence bundle
capstone-incident/
timeline.md
version-topology.txt
client-request-redacted.txt
routing-authz-state.json
artifact-hashes.json
cache-upstream-evidence.txt
db-blob-capacity-evidence.txt
logs-tasks-metrics-redacted.txt
hypothesis-and-rejected-fixes.md
correction.md
verification.md
recovery-rpo-rto.md
Knowledge check
Why is deleting an exposed secret from a log insufficient remediation?
The credential remains usable; rotate/revoke it, investigate use, replace it through the secret path, and then handle evidence safely.
What does “no request in Nexus logs” suggest when a dependency still downloads?
The client may be bypassing Nexus, so inspect client/network routing rather than only Nexus policy.
Why can a database-only restore leave repository content broken?
Database metadata can reference blob content that was not restored to the same consistency point.
What is the safe rollback when an upgrade changed persistent state?
Restore the verified pre-upgrade database/blob/configuration checkpoint using the supported procedure, not a blind binary downgrade.
How did the fixture prove the capstone RPO/RTO assumptions?
The simulated restore completed in 46 minutes and the selected recovery point represented a 22-minute data-loss window, both within the stated synthetic targets.
Summary and next step
The failure drill proved that the capstone can classify incidents by layer, preserve evidence, reject folklore fixes, recover coherent state, verify immutable artifact identity, and measure synthetic RPO/RTO.
Lesson 5 turns that technical proof into a final production-readiness review and operational handoff: ownership, runbooks, evidence retention, change gates, capacity, security, recovery, and course completion.
Official references and version notes
- Sonatype: Nexus Repository documentation — current self-hosted product entry point and supported-format documentation.
- Sonatype: 2026 self-hosted release notes — release-specific changes and known-issue/upgrade guidance.
- Sonatype: Self-Hosted Feature Matrix — Community versus Pro capability boundary.
- Sonatype: System Requirements — Java 21, H2 workload limits, PostgreSQL guidance, sizing, file handles, and deployment constraints.
- Sonatype: Repository Types — hosted, proxy, and group responsibilities.
- Sonatype: Content Selectors — fine-grained content authorization concepts.
- Sonatype: Cleanup Policies — retention criteria, preview/evaluation, deletion semantics, and blob reclamation boundary.
- Sonatype: REST and Integration API — documented REST/OpenAPI automation boundary.
- Sonatype: Webhooks — repository/global events and delivery semantics.
- Sonatype: Staging — current Pro-only staged component movement/promotion capability.
- Sonatype: Repository Firewall — separately licensed IQ-powered policy integration and supported Nexus editions.
- Sonatype: Prepare a Backup — coordinated database/blob backups and node-ID preservation.
- Sonatype: Database migration — current migration tooling and compatibility gates.
- Sonatype: Instance Migrator — current source/target migration workflow and constraints.
- Sonatype: Upgrade Paths — version thresholds and mandatory crossed-version procedures.
- Sonatype: Rolling Upgrades in HA — Pro/HA rolling-upgrade behavior and mixed-mode constraints.
- Sonatype: HA system requirements — Pro-only active-node topology, shared PostgreSQL/blob state, and failure-domain requirements.
- Sonatype: Logging — application/request/outbound/audit/JVM evidence.
- Sonatype: Prometheus — current metrics endpoint and privilege boundary.
- Sonatype: Support Features — support ZIP generation and evidence handling.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.