Capstone: Operate a Governed Production SonarQube Quality Platform: Diagnostics, Failure Modes, and Production Practices
Diagnose capstone incidents by preserving first-failure evidence and separating scanner, compute, policy, trust, persistence, infrastructure and governance layers.
Learning objectives
- Apply the same evidence-first diagnostic sequence across scanner, network/auth, Compute Engine, policy, database/search, host/container, plugin/compatibility and edition layers.
- Diagnose capstone anti-patterns without hiding the original cause through restarts, project-key changes, threshold lowering or direct database/search edits.
- Interpret an intentionally broken Compute Engine-style incident and a runnable least-privilege/auth failure.
- Turn a database backup into an isolated restore/reindex recovery proof.
- Write a post-incident record that preserves residual risk and prevents recurrence.
1. Production anti-patterns become incident multipliers
Capstone failures are often caused by plausible shortcuts: install another plugin, grant more permission, rotate credentials, restart everything, lower a gate, copy an index, change the project key, or downgrade. Each shortcut can erase evidence or move the symptom without repairing the owning layer.
Generation date: 2026-09-08. Mandatory executable
work uses
SonarQube Community Build 26.9.0.129388 and the
official sonarqube:26.9.0.129388-community image. The
recorded scanner baseline is
SonarScanner CLI 8.1.0.6389. When scanner JRE
auto-provisioning is unavailable or disabled, use Java 21 or
newer. The local database family is PostgreSQL 17.x; the 2026.1
Server LTA documentation supports PostgreSQL 14–18. Commercial
references are SonarQube Server 2026 Release 4.1 and 2026.1.5 LTA.
Record the exact image digest, scanner output and database image
digest you actually run.
| Failure mode | Why it fails | Least-destructive response |
|---|---|---|
| More plugins for every gap | Adds executable supply-chain and compatibility risk. | Prove the missing capability and prefer built-in/external-tool boundary first. |
| Administrator token in CI | Expands blast radius and hides actual permission requirements. | Use project-scoped/narrow token and test permission explicitly. |
| Scan success = gate success | Collapses asynchronous and policy states. | Correlate task, result, gate and CI separately. |
| Restart before evidence | Destroys transient causal context. | Preserve scanner/server/CI/API logs/task/config first. |
| Untested backup | Recovery time and completeness are unknown. | Restore to isolated destination, reindex, verify project, run fresh analysis. |
| Unsupported downgrade | May cross incompatible DB/schema/runtime boundaries. | Use supported backup/update recovery path. |
| Vanity metric target | Encourages exclusions/suppressions and developer ranking. | Measure remediation flow, policy outcomes and context. |
| Commercial dependency with no fallback | Core governance may vanish with license/provider availability. | Document the underlying control and free/local simulation/fallback. |
| Undocumented exception | Becomes permanent invisible policy drift. | Owner + rationale + scope + expiry + review evidence. |
2. Evidence-first incident sequence
This sequence is deliberately the same one taught in Chapter 32. Production maturity means the organization follows it even under pressure.
- Preserve scanner, server, CI/provider and API evidence before cleanup or retry.
- Confirm versions: exact Community/Server edition, Java, database, scanner, plugins and integrations.
- Confirm source and effective inputs: revision, workspace state, scope/report paths and parameter precedence.
-
Inspect scanner/index/report: indexed files,
analyzer/sensor/import warnings, upload and
report-task.txt. -
Inspect Compute Engine: task ID/status,
timestamps and
ce.logcorrelation. - Inspect policy/result: profile, gate, New Code, security/review state and durable measures/issues.
- Inspect trust/external state: token permission, TLS/network, CI/provider/IDE integration.
- Inspect persistence/resources: database/search/JVM/host/container only if evidence points there.
- Apply one least-destructive correction at the owning layer.
- Rerun the smallest equivalent scenario with the same project key/revision/workload and verify durable recovery.
Do not delete .scannerwork/logs before preservation,
disable TLS verification, edit database/search internals directly,
lower policy thresholds, mass-suppress findings, replace the
project key, or use unsupported downgrade as “troubleshooting.”
3. Runnable broken example: authentication without privilege escalation
This fault changes exactly one trust input: the bearer token. The expected evidence is an authentication failure response while server health and project identity remain unchanged. Because auth is now proven as the owning layer, rotating unrelated tokens, changing TLS, restarting services or granting administrator privileges would be unjustified.
# Preserve a baseline successful read first.
curl -fsS -D evidence/api/auth-good.headers \
-H "Authorization: Bearer $SONAR_TOKEN" \
"http://localhost:9000/api/qualitygates/project_status?projectKey=sq34%3Acapstone" \
-o evidence/api/auth-good.json
# Intentionally use a fake invalid token. Never overwrite the real lab token value.
BAD_TOKEN='sq34-intentionally-invalid'
curl -sS -D evidence/api/auth-bad.headers \
-H "Authorization: Bearer $BAD_TOKEN" \
"http://localhost:9000/api/qualitygates/project_status?projectKey=sq34%3Acapstone" \
-o evidence/api/auth-bad.body
# Repair: restore the known-good scoped credential; do not promote to admin.
curl -fsS -H "Authorization: Bearer $SONAR_TOKEN" \
"http://localhost:9000/api/qualitygates/project_status?projectKey=sq34%3Acapstone" \
> evidence/api/auth-recovered.json
Recovery is the same read using the known-good least-privilege lab credential. Preserve both responses and record token owner/purpose—never the token value.
4. Intentionally broken Compute Engine incident: interpret before repair
Creating a real incompatible plugin merely to crash a current server is an unnecessary safety risk. The mandatory free path therefore uses a clearly labeled synthetic CE transcript while preserving the same diagnostic reasoning. Optional plugin lifecycle testing belongs only in an isolated extension lab.
# Synthetic CE evidence for interpretation only — do NOT install an incompatible plugin to create this.
2026.09.08 12:15:31 INFO ce[][o.s.c.t.CeWorkerImpl] Execute task | project=sq34:capstone | type=REPORT | id=AY-sq34-demo
2026.09.08 12:15:32 ERROR ce[][o.s.c.t.CeWorkerImpl] Failed to execute task AY-sq34-demo
java.lang.IllegalStateException: SQ34 simulated extension incompatible with current server API
... first causal stack frames preserved ...
Task status: FAILED
Scanner evidence: report uploaded successfully; ceTaskId=AY-sq34-demo
| Evidence | Interpretation | Wrong fix | Correct next action |
|---|---|---|---|
| Scanner upload success + task ID | Scanner/transport reached server hand-off. | Change project key or rerun blindly. | Inspect referenced CE task/ce.log. |
| CE task FAILED | Analysis did not become a successful durable result. | Lower gate threshold. | Preserve stack/version/plugin inventory. |
| Synthetic incompatible-extension stack | Plugin/compatibility layer owns failure in this simulation. | Edit database/search internals. | In isolated lab, remove/replace only the incompatible extension using documented compatibility/rollback path, restart if required, rerun same report/input. |
| Old gate/result still visible | UI may show previous successful analysis. | Call that current analysis success. | Compare analysis/task timestamp and revision before interpreting. |
5. Turn backup creation into recovery proof
The goal is not to overwrite the working lab. Restore to an isolated destination with a fresh search/data volume. SonarQube backup guidance centers the supported database and uses search reindexing during recovery. This exercise therefore refuses to “restore” a copied embedded-search directory as the authoritative state.
# compose.restore.yaml — isolated restore target, different port and volumes.
services:
restore-db:
image: postgres:17
environment:
POSTGRES_USER: sonar
POSTGRES_PASSWORD: sq34-restore-db-only
POSTGRES_DB: sonarqube
volumes: [sq34_restore_db:/var/lib/postgresql/data]
restore-sonar:
image: sonarqube:26.9.0.129388-community
depends_on: [restore-db]
environment:
SONAR_JDBC_URL: jdbc:postgresql://restore-db:5432/sonarqube
SONAR_JDBC_USERNAME: sonar
SONAR_JDBC_PASSWORD: sq34-restore-db-only
ports: ["9001:9000"]
volumes:
- sq34_restore_data:/opt/sonarqube/data
- sq34_restore_logs:/opt/sonarqube/logs
volumes:
sq34_restore_db:
sq34_restore_data:
sq34_restore_logs:
# Verify checksum before restore.
sha256sum -c evidence/backup/sonarqube.dump.sha256
# Start ONLY the isolated database first.
docker compose -f compose.restore.yaml -p sq34restore up -d restore-db
# Wait for pg_isready, then restore the database backup.
docker compose -f compose.restore.yaml -p sq34restore exec -T restore-db \
pg_restore -U sonar -d sonarqube --clean --if-exists \
< evidence/backup/sonarqube.dump
# Start SonarQube with a fresh search/data volume so indexes are built for the restored DB.
docker compose -f compose.restore.yaml -p sq34restore up -d restore-sonar
curl -fsS http://localhost:9001/api/system/status \
| tee evidence/backup/restore-system-status.json
# Verify project/configuration in UI/API, then run a fresh equivalent analysis against port 9001.
# Record the restored project's prior analysis evidence AND the new validation task.
# Cleanup only the restore target after evidence capture.
docker compose -f compose.restore.yaml -p sq34restore down -v
In a real environment, follow the exact current backup/restore procedure for your deployment model, database and SonarQube release. The Compose drill teaches the evidence pattern, not a universal production runbook.
6. Unsupported upgrade/downgrade shortcuts
An update changes multiple compatibility contracts at once. Production recovery must know whether the database schema, plugin API, Java runtime, scanner or automation endpoint changed. “Install the old binaries again” is not a safe generic rollback.
| Failure evidence | Preserve | Decision |
|---|---|---|
| Server fails after update | startup/web/CE/ES logs, exact old/new versions, DB version, plugin inventory, migration state | Stop mutation; consult update notes and supported recovery path. |
| Plugin incompatible | plugin provenance/version/compatibility + startup/CE error | Remove/replace only through documented extension path in controlled environment; do not delete random files/data. |
| Deprecated API breaks automation | HTTP status/body, endpoint/version, response headers, automation revision | Migrate to documented replacement/V2 path and add compatibility test. |
| Scanner runtime mismatch | scanner version + Java/JRE provisioning mode + failure log | Update scanner/runtime per current requirements; do not change server policy. |
2026.1 LTA removes Java 17 for the server and requires a JDK, supports Java 21 or 25, supports PostgreSQL 14–18 and includes Elasticsearch 8.x. Those are precisely the kinds of assumptions an update worksheet must record.
7. Vanity metrics and commercial dependencies as incident causes
Operational incidents can be organizational. A misleading KPI creates pressure to mutate analysis scope or policy; a paid feature can become an undocumented single point of governance. Treat incentives and ownership as causal layers, not “soft” afterthoughts.
| Bad incident response | Hidden incentive/cause | Production correction |
|---|---|---|
| Lower the gate so CI turns green | Optimizes dashboard outcome instead of code/evidence. | Keep policy; remediate or use bounded exception with owner/expiry. |
| Exclude files to improve coverage | Changes analysis scope, not tested behavior. | Fix producer tests/report or justify genuine generated/vendor exclusion. |
| Rank developers by issues | Context-free individual scoring drives suppression/gaming. | Review team/system trends, remediation lead time, New Code policy outcomes. |
| Buy enterprise report and call it compliance | Confuses reporting evidence with control effectiveness/security assurance. | Use report as one artifact in broader control/audit packet with residual risks. |
8. Post-incident record: prove durable recovery
Closing an incident means more than a green rerun. Preserve equivalence, causal evidence, correction, verification and residual risks.
incident_id: SQ34-INC-001
first_failure_utc: 2026-09-08T12:15:32Z
project_key: sq34:capstone
revision: <exact SHA from evidence/00-revision.txt>
classification: authentication | scanner-config | compute/plugin | persistence | policy | provider | infrastructure
first_failure_artifacts:
- evidence/scanner/...
- evidence/server/...
- evidence/api/...
hypothesis: <one falsifiable statement>
change_applied: <smallest reversible correction>
rerun_equivalence:
same_project_key: true
same_revision: true
same_build_reports: true
verification:
scanner: <status>
compute_engine: <status>
gate: <status>
ci_provider: <status or simulated>
residual_risks:
- <risk, owner, due date>
prevention:
- <compatibility test / permission check / runbook / training change>
9. Knowledge check
A CE task failed after a successful upload. Should you inspect the quality gate first?
No. The gate evaluates a successfully processed analysis. Start with the referenced CE task and CE/server evidence; a previous gate visible in the UI may belong to an older analysis.
Why is an invalid-token drill safer than disabling TLS to create a network incident?
It changes one bounded authentication input without weakening the trust boundary. TLS verification disablement teaches an unsafe shortcut and can hide certificate/host problems.
What turns a backup artifact into recovery evidence?
Checksum/tool/version provenance plus an isolated restore, correct search reindex/fresh index behavior, verification of project/configuration and a fresh equivalent analysis with measured recovery time.
Why not change the project key when analysis history looks corrupted?
The project key is identity and history. Changing it abandons evidence instead of repairing the owning persistence/search/configuration issue and makes before/after comparison invalid.
What organizational signal can create technical incident risk?
Metrics/incentives that reward hiding findings or passing gates at any cost. They encourage exclusions, suppressions and policy weakening, so governance design is part of reliability.
10. Summary and bridge
You can now operate the incident side of the platform without destructive shortcuts. Lesson 5 is the final checkpoint: assemble architecture, analysis, policy, security/RBAC, automation, observability, recovery, upgrade, capacity and governance into one production-readiness evidence packet with explicit residual risk.
Official references and version notes
Further reading
Verify version-sensitive behavior against primary documentation before using these patterns outside the disposable lab.
- SonarQube downloads — current Community Build, commercial release and LTA identities
- SonarQube Server documentation — server, administration, security and operations
- SonarQube Community Build documentation — free/local product behavior
- Web API — authentication and the ongoing Web API V2 migration
- Backup and restore — database backup/restore and reindex guidance
- Official SonarQube Docker image — current image tags and deployment notes
- Performance troubleshooting — host/search/database/CE performance guidance
SonarQube product names, editions, release trains, scanner runtimes, APIs, authentication options, and platform prerequisites can change independently. Re-check the linked SonarSource primary documentation for the exact target release before applying version-sensitive commands or operational guidance outside the disposable course environment.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.