Chapter 34Lesson 04~220 minutes

Capstone: Operate a Governed Production SonarQube Quality Platform: Diagnostics, Failure Modes, and Production Practices

Diagnose capstone incidents by preserving first-failure evidence and separating scanner, compute, policy, trust, persistence, infrastructure and governance layers.

Incident responseFirst failureRecoverySafe operationsRoot cause

Learning objectives

  • Apply the same evidence-first diagnostic sequence across scanner, network/auth, Compute Engine, policy, database/search, host/container, plugin/compatibility and edition layers.
  • Diagnose capstone anti-patterns without hiding the original cause through restarts, project-key changes, threshold lowering or direct database/search edits.
  • Interpret an intentionally broken Compute Engine-style incident and a runnable least-privilege/auth failure.
  • Turn a database backup into an isolated restore/reindex recovery proof.
  • Write a post-incident record that preserves residual risk and prevents recurrence.

1. Production anti-patterns become incident multipliers

Capstone failures are often caused by plausible shortcuts: install another plugin, grant more permission, rotate credentials, restart everything, lower a gate, copy an index, change the project key, or downgrade. Each shortcut can erase evidence or move the symptom without repairing the owning layer.

Pinned capstone baseline

Generation date: 2026-09-08. Mandatory executable work uses SonarQube Community Build 26.9.0.129388 and the official sonarqube:26.9.0.129388-community image. The recorded scanner baseline is SonarScanner CLI 8.1.0.6389. When scanner JRE auto-provisioning is unavailable or disabled, use Java 21 or newer. The local database family is PostgreSQL 17.x; the 2026.1 Server LTA documentation supports PostgreSQL 14–18. Commercial references are SonarQube Server 2026 Release 4.1 and 2026.1.5 LTA. Record the exact image digest, scanner output and database image digest you actually run.

Failure mode Why it fails Least-destructive response
More plugins for every gap Adds executable supply-chain and compatibility risk. Prove the missing capability and prefer built-in/external-tool boundary first.
Administrator token in CI Expands blast radius and hides actual permission requirements. Use project-scoped/narrow token and test permission explicitly.
Scan success = gate success Collapses asynchronous and policy states. Correlate task, result, gate and CI separately.
Restart before evidence Destroys transient causal context. Preserve scanner/server/CI/API logs/task/config first.
Untested backup Recovery time and completeness are unknown. Restore to isolated destination, reindex, verify project, run fresh analysis.
Unsupported downgrade May cross incompatible DB/schema/runtime boundaries. Use supported backup/update recovery path.
Vanity metric target Encourages exclusions/suppressions and developer ranking. Measure remediation flow, policy outcomes and context.
Commercial dependency with no fallback Core governance may vanish with license/provider availability. Document the underlying control and free/local simulation/fallback.
Undocumented exception Becomes permanent invisible policy drift. Owner + rationale + scope + expiry + review evidence.

2. Evidence-first incident sequence

This sequence is deliberately the same one taught in Chapter 32. Production maturity means the organization follows it even under pressure.

  1. Preserve scanner, server, CI/provider and API evidence before cleanup or retry.
  2. Confirm versions: exact Community/Server edition, Java, database, scanner, plugins and integrations.
  3. Confirm source and effective inputs: revision, workspace state, scope/report paths and parameter precedence.
  4. Inspect scanner/index/report: indexed files, analyzer/sensor/import warnings, upload and report-task.txt.
  5. Inspect Compute Engine: task ID/status, timestamps and ce.log correlation.
  6. Inspect policy/result: profile, gate, New Code, security/review state and durable measures/issues.
  7. Inspect trust/external state: token permission, TLS/network, CI/provider/IDE integration.
  8. Inspect persistence/resources: database/search/JVM/host/container only if evidence points there.
  9. Apply one least-destructive correction at the owning layer.
  10. Rerun the smallest equivalent scenario with the same project key/revision/workload and verify durable recovery.
Forbidden shortcuts

Do not delete .scannerwork/logs before preservation, disable TLS verification, edit database/search internals directly, lower policy thresholds, mass-suppress findings, replace the project key, or use unsupported downgrade as “troubleshooting.”

3. Runnable broken example: authentication without privilege escalation

This fault changes exactly one trust input: the bearer token. The expected evidence is an authentication failure response while server health and project identity remain unchanged. Because auth is now proven as the owning layer, rotating unrelated tokens, changing TLS, restarting services or granting administrator privileges would be unjustified.

# Preserve a baseline successful read first.
curl -fsS -D evidence/api/auth-good.headers \
  -H "Authorization: Bearer $SONAR_TOKEN" \
  "http://localhost:9000/api/qualitygates/project_status?projectKey=sq34%3Acapstone" \
  -o evidence/api/auth-good.json

# Intentionally use a fake invalid token. Never overwrite the real lab token value.
BAD_TOKEN='sq34-intentionally-invalid'
curl -sS -D evidence/api/auth-bad.headers \
  -H "Authorization: Bearer $BAD_TOKEN" \
  "http://localhost:9000/api/qualitygates/project_status?projectKey=sq34%3Acapstone" \
  -o evidence/api/auth-bad.body

# Repair: restore the known-good scoped credential; do not promote to admin.
curl -fsS -H "Authorization: Bearer $SONAR_TOKEN" \
  "http://localhost:9000/api/qualitygates/project_status?projectKey=sq34%3Acapstone" \
  > evidence/api/auth-recovered.json
Repair interpretation

Recovery is the same read using the known-good least-privilege lab credential. Preserve both responses and record token owner/purpose—never the token value.

4. Intentionally broken Compute Engine incident: interpret before repair

Creating a real incompatible plugin merely to crash a current server is an unnecessary safety risk. The mandatory free path therefore uses a clearly labeled synthetic CE transcript while preserving the same diagnostic reasoning. Optional plugin lifecycle testing belongs only in an isolated extension lab.

# Synthetic CE evidence for interpretation only — do NOT install an incompatible plugin to create this.
2026.09.08 12:15:31 INFO  ce[][o.s.c.t.CeWorkerImpl] Execute task | project=sq34:capstone | type=REPORT | id=AY-sq34-demo
2026.09.08 12:15:32 ERROR ce[][o.s.c.t.CeWorkerImpl] Failed to execute task AY-sq34-demo
java.lang.IllegalStateException: SQ34 simulated extension incompatible with current server API
... first causal stack frames preserved ...
Task status: FAILED
Scanner evidence: report uploaded successfully; ceTaskId=AY-sq34-demo
Evidence Interpretation Wrong fix Correct next action
Scanner upload success + task ID Scanner/transport reached server hand-off. Change project key or rerun blindly. Inspect referenced CE task/ce.log.
CE task FAILED Analysis did not become a successful durable result. Lower gate threshold. Preserve stack/version/plugin inventory.
Synthetic incompatible-extension stack Plugin/compatibility layer owns failure in this simulation. Edit database/search internals. In isolated lab, remove/replace only the incompatible extension using documented compatibility/rollback path, restart if required, rerun same report/input.
Old gate/result still visible UI may show previous successful analysis. Call that current analysis success. Compare analysis/task timestamp and revision before interpreting.

5. Turn backup creation into recovery proof

The goal is not to overwrite the working lab. Restore to an isolated destination with a fresh search/data volume. SonarQube backup guidance centers the supported database and uses search reindexing during recovery. This exercise therefore refuses to “restore” a copied embedded-search directory as the authoritative state.

# compose.restore.yaml — isolated restore target, different port and volumes.
services:
  restore-db:
    image: postgres:17
    environment:
      POSTGRES_USER: sonar
      POSTGRES_PASSWORD: sq34-restore-db-only
      POSTGRES_DB: sonarqube
    volumes: [sq34_restore_db:/var/lib/postgresql/data]
  restore-sonar:
    image: sonarqube:26.9.0.129388-community
    depends_on: [restore-db]
    environment:
      SONAR_JDBC_URL: jdbc:postgresql://restore-db:5432/sonarqube
      SONAR_JDBC_USERNAME: sonar
      SONAR_JDBC_PASSWORD: sq34-restore-db-only
    ports: ["9001:9000"]
    volumes:
      - sq34_restore_data:/opt/sonarqube/data
      - sq34_restore_logs:/opt/sonarqube/logs
volumes:
  sq34_restore_db:
  sq34_restore_data:
  sq34_restore_logs:
# Verify checksum before restore.
sha256sum -c evidence/backup/sonarqube.dump.sha256

# Start ONLY the isolated database first.
docker compose -f compose.restore.yaml -p sq34restore up -d restore-db
# Wait for pg_isready, then restore the database backup.
docker compose -f compose.restore.yaml -p sq34restore exec -T restore-db \
  pg_restore -U sonar -d sonarqube --clean --if-exists \
  < evidence/backup/sonarqube.dump

# Start SonarQube with a fresh search/data volume so indexes are built for the restored DB.
docker compose -f compose.restore.yaml -p sq34restore up -d restore-sonar
curl -fsS http://localhost:9001/api/system/status \
  | tee evidence/backup/restore-system-status.json

# Verify project/configuration in UI/API, then run a fresh equivalent analysis against port 9001.
# Record the restored project's prior analysis evidence AND the new validation task.

# Cleanup only the restore target after evidence capture.
docker compose -f compose.restore.yaml -p sq34restore down -v
Production adaptation

In a real environment, follow the exact current backup/restore procedure for your deployment model, database and SonarQube release. The Compose drill teaches the evidence pattern, not a universal production runbook.

6. Unsupported upgrade/downgrade shortcuts

An update changes multiple compatibility contracts at once. Production recovery must know whether the database schema, plugin API, Java runtime, scanner or automation endpoint changed. “Install the old binaries again” is not a safe generic rollback.

Failure evidence Preserve Decision
Server fails after update startup/web/CE/ES logs, exact old/new versions, DB version, plugin inventory, migration state Stop mutation; consult update notes and supported recovery path.
Plugin incompatible plugin provenance/version/compatibility + startup/CE error Remove/replace only through documented extension path in controlled environment; do not delete random files/data.
Deprecated API breaks automation HTTP status/body, endpoint/version, response headers, automation revision Migrate to documented replacement/V2 path and add compatibility test.
Scanner runtime mismatch scanner version + Java/JRE provisioning mode + failure log Update scanner/runtime per current requirements; do not change server policy.
Current LTA example

2026.1 LTA removes Java 17 for the server and requires a JDK, supports Java 21 or 25, supports PostgreSQL 14–18 and includes Elasticsearch 8.x. Those are precisely the kinds of assumptions an update worksheet must record.

7. Vanity metrics and commercial dependencies as incident causes

Operational incidents can be organizational. A misleading KPI creates pressure to mutate analysis scope or policy; a paid feature can become an undocumented single point of governance. Treat incentives and ownership as causal layers, not “soft” afterthoughts.

Bad incident response Hidden incentive/cause Production correction
Lower the gate so CI turns green Optimizes dashboard outcome instead of code/evidence. Keep policy; remediate or use bounded exception with owner/expiry.
Exclude files to improve coverage Changes analysis scope, not tested behavior. Fix producer tests/report or justify genuine generated/vendor exclusion.
Rank developers by issues Context-free individual scoring drives suppression/gaming. Review team/system trends, remediation lead time, New Code policy outcomes.
Buy enterprise report and call it compliance Confuses reporting evidence with control effectiveness/security assurance. Use report as one artifact in broader control/audit packet with residual risks.

8. Post-incident record: prove durable recovery

Closing an incident means more than a green rerun. Preserve equivalence, causal evidence, correction, verification and residual risks.

incident_id: SQ34-INC-001
first_failure_utc: 2026-09-08T12:15:32Z
project_key: sq34:capstone
revision: <exact SHA from evidence/00-revision.txt>
classification: authentication | scanner-config | compute/plugin | persistence | policy | provider | infrastructure
first_failure_artifacts:
  - evidence/scanner/...
  - evidence/server/...
  - evidence/api/...
hypothesis: <one falsifiable statement>
change_applied: <smallest reversible correction>
rerun_equivalence:
  same_project_key: true
  same_revision: true
  same_build_reports: true
verification:
  scanner: <status>
  compute_engine: <status>
  gate: <status>
  ci_provider: <status or simulated>
residual_risks:
  - <risk, owner, due date>
prevention:
  - <compatibility test / permission check / runbook / training change>

9. Knowledge check

A CE task failed after a successful upload. Should you inspect the quality gate first?

Why is an invalid-token drill safer than disabling TLS to create a network incident?

What turns a backup artifact into recovery evidence?

Why not change the project key when analysis history looks corrupted?

What organizational signal can create technical incident risk?

10. Summary and bridge

You can now operate the incident side of the platform without destructive shortcuts. Lesson 5 is the final checkpoint: assemble architecture, analysis, policy, security/RBAC, automation, observability, recovery, upgrade, capacity and governance into one production-readiness evidence packet with explicit residual risk.

Next lesson

Checkpoint Lab — Capstone: Operate a Governed Production SonarQube Quality Platform

Continue to the next lesson to build on this lesson’s evidence, workflow, and operational practices.

Official references and version notes

Further reading

Verify version-sensitive behavior against primary documentation before using these patterns outside the disposable lab.

Version and compatibility note

SonarQube product names, editions, release trains, scanner runtimes, APIs, authentication options, and platform prerequisites can change independently. Re-check the linked SonarSource primary documentation for the exact target release before applying version-sensitive commands or operational guidance outside the disposable course environment.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.