Chapter 32Lesson 04~175 minutes

Troubleshooting Analysis Failures, Index Problems, CI Drift, and Incidents: Diagnostics, Failure Modes, and Production Practices

Diagnose realistic SonarQube incidents without deleting evidence, disabling TLS, editing database/search internals, changing project identity or masking CI divergence.

Incident responseFailure modesSearch & DBPluginsProduction safety

Learning objectives

  • Preserve .scannerwork, scanner output, task IDs, server logs and CI artifacts before any restart or cleanup.
  • Diagnose a deliberately failed Compute Engine task by correlating report-task.txt, Web API status and ce.log evidence.
  • Reject unsafe shortcuts such as TLS verification disablement, direct database/search edits, broad admin tokens and project-key changes.
  • Separate database/search pressure, plugin compatibility, host/container state and project policy from scanner symptoms.
  • Define a production incident record with ownership, blast radius, correction, rollback and same-input recovery evidence.

1. Production troubleshooting starts by protecting evidence

Incidents create pressure to “make it green.” The dangerous shortcuts in this chapter all remove information or widen blast radius: deleting .scannerwork and logs, restarting every service, editing database/search internals, disabling TLS verification, rotating tokens without proving authentication, changing project keys, or fixing only local execution while CI remains divergent.

The production objective is not merely a successful next run. It is a durable recovery whose cause, correction and unchanged inputs are independently reviewable.

2. Evidence-first diagnostic sequence

Follow ownership from source to infrastructure; stop when evidence localizes the first failed transition
flowchart TD A["Preserve scanner + CI + API evidence"] --> B["Confirm edition/version matrix"] B --> C["Confirm source revision + effective parameters"] C --> D["Inspect indexing + report + task ID"] D --> E["Inspect CE/background task"] E --> F["Inspect project policy/result"] F --> G["Inspect auth/network/provider"] G --> H["Inspect DB/search/JVM/host/container if relevant"] H --> I["Least-destructive correction"] I --> J["Same-input rerun + durable recovery proof"]

The sequence is not rigid bureaucracy. It is a way to avoid skipping from a scanner symptom directly to a database restart. If evidence localizes a TLS failure before report upload, you stop there. If the report uploaded and CE failed, you move to CE/server state.

3. Preserve before cleanup

# Disposable/local example evidence capture. Run from project root.
STAMP=$(date -u +%Y%m%dT%H%M%SZ)
E="incident-$STAMP"
mkdir -p "$E"

git rev-parse HEAD > "$E/revision.txt"
git status --porcelain > "$E/worktree-status.txt"
sonar-scanner -v > "$E/scanner-version.txt" 2>&1 || true

# Copy, do not delete, scanner bridge/workspace evidence.
[ -f .scannerwork/report-task.txt ] && cp .scannerwork/report-task.txt "$E/"
[ -d .scannerwork ] && find .scannerwork -maxdepth 2 -type f -printf '%P
' > "$E/scannerwork-files.txt"

# Preserve server/container logs before restart in this disposable lab.
docker logs sq32-sonarqube > "$E/sonarqube-container.log" 2>&1 || true
docker logs sq32-db > "$E/postgres-container.log" 2>&1 || true
docker stats --no-stream sq32-sonarqube sq32-db > "$E/container-stats.txt" || true

On a ZIP/service deployment, copy the relevant files from <sonarqubeHome>/logs. In container/Kubernetes deployments, preserve logs/events from the correct container/pod before replacement. Never collect secrets indiscriminately; redact tokens, passwords, private keys and sensitive source paths before sharing evidence.

4. Intentionally broken example — upload succeeded, Compute Engine failed

The following is synthetic but mirrors the evidence shape you should recognize. The scanner ended successfully and wrote:

projectKey=sq32:incident-lab
serverUrl=http://localhost:9000
serverVersion=26.9.0.129388
dashboardUrl=http://localhost:9000/dashboard?id=sq32%3Aincident-lab
ceTaskId=AX_SQ32_BROKEN_001
ceTaskUrl=http://localhost:9000/api/ce/task?id=AX_SQ32_BROKEN_001

An authorized read of the task returns:

{
  "task": {
    "id": "AX_SQ32_BROKEN_001",
    "type": "REPORT",
    "componentKey": "sq32:incident-lab",
    "status": "FAILED",
    "executionTimeMs": 832,
    "errorMessage": "Failed to execute project analysis"
  }
}

And the preserved ce.log window contains:

2026.09.08 14:22:31 INFO  ce[][o.s.c.t.CeWorkerImpl] Execute task | project=sq32:incident-lab | id=AX_SQ32_BROKEN_001
2026.09.08 14:22:32 ERROR ce[][o.s.c.t.CeWorkerImpl] Failed to execute task AX_SQ32_BROKEN_001
java.lang.NoClassDefFoundError: example/plugin/RemovedApi
    at example.plugin.CustomSensor.execute(CustomSensor.java:42)

The first failed transition is server-side Compute Engine execution after successful report upload. The classloading signature makes plugin/runtime compatibility a stronger hypothesis than token, URL or scanner source scope.

5. Repair the broken example without hiding the cause

Preserve the installed plugin inventory, SonarQube version, Java runtime and startup logs. Check the plugin's maintainer compatibility guidance for the exact target release. In a disposable lab, the reversible options are: remove the incompatible third-party plugin through the supported plugin lifecycle and restart once as required, or move to a plugin version explicitly compatible with the current SonarQube release. Record the artifact/checksum/source and rollback package.

Disruptive action

Plugin install/update/removal and server restart affect availability. They belong in an authorized maintenance window with a tested rollback. Do not manipulate plugin records in the database or delete search indexes to “fix” a classloading failure.

After repair, rerun the same commit with the same project key and analysis parameters. Compare the new task ID to the original failed task, but do not replace the original evidence. Recovery is proven when the new CE task succeeds and the durable project result updates as expected.

6. Seven failure modes and their causal correction

Anti-pattern What it destroys or risks Causal alternative
Delete .scannerwork/logs first Task bridge, indexing evidence, first stacktrace Copy/redact evidence first; cleanup only after diagnosis.
Restart every service Process state and timing; broad availability impact Restart only the proven component when documented/necessary.
Edit DB/search internals Supportability and data integrity Use documented admin/recovery paths; treat DB as durable state and search as managed/rebuildable state.
Disable TLS verification Trust boundary and MITM protection Repair CA chain/hostname/truststore configuration.
Rotate token immediately Changes auth variable without proving auth cause Check HTTP status, token type/expiry/permission first; rotate only when evidence supports it or policy requires.
Change project key Creates different project history/identity Keep stable project identity and repair original failure.
Fix local only Leaves delivery path broken and governance evidence divergent Reproduce CI manifest/path/secret/trust/runtime state and repair CI-owned drift.

7. Database/search symptoms: observe before acting

If CE tasks fail with database connectivity/timeout errors, correlate ce.log/web.log with database health and network latency. If project/issues APIs return temporary 503 during indexing or es.log shows search health transitions, preserve that evidence. Do not open the SonarQube database to “fix” rows or manipulate embedded search files as normal troubleshooting.

Database restoration, search reindexing, upgrade recovery and storage repair are operational procedures with version/support boundaries covered in earlier chapters. Incident response should invoke those procedures only when the owning layer is proven.

8. Auth/TLS incidents: distinguish identity, authorization and trust

HTTP 401 suggests authentication failure; 403 suggests an authenticated identity lacks required authorization; a certificate/hostname validation exception is a TLS trust failure. They can look similar in a CI job because all prevent analysis, but they require different corrections.

Use a least-privilege fake/local token in labs. For production, verify token owner/type/expiry and required project/global permission without printing the token. Current API guidance recommends bearer authentication, and responses using tokens can expose an expiration header on supported endpoints. That lets operations detect upcoming expiry without logging secrets.

9. CI drift incidents: preserve the runner as part of the system

A CI runner can change underneath a repository: action/plugin version updates, base image updates, Java changes, working directory changes, cache restore behavior, shallow checkout, proxy changes or secret scope. Capture the job definition revision and runner/action versions alongside SonarQube evidence. Pin dependencies where the provider supports it and review updates intentionally.

If only the local run is repaired, the governed delivery path remains broken. Recovery must include the CI-owned outcome, while still distinguishing a technically successful analysis from gate pass and provider decoration.

10. Production incident record

Field Example content
Incident identity SQ-INC-2026-032; UTC start/end; owner.
Affected scope Project(s), CI pipeline(s), API/UI availability; no developer ranking.
Exact versions Community/Server edition and build, Java, database, scanner, plugins, CI integration.
First failed transition e.g. report upload → CE task execution.
Evidence Revision, redacted scanner log, report-task, CE JSON, log window, health snapshot, CI artifact.
Hypothesis & falsifiers Why plugin compatibility was suspected; what evidence ruled out auth/network/scanner.
Correction Exact reversible change, approver, change window.
Rollback Artifact/config to restore and trigger condition.
Recovery proof Same revision + same effective inputs + new task SUCCESS + durable project result + CI/provider outcome.
Follow-up Compatibility policy, pinning, alert/runbook change; no policy threshold gaming.

11. Knowledge check

Scanner upload succeeded and CE task FAILED with NoClassDefFoundError in a third-party plugin. Should you rotate the analysis token?

Why is deleting the embedded search directory not a normal fix for an index symptom?

A token request gets 403. Does that prove the token is invalid?

After repair, why must you keep the original failed task ID?

Next lesson

Checkpoint Lab

Diagnose three different failures from one frozen fixture, use a written hypothesis/evidence table, repair minimally, and package proof that every rerun used equivalent input.

Official references and version notes

Version and compatibility note

Baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. JRE auto-provisioning is supported by modern scanners; when scanner JRE auto-provisioning is disabled, use Java 21 or newer for the lab baseline. The disposable database path uses PostgreSQL 17.x and records the exact container digest at runtime. Web API V2 is still gradually replacing legacy endpoints; every endpoint used in automation should be checked in the target instance's API documentation.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.