Chapter 32Lesson 01~150 minutes

Troubleshooting Analysis Failures, Index Problems, CI Drift, and Incidents: Core Concepts and Mental Model

Build an evidence-first troubleshooting model that preserves first failure and isolates scanner, network/auth, Compute Engine, database/search, plugin, project and CI layers.

TroubleshootingFirst-failure evidenceLayer isolationCompute EngineCI drift

Learning objectives

  • Explain why a visible symptom is not the same thing as the layer that caused it.
  • Preserve revision, scanner/runtime, effective-parameter, task, server, database/search, plugin and CI evidence before changing state.
  • Separate scanner success, report upload, Compute Engine success, durable analysis completion, gate outcome and CI/provider status.
  • Classify failures across scanner/build/report, network/auth/TLS, Compute Engine, project policy, database/search, plugin/runtime and CI/provider layers.
  • Use a one-hypothesis-at-a-time loop and define what durable recovery must prove.

1. The practical problem: the first visible error is often downstream of the cause

By Chapter 31 you learned to follow work through scanner, queue, Compute Engine, database/search and host resources. Troubleshooting uses the same ownership model under failure. A CI job that says “SonarQube failed” is not a diagnosis. It might mean the scanner could not resolve the server URL, the token lacked Execute Analysis, TLS trust failed, indexing excluded every intended source file, the report uploaded but its Compute Engine task failed, the project policy evaluated as failed, or the CI wrapper interpreted a later gate result as job failure.

The practical discipline is therefore preserve, classify, correlate, test. Preserve the earliest evidence before rerunning. Classify the layer that produced the evidence. Correlate source revision, effective scanner inputs, report/task identifiers and server state. Then test one hypothesis with the least destructive change. If you make several unrelated changes at once, a successful rerun tells you very little about causality.

First-failure rule

Before restart, retry, cache deletion, token rotation, plugin change, project-key change or host cleanup, copy the scanner/CI log, exact revision, command shape with secrets redacted, effective non-secret parameters, .scannerwork/report-task.txt when present, task ID/status, relevant server logs and environment/version manifest.

2. The incident loop: symptom → evidence → hypothesis → same-input proof

A symptom starts the investigation but does not own the fix. “401,” “PKIX path building failed,” “0 files indexed,” “report uploaded but dashboard unchanged,” “CE task FAILED,” “works locally but not in CI,” and “issues API returns 503 while indexing” each point at different boundaries. Classification narrows what evidence can falsify a hypothesis.

Preservation comes before minimal reproduction because a reproduction may warm caches, rotate logs, generate a different task ID or run a different revision. Once the original packet exists, reproduce the smallest equivalent scenario: same project key, same commit, same scanner family/version, same effective analysis scope, same report inputs and—when diagnosing CI drift—the same working-directory/path semantics. Test one hypothesis. Apply the narrowest reversible correction. Finally, rerun the same input and verify both the immediate symptom and the durable server result.

Evidence-first troubleshooting loop — arrows encode causality, not a click sequence
flowchart TD A["Symptom"] --> B["Classify owning layer"] B --> C["Preserve first-failure evidence"] C --> D["Minimal same-input reproduction"] D --> E["One falsifiable hypothesis"] E --> F["Least-destructive correction"] F --> G["Rerun same revision + inputs"] G --> H["Verify task + durable result + CI outcome"] H -->|"not recovered"| B

3. Eight layers to keep separate

The strongest troubleshooting habit is to name the state owner before touching configuration. These layers interact, but they are not interchangeable.

Source / build

Commit SHA, checkout depth, generated files, test/coverage producers, binaries and workspace layout. A wrong revision cannot be repaired by changing the server.

Scanner / indexing

Scanner family/version/runtime, base directory, sources/tests, inclusions/exclusions, report paths, analyzer downloads and indexed-file evidence.

Network / auth / TLS

Server URL, DNS/proxy, certificate trust, token identity/expiry and permission. A 401/403 is not a Compute Engine problem.

Compute Engine

Uploaded report, queue position, CE task ID, task status/duration and ce.log. Upload success does not prove CE success.

Project policy/result

Project key, quality profile, New Code definition, issues/measures and quality gate. A red gate can follow a technically successful analysis.

Database / search

Supported database reachability/latency and embedded search/index health. Durable application state is not the same as rebuildable search state.

Plugin / runtime

Server/plugin versions, analyzer compatibility, Java runtime and startup/classloading evidence. Plugin changes are executable supply-chain changes.

CI / provider

Runner image, working directory, environment, secret injection, caches, checkout, network route and wrapper/action version. Local success does not prove CI parity.

4. State before action: build the incident manifest

An incident manifest is a compact inventory of facts that must remain stable while you test hypotheses. Record secret identifiers or token type/owner—not token values.

State Read-only evidence Why it matters
Source revision git rev-parse HEAD, dirty-state check, checkout metadata Proves what code was actually presented to the scanner.
Scanner/runtime sonar-scanner -v, Java/runtime manifest, scanner debug header Separates scanner drift from server drift.
Effective inputs repository config + CI variables + non-secret command arguments Proves scope/report-path/project identity and precedence.
Upload/task .scannerwork/report-task.txt, CE task ID/status Bridges scanner completion to asynchronous server processing.
Server logs sonar.log, web.log, ce.log, es.log, access/deprecation logs Localizes process ownership and timestamps.
DB/search/host health/monitoring data, container stats, database reachability/latency Tests whether a shared dependency caused the symptom.
Plugin inventory Administration/Marketplace or documented installed-plugin API with appropriate permission Detects compatibility drift without immediately changing plugins.
CI state runner image, cwd, env names, cache key, checkout SHA, action/plugin versions Explains “local green, CI red” without blaming SonarQube generically.

5. Read-only inspection first

The following commands are examples for a disposable lab or authorized workstation. They deliberately avoid printing SONAR_TOKEN. On Windows PowerShell, use the equivalent commands shown below. API endpoints evolve; verify them in the target instance's built-in Web API documentation before operational automation.

# Source and scanner identity
printf 'revision='; git rev-parse HEAD
git status --porcelain
sonar-scanner -v
java -version 2>&1 | head -n 2

# Server identity and basic reachability: no token printed
curl -fsS http://localhost:9000/api/server/version ; echo
curl -fsS http://localhost:9000/api/system/status ; echo

# Preserve the scanner-to-CE bridge if analysis reached report upload
if [ -f .scannerwork/report-task.txt ]; then
  cp .scannerwork/report-task.txt evidence-report-task.txt
  sed -n '1,20p' evidence-report-task.txt
fi

# Container/process evidence in the disposable lab
docker ps --format 'table {{.Names}}	{{.Image}}	{{.Status}}'
docker stats --no-stream sq32-sonarqube sq32-db
$Revision = git rev-parse HEAD
"revision=$Revision"
git status --porcelain
sonar-scanner -v
java -version

Invoke-RestMethod http://localhost:9000/api/server/version
Invoke-RestMethod http://localhost:9000/api/system/status

if (Test-Path .scannerwork/report-task.txt) {
  Copy-Item .scannerwork/report-task.txt evidence-report-task.txt
  Get-Content evidence-report-task.txt
}

docker ps --format "table {{.Names}}	{{.Image}}	{{.Status}}"
docker stats --no-stream sq32-sonarqube sq32-db
Credential handling

For authenticated API checks, inject a fake/local token through SONAR_TOKEN and use an Authorization header without echoing the variable. Never paste a real token into committed scripts, screenshots, shell history or evidence bundles.

6. The scanner-to-task evidence chain

A successful scanner process is only one link. When the report uploads, the scanner writes a small task bridge file in .scannerwork. Preserve it. Its task identifier lets you ask the server what happened after upload. A scanner can finish upload successfully while the Compute Engine later marks the task failed.

Source SHA → scanner/runtime + effective parameters → indexed files → report upload → CE task ID → queue/task status → database/search-backed result → issues/measures/gate → CI/provider outcome

For a task whose ID is known, the documented api/ce/task service returns task details to appropriately authorized callers. Do not assume anonymous access or administrator permission. In a lab, use a project-scoped identity with the minimum permission that can read the task.

CE_TASK_ID="AX_FAKE_TASK_ID_FROM_REPORT_TASK_TXT"
curl -fsS   -H "Authorization: Bearer ${SONAR_TOKEN}"   "http://localhost:9000/api/ce/task?id=${CE_TASK_ID}"   > evidence-ce-task.json
python -m json.tool evidence-ce-task.json | sed -n '1,120p'

If there is no report-task.txt, do not jump to CE logs. First prove whether scanning reached report generation/upload. That distinction saves time and avoids unrelated server changes.

7. Translate common symptoms into first hypotheses

These mappings are starting hypotheses, not final answers. Evidence can contradict them.

Symptom First layer to test Evidence that can falsify the hypothesis
Connection refused / wrong host Network / URL Known-good HTTP reachability from the same runner namespace and exact effective sonar.host.url.
401 / token invalid Authentication Token identity/type/expiry and an authenticated read request from the same environment.
403 / not authorized Authorization Permission check for Execute Analysis/Browse/Administer as required; rotating the token without checking permission is not diagnosis.
PKIX / certificate error TLS trust Certificate chain, hostname and scanner truststore configuration; do not use -k as the “fix”.
0 or unexpected files indexed Scanner scope / workspace Working directory, base dir, sources/tests, inclusions/exclusions and scanner indexing lines.
Upload succeeds, dashboard unchanged CE queue/task Preserved task ID and task status/duration; then ce.log if failed.
Works locally, fails in CI CI/provider drift Same SHA? same cwd? same scanner/action/runtime? same variables? same trust/proxy? same generated reports?
Issues search temporarily 503 Search/index state Server/search evidence and indexing activity; do not edit search indexes directly.

8. What not to do first

Deleting .scannerwork, restarting every service, clearing CI caches, rotating tokens, changing project keys, disabling TLS verification or reinstalling plugins all change evidence. Some can be valid later, but only after a hypothesis points to them and rollback is explicit. Direct edits to the SonarQube database or embedded search indexes are not normal troubleshooting mechanisms and can turn a diagnosable incident into data corruption.

Controlled retry: a retry is useful when it tests the hypothesis “this was transient.” It must preserve the original attempt, use the same input, and be bounded. An unrecorded loop of retries is evidence destruction by noise.

9. Why this matters in DevOps

A SonarQube control is reproducible only when another engineer can independently verify the exact revision, effective analysis inputs, asynchronous task/result, policy context and owning system. Troubleshooting is therefore part of governance. The incident record should show which layer owned the fault, what evidence proved it, what minimal correction changed, what remained unchanged, and what recovery criterion passed.

This prevents two common organizational failures: blaming application developers for platform faults and weakening quality/security policy to hide operational faults. A failed gate is not repaired by changing the scanner. A TLS failure is not repaired by lowering a gate. A CE queue problem is not repaired by changing the project key.

10. Mini lab — classify before you repair

For each observation below, write the first layer you would inspect and one artifact you would preserve. Do not propose a fix yet.

  1. Scanner log ends with HTTP 401 before any “report uploaded” line.
  2. Scanner reports 148 files indexed and gives a task URL, but the project page shows an analysis warning.
  3. Local run succeeds at commit abc123; CI fails and its manifest shows commit abc120.
  4. Both scanner and UI requests slow down while database latency spikes.

The exercise trains sequence: identity and evidence first, repair second.

11. Knowledge check

The scanner prints “EXECUTION SUCCESS.” Can you conclude the quality gate passed?

A CI run fails TLS validation but a teammate suggests adding curl -k. What is the correct layer and safer action?

Why is changing sonar.projectKey a bad troubleshooting shortcut?

What evidence separates a scanner failure from a Compute Engine failure?

Next lesson

Guided Hands-On Workflow

Apply the model to bounded fault drills using one synthetic project and a disposable Community Build instance so every repair has comparable before/after evidence.

Official references and version notes

Version and compatibility note

Baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. JRE auto-provisioning is supported by modern scanners; when scanner JRE auto-provisioning is disabled, use Java 21 or newer for the lab baseline. The disposable database path uses PostgreSQL 17.x and records the exact container digest at runtime. Web API V2 is still gradually replacing legacy endpoints; every endpoint used in automation should be checked in the target instance's API documentation.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.