Chapter 32Lesson 03~155 minutes

Troubleshooting Analysis Failures, Index Problems, CI Drift, and Incidents: Configuration, Design Patterns, and Trade-Offs

Choose troubleshooting patterns by owning layer, determinism, scope, CI drift, edition boundary, rollback and observable evidence instead of habitual retry or restart.

Design patternsConfiguration driftLeast privilegeRollbackEvidence

Learning objectives

  • Decide whether a symptom is scanner-side or server-side before changing either environment.
  • Distinguish transient network failure from deterministic configuration failure using repeatable evidence.
  • Choose project-local corrections before global mutations when scope permits.
  • Prove CI-only drift by comparing revision, working directory, parameters, runtime, cache and trust state with a local reproduction.
  • Explain why retry/restart is sometimes a controlled experiment but never a substitute for root-cause correction.

1. Troubleshooting is a configuration-design problem

Good troubleshooting is not a bag of commands. It is a method for minimizing the number of states changed before causality is understood. The five recurring design choices in this lesson—scanner versus server, transient versus deterministic, project versus global, CI-only versus local, and retry/restart versus root-cause correction—decide who owns the change, how wide the blast radius is, and what evidence can prove rollback.

The maintainable pattern prefers a narrow, observable, reversible change in the layer that owns the failed state. The anti-pattern changes shared server settings because they are easy to reach.

2. Scanner-side versus server-side

Scanner-side state exists before report submission: checkout, build outputs, scanner version/runtime, base directory, scope, report paths, SCM data, credentials and network route. Server-side state begins with Web/Compute Engine processing, project policy, database/search persistence, plugins and host resources. The bridge is the uploaded analysis report and its CE task ID.

Observation Prefer investigating Do not start with
No server connection / 401 / TLS error Scanner runner + network/auth/trust CE workers, database indexes, quality gate.
Unexpected indexed files Workspace + scanner scope/precedence Database/search tuning.
Report uploaded, CE task FAILED CE task + ce.log + server/plugin/DB/search evidence Changing source revision or project key.
CE SUCCESS, gate red Project measures/profile/gate/New Code Restarting scanner/server.
Gate correct, provider decoration missing DevOps integration/provider/CI state Changing gate threshold.

3. Transient network versus deterministic configuration

A transient hypothesis predicts that identical inputs can succeed without a configuration change because an external condition recovered. A deterministic configuration hypothesis predicts repeatable failure at the same stage. The diagnostic difference is not “try again until green.” It is a bounded same-input retry with preserved timestamps, HTTP/error evidence and a stated expectation.

Hypothesis example

H1: the scanner failed because the reverse proxy briefly reset connections. Prediction: DNS, URL, token and configuration remain unchanged; a single bounded rerun from the same runner succeeds, while proxy/access evidence shows a transient reset in the first window. If the failure repeats identically, H1 weakens.

Timeout values such as scanner connect/socket/response timeouts are not first-line cures. Raising them changes failure behavior and can hide a genuinely slow or unreachable dependency. Measure why the timeout occurred first.

4. Project-specific versus global correction

Global settings have broad blast radius. A path problem in one repository usually belongs in that repository/scanner configuration or project-level setting, not a global exclusion. A shared TLS trust issue on every runner may justify a centrally managed trust bundle, but only after the common cause is proven. The same logic applies to permission templates, proxy settings, quality profiles, New Code and CI variables.

Least privilege improves troubleshooting because identities expose only the surfaces needed. A project analysis token with Execute Analysis is easier to reason about than a global administrator token whose excessive rights can mask authorization errors.

5. CI-only drift versus local reproduction

“Works on my machine” is useful only when you compare manifests. Local and CI can differ in source SHA, checkout depth, working directory, case sensitivity, scanner wrapper/action version, Java/runtime, analyzer cache, proxy/TLS trust, environment-variable names, secret availability, generated report timing and filesystem mounts. A local rerun that does not reproduce CI state is not an equivalent test.

# Minimal parity manifest; redact secrets rather than dumping env.
{
  printf 'revision='; git rev-parse HEAD
  printf 'cwd='; pwd
  printf 'os='; uname -a
  sonar-scanner -v
  java -version 2>&1 | head -n 2
  printf 'SONAR_HOST_URL_present='; test -n "${SONAR_HOST_URL:-}" && echo yes || echo no
  printf 'SONAR_TOKEN_present='; test -n "${SONAR_TOKEN:-}" && echo yes || echo no
  find . -maxdepth 3 -type f \( -name 'coverage*.xml' -o -name 'test*.xml' \) -print
} > sonar-parity-manifest.txt

Do not capture the actual token. Presence, owner/type/expiry and permission are enough for most parity investigations.

6. Retry/restart versus root-cause correction

A restart is disruptive because it changes process state, caches, queue timing and logs. It can be appropriate when a documented operation requires it—for example, completing a plugin lifecycle—or when evidence identifies a process that cannot recover. It is not a generic diagnostic step.

A retry is less destructive but can still obscure intermittent behavior. Keep the first attempt, assign attempt IDs, cap retries, and record whether the same task stage failed. If every rerun creates a different source SHA or new project key, you are not testing recovery.

7. Product and edition boundaries

Community Build is the mandatory lab. Commercial Server editions add capabilities and Data Center adds licensed multi-node topology. Horizontal scale, branch/PR behavior, enterprise reporting and some identity/provider features can change which node or integration owns a symptom. SonarQube Cloud is a hosted product with a different operations boundary: you do not inspect its server filesystem, database or embedded search process.

Do not transfer a Data Center incident procedure to Community Build or a self-managed database procedure to SonarQube Cloud. Likewise, scanner/build configuration remains outside the server regardless of edition.

8. Current version-sensitive behavior to verify before diagnosis

As of the chapter recheck, Community Build 26.9 and the 2026 commercial train use modern scanner runtime behavior, and Java 17 support for non-auto-provisioned scanner runtimes has ended. Scanner JRE auto-provisioning can therefore explain why a CI runner succeeds without the Java version you expected. Preserve the scanner debug/runtime header rather than guessing.

Server process logs still have separate ownership: web.log covers web/database migration/request processing, ce.log covers background tasks, and es.log covers search engine operations. Web API V2 migration means an old troubleshooting script can fail independently of SonarQube analysis; check deprecation logs and the target instance API docs.

9. Decision table — choose the narrowest owned correction

Scenario Owning layer Candidate change Observable evidence Prerequisite Rollback
Scanner 401 before upload Auth/permission Project/local credential selection Token owner/type/expiry + permission check + no task ID Community Build+ Restore known-good project-scoped token
CI wrong report path CI/workspace/scanner Repository/CI working directory or path mapping Same SHA + path manifest + scanner import warning Community Build+ Revert CI path change
CE FAILED after upload Server/CE/plugin/DB/search Only the proven server-side cause Task ID + ce.log + inventory/health evidence Community Build+; topology differs in DCE Rollback plugin/config change or resource change
Search-backed API 503 during indexing Search/index state Observe supported recovery path es.log + system/indexing status + API response Community Build+ Rollback triggering deployment/upgrade change; no direct index edits
PR decoration fails, gate correct Provider integration Integration/provider config Gate result + provider/CI status + permission Commercial feature boundary may apply Revert integration change; gate unchanged

10. Worked scenario — local green, CI red

Suppose local and CI both target sq32:incident-lab. Local creates coverage.xml under the repository root. CI runs tests in a container that writes the file to /workspace/out/coverage.xml, but the scanner executes on the host workspace and still expects coverage.xml. The server is healthy and auth succeeds.

The correct pattern is to prove source SHA and scanner version parity, preserve the CI producer artifact and scanner import warning, then fix path publication/mapping in CI. Changing the server's global coverage exclusions, lowering the quality gate or mounting the SonarQube database into the CI runner would all change unrelated state.

Observable recovery

The repaired CI job must upload the producer report before scanning, the scanner must log successful import, the CE task must succeed, and the resulting coverage measure must correspond to the same revision. CI green alone is not enough.

11. Knowledge check

A scanner fails before upload. Which server log should you read first?

When is a retry a valid experiment?

A single project has the wrong source exclusion. Should you change the global exclusion?

A local run and CI run use different scanner versions. Can you conclude CI infrastructure is the cause?

Next lesson

Diagnostics, Failure Modes, and Production Practices

Apply the decision rules to destructive anti-patterns and a deliberately failed Compute Engine example with preserved task and log evidence.

Official references and version notes

Version and compatibility note

Baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. JRE auto-provisioning is supported by modern scanners; when scanner JRE auto-provisioning is disabled, use Java 21 or newer for the lab baseline. The disposable database path uses PostgreSQL 17.x and records the exact container digest at runtime. Web API V2 is still gradually replacing legacy endpoints; every endpoint used in automation should be checked in the target instance's API documentation.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.