Troubleshooting Analysis Failures, Index Problems, CI Drift, and Incidents: Configuration, Design Patterns, and Trade-Offs
Choose troubleshooting patterns by owning layer, determinism, scope, CI drift, edition boundary, rollback and observable evidence instead of habitual retry or restart.
Learning objectives
- Decide whether a symptom is scanner-side or server-side before changing either environment.
- Distinguish transient network failure from deterministic configuration failure using repeatable evidence.
- Choose project-local corrections before global mutations when scope permits.
- Prove CI-only drift by comparing revision, working directory, parameters, runtime, cache and trust state with a local reproduction.
- Explain why retry/restart is sometimes a controlled experiment but never a substitute for root-cause correction.
1. Troubleshooting is a configuration-design problem
Good troubleshooting is not a bag of commands. It is a method for minimizing the number of states changed before causality is understood. The five recurring design choices in this lesson—scanner versus server, transient versus deterministic, project versus global, CI-only versus local, and retry/restart versus root-cause correction—decide who owns the change, how wide the blast radius is, and what evidence can prove rollback.
The maintainable pattern prefers a narrow, observable, reversible change in the layer that owns the failed state. The anti-pattern changes shared server settings because they are easy to reach.
2. Scanner-side versus server-side
Scanner-side state exists before report submission: checkout, build outputs, scanner version/runtime, base directory, scope, report paths, SCM data, credentials and network route. Server-side state begins with Web/Compute Engine processing, project policy, database/search persistence, plugins and host resources. The bridge is the uploaded analysis report and its CE task ID.
| Observation | Prefer investigating | Do not start with |
|---|---|---|
| No server connection / 401 / TLS error | Scanner runner + network/auth/trust | CE workers, database indexes, quality gate. |
| Unexpected indexed files | Workspace + scanner scope/precedence | Database/search tuning. |
| Report uploaded, CE task FAILED |
CE task + ce.log + server/plugin/DB/search
evidence
|
Changing source revision or project key. |
| CE SUCCESS, gate red | Project measures/profile/gate/New Code | Restarting scanner/server. |
| Gate correct, provider decoration missing | DevOps integration/provider/CI state | Changing gate threshold. |
3. Transient network versus deterministic configuration
A transient hypothesis predicts that identical inputs can succeed without a configuration change because an external condition recovered. A deterministic configuration hypothesis predicts repeatable failure at the same stage. The diagnostic difference is not “try again until green.” It is a bounded same-input retry with preserved timestamps, HTTP/error evidence and a stated expectation.
H1: the scanner failed because the reverse proxy briefly reset connections. Prediction: DNS, URL, token and configuration remain unchanged; a single bounded rerun from the same runner succeeds, while proxy/access evidence shows a transient reset in the first window. If the failure repeats identically, H1 weakens.
Timeout values such as scanner connect/socket/response timeouts are not first-line cures. Raising them changes failure behavior and can hide a genuinely slow or unreachable dependency. Measure why the timeout occurred first.
4. Project-specific versus global correction
Global settings have broad blast radius. A path problem in one repository usually belongs in that repository/scanner configuration or project-level setting, not a global exclusion. A shared TLS trust issue on every runner may justify a centrally managed trust bundle, but only after the common cause is proven. The same logic applies to permission templates, proxy settings, quality profiles, New Code and CI variables.
Least privilege improves troubleshooting because identities expose only the surfaces needed. A project analysis token with Execute Analysis is easier to reason about than a global administrator token whose excessive rights can mask authorization errors.
5. CI-only drift versus local reproduction
“Works on my machine” is useful only when you compare manifests. Local and CI can differ in source SHA, checkout depth, working directory, case sensitivity, scanner wrapper/action version, Java/runtime, analyzer cache, proxy/TLS trust, environment-variable names, secret availability, generated report timing and filesystem mounts. A local rerun that does not reproduce CI state is not an equivalent test.
# Minimal parity manifest; redact secrets rather than dumping env.
{
printf 'revision='; git rev-parse HEAD
printf 'cwd='; pwd
printf 'os='; uname -a
sonar-scanner -v
java -version 2>&1 | head -n 2
printf 'SONAR_HOST_URL_present='; test -n "${SONAR_HOST_URL:-}" && echo yes || echo no
printf 'SONAR_TOKEN_present='; test -n "${SONAR_TOKEN:-}" && echo yes || echo no
find . -maxdepth 3 -type f \( -name 'coverage*.xml' -o -name 'test*.xml' \) -print
} > sonar-parity-manifest.txt
Do not capture the actual token. Presence, owner/type/expiry and permission are enough for most parity investigations.
6. Retry/restart versus root-cause correction
A restart is disruptive because it changes process state, caches, queue timing and logs. It can be appropriate when a documented operation requires it—for example, completing a plugin lifecycle—or when evidence identifies a process that cannot recover. It is not a generic diagnostic step.
A retry is less destructive but can still obscure intermittent behavior. Keep the first attempt, assign attempt IDs, cap retries, and record whether the same task stage failed. If every rerun creates a different source SHA or new project key, you are not testing recovery.
7. Product and edition boundaries
Community Build is the mandatory lab. Commercial Server editions add capabilities and Data Center adds licensed multi-node topology. Horizontal scale, branch/PR behavior, enterprise reporting and some identity/provider features can change which node or integration owns a symptom. SonarQube Cloud is a hosted product with a different operations boundary: you do not inspect its server filesystem, database or embedded search process.
Do not transfer a Data Center incident procedure to Community Build or a self-managed database procedure to SonarQube Cloud. Likewise, scanner/build configuration remains outside the server regardless of edition.
8. Current version-sensitive behavior to verify before diagnosis
As of the chapter recheck, Community Build 26.9 and the 2026 commercial train use modern scanner runtime behavior, and Java 17 support for non-auto-provisioned scanner runtimes has ended. Scanner JRE auto-provisioning can therefore explain why a CI runner succeeds without the Java version you expected. Preserve the scanner debug/runtime header rather than guessing.
Server process logs still have separate ownership:
web.log covers web/database migration/request
processing, ce.log covers background tasks, and
es.log covers search engine operations. Web API V2
migration means an old troubleshooting script can fail independently
of SonarQube analysis; check deprecation logs and the target
instance API docs.
9. Decision table — choose the narrowest owned correction
| Scenario | Owning layer | Candidate change | Observable evidence | Prerequisite | Rollback |
|---|---|---|---|---|---|
| Scanner 401 before upload | Auth/permission | Project/local credential selection | Token owner/type/expiry + permission check + no task ID | Community Build+ | Restore known-good project-scoped token |
| CI wrong report path | CI/workspace/scanner | Repository/CI working directory or path mapping | Same SHA + path manifest + scanner import warning | Community Build+ | Revert CI path change |
| CE FAILED after upload | Server/CE/plugin/DB/search | Only the proven server-side cause | Task ID + ce.log + inventory/health evidence | Community Build+; topology differs in DCE | Rollback plugin/config change or resource change |
| Search-backed API 503 during indexing | Search/index state | Observe supported recovery path | es.log + system/indexing status + API response | Community Build+ | Rollback triggering deployment/upgrade change; no direct index edits |
| PR decoration fails, gate correct | Provider integration | Integration/provider config | Gate result + provider/CI status + permission | Commercial feature boundary may apply | Revert integration change; gate unchanged |
10. Worked scenario — local green, CI red
Suppose local and CI both target sq32:incident-lab.
Local creates coverage.xml under the repository root.
CI runs tests in a container that writes the file to
/workspace/out/coverage.xml, but the scanner executes
on the host workspace and still expects coverage.xml.
The server is healthy and auth succeeds.
The correct pattern is to prove source SHA and scanner version parity, preserve the CI producer artifact and scanner import warning, then fix path publication/mapping in CI. Changing the server's global coverage exclusions, lowering the quality gate or mounting the SonarQube database into the CI runner would all change unrelated state.
The repaired CI job must upload the producer report before scanning, the scanner must log successful import, the CE task must succeed, and the resulting coverage measure must correspond to the same revision. CI green alone is not enough.
11. Knowledge check
A scanner fails before upload. Which server log should you read first?
Often none. First inspect scanner/build/network/auth evidence. Server access/web logs can help if a request reached the server, but ce.log is not relevant until a CE task exists.
When is a retry a valid experiment?
When the transient hypothesis is explicit, the original attempt is preserved, the retry is bounded, and revision/effective inputs remain equivalent.
A single project has the wrong source exclusion. Should you change the global exclusion?
Normally no. Prefer project/repository-scoped configuration so the blast radius and rollback are narrow unless evidence proves a global policy is intentionally required.
A local run and CI run use different scanner versions. Can you conclude CI infrastructure is the cause?
Not yet. Scanner version is itself part of the environment difference. Align or deliberately test that variable before attributing the failure elsewhere.
Official references and version notes
- Troubleshooting the analysis — Current scanner-to-server asynchronous troubleshooting model and analysis progress behavior.
- Server logs — Current log split: sonar.log, web.log, ce.log, es.log and access.log, plus API deprecation logging.
- SonarScanner CLI — Scanner invocation, verbose/debug options, runtime requirements and TLS troubleshooting entry points.
- Scanner environment requirements — Current scanner runtime and JRE auto-provisioning requirements.
- TLS certificates on client side — Supported truststore/keystore paths instead of disabling TLS verification.
- Web API — Bearer authentication guidance and Web API V2 migration warning.
- SonarQube downloads — Current Community Build, Server release train and active LTA identities.
- SonarScanner CLI releases — Current scanner release identity and release notes.
- Docker Official Image — Current Community Build image tags for a pinned disposable local fixture.
Baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. JRE auto-provisioning is supported by modern scanners; when scanner JRE auto-provisioning is disabled, use Java 21 or newer for the lab baseline. The disposable database path uses PostgreSQL 17.x and records the exact container digest at runtime. Web API V2 is still gradually replacing legacy endpoints; every endpoint used in automation should be checked in the target instance's API documentation.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.