Language Analysis, Sensors, SCM Data, and Source-Code Indexing: Diagnostics, Failure Modes, and Production Practices
Diagnose scope and indexing failures from the first preserved evidence instead of hiding them with broader exclusions, disabled SCM, different project keys, or lower policy thresholds.
Learning objectives
- Diagnose missing/unexpected files by walking base directory, scope, filters, language recognition and sensor stages in order.
- Repair source/test overlap without suppressing the scanner error.
- Distinguish shallow-SCM symptoms from authentication, analyzer and server failures.
- Preserve and interpret sensor/report warnings before rerunning analysis.
- Apply the least-destructive correction and rerun the smallest equivalent scenario.
1. Evidence-first diagnostic sequence
-
Preserve: scanner debug log,
sonar-project.properties/effective build config, exact revision, working directory,report-task.txtif produced, and server background-task evidence. - Confirm versions: Community Build edition/version, scanner, analyzer/build tool/runtime assumptions.
- Confirm repository state: base directory, tracked/ignored files, full vs shallow checkout.
- Inspect initial scope: source roots, test roots, build-derived source sets.
- Inspect filters: global/project/scanner exclusions/inclusions and SCM ignore behavior.
- Inspect recognition: language suffixes, unsupported/ambiguous files, build prerequisites such as Java bytecode.
- Inspect sensors/report inputs: warnings, missing coverage/external-report files, skipped analyzers.
- Then inspect upload/Compute Engine/policy state if a report was produced.
- Change one causal variable and rerun the smallest equivalent analysis.
2. Failure map: symptom does not identify the owning layer
| Symptom | Likely layer | First evidence | Wrong shortcut |
|---|---|---|---|
| Expected file absent | base dir / initial scope / filters / suffix | debug indexing + scanner context | remove all exclusions globally |
| File indexed twice | source/test classification overlap | resolved source/test roots + patterns | mark tests as source only to pass |
| Unexpected vendor/generated issues | scope/ownership policy | file path + exclusion origin + analyzer | accept all issues or lower gate |
| Missing blame / shallow warning | SCM checkout |
git rev-parse --is-shallow-repository + log
|
sonar.scm.disabled=true |
| Sensor skipped/missing report | build/report input | sensor warning + file timestamp/path | create fake empty report |
| Scanner success but stale dashboard | CE/task/project identity | report-task + CE status + project key | rerun with a new project key |
3. Intentionally broken example: source/test overlap
On a disposable copy of the fixture, configure both source and test roots to the same tree without complementary filters:
sonar.sources=src,tests
sonar.tests=tests
The scanner should refuse to index the same code file twice.
Preserve the exact error. The repair is to make the sets
disjoint—for example sonar.sources=src,infra and
sonar.tests=tests—not to suppress test classification
or invent a new project key.
4. Over-broad exclusion disguised as “metric cleanup”
# Dangerous governance smell:
sonar.exclusions=**/*
A dramatic metric improvement after excluding nearly all source is not a quality improvement. Preserve the before/after indexed-file count and the exclusion change. Restore the narrow intended scope and treat the exclusion proposal as a policy change requiring review.
5. Vendor/generated noise: diagnose ownership first
If generated or vendor files unexpectedly appear, ask: were they
inside sonar.sources? Did Git ignore them? Did a
project/global exclusion disappear? Did a suffix override cause
recognition? If the files are genuinely owned and shipped, excluding
them solely because they create issues can hide risk. If they are
non-owned third-party copies, move the policy boundary to a
documented source root rather than growing an opaque list of
arbitrary suppressions.
6. Shallow checkout: SCM is the failing evidence layer
git rev-parse --is-shallow-repository
git rev-parse HEAD
git remote -v
grep -Ei 'shallow|blame|scm|missing blame|could not find ref' scanner-debug.log
Repair by fetching sufficient/full history and rerun the same
revision/configuration if possible. Do not delete .git,
rewrite commits, or disable TLS/SCM merely to get a green analysis.
7. Sensor warning: missing or stale auxiliary report
Suppose a Python project configures
sonar.python.ruff.reportPaths=reports/ruff.json, but
the file is missing or was generated for a different revision.
Preserve:
- the scanner sensor warning,
- the command/version that should create the report,
- the report path and timestamp if it exists,
- the current Git revision.
Repair the report-generation stage or remove the import configuration if the integration is intentionally retired. Do not create a dummy report or reuse a stale artifact to make the log quiet.
8. Keep downstream failures separate
If indexing and sensors complete but upload fails, the problem has moved to authentication/network/server acceptance. If upload succeeds but Compute Engine fails, inspect the exact task. If CE succeeds but the gate fails, inspect policy and measures. Scope diagnosis should not mutate database/search state, restart the server, lower thresholds, grant admin tokens, or disable TLS.
9. Production practices
- Version-control project-specific scope where practical and capture server-side global/project settings in evidence.
- Keep forced global exclusions few, named, reviewed and testable against representative repositories.
- Run CI with full SCM history when blame/new-code semantics matter.
- Record scanner/analyzer/runtime and external-report generator versions.
- Track generated/vendor ownership and exclusion rationale as governance metadata.
- Alert on unexpected sharp drops in indexed files/LOC; they can indicate scope regression, not miraculous cleanup.
Knowledge check
What should you inspect first when an expected file is absent?
Base directory, initial source/test scope, effective filters/SCM ignore state, then language recognition—in that order.
Why is sonar.exclusions=**/* not a legitimate fix
for failing metrics?
It removes assurance input rather than improving code; the apparent metric improvement is caused by hiding files.
How do you repair a file indexed as both source and test?
Make source/test scopes and filtering rules disjoint; do not suppress the error or reclassify tests arbitrarily.
What evidence proves a missing external report is a sensor/input issue?
The sensor warning plus report configuration, generator command/version, path/timestamp and revision evidence.
When should database/search troubleshooting enter the sequence?
Only after evidence points past scanner/indexing/upload/policy layers toward server infrastructure; it is not a first response to scope errors.
Official references and version notes
- Community Build — Setting initial scope — source/test roots, simple-path rules, and project-base-directory semantics.
- Community Build — Path-based inclusions and exclusions — wildcard filtering after the initial scope.
- Community Build — Verifying analysis scope — debug-log indexing evidence and SonarScanner Context.
-
Community Build — File suffixes
— language recognition through
sonar.<language>.file.suffixes. - Community Build — Checked-out code and SCM integration — full-clone/blame requirements and shallow-clone behavior.
-
Community Build — Other scope adjustments
— SCM ignore behavior and
sonar.scm.exclusions.disabled. - Community Build — Supported languages — current analyzer/language matrix.
- Community Build — External issues and generic report format.
- Scanner environment requirements — current JRE auto-provisioning/runtime boundary.
Rechecked on 2026-09-07. Mandatory labs target local/private SonarQube Community Build 26.9.0.129388 with standalone SonarScanner CLI 8.1.0.6389. JRE auto-provisioning remains enabled; current scanner guidance requires Java 11 to launch CLI 7.2+ when provisioning is enabled, while environments that disable provisioning must supply a currently supported Java runtime (Java 21 is the safe current baseline). The mixed-source fixture uses languages supported by Community Build and does not require commercial analyzers, third-party plugins, CI providers, enterprise identity, branch/PR analysis, or external databases beyond the already running disposable lab server. Re-check language and scanner requirements before future runs because analyzer/runtime support evolves.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.