Language Analysis, Sensors, SCM Data, and Source-Code Indexing: Core Concepts and Mental Model
Understand how a workspace becomes a precise indexed-file set, how analyzers and sensors consume that set, and why SCM/report evidence changes interpretation without changing repository bytes.
Learning objectives
- Explain the difference between repository files, initial scope, filtered/indexed files, and files actually recognized by a language analyzer.
- Separate source/test classification from inclusion/exclusion filtering and language-suffix recognition.
- Explain what scanner-side sensors/analyzers and external-report importers contribute to an analysis report.
- Describe how SCM revision and blame metadata affect attribution and new-code behavior.
- Inspect scope and SCM state before changing any exclusion or suffix rule.
1. Chapter baseline: analysis starts by deciding what exists
Chapter 07 chose the scanner that matches the build ecosystem. Chapter 08 asks the next question: what exact files and metadata does that scanner present for analysis? SonarQube does not analyze “the repository” as an indivisible object. The scanner starts from a project base directory, establishes source and test scopes, applies SCM ignore behavior and configured filters, recognizes analyzable languages/file suffixes, invokes the relevant analyzer sensors, imports optional report data, and builds one analysis report.
2. Mental model: workspace to analysis report
Each arrow is a narrowing or enrichment step; later steps cannot analyze files that earlier scope steps removed.
flowchart TD W[Workspace + exact revision] --> B[Project base directory] B --> S[Initial source/test scope] S --> F[SCM ignore + inclusions/exclusions] F --> L[Language / suffix recognition] L --> A[Language analyzers + sensors] SCM[SCM revision + blame] --> A EXT[Optional external reports] --> A A --> R[Analysis report] R --> CE[Compute Engine] CE --> P[Issues / measures / policy result]
The project base directory anchors relative paths.
sonar.sources and sonar.tests establish
the initial source and test roots. Inclusions and exclusions filter
those roots; they do not add a file that was outside the initial
scope. Language recognition then decides which remaining files
belong to supported analyzers. Scanner-side analyzers/sensors
inspect files and import supporting evidence. SCM metadata enriches
line attribution and new-code decisions. The resulting report is
uploaded for asynchronous server processing.
3. State stores you must not collapse into “analysis configuration”
| State | Owned by | Examples | Why it matters |
|---|---|---|---|
| Workspace/revision | Git/checkout/CI | HEAD, worktree, shallow marker | Analysis evidence must map to exact bytes and history. |
| Base directory | Scanner invocation | current directory, sonar.projectBaseDir |
All relative scope/report paths depend on it. |
| Initial scope | Scanner/build model | sonar.sources, sonar.tests |
Defines files eligible to enter source/test analysis. |
| Filters | Global/project/scanner config | exclusions, inclusions, SCM ignore | Only narrows the candidate set. |
| Language recognition | Installed analyzer + suffix settings |
.py, .xml, configured suffixes
|
A file can be indexed yet not interpreted as the language you expected. |
| Sensor/report input | Analyzer/plugin/scanner | native analyzer data, coverage, external issues | Contributes findings/measures but has its own prerequisites. |
| SCM metadata | Git/SVN checkout | revision, blame, shallow/full history | Affects attribution/new-code behavior without changing file content. |
| Server result | Compute Engine + policy | issues, measures, gate | Exists only after uploaded report is processed. |
4. Initial scope versus filtering
For Scanner CLI, sonar.sources and
sonar.tests accept comma-delimited
simple paths, not wildcard expressions. If
sonar.sources=src,infra, a file under
docs/ cannot be pulled into analysis by setting
sonar.inclusions=docs/**/*.py; it was never in the
initial source set. By contrast, sonar.exclusions,
sonar.inclusions, sonar.test.exclusions,
and sonar.test.inclusions use path-matching patterns to
filter the appropriate set.
5. Language recognition happens after scope
A file may be inside scope yet not be recognized by the language you
intended. Many analyzers expose
sonar.<language>.file.suffixes. Changing a suffix
list changes language recognition, not directory scope. Use
overrides sparingly: ambiguous custom extensions can cause a file to
be claimed by the wrong analyzer or make later developers unable to
reproduce the mapping.
Community Build currently supports a broad language set including Java, JavaScript/TypeScript, Python, C#, Go, Kotlin, PHP, Ruby, Rust, Scala, XML and multiple IaC formats. Support and version ranges are language-specific, so “Community Build supports the language” is not proof that every language version or build context is supported.
6. What “sensor” means in this course
A sensor is scanner-side analyzer/plugin logic that consumes relevant indexed files or auxiliary reports and contributes analysis data. It is not a database table, Quality Gate, or CI job. A language analyzer may parse source and produce issues/measures; another sensor may import coverage or third-party findings. Sensor logs are therefore evidence of what client-side analysis ran and what inputs it consumed, not proof that the server accepted the report or that the Quality Gate passed.
7. SCM metadata is evidence, not source code
Git support is detected from repository metadata and Sonar uses blame/history to support line attribution and SCM-driven new-code behavior. Current guidance requires a full Git clone for reliable blame. A shallow clone can be detected, blame retrieval can be skipped, and analysis may fail. Do not respond by globally disabling SCM just to remove a warning; first verify checkout depth and whether history is an intended analysis input.
8. Read-only preflight before changing scope
pwd
git status --short
git rev-parse --show-toplevel
git rev-parse HEAD
git rev-parse --is-shallow-repository
git ls-files | sort | sed -n '1,120p'
git status --ignored --short | sed -n '1,120p'
find . -maxdepth 3 -type f -not -path './.git/*' | sort | sed -n '1,160p'
Record the repository root, exact revision, shallow/full status, tracked/ignored file sets, and filesystem inventory. These are independent observations. A file can exist on disk but be ignored by Git; a Git-ignored file is excluded by Sonar SCM behavior by default unless that behavior is deliberately disabled.
9. DevOps connection: scope is part of the control definition
A reproducible quality control is not “run SonarQube.” It is “analyze revision X with scanner/runtime Y, base directory Z, source/test scope S, filters F, language mapping L, SCM state M and report inputs R, then correlate the upload to Compute Engine result C and policy outcome Q.” Chapter 08 makes those hidden inputs reviewable.
Knowledge check
Can sonar.inclusions add a file that is outside
sonar.sources?
No. Inclusions and exclusions filter the initial scope; they do not expand it.
Why are source and test sets separate?
They use different rules/metrics and a physical code file cannot be indexed as both source and test.
Does a sensor log prove the Quality Gate passed?
No. It proves scanner-side analysis/import activity; server Compute Engine processing and gate evaluation are separate states.
What is the first response to a shallow-clone SCM warning?
Preserve the warning and inspect checkout depth/history. Restore a full valid checkout if SCM evidence is required instead of disabling SCM blindly.
What does a language suffix override change?
Language recognition for files matching those suffixes; it does not enlarge the initial source/test scope.
Official references and version notes
- Community Build — Setting initial scope — source/test roots, simple-path rules, and project-base-directory semantics.
- Community Build — Path-based inclusions and exclusions — wildcard filtering after the initial scope.
- Community Build — Verifying analysis scope — debug-log indexing evidence and SonarScanner Context.
-
Community Build — File suffixes
— language recognition through
sonar.<language>.file.suffixes. - Community Build — Checked-out code and SCM integration — full-clone/blame requirements and shallow-clone behavior.
-
Community Build — Other scope adjustments
— SCM ignore behavior and
sonar.scm.exclusions.disabled. - Community Build — Supported languages — current analyzer/language matrix.
- Community Build — External issues and generic report format.
- Scanner environment requirements — current JRE auto-provisioning/runtime boundary.
Rechecked on 2026-09-07. Mandatory labs target local/private SonarQube Community Build 26.9.0.129388 with standalone SonarScanner CLI 8.1.0.6389. JRE auto-provisioning remains enabled; current scanner guidance requires Java 11 to launch CLI 7.2+ when provisioning is enabled, while environments that disable provisioning must supply a currently supported Java runtime (Java 21 is the safe current baseline). The mixed-source fixture uses languages supported by Community Build and does not require commercial analyzers, third-party plugins, CI providers, enterprise identity, branch/PR analysis, or external databases beyond the already running disposable lab server. Re-check language and scanner requirements before future runs because analyzer/runtime support evolves.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.