Chapter 08Lesson 01~125 minutes

Language Analysis, Sensors, SCM Data, and Source-Code Indexing: Core Concepts and Mental Model

Understand how a workspace becomes a precise indexed-file set, how analyzers and sensors consume that set, and why SCM/report evidence changes interpretation without changing repository bytes.

IndexingSource vs testSensorsSCM blameAnalysis scope

Learning objectives

  • Explain the difference between repository files, initial scope, filtered/indexed files, and files actually recognized by a language analyzer.
  • Separate source/test classification from inclusion/exclusion filtering and language-suffix recognition.
  • Explain what scanner-side sensors/analyzers and external-report importers contribute to an analysis report.
  • Describe how SCM revision and blame metadata affect attribution and new-code behavior.
  • Inspect scope and SCM state before changing any exclusion or suffix rule.

1. Chapter baseline: analysis starts by deciding what exists

Chapter 07 chose the scanner that matches the build ecosystem. Chapter 08 asks the next question: what exact files and metadata does that scanner present for analysis? SonarQube does not analyze “the repository” as an indivisible object. The scanner starts from a project base directory, establishes source and test scopes, applies SCM ignore behavior and configured filters, recognizes analyzable languages/file suffixes, invokes the relevant analyzer sensors, imports optional report data, and builds one analysis report.

Governance boundary: exclusions are policy inputs. Removing code from scope can make dashboards look better while reducing assurance. Every exclusion, inclusion, suffix override, generated/vendor decision, and intentional SCM disablement needs a reason, owner, evidence, and rollback path.

2. Mental model: workspace to analysis report

Discovery, classification, recognition, sensing, report

Each arrow is a narrowing or enrichment step; later steps cannot analyze files that earlier scope steps removed.

flowchart TD
W[Workspace + exact revision] --> B[Project base directory]
B --> S[Initial source/test scope]
S --> F[SCM ignore + inclusions/exclusions]
F --> L[Language / suffix recognition]
L --> A[Language analyzers + sensors]
SCM[SCM revision + blame] --> A
EXT[Optional external reports] --> A
A --> R[Analysis report]
R --> CE[Compute Engine]
CE --> P[Issues / measures / policy result]

The project base directory anchors relative paths. sonar.sources and sonar.tests establish the initial source and test roots. Inclusions and exclusions filter those roots; they do not add a file that was outside the initial scope. Language recognition then decides which remaining files belong to supported analyzers. Scanner-side analyzers/sensors inspect files and import supporting evidence. SCM metadata enriches line attribution and new-code decisions. The resulting report is uploaded for asynchronous server processing.

3. State stores you must not collapse into “analysis configuration”

State Owned by Examples Why it matters
Workspace/revision Git/checkout/CI HEAD, worktree, shallow marker Analysis evidence must map to exact bytes and history.
Base directory Scanner invocation current directory, sonar.projectBaseDir All relative scope/report paths depend on it.
Initial scope Scanner/build model sonar.sources, sonar.tests Defines files eligible to enter source/test analysis.
Filters Global/project/scanner config exclusions, inclusions, SCM ignore Only narrows the candidate set.
Language recognition Installed analyzer + suffix settings .py, .xml, configured suffixes A file can be indexed yet not interpreted as the language you expected.
Sensor/report input Analyzer/plugin/scanner native analyzer data, coverage, external issues Contributes findings/measures but has its own prerequisites.
SCM metadata Git/SVN checkout revision, blame, shallow/full history Affects attribution/new-code behavior without changing file content.
Server result Compute Engine + policy issues, measures, gate Exists only after uploaded report is processed.

4. Initial scope versus filtering

For Scanner CLI, sonar.sources and sonar.tests accept comma-delimited simple paths, not wildcard expressions. If sonar.sources=src,infra, a file under docs/ cannot be pulled into analysis by setting sonar.inclusions=docs/**/*.py; it was never in the initial source set. By contrast, sonar.exclusions, sonar.inclusions, sonar.test.exclusions, and sonar.test.inclusions use path-matching patterns to filter the appropriate set.

Disjoint-set invariant: one physical code file cannot be both source and test. If scope rules place it in both sets, the scanner reports that the file cannot be indexed twice. Fix the classification instead of suppressing the error.

5. Language recognition happens after scope

A file may be inside scope yet not be recognized by the language you intended. Many analyzers expose sonar.<language>.file.suffixes. Changing a suffix list changes language recognition, not directory scope. Use overrides sparingly: ambiguous custom extensions can cause a file to be claimed by the wrong analyzer or make later developers unable to reproduce the mapping.

Community Build currently supports a broad language set including Java, JavaScript/TypeScript, Python, C#, Go, Kotlin, PHP, Ruby, Rust, Scala, XML and multiple IaC formats. Support and version ranges are language-specific, so “Community Build supports the language” is not proof that every language version or build context is supported.

6. What “sensor” means in this course

A sensor is scanner-side analyzer/plugin logic that consumes relevant indexed files or auxiliary reports and contributes analysis data. It is not a database table, Quality Gate, or CI job. A language analyzer may parse source and produce issues/measures; another sensor may import coverage or third-party findings. Sensor logs are therefore evidence of what client-side analysis ran and what inputs it consumed, not proof that the server accepted the report or that the Quality Gate passed.

7. SCM metadata is evidence, not source code

Git support is detected from repository metadata and Sonar uses blame/history to support line attribution and SCM-driven new-code behavior. Current guidance requires a full Git clone for reliable blame. A shallow clone can be detected, blame retrieval can be skipped, and analysis may fail. Do not respond by globally disabling SCM just to remove a warning; first verify checkout depth and whether history is an intended analysis input.

8. Read-only preflight before changing scope

pwd
git status --short
git rev-parse --show-toplevel
git rev-parse HEAD
git rev-parse --is-shallow-repository
git ls-files | sort | sed -n '1,120p'
git status --ignored --short | sed -n '1,120p'
find . -maxdepth 3 -type f -not -path './.git/*' | sort | sed -n '1,160p'

Record the repository root, exact revision, shallow/full status, tracked/ignored file sets, and filesystem inventory. These are independent observations. A file can exist on disk but be ignored by Git; a Git-ignored file is excluded by Sonar SCM behavior by default unless that behavior is deliberately disabled.

9. DevOps connection: scope is part of the control definition

A reproducible quality control is not “run SonarQube.” It is “analyze revision X with scanner/runtime Y, base directory Z, source/test scope S, filters F, language mapping L, SCM state M and report inputs R, then correlate the upload to Compute Engine result C and policy outcome Q.” Chapter 08 makes those hidden inputs reviewable.

Knowledge check

Can sonar.inclusions add a file that is outside sonar.sources?

Why are source and test sets separate?

Does a sensor log prove the Quality Gate passed?

What is the first response to a shallow-clone SCM warning?

What does a language suffix override change?

Next lesson

Prove the indexed set with a mixed-source fixture

Lesson 2 builds a disposable repository, records verbose indexing/sensor evidence, changes one exclusion and one suffix mapping, then demonstrates shallow-SCM behavior safely.

Official references and version notes

Version and compatibility note

Rechecked on 2026-09-07. Mandatory labs target local/private SonarQube Community Build 26.9.0.129388 with standalone SonarScanner CLI 8.1.0.6389. JRE auto-provisioning remains enabled; current scanner guidance requires Java 11 to launch CLI 7.2+ when provisioning is enabled, while environments that disable provisioning must supply a currently supported Java runtime (Java 21 is the safe current baseline). The mixed-source fixture uses languages supported by Community Build and does not require commercial analyzers, third-party plugins, CI providers, enterprise identity, branch/PR analysis, or external databases beyond the already running disposable lab server. Re-check language and scanner requirements before future runs because analyzer/runtime support evolves.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.