Chapter 08Lesson 03~130 minutes

Language Analysis, Sensors, SCM Data, and Source-Code Indexing: Configuration, Design Patterns, and Trade-Offs

Design analysis scope as a reviewable contract: choose explicit or build-inferred roots, global or project filters, generated/vendor treatment, SCM policy, monorepo boundaries, and report imports deliberately.

Scope designGenerated codeMonoreposSCM policyExternal reports

Learning objectives

  • Compare explicit Scanner CLI scope with build-inferred Maven/Gradle/.NET scope.
  • Choose global, project, and scanner-host exclusions with attention to precedence and auditability.
  • Distinguish excluding a file from analysis from excluding only issues, coverage, or duplication.
  • Design generated/vendor and SCM policies without improving metrics by hiding risk.
  • Define monorepo/project and external-report boundaries that remain reproducible.

1. Scope-design decision table

Decision Safer default When to deviate Evidence / rollback
Explicit vs inferred roots Use build-derived roots for native build scanners; explicit roots for CLI Repository layout/build metadata is unusual scanner context, indexed-file logs, config diff
Global vs project exclusions Keep forced global exclusions rare and documented organization-wide generated/vendor convention global setting owner + affected projects + rollback
Generated code Exclude only when ownership/value justifies it generated code is edited or production-critical generator provenance, path rule, exception owner
Vendor code Usually outside project-owned source scope forked/vendor code is maintained internally ownership/license boundary, source root decision
SCM integration Keep enabled with full checkout controlled case where SCM data is intentionally unavailable checkout evidence, explicit rationale, new-code/attribution impact
Monorepo boundary One project per meaningful ownership/release unit shared release/policy justifies a broader project base dirs, project keys, no overlapping ownership
External reports Import when provenance/format is stable third-party analyzer is authoritative for a specific rule set report generator version, path, timestamp, rule owner

2. Explicit versus inferred source scopes

Scanner CLI typically needs explicit roots when the repository has separate source/test trees. Maven and Gradle derive source/test information from their build model; .NET discovers files through MSBuild integration and does not use user-defined sonar.sources/sonar.tests in the same way. Forcing CLI-style scope parameters into a native build scanner can make the lesson appear portable while actually overriding correct build metadata.

3. Global, project, and scanner-host filters

Some analysis-scope settings can live in the server UI at global/project level and can be overridden by scanner-host parameters; forced global exclusions are intentionally stronger. Use central settings for stable organization-wide policy and repository/CI settings for project-specific, version-controlled decisions. Whichever location you choose, preserve the effective scanner context for the analysis run—declared configuration is not evidence until the scanner processes it.

4. “Exclude” is not one operation

Mechanism Removes file from analysis? Typical intent
sonar.exclusions / test equivalent Yes, from the corresponding analyzed code set File should not participate as source/test analysis input.
Language suffix settings May prevent language analyzer recognition File extension should/should not map to a language.
Ignore Issues on Files / blocks No; other measures can remain Suppress issue generation in known generated/content regions.
sonar.coverage.exclusions No Exclude file from coverage calculation only.
sonar.cpd.exclusions No Exclude file from duplication detection only.
SCM ignore behavior Typically removes SCM-ignored files from analysis Respect repository ignore contract.

If a team wants generated code to remain visible for LOC/duplication but not produce selected issues, a path exclusion is too broad. If a file must not be analyzed at all, an issue-only exclusion is too narrow. Match the mechanism to the assurance question.

5. Generated and vendor code: ownership before metrics

Generated code is not automatically worthless. Ask whether developers edit it, whether the generator is controlled, whether defects in generated output can reach production, and whether scanning it creates actionable signal. Some analyzers have language-specific generated-code behavior; for example C# skips tool-generated code by default unless configured otherwise. Vendor code raises a different question: who can remediate it? If the team does not own it, scanning may create noise; if the team maintains a fork, excluding it as “vendor” can hide owned risk.

6. SCM enabled versus intentionally disabled

SCM metadata supports blame, attribution and new-code reasoning. Disabling SCM changes evidence semantics, not just performance. If an isolated source package truly has no usable history, document the limitation and understand that attribution/new-code behavior differs. Never make “no shallow warning” the objective; make “valid reproducible SCM evidence, or a documented absence of it” the objective.

7. Monorepo boundaries

A monorepo can contain several products with different release owners, quality profiles, test ecosystems and deployment schedules. One giant Sonar project may simplify administration while creating misleading aggregate policy. Multiple projects improve ownership but require non-overlapping base directories, stable project keys, explicit CI triggers and careful shared-code treatment. Treat the project boundary as governance architecture, not a path glob.

One repository, multiple analysis contracts
flowchart TD
R[Monorepo revision] --> A[service-a baseDir]
R --> B[service-b baseDir]
R --> S[shared library]
A --> PA[Project key A / policy A]
B --> PB[Project key B / policy B]
S --> D{Ownership choice}
D -->|owned by A| PA
D -->|separate lifecycle| PS[Project key Shared]

8. Sensors and external reports: provenance matters

Community Build can import findings from supported third-party analyzers and generic/SARIF-style reports. Those findings become part of the analysis report, but the external tool still owns rule activation and its native issue state. A false-positive action in Sonar does not synchronize back to the external analyzer. Record which tool/version generated the report, which revision it analyzed, where the report lives, and which scanner parameter imported it.

# Example generic external-issue import path (optional; not required by the lab)
sonar.externalIssuesReportPaths=reports/external-issues.json

Do not create an empty/stale file merely to make the sensor log quiet. A missing expected report is a build/evidence defect to diagnose.

9. Worked scenario

A monorepo contains services/payments, services/catalog, generated/sdk and vendor/crypto. Payments and Catalog release independently. The SDK is regenerated from an owned schema and never edited. Crypto is an upstream dependency copied for an air-gapped build.

  1. Create separate stable project keys/base directories for Payments and Catalog because ownership/release policies differ.
  2. Do not include vendor/crypto merely to increase “coverage”; track its provenance and vulnerability management outside this source-quality project unless the team maintains a fork.
  3. For generated/sdk, decide whether to exclude whole files or only issues based on whether generated defects are actionable; record generator/schema version.
  4. Keep full Git history in CI. If external analyzer reports are imported, correlate each report to the same revision as the Sonar scan.

Knowledge check

Why can a path exclusion be too broad for generated code?

What is the main risk of a forced global exclusion?

Should Scanner CLI scope parameters be copied mechanically into .NET analysis?

Who owns rule activation for imported external issues?

What should determine a monorepo Sonar project boundary?

Next lesson

Diagnose what the scanner actually did

Lesson 4 turns scope/sensor/SCM mistakes into evidence-first failure trees, including over-broad exclusions, source/test overlap, vendor noise, shallow checkouts, and ignored sensor warnings.

Official references and version notes

Version and compatibility note

Rechecked on 2026-09-07. Mandatory labs target local/private SonarQube Community Build 26.9.0.129388 with standalone SonarScanner CLI 8.1.0.6389. JRE auto-provisioning remains enabled; current scanner guidance requires Java 11 to launch CLI 7.2+ when provisioning is enabled, while environments that disable provisioning must supply a currently supported Java runtime (Java 21 is the safe current baseline). The mixed-source fixture uses languages supported by Community Build and does not require commercial analyzers, third-party plugins, CI providers, enterprise identity, branch/PR analysis, or external databases beyond the already running disposable lab server. Re-check language and scanner requirements before future runs because analyzer/runtime support evolves.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.