Language Analysis, Sensors, SCM Data, and Source-Code Indexing: Guided Hands-On Workflow
Build and analyze a disposable mixed-language repository, prove its indexed-file set from debug logs, change scope deliberately, and reproduce a shallow-SCM warning without touching production data.
Learning objectives
- Create a synthetic mixed-source Git fixture with explicit source/test boundaries.
- Run a baseline analysis with debug logging and preserve indexed-file and sensor evidence.
- Change one exclusion and one language suffix rule, then compare the effective indexed set.
- Reproduce shallow-clone SCM behavior in a disposable local clone and preserve the warning.
- Clean up tokens/projects/fixtures without deleting reusable scanner caches or unrelated server state.
1. Lab assumptions and preflight
Use the disposable local Community Build 26.9.0.129388 instance from earlier chapters and Scanner CLI 8.1.0.6389. The fixture contains Python, XML, YAML, Dockerfile/IaC-style text, tests, generated code, and ignored scratch data. No proprietary source, public endpoint, CI provider, commercial edition, external plugin, or real credential is required.
export SONAR_HOST_URL='http://127.0.0.1:9000'
curl -fsS "$SONAR_HOST_URL/api/system/status"
sonar-scanner --version
git --version
java -version || true
academy-sonarqube-p08 through the local UI/API
using the lab administrator only for setup, then use that narrowly
scoped token for analysis. Never paste the token into
sonar-project.properties, shell history, screenshots,
or the evidence packet.
2. Create the mixed-source fixture
rm -rf sonar-p08 sonar-p08-shallow
mkdir sonar-p08 && cd sonar-p08
git init
git config user.email 'academy@example.invalid'
git config user.name 'DevOps Academy Lab'
mkdir -p src/generated tests infra vendor scratch
cat > src/app.py <<'PY'
def classify(value):
if value < 0:
return "negative"
return "non-negative"
PY
cat > src/helper.academy <<'PY'
def helper(name):
return "hello " + name
PY
cat > src/generated/client.py <<'PY'
# synthetic generated fixture
def generated(x):
return x
PY
cat > tests/test_app.py <<'PY'
from src.app import classify
def test_zero():
assert classify(0) == "non-negative"
PY
cat > infra/service.yaml <<'YAML'
apiVersion: v1
kind: ConfigMap
metadata:
name: academy-p08
data:
mode: lab
YAML
cat > infra/settings.xml <<'XML'
<settings><mode>lab</mode></settings>
XML
cat > vendor/vendor.py <<'PY'
def vendored():
return "outside initial scope"
PY
cat > scratch/ignored.py <<'PY'
print("ignored scratch")
PY
cat > .gitignore <<'EOF'
scratch/
.scannerwork/
*.log
EOF
cat > sonar-project.properties <<'EOF'
sonar.projectKey=academy-sonarqube-p08
sonar.projectName=Academy SonarQube P08 Indexing Lab
sonar.sources=src,infra
sonar.tests=tests
sonar.sourceEncoding=UTF-8
EOF
git add . && git commit -m 'mixed-source indexing fixture'
git rev-parse HEAD | tee revision-baseline.txt
find . -maxdepth 3 -type f -not -path './.git/*' | sort
Predict before scanning: vendor/vendor.py is outside
sonar.sources; scratch/ignored.py is
Git-ignored and outside scope anyway;
tests/test_app.py is test code;
src/generated/client.py is initially eligible source;
and src/helper.academy is inside the source directory
but should not yet be recognized as Python.
3. Baseline run: preserve verbose indexing and sensor evidence
read -rsp 'Project-analysis token: ' SONAR_TOKEN; echo
export SONAR_TOKEN
sonar-scanner -X -Dsonar.host.url="$SONAR_HOST_URL" 2>&1 | tee baseline-debug.log
cp .scannerwork/report-task.txt baseline-report-task.txt
grep -Ei 'indexed|excluded|included|sensor|scm|blame|language|project configuration' baseline-debug.log | tee baseline-scope-evidence.txt
Do not count only log lines that contain the word “indexed”; scanner
versions can format debug output differently. Preserve the complete
debug log, then use the Project configuration/indexing/sensor
sections to reconstruct the effective set. Correlate
ceTaskId from report-task.txt with the
server background task before interpreting measures.
4. Change exactly one path exclusion and compare
Add one source exclusion for the synthetic generated directory:
printf '
sonar.exclusions=src/generated/**
' >> sonar-project.properties
git add sonar-project.properties && git commit -m 'exclude generated fixture'
git rev-parse HEAD | tee revision-exclusion.txt
sonar-scanner -X -Dsonar.host.url="$SONAR_HOST_URL" 2>&1 | tee exclusion-debug.log
cp .scannerwork/report-task.txt exclusion-report-task.txt
grep -Ei 'excluded|indexed|project configuration|sensor' exclusion-debug.log | tee exclusion-scope-evidence.txt
Prediction: the source candidate set becomes smaller because
src/generated/client.py is filtered out. The test root
does not change. vendor/ still cannot appear because it
remains outside the initial scope. Verify from logs and resulting
measures rather than assuming the configuration was honored.
5. Change one language suffix rule
Now teach Sonar that .academy is also a Python suffix
for this disposable project. This is intentionally unusual so the
effect is obvious and reversible.
printf '
sonar.python.file.suffixes=.py,.academy
' >> sonar-project.properties
git add sonar-project.properties && git commit -m 'recognize academy suffix as python'
git rev-parse HEAD | tee revision-suffix.txt
sonar-scanner -X -Dsonar.host.url="$SONAR_HOST_URL" 2>&1 | tee suffix-debug.log
cp .scannerwork/report-task.txt suffix-report-task.txt
grep -Ei 'helper\.academy|python|language|indexed|sensor' suffix-debug.log | tee suffix-recognition-evidence.txt
Prediction: src/helper.academy remains in the same
source directory but changes from unrecognized/custom text to a
Python-analyzer input. That demonstrates why language recognition is
a distinct layer from directory scope.
6. Prove SCM-ignore behavior without bypassing it
git check-ignore -v scratch/ignored.py
git status --ignored --short | grep scratch || true
grep -F 'scratch/ignored.py' baseline-debug.log || true
Sonar respects SCM ignore directives by default. The property
sonar.scm.exclusions.disabled=true can disable that
behavior, but this lab does not use it: forcing ignored
scratch/vendor artifacts into analysis merely to prove a switch
would teach the wrong production habit.
7. Demonstrate shallow SCM metadata safely
Create a local shallow clone of the same synthetic repository. A
file:// URL is used so Git actually honors
--depth 1 for a local source.
cd ..
git clone --depth 1 "file://$PWD/sonar-p08" sonar-p08-shallow
cd sonar-p08-shallow
git rev-parse HEAD | tee shallow-revision.txt
git rev-parse --is-shallow-repository | tee shallow-state.txt
# Same project key/token: this is a controlled comparison of checkout state.
sonar-scanner -X -Dsonar.host.url="$SONAR_HOST_URL" 2>&1 | tee shallow-debug.log || true
grep -Ei 'shallow|blame|scm|missing blame|could not find ref' shallow-debug.log | tee shallow-scm-evidence.txt
Expected diagnostic shape: the scanner detects shallow history, cannot rely on normal blame retrieval, and may fail depending on the exact repository/analyzer state. Preserve the actual result. Do not claim a successful analysis if the scanner stopped.
8. Repair the SCM defect, not the evidence
git fetch --unshallow 2>/dev/null || git fetch --depth=2147483647
git rev-parse --is-shallow-repository | tee repaired-shallow-state.txt
sonar-scanner -X -Dsonar.host.url="$SONAR_HOST_URL" 2>&1 | tee full-history-debug.log
grep -Ei 'shallow|blame|scm' full-history-debug.log | tee full-history-scm-evidence.txt
The desired change is checkout history, not Sonar policy. Compare the exact revision and configuration across the shallow and repaired runs so the causal variable is clear.
9. Small challenge: diagnose a missing file
A teammate expects vendor/vendor.py in the Python
analyzer log and adds sonar.inclusions=vendor/**/*.py.
It still does not appear. Explain why and propose the
least-surprising fix. Correct reasoning starts with the initial
source roots: inclusions cannot expand
sonar.sources=src,infra. Decide whether vendor code
should be analyzed at all before changing
sonar.sources.
10. Guarded cleanup
-
Preserve sanitized debug logs, revisions,
report-task.txtfiles and CE/gate evidence. -
Revoke the disposable project-analysis token and unset
SONAR_TOKEN. -
Delete only
sonar-p08andsonar-p08-shallowafter evidence is copied. - Delete the disposable project only if later lessons will not reuse it.
- Do not delete shared scanner caches, server data/search volumes, global exclusions, or unrelated projects.
Knowledge check
Why did vendor/vendor.py stay absent even after a
matching inclusion pattern?
Because it was outside the initial
sonar.sources roots; inclusions only filter files
already in the candidate set.
What changed when .academy was added to Python
suffixes?
Language recognition changed for an already in-scope file; the directory/source scope itself did not expand.
What is the causal repair for shallow-clone blame problems?
Restore sufficient/full Git history and verify repository metadata, rather than disabling SCM globally.
Why keep three report-task files?
Each scan produces a distinct report/Compute Engine task that must be correlated to its own configuration/revision evidence.
Why is sonar.scm.exclusions.disabled=true omitted
from the mandatory lab?
Because overriding SCM ignore behavior merely to increase analyzed files is usually unnecessary and can pull intentionally ignored artifacts into scope.
Official references and version notes
- Community Build — Setting initial scope — source/test roots, simple-path rules, and project-base-directory semantics.
- Community Build — Path-based inclusions and exclusions — wildcard filtering after the initial scope.
- Community Build — Verifying analysis scope — debug-log indexing evidence and SonarScanner Context.
-
Community Build — File suffixes
— language recognition through
sonar.<language>.file.suffixes. - Community Build — Checked-out code and SCM integration — full-clone/blame requirements and shallow-clone behavior.
-
Community Build — Other scope adjustments
— SCM ignore behavior and
sonar.scm.exclusions.disabled. - Community Build — Supported languages — current analyzer/language matrix.
- Community Build — External issues and generic report format.
- Scanner environment requirements — current JRE auto-provisioning/runtime boundary.
Rechecked on 2026-09-07. Mandatory labs target local/private SonarQube Community Build 26.9.0.129388 with standalone SonarScanner CLI 8.1.0.6389. JRE auto-provisioning remains enabled; current scanner guidance requires Java 11 to launch CLI 7.2+ when provisioning is enabled, while environments that disable provisioning must supply a currently supported Java runtime (Java 21 is the safe current baseline). The mixed-source fixture uses languages supported by Community Build and does not require commercial analyzers, third-party plugins, CI providers, enterprise identity, branch/PR analysis, or external databases beyond the already running disposable lab server. Re-check language and scanner requirements before future runs because analyzer/runtime support evolves.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.