Performance Sizing, Reference Architectures, and Large-Instance Tuning: Guided Hands-On Workflow
Measure repeated local analyses, preserve task evidence, change one workload factor and compare scanner, queue and host behavior in a disposable Community Build lab.
Learning objectives
- Start a disposable Community Build 26.9.0.129388 + PostgreSQL fixture and record exact image/runtime identities.
- Create a deterministic synthetic repository and preserve an exact Git revision for repeated measurements.
- Capture scanner elapsed time, report-task identity, CE status/timing, host/container metrics and read-only PostgreSQL counters.
- Separate warm-up effects from repeated baseline measurements, then change exactly one workload factor.
- Create a controlled burst and decide whether the first limiting layer is scanner, CE queue, database/search, host or network.
1. Scenario and preflight: a bounded lab, not a synthetic “production benchmark”
The goal is not to publish a universal benchmark number. It is to practice the measurement discipline you would use in production while keeping every mutation inside resources you own. The lab uses the exact Community Build release currently published at this chapter recheck, a supported PostgreSQL major version, a synthetic Python repository and a stable project key. No proprietary code, external CI provider, paid feature or real credential is required.
Use a machine with enough free RAM and disk for Docker Desktop/Engine plus the SonarQube and PostgreSQL containers. Record host CPU model/core count, physical RAM, free disk, Docker version and storage location before the run. If the host is already under memory or I/O pressure, mark that limitation rather than pretending the measurements are clean.
| Component | Lab assumption | Why it is recorded |
|---|---|---|
| SonarQube | Community Build 26.9.0.129388 official image | Pins server behavior and keeps mandatory path free |
| Database |
Official postgres:17; record
SELECT version() and image digest
|
PostgreSQL 14–18 is supported in current 2026.1 Server guidance; exact patch/digest is evidence |
| Scanner | SonarScanner CLI 8.1.0.6389 or recorded later compatible version | Scanner runtime and analyzer bootstrap can affect client-side time |
| Java | Scanner embedded/JRE auto-provisioned runtime where applicable; otherwise Java 21+ | Current scanner runtime requirements have moved beyond Java 17 |
| Plugins | No third-party plugins | Removes an uncontrolled performance variable |
| CI/IdP | None; local shell only | Keeps provider/auth latency out of the mandatory baseline |
2. Start the isolated server and prove its identity before analysis
Create a new empty directory such as sq31-capacity-lab.
Save the Compose file below as compose.yaml. The
password is intentionally fake and local-only; do not reuse it
elsewhere. The SonarQube image tag is pinned to the current
Community Build. PostgreSQL is major-version pinned, so the
preflight explicitly records its exact patch and image digest.
name: sq31-capacity-lab
services:
db:
image: postgres:17
container_name: sq31-postgres
environment:
POSTGRES_USER: sonar
POSTGRES_PASSWORD: local-only-sq31-db
POSTGRES_DB: sonar
volumes:
- sq31_db:/var/lib/postgresql/data
sonarqube:
image: sonarqube:26.9.0.129388-community
container_name: sq31-sonarqube
depends_on:
- db
environment:
SONAR_JDBC_URL: jdbc:postgresql://db:5432/sonar
SONAR_JDBC_USERNAME: sonar
SONAR_JDBC_PASSWORD: local-only-sq31-db
ports:
- "9000:9000"
volumes:
- sq31_data:/opt/sonarqube/data
- sq31_logs:/opt/sonarqube/logs
- sq31_extensions:/opt/sonarqube/extensions
volumes:
sq31_db:
sq31_data:
sq31_logs:
sq31_extensions:
docker compose up -d
# Wait until the web process reports UP, then record identities.
curl -fsS http://localhost:9000/api/system/status
curl -fsS http://localhost:9000/api/server/version
docker image inspect sonarqube:26.9.0.129388-community --format '{{json .RepoDigests}}'
docker image inspect postgres:17 --format '{{json .RepoDigests}}'
docker exec sq31-postgres psql -U sonar -d sonar -c 'select version();'
docker stats --no-stream sq31-sonarqube sq31-postgres
Log in only to this disposable instance, change the default
administrator password, create the project(s) used by the lab, and
create a local analysis identity with only the permissions required
to execute analysis on those projects. Export its token through
SONAR_TOKEN; do not place it in
sonar-project.properties, screenshots, Git history or
the evidence packet.
# PowerShell — paste the local token only into this process environment.
$env:SONAR_HOST_URL = "http://localhost:9000"
$env:SONAR_TOKEN = "<LOCAL_LAB_TOKEN>"
sonar-scanner -v
# Bash equivalent
export SONAR_HOST_URL='http://localhost:9000'
read -s -p 'Local Sonar token: ' SONAR_TOKEN; export SONAR_TOKEN; echo
The value represented by
local-only-sq31-health system
passcode only so the read-only
api/system/health call can run without granting the
scanner identity Administer System; in production,
treat any system passcode as a secret and inject it from a
protected secret mechanism.
3. Generate a deterministic repository and freeze the baseline revision
The fixture generator produces many small Python files with predictable structure. We use file count as a controlled workload factor because it affects scanner indexing and the amount of analyzed source without introducing external dependencies. This is not a claim that “N Python files equals N LOC of Java.” It is deliberately one synthetic shape whose limitations remain in the evidence packet.
from pathlib import Path
import sys
count = int(sys.argv[1])
root = Path('src')
root.mkdir(exist_ok=True)
for old in root.glob('module_*.py'):
old.unlink()
for i in range(count):
(root / f'module_{i:04d}.py').write_text(
f"def normalize_{i}(value):\n"
f" if value is None:\n return 0\n"
f" return (value * {i+1}) % 997\n"
f"\ndef bucket_{i}(items):\n"
f" return [normalize_{i}(x) for x in items]\n",
encoding='utf-8'
)
print(f'generated {count} Python files')
mkdir sq31-fixture
cd sq31-fixture
git init
git config user.email "sq31@example.invalid"
git config user.name "SQ31 Local Lab"
python ../make_fixture.py 120
printf "sonar.projectKey=sq31-capacity\nsonar.sources=src\nsonar.sourceEncoding=UTF-8\n" > sonar-project.properties
git add .
git commit -m "baseline 120-file fixture"
git tag sq31-baseline
git rev-parse HEAD
Write the resulting commit SHA into
evidence/assumptions.md. The first run is a warm-up run
because scanner/analyzer caches, JIT compilation, database caches
and OS page cache may not resemble later runs. Do not silently
discard it; label it warm-up and keep it so reviewers can
see that the test design accounted for cache state.
4. Run the baseline and preserve the asynchronous evidence chain
The PowerShell runner below creates a timestamped evidence
directory, records the revision and scanner version, measures
client-side elapsed time, preserves report-task.txt,
polls the exact CE task until terminal state and records a
host/container snapshot. It never treats scanner exit code alone as
“analysis complete.”
$ErrorActionPreference = "Stop"
$ProjectKey = "sq31-capacity"
$Run = Get-Date -Format "yyyyMMdd-HHmmss"
$Evidence = Join-Path $PWD "evidence/$Run"
New-Item -ItemType Directory -Force $Evidence | Out-Null
git rev-parse HEAD | Tee-Object "$Evidence/revision.txt"
sonar-scanner -v | Tee-Object "$Evidence/scanner-version.txt"
$sw = [Diagnostics.Stopwatch]::StartNew()
sonar-scanner -Dsonar.projectKey=$ProjectKey -Dsonar.sources=src -Dsonar.sourceEncoding=UTF-8 2>&1 | Tee-Object "$Evidence/scanner.log"
$sw.Stop()
$sw.Elapsed.TotalSeconds | Set-Content "$Evidence/scanner-seconds.txt"
Copy-Item .scannerwork/report-task.txt "$Evidence/report-task.txt"
$ceTaskId = ((Get-Content .scannerwork/report-task.txt) | Where-Object { $_ -like 'ceTaskId=*' }).Split('=')[1]
$headers = @{ Authorization = "Bearer $env:SONAR_TOKEN" }
do {
$task = Invoke-RestMethod "$env:SONAR_HOST_URL/api/ce/task?id=$ceTaskId" -Headers $headers
$task | ConvertTo-Json -Depth 8 | Set-Content "$Evidence/ce-task-latest.json"
if ($task.task.status -in @('SUCCESS','FAILED','CANCELED')) { break }
Start-Sleep -Seconds 1
} while ($true)
Invoke-RestMethod "$env:SONAR_HOST_URL/api/system/health" -Headers $headers | ConvertTo-Json -Depth 8 | Set-Content "$Evidence/system-health.json"
docker stats --no-stream --format "{{json .}}" sq31-sonarqube sq31-postgres | Set-Content "$Evidence/docker-stats.jsonl"
Write-Host "Evidence: $Evidence ; CE task: $ceTaskId ; status: $($task.task.status)"
Run one warm-up plus at least three measured repetitions on the same
revision. After each task reaches SUCCESS, append a row
to a worksheet. If a task reaches FAILED or
CANCELED, stop the benchmark and diagnose that task; do
not average a failed state into a performance number.
| run | revision | scanner s | queue wait s | CE processing s | SQ CPU/RAM snapshot | DB counters delta | notes |
|---|---|---|---|---|---|---|---|
| warm-up | same SHA | record | record | record | record | record | cache/JIT warm-up |
| B1 | same SHA | record | record | record | record | record | measured |
| B2 | same SHA | record | record | record | record | record | measured |
| B3 | same SHA | record | record | record | record | record | measured |
5. Derive queue wait and service time instead of lumping them together
The CE task response contains timestamps such as submission, start and execution/completion fields depending on the current endpoint representation. Preserve the raw JSON first. Then calculate queue wait from submit-to-start and processing duration from start-to-finish using fields actually present in your baseline. If the endpoint schema differs, use the API documentation embedded in that exact SonarQube instance; do not invent a field name to make a worksheet work.
# Example derivation after inspecting the raw task JSON.
# Pseudocode — map these names to fields actually returned by your version.
queue_wait_s = (started_at - submitted_at).total_seconds()
ce_service_s = (finished_at - started_at).total_seconds()
end_to_end_server_s = queue_wait_s + ce_service_s
The Web API V2 transition is ongoing. Keep the raw response, server version and endpoint documentation together. Read-only task/health calls used in this lab do not justify building long-lived automation that assumes an undocumented schema.
6. Take read-only database, disk and search-adjacent snapshots
We do not query or edit SonarQube’s internal tables. PostgreSQL exposes supported operational statistics that can be read without knowing application schema. Capture the counters before and after a bounded run window. For search, use SonarQube health/system information, container/host disk metrics and SonarQube logs; do not reach into embedded Elasticsearch indices or mutate them directly.
# PostgreSQL operational counters — read-only.
docker exec sq31-postgres psql -U sonar -d sonar -c "
select datname,numbackends,xact_commit,xact_rollback,
blks_read,blks_hit,temp_files,temp_bytes,
tup_returned,tup_fetched,tup_inserted,tup_updated,tup_deleted
from pg_stat_database where datname='sonar';"
# Storage and process snapshots.
docker stats --no-stream sq31-sonarqube sq31-postgres
docker exec sq31-sonarqube sh -lc 'df -h /opt/sonarqube/data; du -sh /opt/sonarqube/data /opt/sonarqube/logs'
docker logs --since 10m sq31-sonarqube > evidence/server-window.log 2>&1
7. Change exactly one workload factor: source-file count
Now change only the source workload. Generate 480 files instead of 120, commit the change and tag it. Keep the same server resources, database, scanner family/version, project key, token permissions and host. This makes the comparison interpretable. The new commit is part of the evidence, not noise to hide.
python ../make_fixture.py 480
git add src
git commit -m "workload factor: 480 files"
git tag sq31-large
git rev-parse HEAD
# Repeat one warm-up and at least three measured runs with the same runner.
pwsh ../run-analysis.ps1
Compare scanner time and CE processing time separately. If scanner time grows much more than CE time, the client-side workload changed more than the server service demand. If CE time grows substantially, correlate it with database, disk and process evidence. Do not conclude that “4× files needs 4× RAM”; the measurement describes this fixture on this host, not a universal scaling law.
8. Create a controlled burst to expose queue behavior
Averages hide bursts, so the last experiment submits several
independent project analyses close together. Pre-create three local
projects sq31-burst-a, sq31-burst-b and
sq31-burst-c and grant the local analysis identity
Execute Analysis on them. Use the same synthetic source revision,
changing only the stable project key per submission. Start them
within a few seconds of each other and preserve each
ceTaskId.
# Conceptual PowerShell burst. Each working copy has the same source revision
# and its own stable project key. Do not reuse one .scannerwork directory concurrently.
$projects = @('sq31-burst-a','sq31-burst-b','sq31-burst-c')
$jobs = foreach ($p in $projects) {
Start-Job -ArgumentList $p,$env:SONAR_HOST_URL,$env:SONAR_TOKEN -ScriptBlock {
param($projectKey,$hostUrl,$token)
$env:SONAR_HOST_URL=$hostUrl; $env:SONAR_TOKEN=$token
Set-Location "../copies/$projectKey"
sonar-scanner -Dsonar.projectKey=$projectKey -Dsonar.sources=src 2>&1
}
}
$jobs | Wait-Job | Receive-Job
Community Build provides the useful learning constraint here: background work serializes, so closely arriving reports should expose queue wait if they reach the server faster than the Compute Engine can finish them. The scanner jobs may all upload successfully while later CE tasks are still pending. That is exactly why upload success, CE success and end-to-end completion are separate states.
Preserve the queue evidence before changing scheduling, memory or edition. The point is to prove the arrival/service imbalance first.
9. Compare your measurements with a reference architecture without normalizing away differences
Create an assumptions-diff table between the current 10M reference architecture and your lab. Your laptop/container results are not expected to match the published host numbers; they are expected to teach you which differences make the comparison invalid as a benchmark.
| Dimension | Reference-architecture assumption | Your lab evidence | Implication |
|---|---|---|---|
| Edition | Developer/Enterprise single node | Community Build | Worker/branch/PR capabilities differ |
| Host | Dedicated VM, published vCPU/RAM/local SSD | Docker on your measured host | Virtualization/contention/storage path differ |
| Database | Dedicated PostgreSQL host | Local PostgreSQL container | Network and DB isolation differ |
| Repository mix | ~50k LOC average; typical daily main + PR work | Synthetic Python files | Analyzer/project-shape mismatch |
| API/users | Occasional API load | Near-zero interactive load | Web path is underrepresented |
| Plugins | No third-party plugins | None | This variable intentionally matches |
10. Challenge: choose the layer, not the command
Your large-fixture scanner time doubles, but CE processing time, queue wait, server CPU, database counters and disk latency remain almost unchanged. Which first action is best?
- A. Increase CE workers.
- B. Increase search heap.
- C. Inspect scanner indexing/analyzer time, runner CPU/RAM/cache and effective source scope.
- D. Move PostgreSQL to a larger machine.
C. The changed evidence is client-side. Server changes would not target the observed bottleneck and would add confounding variables.
11. Cleanup and rollback
First copy the evidence packet outside the lab directory. Revoke/delete the local lab token through the disposable instance UI. Then remove only the named Compose project and volumes created by this lab. Do not use broad Docker prune commands as “cleanup,” because they can delete unrelated development data.
# From the directory that contains this lab's compose.yaml only:
docker compose down -v
cd ..
# Delete sq31-capacity-lab only after confirming evidence was copied out.
# Remove the shell token from the current process.
Remove-Item Env:SONAR_TOKEN -ErrorAction SilentlyContinue
The -v switch deletes the named lab volumes. Verify
the Compose project name and directory before running it. Never
point this command at a production or shared Compose project.
Knowledge check
Why keep the warm-up run instead of simply deleting it?
It proves the test design recognized cache/JIT/bootstrap state. Hiding it can make a warmed benchmark look universally representative.
Three scanners exit 0, but two CE tasks are still PENDING. How many completed analyses do you have?
Only the task(s) that reached terminal SUCCESS count as completed server analyses. Scanner success/report upload does not complete the asynchronous CE work.
Why use separate working copies for concurrent scanner jobs?
The scanner working directory is analysis-local state and should not be raced by multiple concurrent analyses. Separate workspaces preserve evidence and avoid .scannerwork collisions.
What changed in the 120→480-file comparison?
The source workload/revision. Server resources, database, scanner family/version, project identity and other variables should remain controlled.
Why is a local Docker result not a benchmark of the SonarSource 10M architecture?
Edition, host isolation, storage, database topology, project/language mix, users/API load and CI arrival shape differ; it is a learning experiment, not a normalized production comparison.
Official references and version notes
- SonarQube downloads — Current Community Build, commercial Server release train, editions and active LTA.
- Server host requirements — Current disk, memory, CPU, local-storage and production database-host guidance.
- Community Build host requirements — Equivalent Community Build host/search guidance for the mandatory free path.
- Monitoring the instance — Web/CE/search JVMs, read-only JMX MBeans and ComputeEngineTasks/Database signals.
- Reference architecture up to 10 M LOC — Planning anchor and its stated normal-usage assumptions.
- Improving performance — Enterprise+ CE-worker guidance and the requirement to measure external bottlenecks.
- Performance issues — Current troubleshooting starting points for storage, scope, workers and database-related performance.
- Server release notes — 2026.1 runtime/database changes, including JDK 21/25 and PostgreSQL 14–18.
- SonarScanner CLI releases — Current scanner release identity; 8.1.0.6389 is the latest listed release at this chapter recheck.
- Web API — Bearer authentication guidance and the ongoing Web API V2 migration.
Version/compatibility baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. For the 2026.1 LTA, server runtime requires a JDK and supports Java 21 or 25; PostgreSQL 14–18 is supported. Scanner runtimes without JRE auto-provisioning should use Java 21 or newer. Web API V2 is still gradually replacing legacy endpoints. Always re-check the target release before copying any sizing or runtime value.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.