Performance Sizing, Reference Architectures, and Large-Instance Tuning: Diagnostics, Failure Modes, and Production Practices
Diagnose SonarQube performance failures by preserving first-failure evidence and separating scanner, Compute Engine, database, search, JVM, storage, network and CI causes.
Learning objectives
- Apply one evidence-first diagnostic sequence across scanner, CE, project-policy, auth/network, DB/search/JVM/host and CI layers.
- Diagnose why assigning all RAM to Java heap can reduce search responsiveness.
- Recognize worker, network-storage, LOC-only, burst and warmed-benchmark failure patterns.
- Preserve original logs/task IDs/configuration before the least-destructive correction.
- Use rollback criteria instead of tuning production by repeated restarts or destructive cleanup.
1. Diagnostic sequence: preserve first failure, then narrow the owning layer
Performance incidents tempt teams to restart services and erase the evidence that identifies the constrained layer. Use the same sequence whether the symptom is a slow scanner, growing CE queue, failed background task, sluggish UI or unstable search. Each step either rules in or rules out a state owner.
ceTaskId → inspect CE task/queue
→ check project/profile/gate/new-code/security
state → check auth/network/provider
→ check DB/search/JVM/host/container
→ least-destructive correction
→ smallest equivalent rerun
Project policy is included even in a performance incident because an “analysis is slow” report can actually be a quality-gate wait, provider-decoration delay, permission/API failure or scanner-scope drift. Do not assume the symptom names the layer.
2. Intentionally broken example: heap everything, cache nothing
Assume a disposable single-node lab host has 8 GiB available to the SonarQube container. An operator sees search latency and decides to set the search JVM maximum close to the entire container limit while Web and CE remain on the same container. After restart, container memory stays near the limit, the host reclaims aggressively, search queries become erratic and CE tasks slow during indexing. The change looked like “more memory for search,” but it removed memory from the OS cache and competing processes.
# INTENTIONALLY BAD — educational example only. Do not apply to production.
# Container memory: 8 GiB
SONAR_SEARCH_JAVAOPTS=-Xms7g -Xmx7g
# Web + CE + native memory + Lucene page cache now have almost no safe budget.
JVM heap changes require a restart and can prevent a server from starting. Preserve current configuration, system info, logs and host memory evidence first. Use only disposable resources for this broken example.
Interpret the evidence
| Evidence | Interpretation |
|---|---|
| Container memory near hard limit; host available/cache drops | Whole-host/container budget is exhausted |
| Search heap max increased but UI/API p95 worsens | Bigger heap did not solve the causal path |
| Disk read latency / page faults rise | Lucene/OS cache may be starved |
| CE task time also rises | Shared-host memory/I/O pressure affects more than search |
| Scanner time unchanged | Client-side analyzer is not the first limiting layer |
Repair: restore the previous supported heap values first, restart through the normal deployment procedure, and rerun the smallest equivalent workload. Then build a measured process/OS memory budget. If search still needs capacity, change one supported resource at a time while retaining page-cache headroom.
3. Failure mode: increasing CE workers into a constrained database/search path
A classic secondary failure begins with a real queue. An Enterprise operator increases workers from two to six. Pending count briefly falls, but each task becomes slower, database CPU and active connections rise, disk latency increases and the queue soon returns. The worker change did increase concurrency—exactly as configured—but concurrency amplified contention in the actual bottleneck.
Preserve worker count,
PendingCount/PendingTime, task durations,
DB pool/activity and disk/network evidence from both states. The
least-destructive correction is usually to return to the previous
worker count while diagnosing the downstream limit. Do not
“compensate” with progressively more workers or unreviewed database
changes.
4. Failure mode: performance-sensitive search data on remote network storage
Search can look CPU-idle while queries and indexing stall on storage latency. Current SonarSource host guidance warns against remote-mounted NFS, SMB/CIFS and NAS for search-sensitive storage because latency is higher and more variable. A deployment can pass a trivial smoke test and still fail under concurrent indexing/search.
Evidence should include actual mount/storage class, latency/IOPS observations, search/server logs, disk free space and UI/API/CE timing correlation. The correction is a supported migration to suitable low-latency storage—not copying an embedded search directory while the service runs, editing search metadata or deleting indices. Preserve the database as durable application state and follow supported recovery/migration paths.
5. Failure mode: sizing only by LOC
A 20M-LOC portfolio can be low-frequency and mostly small repositories; a 5M-LOC estate can generate continuous PR traffic from a few huge monorepos. LOC matters for stored/analyzed scale, but it is one dimension. When a reference architecture is selected only from LOC, high-frequency analysis, large-repository tails, API automation and plugin/analyzer cost disappear from the model.
| Missing dimension | Symptom when ignored | Evidence to add |
|---|---|---|
| Arrival burstiness | Morning/nightly queue spikes despite low daily average | analyses/5m, peak concurrent uploads, pending time |
| Project-size distribution | Few very large tasks monopolize CE | per-project scanner/CE p50/p95/max |
| Language/analyzer mix | Same LOC produces different scanner/report cost | languages, analyzer versions, file counts |
| API/user concurrency | UI slow while CE appears healthy | Web CPU/heap, API p95/request rate |
| Database/search/storage | Workers do not improve throughput | DB latency/pool, search disk latency/free space |
| CI/network path | Upload/scan time dominates | runner/network timing and scanner logs |
6. Failure mode: ignoring bursty CI
Suppose 120 analyses arrive over an eight-hour day: average 15/hour. That sounds small. If 80 of them start within ten minutes after a scheduled merge train, the relevant arrival rate during the burst is 480/hour. A capacity worksheet that only stores the daily average will systematically underpredict queueing.
Preserve CI scheduling/run-start evidence beside CE pending/time evidence. If the service is healthy outside the burst, test a reversible scheduling/concurrency change before a server expansion. If the burst is legitimate merge-critical traffic, use its peak shape in the scaling model rather than smoothing it away.
7. Failure mode: trusting one warmed benchmark
Caches are real operational mechanisms, but a single warmed run can hide cold-start and change-related cost. Scanner analyzer downloads, filesystem cache, JVM JIT and database buffers can make later repetitions faster. Conversely, one cold run can exaggerate normal steady-state time. A trustworthy test labels cache state and keeps multiple repetitions.
run,revision,cache_state,scanner_s,queue_wait_s,ce_s
warmup,abc123,cold_or_unknown,48.2,0.0,31.4
B1,abc123,warm,31.0,0.0,27.9
B2,abc123,warm,30.7,0.0,28.1
B3,abc123,warm,31.4,0.0,27.6
Do not delete caches before every run unless “cold cache” is the scenario you are explicitly testing. Do not delete them after a failure before preserving the failure. State which scenario the benchmark represents.
8. Failure mode: tuning production without a rollback envelope
Changing heap, workers, container limits, storage, database tier or cluster topology can trade one bottleneck for another. A production tuning plan therefore needs a before snapshot, exact configuration diff, one primary hypothesis, validation window, guardrails and a rollback condition. “We can restart again if it is bad” is not a rollback plan.
Before:
CE task p95: 70 s
pending-time p95: 210 s
DB p95 latency: 8 ms
search disk p95: 4 ms
Change:
one supported variable only
Success:
pending-time p95 <= 60 s AND CE task p95 <= 77 s
Guardrails:
DB p95 <= 12 ms; no health degradation; UI/API p95 <= baseline + 10%
Rollback:
restore exact previous config/deployment value and verify smallest equivalent workload
9. Causal separation checklist
| Symptom | First evidence | Likely layer families to test | Wrong shortcut |
|---|---|---|---|
| Scanner process slow before upload | scanner timestamps/indexed files/runner metrics | scanner/build/analysis scope/cache/CI runner | add server CE workers |
| Upload/auth fails | HTTP/auth/network evidence | credential/trust/proxy/network | disable TLS verification / admin token everywhere |
| Report uploaded; task pending long | ceTaskId + pending time/count | CE arrival vs service capacity | replace project key or rerun blindly |
| Task starts then runs long | CE log/task + DB/search/host metrics | CE heap/CPU, DB, search/disk, project shape | delete search data / edit DB |
| Gate fails after fast SUCCESS | gate/new-code/profile/measures | project policy/result state | lower thresholds to make CI green |
| Provider/CI still red after gate pass | CI/provider artifact/status | provider integration/CI logic | rerun different analysis and call it fixed |
| UI/API slow, queue empty | Web/search/DB + request load | Web heap/CPU, search queries/disk, DB | increase scanner heap |
10. Evidence preservation bundle for a local incident
$Incident = "evidence/incident-$(Get-Date -Format yyyyMMdd-HHmmss)"
New-Item -ItemType Directory -Force $Incident | Out-Null
# Preserve identities before changing anything.
curl.exe -fsS http://localhost:9000/api/server/version | Set-Content "$Incident/server-version.txt"
sonar-scanner -v 2>&1 | Set-Content "$Incident/scanner-version.txt"
git rev-parse HEAD | Set-Content "$Incident/revision.txt"
docker inspect sq31-sonarqube | Set-Content "$Incident/sonarqube-inspect.json"
docker stats --no-stream sq31-sonarqube sq31-postgres | Set-Content "$Incident/docker-stats.txt"
docker logs --since 30m sq31-sonarqube 2>&1 | Set-Content "$Incident/server.log"
# If an analysis ran, preserve its task identity and raw server response.
Copy-Item .scannerwork/report-task.txt "$Incident/report-task.txt" -ErrorAction SilentlyContinue
# Poll api/ce/task for that exact id using a least-privilege bearer token.
# Only after this packet exists should you test the smallest correction.
Do not delete logs/caches, restart repeatedly, purge queues, manipulate database/search internals, disable TLS, lower policy thresholds or mass-suppress findings to make the symptom disappear. Preserve the first failure and correct the layer that owns it.
Knowledge check
After raising CE workers, pending count falls briefly but task p95 and DB latency double. What is the most likely lesson?
Concurrency amplified a downstream bottleneck. Restore the prior worker envelope and diagnose database/disk/network pressure before adding concurrency again.
Why is NFS/NAS search storage a performance risk even when average latency looks acceptable?
Search/indexing is sensitive to latency and variance; SonarSource explicitly warns against remote-mounted NFS/SMB/NAS for this path.
A performance incident disappears after restart. Is the root cause proven?
No. Restart changes process/cache/queue state and can erase evidence. Without preserved first-failure data, the owning layer remains uncertain.
What is wrong with using only total LOC to size an instance?
It omits arrival/burst shape, project/language distribution, CE task times, API/user load, DB/search/storage/network behavior and headroom expectations.
Why keep the same project key when rerunning the smallest equivalent failure?
Changing project identity changes persistent/history/policy state and can hide rather than repair the original cause.
Official references and version notes
- SonarQube downloads — Current Community Build, commercial Server release train, editions and active LTA.
- Server host requirements — Current disk, memory, CPU, local-storage and production database-host guidance.
- Community Build host requirements — Equivalent Community Build host/search guidance for the mandatory free path.
- Monitoring the instance — Web/CE/search JVMs, read-only JMX MBeans and ComputeEngineTasks/Database signals.
- Reference architecture up to 10 M LOC — Planning anchor and its stated normal-usage assumptions.
- Improving performance — Enterprise+ CE-worker guidance and the requirement to measure external bottlenecks.
- Performance issues — Current troubleshooting starting points for storage, scope, workers and database-related performance.
- Server release notes — 2026.1 runtime/database changes, including JDK 21/25 and PostgreSQL 14–18.
- SonarScanner CLI releases — Current scanner release identity; 8.1.0.6389 is the latest listed release at this chapter recheck.
- Web API — Bearer authentication guidance and the ongoing Web API V2 migration.
Version/compatibility baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. For the 2026.1 LTA, server runtime requires a JDK and supports Java 21 or 25; PostgreSQL 14–18 is supported. Scanner runtimes without JRE auto-provisioning should use Java 21 or newer. Web API V2 is still gradually replacing legacy endpoints. Always re-check the target release before copying any sizing or runtime value.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.