Chapter 31Lesson 04~175 minutes

Performance Sizing, Reference Architectures, and Large-Instance Tuning: Diagnostics, Failure Modes, and Production Practices

Diagnose SonarQube performance failures by preserving first-failure evidence and separating scanner, Compute Engine, database, search, JVM, storage, network and CI causes.

DiagnosticsFailure modesJVM & hostStorageRollback

Learning objectives

  • Apply one evidence-first diagnostic sequence across scanner, CE, project-policy, auth/network, DB/search/JVM/host and CI layers.
  • Diagnose why assigning all RAM to Java heap can reduce search responsiveness.
  • Recognize worker, network-storage, LOC-only, burst and warmed-benchmark failure patterns.
  • Preserve original logs/task IDs/configuration before the least-destructive correction.
  • Use rollback criteria instead of tuning production by repeated restarts or destructive cleanup.

1. Diagnostic sequence: preserve first failure, then narrow the owning layer

Performance incidents tempt teams to restart services and erase the evidence that identifies the constrained layer. Use the same sequence whether the symptom is a slow scanner, growing CE queue, failed background task, sluggish UI or unstable search. Each step either rules in or rules out a state owner.

preserve scanner/server/CI/API evidence → confirm edition + exact versions → confirm revision + effective parameters → inspect indexing/report + ceTaskId → inspect CE task/queue → check project/profile/gate/new-code/security state → check auth/network/provider → check DB/search/JVM/host/container → least-destructive correction → smallest equivalent rerun

Project policy is included even in a performance incident because an “analysis is slow” report can actually be a quality-gate wait, provider-decoration delay, permission/API failure or scanner-scope drift. Do not assume the symptom names the layer.

2. Intentionally broken example: heap everything, cache nothing

Assume a disposable single-node lab host has 8 GiB available to the SonarQube container. An operator sees search latency and decides to set the search JVM maximum close to the entire container limit while Web and CE remain on the same container. After restart, container memory stays near the limit, the host reclaims aggressively, search queries become erratic and CE tasks slow during indexing. The change looked like “more memory for search,” but it removed memory from the OS cache and competing processes.

# INTENTIONALLY BAD — educational example only. Do not apply to production.
# Container memory: 8 GiB
SONAR_SEARCH_JAVAOPTS=-Xms7g -Xmx7g
# Web + CE + native memory + Lucene page cache now have almost no safe budget.
Disruptive action

JVM heap changes require a restart and can prevent a server from starting. Preserve current configuration, system info, logs and host memory evidence first. Use only disposable resources for this broken example.

Interpret the evidence

Evidence Interpretation
Container memory near hard limit; host available/cache drops Whole-host/container budget is exhausted
Search heap max increased but UI/API p95 worsens Bigger heap did not solve the causal path
Disk read latency / page faults rise Lucene/OS cache may be starved
CE task time also rises Shared-host memory/I/O pressure affects more than search
Scanner time unchanged Client-side analyzer is not the first limiting layer

Repair: restore the previous supported heap values first, restart through the normal deployment procedure, and rerun the smallest equivalent workload. Then build a measured process/OS memory budget. If search still needs capacity, change one supported resource at a time while retaining page-cache headroom.

3. Failure mode: increasing CE workers into a constrained database/search path

A classic secondary failure begins with a real queue. An Enterprise operator increases workers from two to six. Pending count briefly falls, but each task becomes slower, database CPU and active connections rise, disk latency increases and the queue soon returns. The worker change did increase concurrency—exactly as configured—but concurrency amplified contention in the actual bottleneck.

worker count ↑ → simultaneous CE DB/search work ↑ → DB/disk/network contention ↑ → individual CE service time ↑ → effective throughput fails to improve

Preserve worker count, PendingCount/PendingTime, task durations, DB pool/activity and disk/network evidence from both states. The least-destructive correction is usually to return to the previous worker count while diagnosing the downstream limit. Do not “compensate” with progressively more workers or unreviewed database changes.

4. Failure mode: performance-sensitive search data on remote network storage

Search can look CPU-idle while queries and indexing stall on storage latency. Current SonarSource host guidance warns against remote-mounted NFS, SMB/CIFS and NAS for search-sensitive storage because latency is higher and more variable. A deployment can pass a trivial smoke test and still fail under concurrent indexing/search.

Evidence should include actual mount/storage class, latency/IOPS observations, search/server logs, disk free space and UI/API/CE timing correlation. The correction is a supported migration to suitable low-latency storage—not copying an embedded search directory while the service runs, editing search metadata or deleting indices. Preserve the database as durable application state and follow supported recovery/migration paths.

5. Failure mode: sizing only by LOC

A 20M-LOC portfolio can be low-frequency and mostly small repositories; a 5M-LOC estate can generate continuous PR traffic from a few huge monorepos. LOC matters for stored/analyzed scale, but it is one dimension. When a reference architecture is selected only from LOC, high-frequency analysis, large-repository tails, API automation and plugin/analyzer cost disappear from the model.

Missing dimension Symptom when ignored Evidence to add
Arrival burstiness Morning/nightly queue spikes despite low daily average analyses/5m, peak concurrent uploads, pending time
Project-size distribution Few very large tasks monopolize CE per-project scanner/CE p50/p95/max
Language/analyzer mix Same LOC produces different scanner/report cost languages, analyzer versions, file counts
API/user concurrency UI slow while CE appears healthy Web CPU/heap, API p95/request rate
Database/search/storage Workers do not improve throughput DB latency/pool, search disk latency/free space
CI/network path Upload/scan time dominates runner/network timing and scanner logs

6. Failure mode: ignoring bursty CI

Suppose 120 analyses arrive over an eight-hour day: average 15/hour. That sounds small. If 80 of them start within ten minutes after a scheduled merge train, the relevant arrival rate during the burst is 480/hour. A capacity worksheet that only stores the daily average will systematically underpredict queueing.

Preserve CI scheduling/run-start evidence beside CE pending/time evidence. If the service is healthy outside the burst, test a reversible scheduling/concurrency change before a server expansion. If the burst is legitimate merge-critical traffic, use its peak shape in the scaling model rather than smoothing it away.

7. Failure mode: trusting one warmed benchmark

Caches are real operational mechanisms, but a single warmed run can hide cold-start and change-related cost. Scanner analyzer downloads, filesystem cache, JVM JIT and database buffers can make later repetitions faster. Conversely, one cold run can exaggerate normal steady-state time. A trustworthy test labels cache state and keeps multiple repetitions.

run,revision,cache_state,scanner_s,queue_wait_s,ce_s
warmup,abc123,cold_or_unknown,48.2,0.0,31.4
B1,abc123,warm,31.0,0.0,27.9
B2,abc123,warm,30.7,0.0,28.1
B3,abc123,warm,31.4,0.0,27.6

Do not delete caches before every run unless “cold cache” is the scenario you are explicitly testing. Do not delete them after a failure before preserving the failure. State which scenario the benchmark represents.

8. Failure mode: tuning production without a rollback envelope

Changing heap, workers, container limits, storage, database tier or cluster topology can trade one bottleneck for another. A production tuning plan therefore needs a before snapshot, exact configuration diff, one primary hypothesis, validation window, guardrails and a rollback condition. “We can restart again if it is bad” is not a rollback plan.

Before:
  CE task p95: 70 s
  pending-time p95: 210 s
  DB p95 latency: 8 ms
  search disk p95: 4 ms
Change:
  one supported variable only
Success:
  pending-time p95 <= 60 s AND CE task p95 <= 77 s
Guardrails:
  DB p95 <= 12 ms; no health degradation; UI/API p95 <= baseline + 10%
Rollback:
  restore exact previous config/deployment value and verify smallest equivalent workload

9. Causal separation checklist

Symptom First evidence Likely layer families to test Wrong shortcut
Scanner process slow before upload scanner timestamps/indexed files/runner metrics scanner/build/analysis scope/cache/CI runner add server CE workers
Upload/auth fails HTTP/auth/network evidence credential/trust/proxy/network disable TLS verification / admin token everywhere
Report uploaded; task pending long ceTaskId + pending time/count CE arrival vs service capacity replace project key or rerun blindly
Task starts then runs long CE log/task + DB/search/host metrics CE heap/CPU, DB, search/disk, project shape delete search data / edit DB
Gate fails after fast SUCCESS gate/new-code/profile/measures project policy/result state lower thresholds to make CI green
Provider/CI still red after gate pass CI/provider artifact/status provider integration/CI logic rerun different analysis and call it fixed
UI/API slow, queue empty Web/search/DB + request load Web heap/CPU, search queries/disk, DB increase scanner heap

10. Evidence preservation bundle for a local incident

$Incident = "evidence/incident-$(Get-Date -Format yyyyMMdd-HHmmss)"
New-Item -ItemType Directory -Force $Incident | Out-Null

# Preserve identities before changing anything.
curl.exe -fsS http://localhost:9000/api/server/version | Set-Content "$Incident/server-version.txt"
sonar-scanner -v 2>&1 | Set-Content "$Incident/scanner-version.txt"
git rev-parse HEAD | Set-Content "$Incident/revision.txt"
docker inspect sq31-sonarqube | Set-Content "$Incident/sonarqube-inspect.json"
docker stats --no-stream sq31-sonarqube sq31-postgres | Set-Content "$Incident/docker-stats.txt"
docker logs --since 30m sq31-sonarqube 2>&1 | Set-Content "$Incident/server.log"

# If an analysis ran, preserve its task identity and raw server response.
Copy-Item .scannerwork/report-task.txt "$Incident/report-task.txt" -ErrorAction SilentlyContinue
# Poll api/ce/task for that exact id using a least-privilege bearer token.

# Only after this packet exists should you test the smallest correction.
Never use evidence destruction as diagnosis

Do not delete logs/caches, restart repeatedly, purge queues, manipulate database/search internals, disable TLS, lower policy thresholds or mass-suppress findings to make the symptom disappear. Preserve the first failure and correct the layer that owns it.

Knowledge check

After raising CE workers, pending count falls briefly but task p95 and DB latency double. What is the most likely lesson?

Why is NFS/NAS search storage a performance risk even when average latency looks acceptable?

A performance incident disappears after restart. Is the root cause proven?

What is wrong with using only total LOC to size an instance?

Why keep the same project key when rerunning the smallest equivalent failure?

Next lesson

Turn diagnosis into a governed capacity decision

Lesson 5 is the checkpoint: build a measured worksheet from several bounded analyses, identify the first limiting layer, propose one reversible change and define evidence-based acceptance/rollback criteria.

Official references and version notes

Version and compatibility note

Version/compatibility baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. For the 2026.1 LTA, server runtime requires a JDK and supports Java 21 or 25; PostgreSQL 14–18 is supported. Scanner runtimes without JRE auto-provisioning should use Java 21 or newer. Web API V2 is still gradually replacing legacy endpoints. Always re-check the target release before copying any sizing or runtime value.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.