Performance Sizing, Reference Architectures, and Large-Instance Tuning: Core Concepts and Mental Model
Build a causal sizing model from analysis arrivals, scanner work, Compute Engine queueing and host/database/search demand before changing resources.
Learning objectives
- Explain why total LOC alone cannot predict SonarQube capacity or user experience.
- Trace one analysis from exact revision and effective scanner inputs through report upload, Compute Engine queueing and durable/search-backed results.
- Separate scanner duration, queue wait, Compute Engine service time and UI/API response time as different performance signals.
- Identify read-only evidence for CPU, RAM/heap/GC, database/search, disk, network, concurrent API use and capacity headroom.
- Use reference architectures as planning anchors while preserving version, workload and hardware assumptions.
1. The practical problem: “How big a server?” is the wrong first question
By Chapter 30 you have already separated a single-node SonarQube deployment from Data Center Edition and learned that adding nodes only helps when the added role owns the constrained work. Performance sizing applies the same causal discipline to every edition. The operational question is not “How many gigabytes of RAM does 20 million LOC require?” It is “What work arrives, where does that work execute, how long does each stage take, what shared resource becomes saturated first, and how much headroom do we need for bursts and failures?”
Two instances with the same total lines of code can behave very differently. One may analyze a stable set of small Java repositories once nightly. Another may analyze large JavaScript monorepos on every pull request while hundreds of developers query dashboards and automation calls the Web API. Total LOC describes stored and analyzed scope, but it does not encode arrival rate, language/analyzer cost, report size, queue burstiness, database churn, search traffic, API concurrency or CI expectations.
A capacity claim is useful only when it names the exact workload window, project/language mix, source revisions, scanner/runtime versions, server edition/version, database, host/storage class and observed evidence. “Our SonarQube handles 30M LOC” without those facts is not a reproducible statement.
2. Mental model: demand enters as analyses and leaves as responsive results
The causal chain starts outside SonarQube. CI jobs and developers create an analysis arrival rate: how many reports reach the server per unit time and whether they arrive smoothly or in bursts. Each arrival has a project and language shape. File count, LOC, generated-code scope, analyzer set, external reports and build/scanner behavior determine how much work happens before upload and how large/complex the resulting analysis report is.
The scanner indexes files and runs client-side sensors/analyzers,
then uploads a report. Upload success is only a handoff. SonarQube
records a background-task identity (ceTaskId in
.scannerwork/report-task.txt for Scanner CLI) and the
Compute Engine processes the report asynchronously. Queue wait and
task processing time are distinct. The Compute Engine then drives
persistence and indexing work against the supported database and
embedded search engine. Those operations consume CPU, JVM heap, OS
page cache, disk I/O and network capacity. Meanwhile the Web process
serves users, CI polling and API automation.
Every arrow matters. A slow scanner with an empty CE queue points toward build-runner, source-scope, analyzer/cache or scanner-network work—not automatically toward more server workers. A growing CE queue with short scanner times points to arrival/service imbalance. Long CE tasks while database latency or storage latency rises point to a shared downstream resource. Fast analyses with sluggish dashboards during API bursts point toward Web/search/database responsiveness rather than scanner capacity.
3. The state you must name before you change anything
Source / revision state
Commit SHA, branch identity where supported, generated source, repository size and language/file mix. A benchmark without an exact revision cannot be repeated.
Scanner / runtime state
Scanner family/version, Java/JRE provisioning mode, cache state, build tool, runner CPU/RAM and scanner duration. Scanner time is not CE time.
Analysis-parameter state
Project key, base directory, sources/tests, inclusions/exclusions and report paths plus the effective precedence source. Command-line overrides can make one run differ from persisted project settings.
Compute Engine state
Queue depth/wait, task ID, worker count where edition permits, task duration/status and CE JVM/CPU. Community Build is the mandatory single-instance learning path; additional worker controls are commercial.
Database / search state
Supported DB version, connection/pool pressure, transaction/read activity, DB latency/space, embedded-search disk latency/index size/health and OS page-cache availability.
Host / network state
vCPU, RAM, per-process heap, swap/memory pressure, disk free space/IOPS/latency, container limits and latency/bandwidth among scanner, server and database.
User / API state
Concurrent interactive users, automation request rate, expensive report/dashboard patterns and Web process responsiveness.
Governance state
Sizing window, owner, SLO/validation criteria, headroom policy, reference-architecture assumptions, change record and rollback trigger.
4. What to measure: workload, service, resource and experience signals
A useful baseline combines four families of evidence. Workload signals describe what arrives: analyses per hour, peak burst size, project/file/LOC distribution and language mix. Service signals describe the analysis lifecycle: scanner elapsed time, queue wait, CE processing time and completion rate. Resource signals describe what the platform consumes: process CPU/RAM, configured heap, GC symptoms, disk space/latency, database activity and network behavior. Experience signals describe whether the service still meets its purpose: time from CI start to authoritative analysis, UI/API latency, time to quality-gate evidence and queue age.
| Signal | Layer / owner | What it can prove | What it cannot prove alone |
|---|---|---|---|
| LOC / file / language mix | Source + analysis scope | Size and analyzer mix of the work submitted | How often work arrives or which resource saturates |
| Scanner elapsed time | Scanner / CI runner | Client-side indexing/analyzer/report time | CE queue or server persistence time |
PendingCount / PendingTime |
Compute Engine | Backlog and oldest wait when observed through supported monitoring | Why the CE is slow without DB/disk/CPU evidence |
| CE task processing time | Compute Engine | Service time of completed reports | Scanner time or UI responsiveness |
| CPU / RSS / heap / GC | Host + JVM process | Compute/memory pressure and process-specific trends | Database or network root cause by itself |
| DB pool/activity + latency | Database | Connection pressure and database contribution | Search-index or scanner bottlenecks |
| Disk latency/free space | Host/search/database storage | I/O pressure and capacity risk | Which caller created the I/O without correlation |
| API/UI latency | Web/search/DB path | Developer/automation experience | Whether analysis throughput itself is constrained |
5. A small queueing model that helps without pretending reality is simple
For a bounded window, let λ be analyses arriving per
hour, S the mean Compute Engine processing time in
seconds, and W the number of effective CE workers. A
useful first approximation of offered CE utilization is
ρ ≈ λ × S / (3600 × W). This is not a SonarQube
guarantee and not a complete queueing model: tasks have different
sizes, arrivals are bursty, workers can contend for downstream
resources and edition rules affect concurrency. Its value is
diagnostic. If offered utilization is already close to one during an
average hour, bursts will create queueing unless service capacity or
scheduling changes.
Use percentiles and windows as well as averages. Ten tasks that each need one minute but arrive together can create visible queue wait on a one-worker Community Build instance even though their hourly average looks tiny. Conversely, a long single task can dominate one worker while the rest of the host looks quiet. Record median and slower-tail task times, peak pending count and oldest pending time rather than averaging away the workload shape.
# Worksheet approximation — not a sizing guarantee
arrival_per_hour = 24
mean_ce_seconds = 95
workers = 1
rho = arrival_per_hour * mean_ce_seconds / (3600 * workers)
# rho ~= 0.63; now examine bursts, p95 task time and downstream saturation.
6. Memory is a budget: heap is not the same thing as usable RAM
SonarQube has separate Web, Compute Engine and embedded-search Java
processes on a single node. Heap settings such as
sonar.web.javaOpts, sonar.ce.javaOpts and
sonar.search.javaOpts cap JVM heaps; they do not
reserve all native memory the processes and operating system may
need. Search uses Lucene, which benefits strongly from the
operating-system page cache. Current SonarSource host guidance
recommends leaving substantial memory outside the Elasticsearch
heap—described as roughly a 50/50 search-heap versus
available-memory split—and cautions against oversized search heaps.
Therefore “the host has 16 GB, set Java -Xmx16G” is a
causal error. On a single-node installation that same host also
needs Web and CE heaps, native memory, container/runtime overhead
and filesystem cache. When memory pressure causes reclaim or swap, a
larger heap can make search and overall response time worse even
though the JVM reports a larger maximum.
Treat SonarSource’s search-memory guidance as a relationship: Lucene needs OS cache. First budget memory across all SonarQube processes and the operating system; then validate with actual memory pressure, search latency and GC evidence.
7. Disk and database behavior can cap throughput before CPU does
Search indexing and queries are sensitive to disk latency. Current SonarSource guidance explicitly warns against remote-mounted NFS, SMB/CIFS and NAS for performance-sensitive search storage and favors SSD/local low-latency storage. Free space is also a correctness concern: search engines protect themselves near high disk-watermark conditions, so “capacity” includes space headroom, not only IOPS.
The database is durable application state; search indexes are a different operational layer. CE workers read/write through database and search paths, so raising concurrency can increase database connections, transaction load, disk activity and network traffic. That is why an empty CPU graph does not prove the instance can safely accept more workers. Inspect supported database metrics and SonarQube’s read-only Database MBean/pool signals rather than editing database tables or embedded search internals.
8. Reference architectures are hypotheses with stated assumptions
SonarSource’s “up to 10 M LOC” architecture is useful because it publishes assumptions: a 4-vCPU/8-GB/50-GB-local-SSD SonarQube host, a separate PostgreSQL host, normal daily main-branch analysis plus several pull-request analyses, an average repository around 50k LOC, occasional API use and no third-party plugin cost. Those assumptions are more important than the headline “10 M LOC.” The same page explicitly says high-frequency analysis, very large repositories, heavy API use and plugins may require more capacity.
There is also a versioning lesson hidden in the page itself. At this chapter’s 2026-09-08 recheck, the generic reference page still prints an OpenJDK 17 software line, while the active 2026.1 LTA release notes require a JDK and remove Java 17 in favor of Java 21 or 25. Therefore copy the architecture reasoning, not stale software prerequisites. Hardware examples are anchors; the target release’s compatibility matrix wins.
| Reference-architecture statement | Use it as | Do not turn it into |
|---|---|---|
| 4 vCPU / 8 GB / local SSD for the 10M scenario | A starting hardware hypothesis with its workload assumptions | A universal minimum or guarantee for every 10M-LOC estate |
| Separate supported database host | A latency/failure-isolation production pattern | Permission to ignore measured DB latency or DB sizing |
| Daily main + several PR analyses | An explicit arrival-rate assumption | A model for bursty monorepo CI without remeasurement |
| Average repo ~50k LOC | A project-mix assumption | Evidence that one 2M-LOC repo behaves like forty 50k repos |
| Enterprise CE worker examples | An edition-dependent lever | A Community Build setting or a reason to add workers before measuring DB/search |
9. Read-only inspection before any sizing change
Capture a baseline before touching heap, workers, storage or CI scheduling. The exact tools vary by deployment, but the evidence should be exportable and timestamped. A local Docker fixture can use the following read-only commands; production teams should map them to their monitoring system and least-privilege API access.
# Server identity and health
curl -fsS http://localhost:9000/api/server/version
curl -fsS http://localhost:9000/api/system/status
# Host/container snapshot
docker stats --no-stream sq31-sonarqube sq31-postgres
docker inspect sq31-sonarqube --format '{{json .HostConfig.Memory}}'
docker exec sq31-sonarqube sh -lc 'df -h /opt/sonarqube/data && du -sh /opt/sonarqube/data /opt/sonarqube/logs'
# PostgreSQL identity and read-only workload counters
docker exec sq31-postgres psql -U sonar -d sonar -c 'select version();'
docker exec sq31-postgres psql -U sonar -d sonar -c "select datname,numbackends,xact_commit,blks_read,blks_hit,temp_files,temp_bytes from pg_stat_database where datname='sonar';"
For each scanner run, preserve the scanner log and
.scannerwork/report-task.txt before cleanup. The task
ID is the join key from client-side evidence to server-side Compute
Engine evidence. If you use JMX, the SonarQube-specific MBeans are
read-only; PendingCount, PendingTime,
ProcessingTime, SuccessCount,
ErrorCount and WorkerCount are designed
precisely for this style of diagnosis.
10. DevOps connection: size the service around delivery demand, not vanity capacity
In DevOps, SonarQube is part of a delivery feedback loop. Capacity is adequate when scanner work, server processing and policy evidence arrive predictably enough for developers and automation, with explicit headroom for routine bursts and maintenance—not when a dashboard shows a large LOC number. The reproducible unit is the evidence chain: exact revision + effective analysis inputs + scanner result + task ID + CE result + resource observations + policy context + owner.
The chapter will keep that chain intact. Lesson 2 measures it locally; Lesson 3 uses the measurements to choose between scaling and scheduling patterns; Lesson 4 deliberately breaks the model; Lesson 5 turns the measurements into a governed capacity worksheet.
Knowledge check
Two SonarQube estates each contain 12M LOC. Why can their required capacity differ dramatically?
LOC does not encode project/language mix, analysis arrival rate and burstiness, report/CE service time, API concurrency, database/search behavior or CI latency expectations.
A scanner takes 11 minutes but its uploaded CE task waits 0 seconds and finishes in 20 seconds. Should you add CE workers first?
No. The evidence points primarily to scanner/build-runner, analysis scope, analyzer/cache or scanner-to-server behavior. More CE workers target queue/CE throughput, not an 11-minute client-side scan.
Why can allocating every spare gigabyte to search heap hurt performance?
Lucene relies on OS page cache, and Web/CE/native memory also share the host. Oversizing heap can create memory pressure, reclaim/swap and worse search latency.
What is the most important thing to copy from a reference architecture?
Its assumptions and causal topology. Hardware numbers are planning anchors; target-release software requirements and measured workload must be revalidated.
Which identifier joins scanner upload evidence to Compute Engine processing evidence?
The Compute Engine task ID, commonly recorded as ceTaskId in the scanner report-task file.
Official references and version notes
- SonarQube downloads — Current Community Build, commercial Server release train, editions and active LTA.
- Server host requirements — Current disk, memory, CPU, local-storage and production database-host guidance.
- Community Build host requirements — Equivalent Community Build host/search guidance for the mandatory free path.
- Monitoring the instance — Web/CE/search JVMs, read-only JMX MBeans and ComputeEngineTasks/Database signals.
- Reference architecture up to 10 M LOC — Planning anchor and its stated normal-usage assumptions.
- Improving performance — Enterprise+ CE-worker guidance and the requirement to measure external bottlenecks.
- Performance issues — Current troubleshooting starting points for storage, scope, workers and database-related performance.
- Server release notes — 2026.1 runtime/database changes, including JDK 21/25 and PostgreSQL 14–18.
- SonarScanner CLI releases — Current scanner release identity; 8.1.0.6389 is the latest listed release at this chapter recheck.
- Web API — Bearer authentication guidance and the ongoing Web API V2 migration.
Version/compatibility baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. For the 2026.1 LTA, server runtime requires a JDK and supports Java 21 or 25; PostgreSQL 14–18 is supported. Scanner runtimes without JRE auto-provisioning should use Java 21 or newer. Web API V2 is still gradually replacing legacy endpoints. Always re-check the target release before copying any sizing or runtime value.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.