Performance Sizing, Reference Architectures, and Large-Instance Tuning: Configuration, Design Patterns, and Trade-Offs
Choose vertical or edition-dependent horizontal scaling, memory and Compute Engine strategies from observed bottlenecks, reversibility and operational evidence.
Learning objectives
- Choose vertical versus edition-dependent horizontal scaling from the constrained SonarQube role.
- Budget Web/CE/search heaps without starving the operating system and Lucene page cache.
- Explain why additional CE workers can increase database, disk and network contention.
- Compare scheduled/staggered CI with bursty arrivals as a capacity-control pattern.
- Evaluate reference architectures with an explicit assumption delta, validation plan and rollback.
1. Configuration is a hypothesis about the bottleneck
Lesson 2 produced measurements rather than a magical server size. Lesson 3 asks what change those measurements justify. Every performance change should have the form: evidence says layer X is limiting; change Y should affect X; metric Z should improve without violating guardrails A/B; rollback is R if it does not. That structure prevents “tuning by folklore.”
Performance choices also change operational state. Increasing a JVM heap changes host memory pressure and restart requirements. Increasing Enterprise CE workers changes database/search concurrency. Data Center horizontal scaling changes licensed topology, node count and failure domains. Staggering CI changes provider scheduling but no SonarQube persistent state. Treat these as different changes with different owners and rollback mechanisms.
2. Vertical versus horizontal scaling: first name the role
| Observed limiting signal | Candidate change | Edition / prerequisite | Evidence after change |
|---|---|---|---|
| Web CPU/heap pressure during API/user peak; CE/search healthy | Increase Web/host capacity | Single-node editions; supported host limits | Web latency/CPU/GC improve; CE timings remain stable |
| CE CPU/heap pressure on large tasks; external resources healthy | Increase CE heap/host CPU; optionally workers if queue-bound | Worker count management: Enterprise+ | Task time/queue age improves without DB/disk regression |
| Search disk latency/page-cache pressure | Faster local SSD / correct memory budget / search-role capacity | All single-node editions for vertical; DCE for search-node horizontal role | Index/query latency and UI/API response improve |
| Database saturation/latency | Scale/tune supported DB service and connection path | All editions use supported external production DB | DB latency/pool pressure falls and CE/Web recover |
| Node-level capacity/availability beyond one server | DCE application/search horizontal scaling | Licensed Data Center Edition; topology rules from Chapter 30 | Role-specific throughput/HA improves; cluster remains healthy |
| Bursty CI creates short queue while daily total is modest | Stagger/schedule analyses | CI/provider capability; edition-neutral | Peak pending count/time falls with same server resources |
“Horizontal scaling” is not a generic Community Build switch. Data Center Edition is the licensed distributed architecture that separates application and search roles and permits horizontal scaling. Enterprise Edition can provide multiple CE workers and other scale-oriented capabilities inside a single-node architecture, but that does not turn several independent Enterprise servers into one cluster.
3. Heap versus OS page cache: optimize the whole memory system
Memory decisions are process-specific. Web heap supports interactive/API work. CE heap supports background report processing and scales with project/task complexity and worker concurrency. Search heap supports Elasticsearch structures, while Lucene relies heavily on filesystem cache outside the JVM. Native allocations and the operating system also consume memory.
Start with a host memory budget, not a heap target. On a single node, list observed peaks for Web, CE and search RSS/heap, add OS/native/container overhead, then preserve page-cache headroom. SonarSource’s current host guidance says roughly half of available memory should remain outside Elasticsearch heap for Lucene caching and warns not to make search heap excessively large. That search-specific principle is incompatible with “set every Java process to the maximum.”
# Capacity worksheet, not sonar.properties syntax
host_ram_gib = 32
web_heap_gib = 2 # hypothesis from observed Web pressure
ce_heap_gib = 6 # hypothesis from observed large CE tasks
search_heap_gib = 8 # preserve substantial non-heap cache
known_heap_total = web_heap_gib + ce_heap_gib + search_heap_gib
non_heap_budget = host_ram_gib - known_heap_total
# Validate native/RSS + page-cache + OS margin before approving.
Changing Java options is disruptive and must be done through supported SonarQube configuration/deployment mechanisms with a rollback value and maintenance plan. Do not live-edit JVM internals or delete search data to “apply” memory tuning.
4. More Compute Engine workers can make the wrong bottleneck worse
Worker concurrency helps only when reports are waiting and the infrastructure can absorb parallel work. SonarSource’s performance guidance is explicit that CE workers increase stress on the database, disk I/O, network, heap and CPU. If database latency, storage latency or network is already the bottleneck, more workers create more contenders for the same constrained door and may lengthen each task.
Worker count is also an edition boundary. Current guidance places CE worker management in Enterprise Edition and above. In Data Center Edition, a configured worker count is replicated across application nodes; for example, four configured workers across two app nodes becomes eight effective workers after the required node restarts. That multiplication makes capacity and database impact a cluster-wide calculation, not a per-node tweak.
PendingTime ↑ + PendingCount ↑
→ individual CE time stable
→ DB/disk/network healthy → worker
increase is a plausible hypothesis
individual CE time ↑ → DB latency /
disk wait ↑ → queue grows secondarily
→ fix downstream resource before worker count
5. Scheduled versus bursty CI: smooth demand before buying capacity
CI is an owning system in the sizing model. If nightly jobs for hundreds of repositories all begin at 02:00, the average analyses-per-day metric hides a sharp arrival spike. A one-hour queue after 02:00 may disappear for the rest of the day, yet developers still experience delayed policy evidence. One reversible fix is to stagger non-urgent analysis windows or constrain batch concurrency in the CI orchestrator.
Scheduling can be cheaper and safer than server changes because it alters arrival shape rather than SonarQube persistent state. It does not solve genuine all-day saturation, and it must not postpone required quality/security analysis beyond governance needs. Document which jobs can move, which are merge-critical, expected peak reduction and a rollback to the original schedule.
| Pattern | Benefit | Trade-off / risk | Evidence |
|---|---|---|---|
| All jobs at a fixed minute | Simple orchestration | Large burst; queue wait and DB/search spikes | PendingCount/PendingTime peak |
| Jitter/stagger non-urgent jobs | Reduces peak arrival rate without server mutation | Longer batch window; schedule governance needed | Peak queue falls; same total analyses |
| Concurrency cap at CI layer | Protects SonarQube and DB from fan-in | Jobs wait in CI rather than CE; distinguish wait locations | CI queue vs CE queue evidence |
| Scale platform | Preserves fast arrivals when demand is legitimate | Cost, edition, ops complexity; may shift bottleneck | Role/resource metrics after change |
6. Reference architecture versus measured sizing: build an assumption delta
A reference architecture is most useful as a review checklist. Copy its topology into one column and your measured environment into another. Every mismatch becomes a question, not an automatic defect. The current 10M scenario assumes “normal” daily-main/several-PR analysis, average repositories around 50k LOC, occasional API use and no third-party plugins. The older/currently indexed 50M example uses a larger host and multiple workers, but the same lesson applies: the headline LOC number is conditional on workload shape.
Version prerequisites must be evaluated separately from hardware. The current 10M page still contains an OpenJDK 17 line while the 2026.1 LTA runtime contract requires JDK 21 or 25. This is a concrete example of why an architecture document can remain useful while one software row becomes stale. Timestamp your design review and link the target release’s software/database/scanner requirements.
7. Worked scenarios: select a pattern and state the observable proof
| Scenario | Choice | Prerequisite / boundary | Predicted state change | Validation / rollback |
|---|---|---|---|---|
| CE queue builds only during 02:00 burst; CPU/DB/search quiet outside burst | Stagger batch analyses first | CI scheduler control; no paid edition required | Lower peak arrivals and pending time; same total daily analyses | Compare same workload window; revert schedule if governance window missed |
| CE queue grows all day; tasks stable; DB/disk/network healthy; CE host has CPU/RAM headroom | Optional Enterprise+ CE worker increase | Enterprise+ license; restart/config rules; CE heap capacity | More concurrent CE tasks; more DB/network demand | Queue age drops, task p95/DB latency stay within guardrails; restore worker count if not |
| CE tasks slow when DB CPU/latency rises; more workers already added once | Scale/fix database path before workers | Supported DB administration and rollback | DB service latency falls; task time recovers | DB + CE p95 improve; no direct SonarQube table edits |
| UI/API slow while CE queue is empty; search disk latency high | Move search data to suitable local low-latency storage / adjust search role resources | Supported storage migration/deployment procedure | Search latency lower; UI/API response improves | Health + UI/API p95 + disk latency; rollback storage mapping if regression |
| Single node cannot meet required HA and role-specific growth | Evaluate Data Center Edition | Commercial DCE license; Chapter30 topology/network/DB prerequisites | App/search roles become horizontally scalable and redundant | Cluster health, failure test, throughput; architecture rollback plan |
8. Maintainability, least privilege and operational cost are sizing constraints too
A performance design that requires permanent administrator tokens, undocumented JVM flags or manual index surgery is not production-ready even if one benchmark improves. Keep tuning inputs as versioned deployment configuration, record owners and before/after evidence, and prefer supported metrics/interfaces. Host and database monitoring identities should be read-only. Scanner tokens should retain only analysis permissions. CI scheduling changes belong in the provider/repository configuration that owns the schedule.
Portability also means recording hardware characteristics rather than cloud marketing labels alone: vCPU model/family, memory, storage type/latency, database tier, region/latency, container limits and noisy-neighbor assumptions. “Same 8 vCPU” across generations is not guaranteed equivalent capacity.
9. A reversible sizing change record
change_id: SQ-PERF-031-CE-01
observed_window: 2026-09-08T10:00Z/2026-09-08T11:00Z
exact_release: "SonarQube Server <record exact build>"
workload:
analyses_per_hour: 42
burst_peak_5m: 9
project_mix: "recorded separately"
evidence:
ce_pending_p95: 6
ce_oldest_wait_p95_seconds: 190
ce_task_p95_seconds: 72
database_latency: "healthy in baseline"
search_disk_latency: "healthy in baseline"
hypothesis: "queue is worker-concurrency limited during sustained window"
change: "example only: Enterprise worker count 1 -> 2"
validation:
- "pending p95 <= 2"
- "CE task p95 does not regress >10%"
- "DB latency and pool pressure remain inside baseline guardrail"
rollback: "restore worker count and CE heap configuration; restart per supported procedure"
This is deliberately not a command recipe. The exact UI/config action depends on edition/version and deployment model. The record makes the reasoning reviewable before any disruptive change.
Knowledge check
When is a CE-worker increase a plausible first change?
When queue depth/age is the problem, individual task times are stable, and database, disk, network, CPU and heap have demonstrated headroom. Worker management also requires Enterprise Edition or above.
Why can CI staggering be a capacity solution?
It changes arrival burstiness and can reduce peak queue demand without changing SonarQube resources, provided governance allows the shifted schedule.
What makes horizontal scaling an edition decision?
SonarQube Data Center Edition owns the supported multi-node application/search cluster model. Multiple independent Community/Server nodes are not a single horizontally scaled instance.
A reference page says OpenJDK 17 but the target 2026.1 LTA release notes say JDK 21/25. Which wins?
The target release compatibility contract wins. Keep the reference architecture’s workload/topology reasoning but update version-sensitive prerequisites.
What must accompany a performance configuration change?
A causal hypothesis, before evidence, predicted affected layer, validation guardrails, owner/version assumptions and an explicit rollback.
Official references and version notes
- SonarQube downloads — Current Community Build, commercial Server release train, editions and active LTA.
- Server host requirements — Current disk, memory, CPU, local-storage and production database-host guidance.
- Community Build host requirements — Equivalent Community Build host/search guidance for the mandatory free path.
- Monitoring the instance — Web/CE/search JVMs, read-only JMX MBeans and ComputeEngineTasks/Database signals.
- Reference architecture up to 10 M LOC — Planning anchor and its stated normal-usage assumptions.
- Improving performance — Enterprise+ CE-worker guidance and the requirement to measure external bottlenecks.
- Performance issues — Current troubleshooting starting points for storage, scope, workers and database-related performance.
- Server release notes — 2026.1 runtime/database changes, including JDK 21/25 and PostgreSQL 14–18.
- SonarScanner CLI releases — Current scanner release identity; 8.1.0.6389 is the latest listed release at this chapter recheck.
- Web API — Bearer authentication guidance and the ongoing Web API V2 migration.
- DCE performance — Current worker replication behavior and up-to-10 configured workers per application node context.
Version/compatibility baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. For the 2026.1 LTA, server runtime requires a JDK and supports Java 21 or 25; PostgreSQL 14–18 is supported. Scanner runtimes without JRE auto-provisioning should use Java 21 or newer. Web API V2 is still gradually replacing legacy endpoints. Always re-check the target release before copying any sizing or runtime value.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.