Chapter 31Lesson 03~165 minutes

Performance Sizing, Reference Architectures, and Large-Instance Tuning: Configuration, Design Patterns, and Trade-Offs

Choose vertical or edition-dependent horizontal scaling, memory and Compute Engine strategies from observed bottlenecks, reversibility and operational evidence.

ScalingHeap & page cacheCE workersReference architecturesTrade-offs

Learning objectives

  • Choose vertical versus edition-dependent horizontal scaling from the constrained SonarQube role.
  • Budget Web/CE/search heaps without starving the operating system and Lucene page cache.
  • Explain why additional CE workers can increase database, disk and network contention.
  • Compare scheduled/staggered CI with bursty arrivals as a capacity-control pattern.
  • Evaluate reference architectures with an explicit assumption delta, validation plan and rollback.

1. Configuration is a hypothesis about the bottleneck

Lesson 2 produced measurements rather than a magical server size. Lesson 3 asks what change those measurements justify. Every performance change should have the form: evidence says layer X is limiting; change Y should affect X; metric Z should improve without violating guardrails A/B; rollback is R if it does not. That structure prevents “tuning by folklore.”

Performance choices also change operational state. Increasing a JVM heap changes host memory pressure and restart requirements. Increasing Enterprise CE workers changes database/search concurrency. Data Center horizontal scaling changes licensed topology, node count and failure domains. Staggering CI changes provider scheduling but no SonarQube persistent state. Treat these as different changes with different owners and rollback mechanisms.

2. Vertical versus horizontal scaling: first name the role

Observed limiting signal Candidate change Edition / prerequisite Evidence after change
Web CPU/heap pressure during API/user peak; CE/search healthy Increase Web/host capacity Single-node editions; supported host limits Web latency/CPU/GC improve; CE timings remain stable
CE CPU/heap pressure on large tasks; external resources healthy Increase CE heap/host CPU; optionally workers if queue-bound Worker count management: Enterprise+ Task time/queue age improves without DB/disk regression
Search disk latency/page-cache pressure Faster local SSD / correct memory budget / search-role capacity All single-node editions for vertical; DCE for search-node horizontal role Index/query latency and UI/API response improve
Database saturation/latency Scale/tune supported DB service and connection path All editions use supported external production DB DB latency/pool pressure falls and CE/Web recover
Node-level capacity/availability beyond one server DCE application/search horizontal scaling Licensed Data Center Edition; topology rules from Chapter 30 Role-specific throughput/HA improves; cluster remains healthy
Bursty CI creates short queue while daily total is modest Stagger/schedule analyses CI/provider capability; edition-neutral Peak pending count/time falls with same server resources

“Horizontal scaling” is not a generic Community Build switch. Data Center Edition is the licensed distributed architecture that separates application and search roles and permits horizontal scaling. Enterprise Edition can provide multiple CE workers and other scale-oriented capabilities inside a single-node architecture, but that does not turn several independent Enterprise servers into one cluster.

3. Heap versus OS page cache: optimize the whole memory system

Memory decisions are process-specific. Web heap supports interactive/API work. CE heap supports background report processing and scales with project/task complexity and worker concurrency. Search heap supports Elasticsearch structures, while Lucene relies heavily on filesystem cache outside the JVM. Native allocations and the operating system also consume memory.

Start with a host memory budget, not a heap target. On a single node, list observed peaks for Web, CE and search RSS/heap, add OS/native/container overhead, then preserve page-cache headroom. SonarSource’s current host guidance says roughly half of available memory should remain outside Elasticsearch heap for Lucene caching and warns not to make search heap excessively large. That search-specific principle is incompatible with “set every Java process to the maximum.”

# Capacity worksheet, not sonar.properties syntax
host_ram_gib = 32
web_heap_gib = 2       # hypothesis from observed Web pressure
ce_heap_gib = 6        # hypothesis from observed large CE tasks
search_heap_gib = 8    # preserve substantial non-heap cache
known_heap_total = web_heap_gib + ce_heap_gib + search_heap_gib
non_heap_budget = host_ram_gib - known_heap_total
# Validate native/RSS + page-cache + OS margin before approving.
Restart and rollback boundary

Changing Java options is disruptive and must be done through supported SonarQube configuration/deployment mechanisms with a rollback value and maintenance plan. Do not live-edit JVM internals or delete search data to “apply” memory tuning.

4. More Compute Engine workers can make the wrong bottleneck worse

Worker concurrency helps only when reports are waiting and the infrastructure can absorb parallel work. SonarSource’s performance guidance is explicit that CE workers increase stress on the database, disk I/O, network, heap and CPU. If database latency, storage latency or network is already the bottleneck, more workers create more contenders for the same constrained door and may lengthen each task.

Worker count is also an edition boundary. Current guidance places CE worker management in Enterprise Edition and above. In Data Center Edition, a configured worker count is replicated across application nodes; for example, four configured workers across two app nodes becomes eight effective workers after the required node restarts. That multiplication makes capacity and database impact a cluster-wide calculation, not a per-node tweak.

PendingTime ↑ + PendingCount ↑ → individual CE time stable → DB/disk/network healthy → worker increase is a plausible hypothesis
individual CE time ↑ → DB latency / disk wait ↑ → queue grows secondarily → fix downstream resource before worker count

5. Scheduled versus bursty CI: smooth demand before buying capacity

CI is an owning system in the sizing model. If nightly jobs for hundreds of repositories all begin at 02:00, the average analyses-per-day metric hides a sharp arrival spike. A one-hour queue after 02:00 may disappear for the rest of the day, yet developers still experience delayed policy evidence. One reversible fix is to stagger non-urgent analysis windows or constrain batch concurrency in the CI orchestrator.

Scheduling can be cheaper and safer than server changes because it alters arrival shape rather than SonarQube persistent state. It does not solve genuine all-day saturation, and it must not postpone required quality/security analysis beyond governance needs. Document which jobs can move, which are merge-critical, expected peak reduction and a rollback to the original schedule.

Pattern Benefit Trade-off / risk Evidence
All jobs at a fixed minute Simple orchestration Large burst; queue wait and DB/search spikes PendingCount/PendingTime peak
Jitter/stagger non-urgent jobs Reduces peak arrival rate without server mutation Longer batch window; schedule governance needed Peak queue falls; same total analyses
Concurrency cap at CI layer Protects SonarQube and DB from fan-in Jobs wait in CI rather than CE; distinguish wait locations CI queue vs CE queue evidence
Scale platform Preserves fast arrivals when demand is legitimate Cost, edition, ops complexity; may shift bottleneck Role/resource metrics after change

6. Reference architecture versus measured sizing: build an assumption delta

A reference architecture is most useful as a review checklist. Copy its topology into one column and your measured environment into another. Every mismatch becomes a question, not an automatic defect. The current 10M scenario assumes “normal” daily-main/several-PR analysis, average repositories around 50k LOC, occasional API use and no third-party plugins. The older/currently indexed 50M example uses a larger host and multiple workers, but the same lesson applies: the headline LOC number is conditional on workload shape.

Version prerequisites must be evaluated separately from hardware. The current 10M page still contains an OpenJDK 17 line while the 2026.1 LTA runtime contract requires JDK 21 or 25. This is a concrete example of why an architecture document can remain useful while one software row becomes stale. Timestamp your design review and link the target release’s software/database/scanner requirements.

7. Worked scenarios: select a pattern and state the observable proof

Scenario Choice Prerequisite / boundary Predicted state change Validation / rollback
CE queue builds only during 02:00 burst; CPU/DB/search quiet outside burst Stagger batch analyses first CI scheduler control; no paid edition required Lower peak arrivals and pending time; same total daily analyses Compare same workload window; revert schedule if governance window missed
CE queue grows all day; tasks stable; DB/disk/network healthy; CE host has CPU/RAM headroom Optional Enterprise+ CE worker increase Enterprise+ license; restart/config rules; CE heap capacity More concurrent CE tasks; more DB/network demand Queue age drops, task p95/DB latency stay within guardrails; restore worker count if not
CE tasks slow when DB CPU/latency rises; more workers already added once Scale/fix database path before workers Supported DB administration and rollback DB service latency falls; task time recovers DB + CE p95 improve; no direct SonarQube table edits
UI/API slow while CE queue is empty; search disk latency high Move search data to suitable local low-latency storage / adjust search role resources Supported storage migration/deployment procedure Search latency lower; UI/API response improves Health + UI/API p95 + disk latency; rollback storage mapping if regression
Single node cannot meet required HA and role-specific growth Evaluate Data Center Edition Commercial DCE license; Chapter30 topology/network/DB prerequisites App/search roles become horizontally scalable and redundant Cluster health, failure test, throughput; architecture rollback plan

8. Maintainability, least privilege and operational cost are sizing constraints too

A performance design that requires permanent administrator tokens, undocumented JVM flags or manual index surgery is not production-ready even if one benchmark improves. Keep tuning inputs as versioned deployment configuration, record owners and before/after evidence, and prefer supported metrics/interfaces. Host and database monitoring identities should be read-only. Scanner tokens should retain only analysis permissions. CI scheduling changes belong in the provider/repository configuration that owns the schedule.

Portability also means recording hardware characteristics rather than cloud marketing labels alone: vCPU model/family, memory, storage type/latency, database tier, region/latency, container limits and noisy-neighbor assumptions. “Same 8 vCPU” across generations is not guaranteed equivalent capacity.

9. A reversible sizing change record

change_id: SQ-PERF-031-CE-01
observed_window: 2026-09-08T10:00Z/2026-09-08T11:00Z
exact_release: "SonarQube Server <record exact build>"
workload:
  analyses_per_hour: 42
  burst_peak_5m: 9
  project_mix: "recorded separately"
evidence:
  ce_pending_p95: 6
  ce_oldest_wait_p95_seconds: 190
  ce_task_p95_seconds: 72
  database_latency: "healthy in baseline"
  search_disk_latency: "healthy in baseline"
hypothesis: "queue is worker-concurrency limited during sustained window"
change: "example only: Enterprise worker count 1 -> 2"
validation:
  - "pending p95 <= 2"
  - "CE task p95 does not regress >10%"
  - "DB latency and pool pressure remain inside baseline guardrail"
rollback: "restore worker count and CE heap configuration; restart per supported procedure"

This is deliberately not a command recipe. The exact UI/config action depends on edition/version and deployment model. The record makes the reasoning reviewable before any disruptive change.

Knowledge check

When is a CE-worker increase a plausible first change?

Why can CI staggering be a capacity solution?

What makes horizontal scaling an edition decision?

A reference page says OpenJDK 17 but the target 2026.1 LTA release notes say JDK 21/25. Which wins?

What must accompany a performance configuration change?

Next lesson

Break the sizing model on purpose

Lesson 4 engineers realistic failure modes—heap starvation of page cache, worker fan-out into a constrained database, unsuitable search storage, LOC-only sizing and burst blindness—then repairs the causal layer without erasing evidence.

Official references and version notes

Version and compatibility note

Version/compatibility baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. For the 2026.1 LTA, server runtime requires a JDK and supports Java 21 or 25; PostgreSQL 14–18 is supported. Scanner runtimes without JRE auto-provisioning should use Java 21 or newer. Web API V2 is still gradually replacing legacy endpoints. Always re-check the target release before copying any sizing or runtime value.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.