Checkpoint Lab — Performance Sizing, Reference Architectures, and Large-Instance Tuning
Create a measured capacity worksheet, identify the first limiting layer and propose one reversible sizing change with explicit validation and rollback criteria.
Learning objectives
- Produce a capacity worksheet from repeated same-revision baselines, one changed workload factor and one controlled burst.
- Predict at least two source / indexing / task / database / search / CI / governance state changes and verify each independently.
- Identify the first limiting layer using correlated lifecycle and resource evidence rather than a single utilization graph.
- Propose exactly one reversible sizing or demand-shaping change with explicit acceptance and rollback criteria.
- Package revision, parameters, scanner/task/server/DB/host evidence and limitations for independent review.
1. Checkpoint mission and acceptance boundary
You are the capacity reviewer for a team about to move SonarQube from occasional developer analysis into regular CI. Management asks whether the current instance “needs a bigger architecture.” You are not allowed to answer from LOC alone or from a vendor diagram. Your job is to run several bounded local analyses, identify the first constrained layer in the scenario, recommend one reversible change, and state exactly what evidence would accept or reject it.
The mandatory executable path remains free: Community Build 26.9.0.129388 on a disposable local Docker fixture with PostgreSQL and synthetic source. Commercial options such as additional CE workers or Data Center horizontal scaling appear only as design alternatives/simulations; a paid license is not required to complete the checkpoint.
A correct lab can conclude “no resource change is justified yet” if the evidence does not identify a limiting layer. The point is causal sizing, not forcing a tuning action.
2. Exact assumptions and resource/credential preflight
| Item | Required evidence |
|---|---|
| Server |
api/server/version showing Community Build
26.9.0.129388; official image tag/digest recorded
|
| Scanner |
sonar-scanner -v; target 8.1.0.6389 for this
checkpoint, or exact later compatible version documented
|
| Java/JRE | Scanner runtime/provisioning mode recorded; Java 21+ if auto-provisioning is disabled/not supported |
| Database |
PostgreSQL 17 lab container; exact
SELECT version() and image digest; no
application-table edits
|
| Plugins/integrations | No third-party plugins, IdP, cloud CI or provider decoration in mandatory path |
| Host | CPU model/core count, RAM, Docker/container limits, storage type/free space, OS/Docker version |
| Credential | Local-only least-privilege analysis token supplied through process environment; token value excluded from evidence |
| Source | Synthetic fixture only; Git SHA/tag for each workload state |
If you cannot meet the host preflight without competing workloads, continue only if you mark that contention as an explicit limitation. Do not close unrelated programs or alter host security merely to manufacture a clean benchmark.
3. Write predictions before you run anything
Predictions stop the lab becoming a post-hoc story. Write at least two; four are recommended. Each prediction names a state owner and an independent verification source.
| Prediction before action | State owner | Independent verification |
|---|---|---|
| 120→480 generated files changes revision and increases scanner indexed-file workload | Git + scanner |
git rev-parse plus scanner indexed-file/log
evidence
|
| A burst of three independent project reports creates CE waiting on Community Build | Compute Engine |
three ceTaskId records + task
timestamps/pending evidence
|
| Changing source-file count does not change the project key or quality-policy assignment | Project policy | effective scanner config + project UI/API identity before/after |
| Read-only PostgreSQL counters change during CE processing, but no schema/config mutation occurs | Database |
pg_stat_database snapshots + change log showing
no DB mutation
|
Do not predict a gate pass/fail unless the fixture and policy are intentionally designed for that result. This checkpoint measures capacity; changing policy thresholds to create a convenient result would contaminate the experiment.
4. Recreate the disposable fixture and evidence directory
Use the Compose file and deterministic fixture generator from Lesson
2. Create an evidence directory outside the Compose volume tree so
docker compose down -v cannot delete it. Record
environment.md before the first scanner run.
mkdir -p ../sq31-evidence
cat > ../sq31-evidence/environment.md <<'EOF'
# SQ31 capacity checkpoint
Recheck date: 2026-09-08
Server: Community Build 26.9.0.129388
Scanner: record sonar-scanner -v
Database: record SELECT version()
Host CPU/RAM/storage: record here
Docker + image digests: record here
Plugins: none
CI/provider/IdP: none (local simulation)
Known limitations: record here
EOF
git log -1 --format=fuller > ../sq31-evidence/revision-baseline.txt
sonar-scanner -v > ../sq31-evidence/scanner-version.txt 2>&1
5. Execute the measurement matrix
Run the same matrix for which Lesson 2 supplied the runner. Preserve every raw run directory; the CSV is a summary, never a replacement for raw logs and task JSON.
| Class | Workload | Runs | Purpose |
|---|---|---|---|
| warm-up | 120-file baseline, exact baseline SHA | 1 | Expose cold/warm effects; retain but do not mix into measured baseline |
| measured | 120-file baseline, same SHA | 3–5 | Estimate repeatability and scanner/CE typical/tail behavior |
| warm-up | 480-file changed revision | 1 | Warm changed workload without hiding it |
| measured | 480-file changed revision | 3–5 | Measure one controlled workload-factor change |
| burst | three stable project keys, same source revision, close submission times | 1 bounded burst | Expose queue wait/arrival behavior |
For every run, capture scanner elapsed time, revision, effective
source/project inputs, indexed-file evidence,
report-task.txt, raw CE task JSON/status, and a
host/container snapshot. For each window, capture DB operational
counters and disk/free-space observations. If any run fails, branch
into diagnosis; do not silently rerun until it succeeds.
6. Build the capacity worksheet without hiding raw evidence
class,run,revision,project_key,files,scanner_s,queue_wait_s,ce_s,ce_status,sq_cpu_pct,sq_mem_mib,db_numbackends,db_blks_read_delta,db_blks_hit_delta,disk_free_gib,notes
warmup,W0,<sha>,sq31-capacity,120,,,,,,,,,,,
measured,B1,<same-sha>,sq31-capacity,120,,,,,,,,,,,
measured,B2,<same-sha>,sq31-capacity,120,,,,,,,,,,,
measured,B3,<same-sha>,sq31-capacity,120,,,,,,,,,,,
warmup,L0,<large-sha>,sq31-capacity,480,,,,,,,,,,,
measured,L1,<large-sha>,sq31-capacity,480,,,,,,,,,,,
measured,L2,<large-sha>,sq31-capacity,480,,,,,,,,,,,
measured,L3,<large-sha>,sq31-capacity,480,,,,,,,,,,,
burst,A,<burst-sha>,sq31-burst-a,480,,,,,,,,,,,
burst,B,<burst-sha>,sq31-burst-b,480,,,,,,,,,,,
burst,C,<burst-sha>,sq31-burst-c,480,,,,,,,,,,,
Add columns if your evidence justifies them—API response p95, process-specific RSS, disk latency, GC pause, network RTT—but never invent numbers because a template has a column. “Not collected” is valid evidence of a limitation.
from __future__ import annotations
import csv, statistics, sys
from pathlib import Path
p = Path(sys.argv[1])
rows = list(csv.DictReader(p.open(encoding='utf-8')))
measured = [r for r in rows if r['class'] == 'measured']
def nums(key):
return [float(r[key]) for r in measured if r.get(key)]
def pct(values, q):
if not values: return None
v = sorted(values)
idx = min(len(v)-1, max(0, round((len(v)-1)*q)))
return v[idx]
for key in ['scanner_s','queue_wait_s','ce_s']:
v = nums(key)
print(key, {
'n': len(v),
'median': statistics.median(v) if v else None,
'p95_nearest_rankish': pct(v, 0.95),
'max': max(v) if v else None,
})
# Offered CE utilization approximation for the selected window.
window_s = float(sys.argv[2])
workers = int(sys.argv[3])
ce_total = sum(nums('ce_s'))
print('ce_busy_ratio', ce_total / (window_s * workers))
print('NOTE: this is a bounded-workload diagnostic ratio, not a queueing guarantee.')
The script’s ce_busy_ratio is a bounded diagnostic
ratio: sum of measured CE service seconds divided by worker-seconds
available in the selected window. It is not an SLO and not a
guarantee of queueing behavior. Use the actual burst/pending
evidence for tail behavior.
7. Identify the first limiting layer
Use an elimination table. A limiting layer is not simply the metric with the largest percentage; it is the earliest constrained owner whose pressure causally explains the observed experience and changes with the workload.
| Question | If yes | If no |
|---|---|---|
| Does scanner time dominate while CE wait/service stay low? | Scanner/build-runner/scope/cache/network is first candidate | Continue server lifecycle |
| Does queue wait grow while individual CE time stays stable? | Arrival/service concurrency is first candidate | Continue |
| Does CE time grow with CE CPU/heap pressure but DB/search healthy? | CE process/host is first candidate | Continue |
| Does CE time grow with DB latency/pool/CPU pressure? | Database path is first candidate | Continue |
| Do CE/UI symptoms correlate with search disk/page-cache pressure? | Search/storage/memory budget is first candidate | Continue |
| Is UI/API slow while analysis path is healthy? | Web/search/query/API concurrency is first candidate | Continue |
| Are symptoms only during a CI burst? | Arrival-shaping may be first reversible control | Consider sustained capacity only if demand remains high |
State uncertainty. A Docker stats snapshot cannot prove
p95 disk latency, and pg_stat_database counters do not
prove query-level latency. If the evidence packet lacks the signal
needed to distinguish two hypotheses, recommend better measurement
before infrastructure change.
8. Propose exactly one reversible change
Choose one change that targets your observed first limit. Examples: increase local lab container memory modestly if CE/search memory pressure is proven; move the lab’s search data from a deliberately slow/simulated path back to local SSD; stagger the three analysis submissions if burst wait is the only problem; or, in a documented commercial simulation, propose an Enterprise CE-worker increase only after proving DB/disk/network headroom.
proposal_id: SQ31-CHECKPOINT-CHANGE-01
first_limiting_layer: "<scanner | CE queue | CE process | DB | search/storage | Web/API | CI arrival>"
baseline_evidence:
- "<raw artifact + observation>"
hypothesis: "<why this one change affects the limiting layer>"
change: "<one reversible setting/resource/schedule change>"
edition_prerequisite: "Community/free OR simulated Enterprise/DCE"
expected_effect:
- "<metric should improve>"
guardrails:
- "CE task p95 must not regress by >10%"
- "DB/search/health remains inside recorded baseline envelope"
validation_window: "same bounded workload matrix"
rollback: "<exact previous value/schedule/deployment mapping>"
owner: "<team/role>"
If your proposal is “increase CE workers” or “horizontal scale app/search nodes,” do not attempt it on Community Build. Document the Enterprise/Data Center prerequisite and validate the reasoning with the free queue/capacity evidence. Real execution requires the matching commercial license and current deployment procedure.
9. Rerun the smallest equivalent scenario and decide
Apply the one change only in the disposable lab if it is free/local and safe. If it is commercial-only, simulate the expected concurrency in the worksheet rather than bypassing edition controls. Rerun the same revision/workload, run count and burst shape. Acceptance requires the target metric to improve while guardrails do not regress. A faster one-off run is insufficient.
| Decision | Required statement |
|---|---|
| Accept | Target metric improved across comparable runs; no guardrail or correctness/security state regressed; assumptions remain valid |
| Reject / rollback | Target did not improve enough, tail worsened, DB/search/health pressure increased, or new instability appeared |
| Inconclusive | Run variance or missing metrics prevent attribution; restore baseline and improve instrumentation before another change |
10. Required evidence packet
Package the following without the token value. The packet should let another engineer trace a row in the worksheet back to an exact revision, scanner run and CE task, and then correlate that task with server/database/host evidence.
- revision/manifest: Git SHA/tag, generated-file count and fixture-generator checksum or source;
-
effective analysis inputs: project key,
sonar.sources, scanner/runtime version, JRE provisioning note and redacted environment; -
scanner/index/report evidence: scanner log,
indexed-file evidence, elapsed time and preserved
report-task.txt; -
task/server evidence: raw
api/ce/taskresponse for each task, health/status response and relevant server log window; - resource evidence: CPU/RAM/container limits, heap/config assumptions, disk space/storage type, database operational snapshots and any network/latency observations actually collected;
- capacity worksheet: raw rows plus calculations, with warm-up clearly separated;
- governance: predictions, first-limit conclusion, proposal, edition boundary, validation/rollback criteria and assumptions/limitations.
sq31-evidence/
environment.md
predictions.md
revision-baseline.txt
revision-large.txt
scanner-version.txt
runs/
W0/ B1/ B2/ B3/ L0/ L1/ L2/ L3/ burst-a/ burst-b/ burst-c/
scanner.log
report-task.txt
ce-task.json
host-snapshot.txt
database-before.txt
database-after.txt
server-window.log
capacity.csv
capacity-summary.txt
proposal.yaml
validation.md
assumptions-limitations.md
11. Verification checklist
- ☐ Server, scanner, Java/JRE mode, database and container/image identities are exact and timestamped.
- ☐ No real credential, proprietary source, PII or production endpoint appears in the packet.
- ☐ Warm-up runs are retained and labeled; at least three measured repetitions exist per steady workload.
-
☐ Every scanner run maps to a preserved
ceTaskIdand terminal task status. - ☐ Scanner success, upload, CE success, analysis completion, gate result and any CI/provider state are not collapsed.
- ☐ The first limiting layer is supported by at least two correlated evidence sources or explicitly marked uncertain.
- ☐ The proposal changes one primary factor, names edition prerequisites, guardrails, owner and rollback.
- ☐ Reference-architecture comparisons preserve their workload assumptions and do not import stale Java/runtime lines.
- ☐ Cleanup is scoped only to the disposable lab resources.
12. Cleanup / rollback after the checkpoint
Export the evidence packet first. Revoke the local analysis token. If you changed a free/local parameter, restore its exact baseline value and capture one health check. Then destroy only this Compose project and its named volumes. The evidence packet should survive the teardown.
# PowerShell
Remove-Item Env:SONAR_TOKEN -ErrorAction SilentlyContinue
docker compose down -v
# Verify no lab containers remain.
docker ps -a --filter name=sq31-
# Keep ../sq31-evidence for review; delete it only when your retention policy allows.
Knowledge check
Your three burst tasks wait, but the steady measured runs have almost zero queue wait and healthy DB/search. What is the strongest first proposal?
Smooth or cap the burst at the CI/provider layer if governance permits. The evidence points to arrival shape rather than sustained server saturation.
The scanner time rises 3× after the larger revision, but CE service time and host/DB/search signals barely move. Which layer is first limiting?
The scanner/build-runner/analysis-scope path is the first candidate; a server worker or DB change is not supported by the evidence.
An Enterprise design proposes four CE workers on two DCE application nodes. What hidden multiplication must the reviewer account for?
DCE worker configuration is replicated across app nodes; four configured workers on two application nodes means eight effective workers, with corresponding downstream DB/search/network/heap implications.
Why is “the 10M reference architecture says 8 GB” not sufficient validation for your 8-GB host?
The architecture assumes a particular edition/topology, project mix, analysis frequency, API load and plugin profile, and even its software row can lag current runtime requirements. Your measured workload and target-release compatibility must validate it.
A proposed tuning change improves average CE time but makes p95 DB latency and UI response much worse. Do you accept it?
No. A capacity change must meet its guardrails. Roll back or reject it; shifting the bottleneck to shared DB/Web experience is not a successful improvement.
What makes the evidence packet governable rather than just a benchmark spreadsheet?
It links exact revision and effective inputs to scanner/report/task/server/resource evidence, records edition/version assumptions, predictions, limitations, owner, causal proposal, validation criteria and rollback.
13. What Chapter 31 adds to the governed SonarQube operating model
Earlier chapters made analysis reproducible, policy auditable, operations observable, recovery testable and Data Center topology explicit. Chapter 31 adds capacity governance: a repeatable way to translate measured workload arrivals and service/resource evidence into a reversible infrastructure or scheduling decision. The operating model now has a capacity baseline, headroom assumptions, workload-shape record, role-specific bottleneck diagnosis, edition-aware scaling rules and change-validation envelope.
That does not eliminate incidents. Real systems still fail from scanner drift, damaged indexes, CI/provider changes, authentication/network faults and incompatible upgrades. Chapter 32 takes the evidence discipline you just practiced and applies it to end-to-end troubleshooting of analysis failures, index problems, CI drift and production incidents.
Official references and version notes
- SonarQube downloads — Current Community Build, commercial Server release train, editions and active LTA.
- Server host requirements — Current disk, memory, CPU, local-storage and production database-host guidance.
- Community Build host requirements — Equivalent Community Build host/search guidance for the mandatory free path.
- Monitoring the instance — Web/CE/search JVMs, read-only JMX MBeans and ComputeEngineTasks/Database signals.
- Reference architecture up to 10 M LOC — Planning anchor and its stated normal-usage assumptions.
- Improving performance — Enterprise+ CE-worker guidance and the requirement to measure external bottlenecks.
- Performance issues — Current troubleshooting starting points for storage, scope, workers and database-related performance.
- Server release notes — 2026.1 runtime/database changes, including JDK 21/25 and PostgreSQL 14–18.
- SonarScanner CLI releases — Current scanner release identity; 8.1.0.6389 is the latest listed release at this chapter recheck.
- Web API — Bearer authentication guidance and the ongoing Web API V2 migration.
Version/compatibility baseline rechecked 2026-09-08: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389; commercial Server current train 2026 Release 4.1 / 2026.4.1; active LTA 2026.1.5 LTA. For the 2026.1 LTA, server runtime requires a JDK and supports Java 21 or 25; PostgreSQL 14–18 is supported. Scanner runtimes without JRE auto-provisioning should use Java 21 or newer. Web API V2 is still gradually replacing legacy endpoints. Always re-check the target release before copying any sizing or runtime value.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.