Data Center Edition, Clustering, High Availability, and Scale: Diagnostics, Failure Modes, and Production Practices
Diagnose clustered SonarQube failures by separating load-balancer, application/Hazelcast, Compute Engine, search-cluster, database, network, version/plugin, and licensing layers before changing anything.
Learning objectives
- Apply an evidence-first diagnostic sequence to DCE incidents without flattening every failure into “cluster problem.”
- Diagnose mixed versions/plugins, bad backplane networking, unhealthy search storage, database bottlenecks, and load-balancer routing separately.
- Explain why several standalone servers are not a cluster and why sticky sessions do not repair cluster membership.
- Reject direct Elasticsearch edits, Internet-exposed inter-node ports, blanket restarts, and capacity changes without measurements.
- Repair an intentionally broken mixed-version topology while preserving the original evidence.
1. Evidence-first DCE incident sequence
-
Preserve scanner/CI evidence,
report-task.txt,ceTaskId, load-balancer status, node logs, cluster/search health, and timestamps. - Confirm DCE license, exact server release, Java/database versions, plugins, Helm/chart/image versions, and node role/count.
- Confirm the exact source revision/effective analysis parameters before blaming infrastructure for a project-specific error.
- Inspect scanner upload and specific CE task state.
- Inspect CE queue/workers and application-node Web/CE logs.
-
Inspect search-cluster health, disk, heap/GC, and
es.log. - Inspect database connectivity, latency, pool pressure, storage, and locks.
- Inspect load-balancer health/backends and private 9000/9001/9002/9003 network reachability.
- Apply the least-destructive correction to the owning layer.
- Rerun the smallest equivalent scenario and compare evidence at the same source revision when possible.
2. Failure: “we have five servers, therefore we have HA”
Five independent SonarQube instances do not share cluster membership, application coordination, or one search cluster. Putting a load balancer in front of them can route users to inconsistent state and does not create a supported DCE architecture. The correction is architectural: deploy licensed DCE with explicit node roles and shared database/search membership, or operate a single supported server.
3. Intentionally broken example: mixed versions/plugins
Start from the Lesson 2 manifest and make two changes:
python - <<'PY'
import json
p='dce-topology.json'; d=json.load(open(p))
for n in d['nodes']:
if n['name']=='app-2':
n['version']='2026.3.1'
n['plugins_hash']='different-plugin-set'
json.dump(d, open('dce-broken-mixed.json','w'), indent=2)
PY
python dce_sim.py dce-broken-mixed.json \
> evidence/mixed-version/validator.stdout 2> evidence/mixed-version/validator.stderr || true
Preserve before repair. The validator should report mixed SonarQube versions and divergent application plugin state. In a real DCE incident, preserve exact binaries/image digests, plugin manifests/checksums, node startup logs, and the deployment event that created drift.
Repair: return app-2 to the exact
cluster release and approved plugin artifact set. Do not “fix” the
symptom by removing the app node from the load balancer forever or
by downgrading the rest of the cluster.
4. Failure: exposing or blocking the wrong cluster port
| Symptom | Likely path | Evidence | Repair |
|---|---|---|---|
| LB reports app unhealthy | LB→app 9000 / proxy health | LB probe + web status/log | Repair app health/proxy route; do not open 9003 publicly |
| App nodes cannot coordinate | app↔app 9003/Hazelcast | cluster logs + private firewall tests | Restore private app-to-app reachability |
| Application cannot query/index search | app→search 9001 | web/CE/es logs + private connectivity | Repair application-to-search path |
| Search cluster cannot form | search↔search 9002 | es.log + host list/version |
Repair private search transport/host inventory |
5. Failure: unsuitable search storage
Search nodes are latency-sensitive and current Sonar guidance says
SSDs perform significantly better than HDDs. If search heap is
healthy but disk latency or capacity pressure is high, adding
application nodes increases incoming work without repairing the
bottleneck. Preserve es.log, disk latency/utilization,
heap/GC, and search health first. Never edit Elasticsearch indexes
directly or move live search data with generic file-copy commands as
a repair.
6. Failure: shared database is the actual bottleneck
Every application node depends on the same supported database. Adding application nodes can increase database concurrency. If JDBC pool waits, DB CPU/IO/locks, or network latency are already saturated, horizontal app scaling can worsen the incident. Use database/vendor metrics and Sonar Web/CE evidence together. Current server support includes PostgreSQL 14–18, SQL Server 2017/2019/2022, and supported Oracle versions; verify the exact release matrix before changing engines or versions.
7. Failure: treating sticky sessions as the HA fix
Sticky sessions can hide routing symptoms but do not repair a broken application cluster, inconsistent JWT secret, mixed node version, or unhealthy backend. Current DCE explicitly does not require session affinity. Preserve the load-balancer backend/health evidence, then repair the owning cluster or network layer.
8. Failure: adding app nodes when search/database is saturated
Capacity management is causal. If CE queue/CPU is high and
downstream tiers are healthy, app scaling may help. If
es.log, search heap, or disk is the limiting layer, fix
search. If JDBC/DB latency is limiting, fix the database path. Node
count is not a generic “performance” knob.
9. Production shortcuts to reject
- Do not expose 9001/9002/9003 to the public Internet.
- Do not direct-edit Elasticsearch state or the SonarQube database.
- Do not repeatedly restart the whole cluster before preserving first-failure evidence.
- Do not run mixed server/plugin versions as a tolerated steady state.
- Do not bypass TLS verification or place cluster/shared secrets in source.
- Do not add nodes before identifying the bottleneck.
- Do not treat node HA as database/region disaster recovery.
- Do not change project keys, gates, or profiles to hide infrastructure failures.
Knowledge check
App nodes are healthy, but search nodes cannot reach each other. Which default path is suspect?
Search-to-search Elasticsearch transport, default port 9002, plus search host/version configuration.
Why is adding an app node dangerous when the database is saturated?
It can increase concurrent DB load while leaving the actual bottleneck unchanged.
What should you preserve before fixing a mixed-version node?
Exact node/image versions, plugin manifests/checksums, startup/cluster logs, deployment event, and current health evidence.
Should you copy Elasticsearch data directories to another search node as a first repair?
No. Use supported SonarQube/search recovery behavior; do not manipulate internal indexes directly.
A sticky-session rule makes logins look stable. Does that prove the DCE cluster is healthy?
No. DCE does not require sticky sessions; inspect node/JWT/network/search/database health instead.
Official references and version notes
- DCE topology — current role model and minimum/default 2 application + 3 search topology.
- DCE installation requirements — node identity, search-node placement, hardware guidance, Docker/network/load-balancer requirements.
- DCE-specific system properties — current cluster roles, host lists and default application-cluster port.
- DCE network rules — default 9000/9001/9002/9003 trust paths.
- DCE scaling — application-node scaling, currently up to ten application nodes.
- Installing the database — current supported database engines/versions including PostgreSQL 14–18.
- Official SonarQube DCE Helm chart — current DCE chart/image structure and application/search node configuration.
- Performing an update — coordinated DCE Helm/database migration behavior and post-update reanalysis.
Baseline: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389. Evidence continuity preserves ceTaskId as the Compute Engine task identity. Rechecked 2026-09-08. Mandatory executable examples remain Community Build 26.9.0.129388 with SonarScanner CLI 8.1.0.6389. DCE examples model the current SonarQube Server 2026 Release 4.1 / 2026.4.1 commercial architecture and require a valid Data Center Edition license for real execution. Current minimum/default DCE topology is two application plus three search nodes; one application and one search node may be lost without user impact when remaining dependencies are healthy. Application nodes can currently scale up to ten. Default private network paths are LB→app 9000, app→search 9001, search→search 9002 and app→app/Hazelcast 9003. Current supported PostgreSQL versions are 14–18. Recheck the exact target release, DCE license, Java/database matrix, Helm/chart version, plugin matrix, Kubernetes/OpenShift matrix, network properties, and update notes immediately before any real deployment.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.