Chapter 30Lesson 04~150 minutes

Data Center Edition, Clustering, High Availability, and Scale: Diagnostics, Failure Modes, and Production Practices

Diagnose clustered SonarQube failures by separating load-balancer, application/Hazelcast, Compute Engine, search-cluster, database, network, version/plugin, and licensing layers before changing anything.

Data Center EditionClusteringHigh AvailabilityScaleOperations

Learning objectives

  • Apply an evidence-first diagnostic sequence to DCE incidents without flattening every failure into “cluster problem.”
  • Diagnose mixed versions/plugins, bad backplane networking, unhealthy search storage, database bottlenecks, and load-balancer routing separately.
  • Explain why several standalone servers are not a cluster and why sticky sessions do not repair cluster membership.
  • Reject direct Elasticsearch edits, Internet-exposed inter-node ports, blanket restarts, and capacity changes without measurements.
  • Repair an intentionally broken mixed-version topology while preserving the original evidence.

1. Evidence-first DCE incident sequence

  1. Preserve scanner/CI evidence, report-task.txt, ceTaskId, load-balancer status, node logs, cluster/search health, and timestamps.
  2. Confirm DCE license, exact server release, Java/database versions, plugins, Helm/chart/image versions, and node role/count.
  3. Confirm the exact source revision/effective analysis parameters before blaming infrastructure for a project-specific error.
  4. Inspect scanner upload and specific CE task state.
  5. Inspect CE queue/workers and application-node Web/CE logs.
  6. Inspect search-cluster health, disk, heap/GC, and es.log.
  7. Inspect database connectivity, latency, pool pressure, storage, and locks.
  8. Inspect load-balancer health/backends and private 9000/9001/9002/9003 network reachability.
  9. Apply the least-destructive correction to the owning layer.
  10. Rerun the smallest equivalent scenario and compare evidence at the same source revision when possible.

2. Failure: “we have five servers, therefore we have HA”

Five independent SonarQube instances do not share cluster membership, application coordination, or one search cluster. Putting a load balancer in front of them can route users to inconsistent state and does not create a supported DCE architecture. The correction is architectural: deploy licensed DCE with explicit node roles and shared database/search membership, or operate a single supported server.

3. Intentionally broken example: mixed versions/plugins

Start from the Lesson 2 manifest and make two changes:

python - <<'PY'
import json
p='dce-topology.json'; d=json.load(open(p))
for n in d['nodes']:
    if n['name']=='app-2':
        n['version']='2026.3.1'
        n['plugins_hash']='different-plugin-set'
json.dump(d, open('dce-broken-mixed.json','w'), indent=2)
PY
python dce_sim.py dce-broken-mixed.json \
  > evidence/mixed-version/validator.stdout 2> evidence/mixed-version/validator.stderr || true

Preserve before repair. The validator should report mixed SonarQube versions and divergent application plugin state. In a real DCE incident, preserve exact binaries/image digests, plugin manifests/checksums, node startup logs, and the deployment event that created drift.

Repair: return app-2 to the exact cluster release and approved plugin artifact set. Do not “fix” the symptom by removing the app node from the load balancer forever or by downgrading the rest of the cluster.

4. Failure: exposing or blocking the wrong cluster port

Symptom Likely path Evidence Repair
LB reports app unhealthy LB→app 9000 / proxy health LB probe + web status/log Repair app health/proxy route; do not open 9003 publicly
App nodes cannot coordinate app↔app 9003/Hazelcast cluster logs + private firewall tests Restore private app-to-app reachability
Application cannot query/index search app→search 9001 web/CE/es logs + private connectivity Repair application-to-search path
Search cluster cannot form search↔search 9002 es.log + host list/version Repair private search transport/host inventory

5. Failure: unsuitable search storage

Search nodes are latency-sensitive and current Sonar guidance says SSDs perform significantly better than HDDs. If search heap is healthy but disk latency or capacity pressure is high, adding application nodes increases incoming work without repairing the bottleneck. Preserve es.log, disk latency/utilization, heap/GC, and search health first. Never edit Elasticsearch indexes directly or move live search data with generic file-copy commands as a repair.

6. Failure: shared database is the actual bottleneck

Every application node depends on the same supported database. Adding application nodes can increase database concurrency. If JDBC pool waits, DB CPU/IO/locks, or network latency are already saturated, horizontal app scaling can worsen the incident. Use database/vendor metrics and Sonar Web/CE evidence together. Current server support includes PostgreSQL 14–18, SQL Server 2017/2019/2022, and supported Oracle versions; verify the exact release matrix before changing engines or versions.

7. Failure: treating sticky sessions as the HA fix

Sticky sessions can hide routing symptoms but do not repair a broken application cluster, inconsistent JWT secret, mixed node version, or unhealthy backend. Current DCE explicitly does not require session affinity. Preserve the load-balancer backend/health evidence, then repair the owning cluster or network layer.

8. Failure: adding app nodes when search/database is saturated

Capacity management is causal. If CE queue/CPU is high and downstream tiers are healthy, app scaling may help. If es.log, search heap, or disk is the limiting layer, fix search. If JDBC/DB latency is limiting, fix the database path. Node count is not a generic “performance” knob.

9. Production shortcuts to reject

  • Do not expose 9001/9002/9003 to the public Internet.
  • Do not direct-edit Elasticsearch state or the SonarQube database.
  • Do not repeatedly restart the whole cluster before preserving first-failure evidence.
  • Do not run mixed server/plugin versions as a tolerated steady state.
  • Do not bypass TLS verification or place cluster/shared secrets in source.
  • Do not add nodes before identifying the bottleneck.
  • Do not treat node HA as database/region disaster recovery.
  • Do not change project keys, gates, or profiles to hide infrastructure failures.

Knowledge check

App nodes are healthy, but search nodes cannot reach each other. Which default path is suspect?

Why is adding an app node dangerous when the database is saturated?

What should you preserve before fixing a mixed-version node?

Should you copy Elasticsearch data directories to another search node as a first repair?

A sticky-session rule makes logins look stable. Does that prove the DCE cluster is healthy?

Next lesson

Produce the DCE architecture and failure checkpoint

Lesson 5 packages node roles, private trust boundaries, failure predictions, capacity limits, real ceTaskId evidence, and explicit commercial boundaries.

Official references and version notes

Version and compatibility note

Baseline: Community Build 26.9.0.129388; SonarScanner CLI 8.1.0.6389. Evidence continuity preserves ceTaskId as the Compute Engine task identity. Rechecked 2026-09-08. Mandatory executable examples remain Community Build 26.9.0.129388 with SonarScanner CLI 8.1.0.6389. DCE examples model the current SonarQube Server 2026 Release 4.1 / 2026.4.1 commercial architecture and require a valid Data Center Edition license for real execution. Current minimum/default DCE topology is two application plus three search nodes; one application and one search node may be lost without user impact when remaining dependencies are healthy. Application nodes can currently scale up to ten. Default private network paths are LB→app 9000, app→search 9001, search→search 9002 and app→app/Hazelcast 9003. Current supported PostgreSQL versions are 14–18. Recheck the exact target release, DCE license, Java/database matrix, Helm/chart version, plugin matrix, Kubernetes/OpenShift matrix, network properties, and update notes immediately before any real deployment.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.