Server Logs, Monitoring, Health, Elasticsearch, and Operational Diagnostics: Diagnostics, Failure Modes, and Production Practices
Diagnose realistic web/CE/search/database/JVM/disk failures without deleting logs, repeatedly restarting, editing Elasticsearch directly, or mistaking UI availability for analysis health.
Learning objectives
- Apply a repeatable evidence-first sequence to scanner, CE, web, DB, search, JVM, disk, network, and policy incidents.
- Reject destructive shortcuts that erase first-failure evidence or mutate SonarQube-owned state.
- Diagnose the difference between UI availability and end-to-end analysis health.
- Recognize disk-watermark, GC/memory, DB-pool, queue, and project-topology signatures.
- Use a deliberately broken CE-task example without turning troubleshooting into a production attack.
1. Evidence-first diagnostic sequence
-
Preserve first failure: scanner/CI log,
report-task.txt, task JSON/error, relevant process logs, timestamps, and resource snapshots. - Confirm versions: Community Build/Server edition, exact release, Java, database, scanner, plugins/integrations, container image.
- Confirm source and scanner context: Git SHA, base directory, indexed files, effective parameters.
-
Confirm report/task transition: was a report
uploaded? What is the
ceTaskId? PENDING, IN_PROGRESS, SUCCESS, FAILED? - Confirm project/policy state: gate/new-code/profile/security evidence only after analysis completion.
- Confirm shared dependencies: auth/network, DB pool/latency, search health/disk, JVM/GC, host/container resources.
- Apply the least destructive correction at the owning layer.
- Rerun the smallest equivalent scenario and independently verify recovery.
2. Failure: deleting or rotating logs before diagnosis
Deleting logs can erase the only copy of a startup stack trace, CE failure, disk-watermark transition, or DB timeout. Copy the incident window first. If storage pressure requires cleanup, move/compress already-preserved logs according to retention policy; do not destroy the root-cause evidence and then ask why the service failed.
3. Failure: leaving DEBUG/TRACE enabled indefinitely
DEBUG can log personal user information; TRACE adds SQL/Elasticsearch request detail and can materially increase volume and overhead. A useful runbook contains four fields: process, temporary level, start time, stop condition. “Turn on debug and see what happens” is not a controlled diagnostic action.
4. Failure: assuming UI availability means CE/search are healthy
The web server can serve cached/previous project state while the CE
queue is delayed or a new task has failed. Search may be degraded
while some pages still load. Always compare the newest analysis
timestamp to the task submission/completion time and confirm the
exact ceTaskId.
5. Failure: repeated restart as a queue strategy
Restarting can interrupt in-flight work, add a new startup event to the evidence stream, and leave the causal defect unchanged. A component-key clash, malformed scanner context, project permission issue, or report-specific failure will return after restart. Use task evidence first; restart only for a supported recovery reason.
6. Failure: touching Elasticsearch internals directly
Do not use Elasticsearch APIs, delete indices, or modify search
settings directly to “unstick” SonarQube. Diagnose search via
es.log, system health, host disk/IO, and documented
SonarQube recovery. For disk-watermark read-only behavior, free disk
and follow Sonar’s recovery procedure instead of overriding ES
safety settings.
7. Failure: ignoring DB saturation
Errors such as
HikariPool-1 - Connection is not available or
Cannot acquire connection from data source point toward
the DB connection path/pool. Correlate web/CE logs with database
connection count, CPU, latency, firewall timeouts, and the current
JDBC pool configuration. Do not edit SonarQube tables.
8. Failure: treating memory pressure as “just slow”
Monitor the Web, CE, and Elasticsearch JVMs independently. Sustained high heap, long GC pauses, or host swapping can manifest as request latency, queue delay, or search instability. The correct response is capacity/evidence-driven tuning—not one global heap change copied from a forum post.
9. Intentionally broken example: scanner succeeds, CE fails
Use the disposable key-clash fixture from Lesson 2. Preserve the contradiction rather than “fixing” it immediately:
scanner process: SUCCESS / report uploaded
report-task.txt: contains ceTaskId=AX...
api/ce/task: status=FAILED
ce.log: task AX... reports project/component ownership clash
api/system/health: GREEN
UI: last successful analysis still visible
Quality Gate: belongs to last completed analysis, not the failed task
This is exactly why operational state must remain multi-dimensional. The server is globally healthy enough to reject an invalid topology correctly. Restarting the server would only replay the same invalid ownership relationship.
10. Failure-layer matrix
| First durable evidence | Likely owner | Smallest correction |
|---|---|---|
No report-task.txt |
Scanner/build/config/auth/network before upload | Fix scanner/build path and rerun. |
ceTaskId FAILED; one project only |
Report/project processing | Fix task-specific cause; keep server running if healthy. |
| Many CE tasks delayed + high CE CPU/heap | CE capacity/workload | Characterize task durations/resources before scaling. |
| Web + CE Hikari errors | DB path/pool | Repair DB/network/pool cause with supported config. |
es.log watermark/read-only + disk 95% used
|
Search/host disk | Preserve, free space, documented recovery. |
Health RED at startup + sonar.log child failure
|
Server process startup | Follow owning child log; do not scan yet. |
11. Security-sensitive actions
12. Production operating practices
- Centralize logs with timestamp normalization and retention; restrict DEBUG/TRACE access.
- Alert on system health plus CE queue age/depth, host disk, JVM memory/GC, DB pool/latency, and search health.
-
Correlate every analysis incident with
ceTaskIdand revision. - Subscribe responsible project/admin users to background-task failure notifications where appropriate.
- Keep API deprecation logs in upgrade readiness reviews.
- Document restart prerequisites and reasons; count restarts as operational events, not fixes.
- Never modify SonarQube database or embedded Elasticsearch state directly.
Knowledge check
Why is a GREEN health response compatible with one failed CE task?
Global application health and individual report validity are different states. A healthy CE can correctly reject one invalid report.
What should happen before increasing log verbosity?
Preserve the first incident evidence and identify the process to target, then set a bounded diagnostic window.
Which evidence suggests a shared DB problem rather than one project defect?
Web/CE Hikari/connection errors across unrelated requests/tasks correlated with DB-side resource/latency/connection abnormalities.
What is wrong with restarting SonarQube repeatedly when a key-clash CE task fails?
The topology conflict remains, first-failure evidence becomes noisier, and availability is disrupted without addressing the owner.
What is the supported operational stance toward embedded Elasticsearch?
Observe through SonarQube logs/health/metrics and host resources, and use SonarSource recovery procedures; do not mutate ES internals directly.
Official references and version notes
- Community Build — Server logs — current process-specific files, log levels, rotation, JSON output, DEBUG/TRACE privacy/performance cautions.
-
Community Build — Monitoring the instance
—
api/system/health, Web/CE/Elasticsearch JVMs, CPU/RAM/disk monitoring. -
Community Build — Monitoring on Kubernetes
—
/api/monitoring/metrics, OpenMetrics, and CE/Web JMX exporter model. - Community Build — Web API — bearer authentication, system-passcode use for monitoring, and continuing Web API V2 transition.
-
Current Web API reference —
api/system/health— GREEN/YELLOW/RED semantics and Administer System/system-passcode requirement. - Current Web API reference — Compute Engine — CE activity/task APIs and permission boundaries.
- Background tasks — scanner success vs CE completion, task failure diagnosis, and project/module key-clash example.
- Community Build — Elasticsearch-related issues — disk-watermark/read-only behavior and supported recovery.
- Community Build — Database-related issues — HikariCP timeout/exhaustion evidence and supported connection-pool tuning.
- Community Build — Installing database — supported production DB engines/versions and evaluation-H2 boundary.
- SonarScanner CLI 8.1.0.6389 — current CLI baseline.
- SonarScanner for Maven 5.7.0.6970 — current Maven scanner baseline used by the optional key-clash failure fixture.
Rechecked 2026-09-08. Mandatory examples target SonarQube
Community Build 26.9.0.129388 and SonarScanner
CLI 8.1.0.6389; the Maven failed-task fixture
uses SonarScanner for Maven 5.7.0.6970. Current
commercial reference is SonarQube Server 2026 Release 4.1 with
2026.1.5 LTA. Current Community Build logs are
sonar.log, web.log, ce.log,
es.log, access and API-deprecation logs.
api/system/health requires Administer System or
system passcode; /api/monitoring/metrics exposes
OpenMetrics under its monitoring-auth boundary. Community Build
has Web, CE and embedded Elasticsearch JVMs; production databases
must use supported PostgreSQL/SQL Server/Oracle rather than the
evaluation H2 database. The project/component key-clash failure is
documented as a CE failure class, but exact scanner versions may
prevalidate the conflict before upload; the lessons explicitly
preserve that layer distinction rather than forcing unsafe
database/search failures.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.