Server Logs, Monitoring, Health, Elasticsearch, and Operational Diagnostics: Configuration, Design Patterns, and Trade-Offs
Choose intentionally among health checks, queue measures, logs, metrics, temporary debug logging, task remediation, and process restarts while keeping project-specific and infrastructure failures separate.
Learning objectives
- Choose the right evidence surface for application health, queue health, request latency, task failure, search pressure, and database saturation.
- Use logs for event causality and metrics for trends/capacity rather than treating either as complete observability.
- Escalate logging from INFO to process-specific DEBUG/TRACE only for a bounded reproduction window.
- Distinguish task remediation from process restart and project-specific failures from infrastructure failures.
- Design monitoring that remains Community Build/local-compatible while recognizing optional Kubernetes/Prometheus and Data Center capabilities.
1. Design principle: monitor states that fail independently
A useful SonarQube dashboard should not collapse the service into one “up/down” light. Monitor at least four independent planes: request serving, background analysis processing, persistence/search dependencies, and host/JVM capacity. Then link every alert to the first evidence an operator should inspect.
symptom → owning state → minimal evidence → threshold/trend →
least-destructive action → independent verification
2. Application health versus analysis queue health
api/system/health is a coarse global signal: GREEN
means operational, YELLOW means usable but needs attention, RED
means not operational. It does not tell you that every queued
analysis is succeeding or that queue latency is acceptable. Queue
state belongs to CE activity and CE metrics.
| Question | Primary evidence | Secondary evidence |
|---|---|---|
| Can the service answer health requests? | api/system/health |
sonar.log, web.log |
| Are analyses waiting too long? | CE activity/queue metrics | ce.log, CPU/heap |
| Why did one report fail? | ceTaskId task detail/scanner context |
ce.log, then DB/search logs if referenced
|
| Is search under pressure? | es.log, disk/IO/health |
system health, host metrics |
| Is DB pool exhausted? | Hikari/DB errors in web/CE logs, DB metrics | pool settings, firewall/latency evidence |
3. Logs versus metrics
Logs preserve discrete events with identities: task X failed at 12:04:11 because of Y. Metrics compress repeated observations into a time series: CE pending count rose from 2 to 80 over 20 minutes while heap and CPU changed. Both are necessary for reliable incident analysis.
Logs
Best for stack traces, exact task IDs, request failures, startup transitions, database/search error messages, and change chronology.
Metrics
Best for alert thresholds, queue age, throughput, memory/GC, CPU, disk, connection counts, and capacity trends.
System info
Best for immutable/slow-changing context: exact version, JVM, plugin set, DB engine, configuration facts.
Project/task APIs
Best for connecting infrastructure observations back to one analysis revision and policy outcome.
4. DEBUG/TRACE versus normal logging
Use INFO continuously and collect it centrally with retention
appropriate to your environment. Increase verbosity only after the
first incident window has been preserved and only for the process
implicated by evidence. Current properties allow process-specific
levels: sonar.log.level.app, .web,
.ce, and .es.
| Level | When | Risk | Exit criterion |
|---|---|---|---|
| INFO | Normal operation | Lowest detail; may not explain rare performance issue | Default |
| DEBUG | Short reproduction when INFO is insufficient | More personal/request information and volume | Immediately after evidence is captured |
| TRACE | Narrow web/request-performance diagnosis | SQL/Elasticsearch request detail, performance overhead, very large logs | Minutes, not days; revert to INFO |
5. Process restart versus task remediation
A restart is appropriate only when the owning subsystem’s documented recovery requires it—for example, after resolving the non-DCE Elasticsearch disk watermark/read-only condition. It is not a universal task-retry mechanism.
| Failure | Correct first action | Restart? |
|---|---|---|
| Project/component key clash | Repair topology after preserving task evidence | No |
| Invalid scanner/report parameter | Correct scanner/build config and rescan | No |
| DB connection pool timeout | Prove DB/firewall/pool cause; change supported DB/pool config if needed | Only if configuration/application restart requires it |
| ES indices read-only after 95% disk use | Free disk and follow documented recovery | Yes for non-DCE, per current guidance |
| One pending task that is no longer wanted | Cancel only while pending, with proper permission | No |
6. Infrastructure failure versus project-specific failure
Project-specific failures tend to cluster around one
ceTaskId, one scanner context, one project key, or one
report shape. Infrastructure failures recur across unrelated
projects and correlate with host/DB/search/JVM signals. Do not infer
either class from a single error message—compare neighboring tasks
and time windows.
incident triage questions
1. Did unrelated projects fail in the same minute?
2. Did system health or search health change?
3. Did CE queue age grow globally?
4. Is the error tied to one project key/report context?
5. Are DB pool / disk / GC / host metrics abnormal?
6. What changed immediately before the first failure?
7. Database operational evidence
Production SonarQube uses a supported external database. Current Community Build supports PostgreSQL 14–18, SQL Server, and Oracle under documented requirements. The application uses a JDBC/HikariCP connection pool; timeouts or exhausted connections appear in web/CE logs and must be correlated with DB-side connections/CPU/latency and firewall behavior.
Do not tune sonar.jdbc.maxActive, keepalive, lifetime,
or timeout values because “the UI feels slow.” Capture pool errors
and DB metrics first. Directly editing SonarQube tables is outside
supported incident repair.
8. Search operational evidence
Monitor disk space and IO latency continuously. Search degradation
can surface in web and CE because both depend on indexed state.
Current SonarQube guidance documents ES warnings at disk pressure
and a read-only safety action around the 95% used watermark.
Preserve es.log before freeing space so the root cause
remains auditable.
9. Worked decision table
| Observed evidence | Diagnosis | Next action |
|---|---|---|
| UI 200 OK, CE pending count rising, CE CPU saturated | Analysis capacity/throughput issue | Inspect task durations and CE resources before scaling/tuning. |
| UI 200 OK, one CE task FAILED, neighbors SUCCESS | Likely project/report-specific | Inspect that task/scanner context first. |
| Web/CE both show Hikari connection errors | Shared DB path/pool problem | Check DB health, latency, firewall and pool evidence. |
es.log reports flood/read-only after high disk
usage
|
Search disk safety mechanism | Preserve evidence, free disk, follow supported restart recovery. |
Only scanner fails before report-task.txt
|
Scanner/build layer | Do not restart the server; fix scanner/build inputs. |
10. Preserve analysis evidence across operational decisions
Even in an infrastructure incident, keep the course’s established
chain: scanner process result → report upload →
report-task.txt/ceTaskId → CE terminal
status → analysis result → Quality Gate → CI job → provider
decoration. A restart or metric alert never substitutes for this
project-level provenance.
Knowledge check
Which is more useful for proving queue growth over an hour: logs or metrics?
Metrics. Logs then explain the specific events/tasks behind the trend.
When is TRACE logging acceptable?
For a narrow, time-bounded diagnostic reproduction after preserving existing evidence, with restricted access and immediate reversion to INFO.
A project/component key clash is fixed by what class of action?
Project/report topology remediation, not a server restart or infrastructure tuning.
Why compare neighboring projects/tasks during an incident?
To distinguish one project/report failure from an infrastructure-wide dependency/resource problem.
Can you change Elasticsearch indices directly to clear a SonarQube alert?
No. Use SonarQube-supported diagnosis and recovery; direct search-state edits bypass application consistency.
Official references and version notes
- Community Build — Server logs — current process-specific files, log levels, rotation, JSON output, DEBUG/TRACE privacy/performance cautions.
-
Community Build — Monitoring the instance
—
api/system/health, Web/CE/Elasticsearch JVMs, CPU/RAM/disk monitoring. -
Community Build — Monitoring on Kubernetes
—
/api/monitoring/metrics, OpenMetrics, and CE/Web JMX exporter model. - Community Build — Web API — bearer authentication, system-passcode use for monitoring, and continuing Web API V2 transition.
-
Current Web API reference —
api/system/health— GREEN/YELLOW/RED semantics and Administer System/system-passcode requirement. - Current Web API reference — Compute Engine — CE activity/task APIs and permission boundaries.
- Background tasks — scanner success vs CE completion, task failure diagnosis, and project/module key-clash example.
- Community Build — Elasticsearch-related issues — disk-watermark/read-only behavior and supported recovery.
- Community Build — Database-related issues — HikariCP timeout/exhaustion evidence and supported connection-pool tuning.
- Community Build — Installing database — supported production DB engines/versions and evaluation-H2 boundary.
- SonarScanner CLI 8.1.0.6389 — current CLI baseline.
- SonarScanner for Maven 5.7.0.6970 — current Maven scanner baseline used by the optional key-clash failure fixture.
Rechecked 2026-09-08. Mandatory examples target SonarQube
Community Build 26.9.0.129388 and SonarScanner
CLI 8.1.0.6389; the Maven failed-task fixture
uses SonarScanner for Maven 5.7.0.6970. Current
commercial reference is SonarQube Server 2026 Release 4.1 with
2026.1.5 LTA. Current Community Build logs are
sonar.log, web.log, ce.log,
es.log, access and API-deprecation logs.
api/system/health requires Administer System or
system passcode; /api/monitoring/metrics exposes
OpenMetrics under its monitoring-auth boundary. Community Build
has Web, CE and embedded Elasticsearch JVMs; production databases
must use supported PostgreSQL/SQL Server/Oracle rather than the
evaluation H2 database. The project/component key-clash failure is
documented as a CE failure class, but exact scanner versions may
prevalidate the conflict before upload; the lessons explicitly
preserve that layer distinction rather than forcing unsafe
database/search failures.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.