Chapter 27Lesson 03~135 minutes

Server Logs, Monitoring, Health, Elasticsearch, and Operational Diagnostics: Configuration, Design Patterns, and Trade-Offs

Choose intentionally among health checks, queue measures, logs, metrics, temporary debug logging, task remediation, and process restarts while keeping project-specific and infrastructure failures separate.

OperationsLogsCompute EngineMonitoringElasticsearch

Learning objectives

  • Choose the right evidence surface for application health, queue health, request latency, task failure, search pressure, and database saturation.
  • Use logs for event causality and metrics for trends/capacity rather than treating either as complete observability.
  • Escalate logging from INFO to process-specific DEBUG/TRACE only for a bounded reproduction window.
  • Distinguish task remediation from process restart and project-specific failures from infrastructure failures.
  • Design monitoring that remains Community Build/local-compatible while recognizing optional Kubernetes/Prometheus and Data Center capabilities.

1. Design principle: monitor states that fail independently

A useful SonarQube dashboard should not collapse the service into one “up/down” light. Monitor at least four independent planes: request serving, background analysis processing, persistence/search dependencies, and host/JVM capacity. Then link every alert to the first evidence an operator should inspect.

symptom → owning state → minimal evidence → threshold/trend → least-destructive action → independent verification

2. Application health versus analysis queue health

api/system/health is a coarse global signal: GREEN means operational, YELLOW means usable but needs attention, RED means not operational. It does not tell you that every queued analysis is succeeding or that queue latency is acceptable. Queue state belongs to CE activity and CE metrics.

Question Primary evidence Secondary evidence
Can the service answer health requests? api/system/health sonar.log, web.log
Are analyses waiting too long? CE activity/queue metrics ce.log, CPU/heap
Why did one report fail? ceTaskId task detail/scanner context ce.log, then DB/search logs if referenced
Is search under pressure? es.log, disk/IO/health system health, host metrics
Is DB pool exhausted? Hikari/DB errors in web/CE logs, DB metrics pool settings, firewall/latency evidence

3. Logs versus metrics

Logs preserve discrete events with identities: task X failed at 12:04:11 because of Y. Metrics compress repeated observations into a time series: CE pending count rose from 2 to 80 over 20 minutes while heap and CPU changed. Both are necessary for reliable incident analysis.

Logs

Best for stack traces, exact task IDs, request failures, startup transitions, database/search error messages, and change chronology.

Metrics

Best for alert thresholds, queue age, throughput, memory/GC, CPU, disk, connection counts, and capacity trends.

System info

Best for immutable/slow-changing context: exact version, JVM, plugin set, DB engine, configuration facts.

Project/task APIs

Best for connecting infrastructure observations back to one analysis revision and policy outcome.

4. DEBUG/TRACE versus normal logging

Use INFO continuously and collect it centrally with retention appropriate to your environment. Increase verbosity only after the first incident window has been preserved and only for the process implicated by evidence. Current properties allow process-specific levels: sonar.log.level.app, .web, .ce, and .es.

Level When Risk Exit criterion
INFO Normal operation Lowest detail; may not explain rare performance issue Default
DEBUG Short reproduction when INFO is insufficient More personal/request information and volume Immediately after evidence is captured
TRACE Narrow web/request-performance diagnosis SQL/Elasticsearch request detail, performance overhead, very large logs Minutes, not days; revert to INFO

5. Process restart versus task remediation

A restart is appropriate only when the owning subsystem’s documented recovery requires it—for example, after resolving the non-DCE Elasticsearch disk watermark/read-only condition. It is not a universal task-retry mechanism.

Failure Correct first action Restart?
Project/component key clash Repair topology after preserving task evidence No
Invalid scanner/report parameter Correct scanner/build config and rescan No
DB connection pool timeout Prove DB/firewall/pool cause; change supported DB/pool config if needed Only if configuration/application restart requires it
ES indices read-only after 95% disk use Free disk and follow documented recovery Yes for non-DCE, per current guidance
One pending task that is no longer wanted Cancel only while pending, with proper permission No

6. Infrastructure failure versus project-specific failure

Project-specific failures tend to cluster around one ceTaskId, one scanner context, one project key, or one report shape. Infrastructure failures recur across unrelated projects and correlate with host/DB/search/JVM signals. Do not infer either class from a single error message—compare neighboring tasks and time windows.

incident triage questions
1. Did unrelated projects fail in the same minute?
2. Did system health or search health change?
3. Did CE queue age grow globally?
4. Is the error tied to one project key/report context?
5. Are DB pool / disk / GC / host metrics abnormal?
6. What changed immediately before the first failure?

7. Database operational evidence

Production SonarQube uses a supported external database. Current Community Build supports PostgreSQL 14–18, SQL Server, and Oracle under documented requirements. The application uses a JDBC/HikariCP connection pool; timeouts or exhausted connections appear in web/CE logs and must be correlated with DB-side connections/CPU/latency and firewall behavior.

Do not tune sonar.jdbc.maxActive, keepalive, lifetime, or timeout values because “the UI feels slow.” Capture pool errors and DB metrics first. Directly editing SonarQube tables is outside supported incident repair.

9. Worked decision table

Observed evidence Diagnosis Next action
UI 200 OK, CE pending count rising, CE CPU saturated Analysis capacity/throughput issue Inspect task durations and CE resources before scaling/tuning.
UI 200 OK, one CE task FAILED, neighbors SUCCESS Likely project/report-specific Inspect that task/scanner context first.
Web/CE both show Hikari connection errors Shared DB path/pool problem Check DB health, latency, firewall and pool evidence.
es.log reports flood/read-only after high disk usage Search disk safety mechanism Preserve evidence, free disk, follow supported restart recovery.
Only scanner fails before report-task.txt Scanner/build layer Do not restart the server; fix scanner/build inputs.

10. Preserve analysis evidence across operational decisions

Even in an infrastructure incident, keep the course’s established chain: scanner process result → report upload → report-task.txt/ceTaskId → CE terminal status → analysis result → Quality Gate → CI job → provider decoration. A restart or metric alert never substitutes for this project-level provenance.

Knowledge check

Which is more useful for proving queue growth over an hour: logs or metrics?

When is TRACE logging acceptable?

A project/component key clash is fixed by what class of action?

Why compare neighboring projects/tasks during an incident?

Can you change Elasticsearch indices directly to clear a SonarQube alert?

Next lesson

Diagnose failures without destroying evidence

Lesson 4 engineers realistic log, search, DB, disk, GC, queue, and task failure modes using an evidence-first diagnostic sequence.

Official references and version notes

Version and compatibility note

Rechecked 2026-09-08. Mandatory examples target SonarQube Community Build 26.9.0.129388 and SonarScanner CLI 8.1.0.6389; the Maven failed-task fixture uses SonarScanner for Maven 5.7.0.6970. Current commercial reference is SonarQube Server 2026 Release 4.1 with 2026.1.5 LTA. Current Community Build logs are sonar.log, web.log, ce.log, es.log, access and API-deprecation logs. api/system/health requires Administer System or system passcode; /api/monitoring/metrics exposes OpenMetrics under its monitoring-auth boundary. Community Build has Web, CE and embedded Elasticsearch JVMs; production databases must use supported PostgreSQL/SQL Server/Oracle rather than the evaluation H2 database. The project/component key-clash failure is documented as a CE failure class, but exact scanner versions may prevalidate the conflict before upload; the lessons explicitly preserve that layer distinction rather than forcing unsafe database/search failures.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.