Chapter 27Lesson 04~145 minutes

Server Logs, Monitoring, Health, Elasticsearch, and Operational Diagnostics: Diagnostics, Failure Modes, and Production Practices

Diagnose realistic web/CE/search/database/JVM/disk failures without deleting logs, repeatedly restarting, editing Elasticsearch directly, or mistaking UI availability for analysis health.

OperationsLogsCompute EngineMonitoringElasticsearch

Learning objectives

  • Apply a repeatable evidence-first sequence to scanner, CE, web, DB, search, JVM, disk, network, and policy incidents.
  • Reject destructive shortcuts that erase first-failure evidence or mutate SonarQube-owned state.
  • Diagnose the difference between UI availability and end-to-end analysis health.
  • Recognize disk-watermark, GC/memory, DB-pool, queue, and project-topology signatures.
  • Use a deliberately broken CE-task example without turning troubleshooting into a production attack.

1. Evidence-first diagnostic sequence

  1. Preserve first failure: scanner/CI log, report-task.txt, task JSON/error, relevant process logs, timestamps, and resource snapshots.
  2. Confirm versions: Community Build/Server edition, exact release, Java, database, scanner, plugins/integrations, container image.
  3. Confirm source and scanner context: Git SHA, base directory, indexed files, effective parameters.
  4. Confirm report/task transition: was a report uploaded? What is the ceTaskId? PENDING, IN_PROGRESS, SUCCESS, FAILED?
  5. Confirm project/policy state: gate/new-code/profile/security evidence only after analysis completion.
  6. Confirm shared dependencies: auth/network, DB pool/latency, search health/disk, JVM/GC, host/container resources.
  7. Apply the least destructive correction at the owning layer.
  8. Rerun the smallest equivalent scenario and independently verify recovery.

2. Failure: deleting or rotating logs before diagnosis

Deleting logs can erase the only copy of a startup stack trace, CE failure, disk-watermark transition, or DB timeout. Copy the incident window first. If storage pressure requires cleanup, move/compress already-preserved logs according to retention policy; do not destroy the root-cause evidence and then ask why the service failed.

3. Failure: leaving DEBUG/TRACE enabled indefinitely

DEBUG can log personal user information; TRACE adds SQL/Elasticsearch request detail and can materially increase volume and overhead. A useful runbook contains four fields: process, temporary level, start time, stop condition. “Turn on debug and see what happens” is not a controlled diagnostic action.

4. Failure: assuming UI availability means CE/search are healthy

The web server can serve cached/previous project state while the CE queue is delayed or a new task has failed. Search may be degraded while some pages still load. Always compare the newest analysis timestamp to the task submission/completion time and confirm the exact ceTaskId.

5. Failure: repeated restart as a queue strategy

Restarting can interrupt in-flight work, add a new startup event to the evidence stream, and leave the causal defect unchanged. A component-key clash, malformed scanner context, project permission issue, or report-specific failure will return after restart. Use task evidence first; restart only for a supported recovery reason.

6. Failure: touching Elasticsearch internals directly

Do not use Elasticsearch APIs, delete indices, or modify search settings directly to “unstick” SonarQube. Diagnose search via es.log, system health, host disk/IO, and documented SonarQube recovery. For disk-watermark read-only behavior, free disk and follow Sonar’s recovery procedure instead of overriding ES safety settings.

7. Failure: ignoring DB saturation

Errors such as HikariPool-1 - Connection is not available or Cannot acquire connection from data source point toward the DB connection path/pool. Correlate web/CE logs with database connection count, CPU, latency, firewall timeouts, and the current JDBC pool configuration. Do not edit SonarQube tables.

8. Failure: treating memory pressure as “just slow”

Monitor the Web, CE, and Elasticsearch JVMs independently. Sustained high heap, long GC pauses, or host swapping can manifest as request latency, queue delay, or search instability. The correct response is capacity/evidence-driven tuning—not one global heap change copied from a forum post.

9. Intentionally broken example: scanner succeeds, CE fails

Use the disposable key-clash fixture from Lesson 2. Preserve the contradiction rather than “fixing” it immediately:

scanner process: SUCCESS / report uploaded
report-task.txt: contains ceTaskId=AX...
api/ce/task: status=FAILED
ce.log: task AX... reports project/component ownership clash
api/system/health: GREEN
UI: last successful analysis still visible
Quality Gate: belongs to last completed analysis, not the failed task

This is exactly why operational state must remain multi-dimensional. The server is globally healthy enough to reject an invalid topology correctly. Restarting the server would only replay the same invalid ownership relationship.

10. Failure-layer matrix

First durable evidence Likely owner Smallest correction
No report-task.txt Scanner/build/config/auth/network before upload Fix scanner/build path and rerun.
ceTaskId FAILED; one project only Report/project processing Fix task-specific cause; keep server running if healthy.
Many CE tasks delayed + high CE CPU/heap CE capacity/workload Characterize task durations/resources before scaling.
Web + CE Hikari errors DB path/pool Repair DB/network/pool cause with supported config.
es.log watermark/read-only + disk 95% used Search/host disk Preserve, free space, documented recovery.
Health RED at startup + sonar.log child failure Server process startup Follow owning child log; do not scan yet.

11. Security-sensitive actions

High-impact controls. Log-level changes, system passcodes, database/JDBC settings, plugin changes, TLS/proxy changes, project deletion, host limits, container restarts, and Data Center operations affect trust or availability. Use fake/local credentials, least privilege, explicit change records, and rollback. Never disable TLS verification, expose system passcodes in command history/artifacts, or use administrator tokens in routine scanners.

12. Production operating practices

  • Centralize logs with timestamp normalization and retention; restrict DEBUG/TRACE access.
  • Alert on system health plus CE queue age/depth, host disk, JVM memory/GC, DB pool/latency, and search health.
  • Correlate every analysis incident with ceTaskId and revision.
  • Subscribe responsible project/admin users to background-task failure notifications where appropriate.
  • Keep API deprecation logs in upgrade readiness reviews.
  • Document restart prerequisites and reasons; count restarts as operational events, not fixes.
  • Never modify SonarQube database or embedded Elasticsearch state directly.

Knowledge check

Why is a GREEN health response compatible with one failed CE task?

What should happen before increasing log verbosity?

Which evidence suggests a shared DB problem rather than one project defect?

What is wrong with restarting SonarQube repeatedly when a key-clash CE task fails?

What is the supported operational stance toward embedded Elasticsearch?

Next lesson

Produce the operational-diagnostics checkpoint

Lesson 5 packages a successful task, failed task, recovery, timestamps, health, and process evidence into one auditable checkpoint.

Official references and version notes

Version and compatibility note

Rechecked 2026-09-08. Mandatory examples target SonarQube Community Build 26.9.0.129388 and SonarScanner CLI 8.1.0.6389; the Maven failed-task fixture uses SonarScanner for Maven 5.7.0.6970. Current commercial reference is SonarQube Server 2026 Release 4.1 with 2026.1.5 LTA. Current Community Build logs are sonar.log, web.log, ce.log, es.log, access and API-deprecation logs. api/system/health requires Administer System or system passcode; /api/monitoring/metrics exposes OpenMetrics under its monitoring-auth boundary. Community Build has Web, CE and embedded Elasticsearch JVMs; production databases must use supported PostgreSQL/SQL Server/Oracle rather than the evaluation H2 database. The project/component key-clash failure is documented as a CE failure class, but exact scanner versions may prevalidate the conflict before upload; the lessons explicitly preserve that layer distinction rather than forcing unsafe database/search failures.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.