Chapter 05Lesson 04~145 minutes

Kubernetes and OpenShift Deployment Patterns: Diagnostics, Failure Modes, and Production Practices

Diagnose Kubernetes/OpenShift SonarQube failures by preserving events, logs, rendered values, PVC/database/network/security evidence, and fixing the owning layer rather than restarting Pods or weakening policies.

DiagnosticsEventsPVC PendingOOMKilledRollback

Learning objectives

  • Use an evidence-first sequence that separates Helm/rendering, admission/security, scheduling/resources, storage, network/database, SonarQube runtime and edition/version failures.
  • Diagnose restricted-policy failures without granting broad privilege or creating custom SCC shortcuts.
  • Interpret Pending PVC, OOMKilled, readiness/liveness, DNS and JDBC symptoms causally.
  • Preserve failed release evidence before retries, upgrades, rollbacks or namespace deletion.
  • Repair one intentionally broken chart value with the smallest reversible change and prove the recovered state independently.

1. Evidence-first diagnostic sequence

  1. Freeze the first failure: chart version, effective values, rendered manifest, namespace, current context and time.
  2. Read helm status/helm history and Kubernetes events before restarting anything.
  3. Identify the failed lifecycle stage: render → admission → scheduling → init → application startup → readiness → Service/routing.
  4. Inspect PVC/StorageClass, Secret references, Service/DNS/network policy and database readiness as independent owners.
  5. Inspect SonarQube logs only after you know the Pod actually started.
  6. Confirm chart/app/edition/platform compatibility.
  7. Change one causal input, rerender, and reproduce the smallest scenario.

2. Preserve evidence before the cluster garbage-collects context

helm -n sq-ch05-lab status sq-ch05 > evidence-helm-status.txt 2>&1 || true
helm -n sq-ch05-lab get values sq-ch05 --all > evidence-values.yaml 2>&1 || true
helm -n sq-ch05-lab get manifest sq-ch05 > evidence-manifest.yaml 2>&1 || true
helm -n sq-ch05-lab history sq-ch05 > evidence-history.txt 2>&1 || true
kubectl -n sq-ch05-lab get events --sort-by=.lastTimestamp > evidence-events.txt
kubectl -n sq-ch05-lab get pod,pvc,svc -o wide > evidence-inventory.txt
kubectl -n sq-ch05-lab get pod -o yaml > evidence-pods.yaml
kubectl -n sq-ch05-lab logs -l app=sonarqube --all-containers=true --tail=500 > evidence-logs.txt 2>&1 || true

Sanitize Secrets and private endpoints before sharing. Preserve raw sensitive evidence only in an appropriately protected location.

3. Failure mode: restricted policy blocks root helper containers

Symptom: admission rejects init-sysctl or init-fs. The wrong response is to grant a privileged SCC/namespace exception reflexively. First confirm whether you intentionally chose a full restricted namespace.

Repair: if the platform team already configures node sysctls and storage ownership correctly, disable initSysctl.enabled and initFs.enabled, rerender, and verify the application containers remain restricted. On OpenShift, use the supported OpenShift.enabled=true path rather than fixed UID/GID or custom SCC hacks.

4. Failure mode: PVC stays Pending

A Pending claim is a storage provisioning problem until evidence says otherwise. Inspect:

kubectl -n sq-ch05-lab describe pvc
kubectl get storageclass
kubectl -n sq-ch05-lab get events --sort-by=.lastTimestamp

Check whether a default StorageClass exists, whether requested access mode/size is supported, whether the provisioner is healthy, and whether topology constraints can satisfy the claim. Restarting the SonarQube Pod cannot provision missing storage.

5. Failure mode: database endpoint or Secret is wrong

Symptom: Pod starts but SonarQube logs show JDBC connection/authentication/DNS errors. Preserve logs, then inspect the effective JDBC URL, Secret reference names/keys, Service/DNS/network policy and external DB readiness. Do not print the Secret value into a troubleshooting ticket.

An intentionally broken render can use jdbc:postgresql://localhost:5432/sonar. In Kubernetes, localhost means the SonarQube Pod itself, not a sibling/external database. Repair only the endpoint to the correct Service/FQDN after the DB is independently verified.

6. Failure mode: OOMKilled or repeated probe failures

OOMKilled means the container exceeded its memory cgroup limit; it is not proof of a SonarQube defect. Compare container termination reason, memory request/limit, node pressure, JVM settings and workload. A too-low request can also cause poor placement/eviction risk.

If readiness/liveness fails, inspect application logs and dependency latency before relaxing probes. Readiness failure may correctly protect traffic from an unready application. Repeated liveness restarts can erase the timing context you need, so preserve logs/events first.

7. Failure mode: public port 9000 or edge misconfiguration

Changing a Service to NodePort/LoadBalancer or exposing port 9000 directly can bypass intended TLS, authentication proxy, IP policy and observability. Keep the Service internal and configure a reviewed Route/Ingress/Gateway path with certificate ownership. Do not fix certificate errors by disabling verification.

For local diagnostics, use kubectl port-forward rather than public exposure.

8. Failure mode: “Pod restarted, therefore recovered”

A new Pod UID only proves Kubernetes created a new runtime object. Verify the external database is the same intended database, PVC is correctly attached, chart values did not drift, search state is healthy, and the server reaches a stable readiness state. A restart can temporarily hide memory pressure or race conditions without fixing them.

9. Failure mode: obsolete chart mixed with newer server assumptions

The chart and SonarQube application are versioned separately. Current Community Build is 26.9.0.129388; this chapter pins released chart 2026.4.1 and overrides its Community Build number explicitly. If you copy values from a future/master chart into an older released chart, keys may be missing or semantics may differ.

Before an upgrade, compare chart changelog, supported Kubernetes/OpenShift matrix, SonarQube upgrade path, plugins, database version and deprecations. Preserve the previous values and database backup.

10. Intentionally broken example: wrong JDBC host

Render the following production-shaped fragment only; do not deploy it against a real DB:

jdbcOverwrite:
  enabled: true
  jdbcUrl: "jdbc:postgresql://localhost:5432/sonar"
  jdbcUsername: "sonar_lab"
  jdbcSecretName: "sq-db-credentials"
  jdbcSecretPasswordKey: "password"

Diagnosis: if SonarQube and PostgreSQL are separate Pods/services, localhost cannot reach the database. The preserved evidence should show the wrong effective JDBC URL while the DB can be healthy independently. Repair the hostname to the database Service/FQDN; do not restart the DB, delete PVCs, change project keys, or lower SonarQube policies.

11. Production operating practices

  • Pin chart/application versions and review rendered diffs.
  • Use external supported DB with tested backup/restore.
  • Run application containers under restricted security policy and manage node prerequisites centrally.
  • Use approved StorageClasses and understand reclaim/snapshot behavior.
  • Keep public routing behind managed TLS/edge infrastructure.
  • Set resources from measured workload and monitor termination/events.
  • Retain Helm values/history, Pod events/logs and upgrade evidence.
  • Test recovery separately from “Pod restart.”

Knowledge check

A Pod is rejected because init-sysctl is privileged. What should you do first?

A PVC is Pending. Which SonarQube setting should you lower first?

Why is localhost usually wrong for an external DB from a Pod?

What does OOMKilled prove?

Why preserve Helm values and manifest before rollback?

Next lesson

From diagnosis to a governed checkpoint

Lesson 5 assembles a namespace-scoped architecture dossier, renders both Kubernetes and OpenShift variants, validates persistence/security/DB/routing assumptions, injects one reversible fault, and records rollback evidence.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current SonarSource, Helm-chart repository, and Kubernetes primary material on 2026-09-07. Mandatory examples use SonarQube Community Build 26.9.0.129388 with the latest released SonarQube Helm chart 2026.4.1. That chart release originally defaults Community Build to 26.7.0.124771, so examples deliberately set community.buildNumber=26.9.0.129388. The released chart documents non-OpenShift Kubernetes support for 1.32–1.35 and OpenShift support for 4.17–4.20. SonarQube Server Developer/Enterprise and the separate Data Center chart are commercial boundaries and are not required. No SonarScanner execution, CI provider, third-party plugin, enterprise identity provider, managed Kubernetes service, public DNS, or paid infrastructure is required by the mandatory Chapter 05 path. Re-check the chart release, supported platform matrix, Community Build number, database compatibility, and upgrade notes before using these examples later.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.