Chapter 28Lesson 01~125 minutes

Database Backup, Restore, Disaster Recovery, and Data Protection: Core Concepts and Mental Model

Treat the database as the primary durable SonarQube state, separate reconstructable search indexes from configuration and secrets, and define recovery through versioned, testable evidence.

Backup & RestorePostgreSQLDisaster RecoveryRPO / RTOData Protection

Learning objectives

  • Identify the SonarQube state that must survive a disaster and separate it from reconstructable search/cache/runtime state.
  • Explain database consistency, backup point, RPO, RTO, restore target, and rollback boundary before executing a recovery.
  • Build the evidence chain from server/database/configuration inventory to protected backup, isolated restore, reindex, and validation.
  • Explain why Docker volumes and Elasticsearch directories are persistence mechanisms but not, by themselves, a complete supported disaster-recovery backup.
  • Record configuration, plugin, secret, and external-system dependencies that are not reconstructed merely by restoring database bytes.
  • Keep recovery and upgrade as separate changes unless a documented update/migration plan intentionally combines them.

1. The practical problem: persistence is not the same as recoverability

Chapter 27 taught you to operate SonarQube by following evidence across the scanner, web process, Compute Engine, database, and embedded Elasticsearch. Chapter 28 asks a different question: which of those states must be protected so the service can be reconstructed after loss?

A mounted Docker volume can survive a container replacement. An Elasticsearch data directory can survive a process restart. Neither statement proves that you have a recoverable SonarQube backup. A disaster-recovery design must identify the authoritative durable state, record the exact server/database/configuration assumptions, protect the backup outside the failure domain, and repeatedly prove that the backup can start an isolated replacement service.

Production invariant. Never call a copied sonarqube_data volume, container filesystem, or Elasticsearch directory a complete backup unless your current supported procedure explicitly defines it that way. Current SonarQube backup guidance centers recovery on the database vendor's backup tooling.

2. Mental model: durable state → protected backup → isolated reconstruction

Start with a known server release, database engine/version, system configuration, plugin inventory, and project state. The database contains the primary durable application data. Startup properties and secret material may live outside the database. Elasticsearch indexes are derived from database-backed application state and are rebuilt after restore. A valid recovery therefore combines these inputs deliberately instead of copying an installation directory wholesale.

Supported recovery causality
flowchart TD
  A[Server/version + config inventory] --> B[Primary SonarQube database]
  A --> C[Plugin + secret references]
  B --> D[Vendor-supported protected backup]
  C --> E[Protected configuration inventory]
  D --> F[Isolated compatible database restore]
  E --> G[Fresh compatible SonarQube instance]
  F --> G
  G --> H[Fresh / cleared Elasticsearch indexes]
  H --> I[Automatic reindexing]
  I --> J[UI/API + project history validation]
  J --> K[Fresh scanner analysis + ceTaskId]
  K --> L[Recovery accepted or rollback retained]

Every arrow is an ownership boundary. Database backup is owned by the database platform. SonarQube owns its schema and the logic that rebuilds search indexes. Configuration may be owned by deployment code, a secret manager, or a host file. The recovery operator owns the evidence that proves the pieces belong to the same intended recovery point.

3. State map: what is authoritative, reconstructable, or external?

State Owner Recovery treatment Common mistake
Projects, issues, measures, profiles, gates, users, tokens, analysis history SonarQube application database Protect with supported database backup; validate after restore. Assuming an Elasticsearch copy is authoritative.
Elasticsearch indexes under data/es8 SonarQube embedded search subsystem Rebuild from restored application state; current restore guidance clears indexes before restart. Restoring stale indexes as the primary data source.
Startup/system properties Deployment configuration Inventory from environment, Helm, command line, or sonar.properties; restore through reviewed configuration. Assuming every server setting is stored in the database.
Third-party plugin artifacts Extension supply chain Record exact plugin version/source/checksum/license; reinstall a compatible artifact. Blindly copying an old extensions directory into a different release.
Encryption key / external secret references Secret-management boundary Protect separately with least privilege and tested retrieval. Backing up encrypted values but losing the key, or storing key and backup together without controls.
SCM, CI artifacts, IdP, external database service, reverse proxy External systems Record integration dependencies and recovery ordering; protect in their own systems. Calling the SonarQube database backup a backup of GitHub, CI, LDAP, or PostgreSQL infrastructure.

4. RPO and RTO turn “we have backups” into measurable objectives

Recovery Point Objective (RPO) is the acceptable amount of data loss measured in time. If you back up nightly and the database fails at 16:00, a worst-case daily schedule might lose nearly a day of SonarQube state. Recovery Time Objective (RTO) is the acceptable time to restore service to a validated operating state.

Neither number comes from SonarQube automatically. They are governance choices that drive database backup frequency, retention, storage location, restore automation, reindex capacity, and drill cadence. A five-minute RPO with a four-hour manual dump schedule is not a plan; it is a contradiction.

business RPO/RTO → backup cadence + retention + off-host protection → restore capacity → reindex duration → validation duration → observed recovery performance

5. Database consistency: use the vendor's backup primitive

Current SonarQube documentation recommends using the database engine's backup tools and explicitly supports hot database backups. For PostgreSQL, pg_dump is a logical-backup utility that takes a consistent snapshot even while the database is in use. That is materially different from copying PostgreSQL data files out from under a running server.

Lab choice. The mandatory exercise uses PostgreSQL 17.11 and pg_dump --format=custom, then verifies the archive with pg_restore --list and a SHA-256 digest before any restore is attempted.

Logical versus physical backup is a database-platform decision. SonarQube does not make every physical snapshot valid. If you use managed snapshots, filesystem snapshots, WAL/PITR, SQL Server backup, or Oracle RMAN in production, the database team's documented consistency and restore procedure remains the authority.

6. Restore compatibility: reconstruct first, migrate second

A recovery drill should minimize simultaneous changes. The safest operational pattern is to restore a known backup into a SonarQube release and edition whose configuration assumptions match the backed-up instance, validate that reconstruction, and only then plan a separate documented upgrade if needed. This is why the lab restores Community Build 26.9 into Community Build 26.9 with the same PostgreSQL major/minor image family.

Do not improvise by pointing an arbitrary older database at a newer server and hoping startup migration will repair everything. SonarQube updates have explicit paths, database requirements, plugin compatibility checks, and sometimes schema migrations. Recovery evidence becomes ambiguous if “restore worked” and “upgrade worked” are tested in one uncontrolled step.

8. Configuration and secret recovery are separate evidence streams

Many startup settings are system properties. They can come from command-line options, environment variables, Helm values/Secrets/ConfigMaps, or sonar.properties. Those values are not all recovered by restoring the SonarQube database. The recovery manifest therefore records where each setting is owned without copying secret values into the evidence packet.

# recovery-manifest.example.yml — no secret values
sonarqube:
  version: 26.9.0.129388
  edition: Community Build
  image: sonarqube:26.9.0.129388-community
database:
  engine: PostgreSQL
  version: 17.11
  backup_method: pg_dump custom archive
configuration:
  source: docker-compose.yml + environment secret injection
  secret_values_in_manifest: false
plugins:
  third_party: []
search:
  recovery: rebuild from restored database
rpo_minutes: 60
rto_minutes: 120

9. Read-only preflight before touching recovery state

Capture the recovery contract before taking or restoring a backup. These are inspection commands; they do not mutate SonarQube:

date -u +%FT%TZ
curl -fsS http://localhost:9000/api/system/status
sonar-scanner --version
git rev-parse HEAD

# PostgreSQL server identity (example local lab container):
docker exec sq28-pg-source psql -U sonar -d sonar -Atc 'select version();'

# Container/image identity and mounts:
docker inspect sq28-sonar-source --format '{{.Config.Image}}'
docker inspect sq28-sonar-source --format '{{json .Mounts}}'

# Record, do not print, secret ownership:
printf '%s\n' 'SONAR_TOKEN: environment-only, value omitted'
printf '%s\n' 'DB password: compose secret/env, value omitted'

The manifest should also capture active plugins, base URL/proxy assumptions, external IdP ownership, database host, backup destination, retention policy, encryption/access controls, and who is authorized to invoke destructive restore or cleanup steps.

Knowledge check

Why is a persisted sonarqube_data volume not automatically a complete disaster-recovery backup?

What does RPO measure?

Why should a DR drill normally restore the same SonarQube release before testing an upgrade?

After restoring the database, what is the supported role of Elasticsearch indexes?

Your database dump restored, but LDAP login fails. Which layer should you investigate?

Next lesson

Build a real isolated PostgreSQL recovery workflow

Lesson 2 turns the durable-state model into a source/restore Docker lab with a real database backup, restore, reindex, and validation analysis.

Official references and version notes

Version and compatibility note

Rechecked 2026-09-08. Mandatory examples target SonarQube Community Build 26.9.0.129388, the official image sonarqube:26.9.0.129388-community, PostgreSQL 17.11 via postgres:17.11-alpine, and SonarScanner CLI 8.1.0.6389. Community Build 26.9 supports PostgreSQL 14–18. The lab restores a 26.9 database into a fresh 26.9 instance first; it does not combine disaster recovery with a SonarQube upgrade or database-vendor migration. Search indexes are treated as rebuildable Sonar-owned state, not the backup authority. Recheck current database requirements, update path, plugins, Java/runtime requirements, and restore instructions before applying these procedures to another release.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.