Database Backup, Restore, Disaster Recovery, and Data Protection: Core Concepts and Mental Model
Treat the database as the primary durable SonarQube state, separate reconstructable search indexes from configuration and secrets, and define recovery through versioned, testable evidence.
Learning objectives
- Identify the SonarQube state that must survive a disaster and separate it from reconstructable search/cache/runtime state.
- Explain database consistency, backup point, RPO, RTO, restore target, and rollback boundary before executing a recovery.
- Build the evidence chain from server/database/configuration inventory to protected backup, isolated restore, reindex, and validation.
- Explain why Docker volumes and Elasticsearch directories are persistence mechanisms but not, by themselves, a complete supported disaster-recovery backup.
- Record configuration, plugin, secret, and external-system dependencies that are not reconstructed merely by restoring database bytes.
- Keep recovery and upgrade as separate changes unless a documented update/migration plan intentionally combines them.
1. The practical problem: persistence is not the same as recoverability
Chapter 27 taught you to operate SonarQube by following evidence across the scanner, web process, Compute Engine, database, and embedded Elasticsearch. Chapter 28 asks a different question: which of those states must be protected so the service can be reconstructed after loss?
A mounted Docker volume can survive a container replacement. An Elasticsearch data directory can survive a process restart. Neither statement proves that you have a recoverable SonarQube backup. A disaster-recovery design must identify the authoritative durable state, record the exact server/database/configuration assumptions, protect the backup outside the failure domain, and repeatedly prove that the backup can start an isolated replacement service.
sonarqube_data volume, container filesystem, or
Elasticsearch directory a complete backup unless your current
supported procedure explicitly defines it that way. Current
SonarQube backup guidance centers recovery on the database vendor's
backup tooling.
2. Mental model: durable state → protected backup → isolated reconstruction
Start with a known server release, database engine/version, system configuration, plugin inventory, and project state. The database contains the primary durable application data. Startup properties and secret material may live outside the database. Elasticsearch indexes are derived from database-backed application state and are rebuilt after restore. A valid recovery therefore combines these inputs deliberately instead of copying an installation directory wholesale.
flowchart TD A[Server/version + config inventory] --> B[Primary SonarQube database] A --> C[Plugin + secret references] B --> D[Vendor-supported protected backup] C --> E[Protected configuration inventory] D --> F[Isolated compatible database restore] E --> G[Fresh compatible SonarQube instance] F --> G G --> H[Fresh / cleared Elasticsearch indexes] H --> I[Automatic reindexing] I --> J[UI/API + project history validation] J --> K[Fresh scanner analysis + ceTaskId] K --> L[Recovery accepted or rollback retained]
Every arrow is an ownership boundary. Database backup is owned by the database platform. SonarQube owns its schema and the logic that rebuilds search indexes. Configuration may be owned by deployment code, a secret manager, or a host file. The recovery operator owns the evidence that proves the pieces belong to the same intended recovery point.
3. State map: what is authoritative, reconstructable, or external?
| State | Owner | Recovery treatment | Common mistake |
|---|---|---|---|
| Projects, issues, measures, profiles, gates, users, tokens, analysis history | SonarQube application database | Protect with supported database backup; validate after restore. | Assuming an Elasticsearch copy is authoritative. |
Elasticsearch indexes under data/es8 |
SonarQube embedded search subsystem | Rebuild from restored application state; current restore guidance clears indexes before restart. | Restoring stale indexes as the primary data source. |
| Startup/system properties | Deployment configuration |
Inventory from environment, Helm, command line, or
sonar.properties; restore through reviewed
configuration.
|
Assuming every server setting is stored in the database. |
| Third-party plugin artifacts | Extension supply chain | Record exact plugin version/source/checksum/license; reinstall a compatible artifact. | Blindly copying an old extensions directory into a different release. |
| Encryption key / external secret references | Secret-management boundary | Protect separately with least privilege and tested retrieval. | Backing up encrypted values but losing the key, or storing key and backup together without controls. |
| SCM, CI artifacts, IdP, external database service, reverse proxy | External systems | Record integration dependencies and recovery ordering; protect in their own systems. | Calling the SonarQube database backup a backup of GitHub, CI, LDAP, or PostgreSQL infrastructure. |
4. RPO and RTO turn “we have backups” into measurable objectives
Recovery Point Objective (RPO) is the acceptable amount of data loss measured in time. If you back up nightly and the database fails at 16:00, a worst-case daily schedule might lose nearly a day of SonarQube state. Recovery Time Objective (RTO) is the acceptable time to restore service to a validated operating state.
Neither number comes from SonarQube automatically. They are governance choices that drive database backup frequency, retention, storage location, restore automation, reindex capacity, and drill cadence. A five-minute RPO with a four-hour manual dump schedule is not a plan; it is a contradiction.
business RPO/RTO → backup cadence + retention + off-host
protection → restore capacity → reindex duration → validation
duration → observed recovery performance
5. Database consistency: use the vendor's backup primitive
Current SonarQube documentation recommends using the database
engine's backup tools and explicitly supports hot database backups.
For PostgreSQL, pg_dump is a logical-backup utility
that takes a consistent snapshot even while the database is in use.
That is materially different from copying PostgreSQL data files out
from under a running server.
pg_dump --format=custom, then verifies the
archive with pg_restore --list and a SHA-256 digest
before any restore is attempted.
Logical versus physical backup is a database-platform decision. SonarQube does not make every physical snapshot valid. If you use managed snapshots, filesystem snapshots, WAL/PITR, SQL Server backup, or Oracle RMAN in production, the database team's documented consistency and restore procedure remains the authority.
6. Restore compatibility: reconstruct first, migrate second
A recovery drill should minimize simultaneous changes. The safest operational pattern is to restore a known backup into a SonarQube release and edition whose configuration assumptions match the backed-up instance, validate that reconstruction, and only then plan a separate documented upgrade if needed. This is why the lab restores Community Build 26.9 into Community Build 26.9 with the same PostgreSQL major/minor image family.
Do not improvise by pointing an arbitrary older database at a newer server and hoping startup migration will repair everything. SonarQube updates have explicit paths, database requirements, plugin compatibility checks, and sometimes schema migrations. Recovery evidence becomes ambiguous if “restore worked” and “upgrade worked” are tested in one uncontrolled step.
7. Elasticsearch: important operational state, but reconstructable after restore
SonarQube's embedded Elasticsearch makes issues and measures
searchable. It is critical to runtime behavior, but current restore
guidance does not tell you to restore search indexes as the backup
authority. Instead the supported sequence is: stop SonarQube,
restore the database, remove the contents of
<sonarqubeHome>/data/es8, and restart. Startup
then rebuilds indexes from restored application state.
During reindexing, some global issue views and filters may be temporarily unavailable even though the server is otherwise operating. That means HTTP 200 is not equivalent to recovery complete. Validation must include reindex progress and project-level evidence.
8. Configuration and secret recovery are separate evidence streams
Many startup settings are system properties. They can come from
command-line options, environment variables, Helm
values/Secrets/ConfigMaps, or sonar.properties. Those
values are not all recovered by restoring the SonarQube database.
The recovery manifest therefore records where each setting
is owned without copying secret values into the evidence packet.
# recovery-manifest.example.yml — no secret values
sonarqube:
version: 26.9.0.129388
edition: Community Build
image: sonarqube:26.9.0.129388-community
database:
engine: PostgreSQL
version: 17.11
backup_method: pg_dump custom archive
configuration:
source: docker-compose.yml + environment secret injection
secret_values_in_manifest: false
plugins:
third_party: []
search:
recovery: rebuild from restored database
rpo_minutes: 60
rto_minutes: 120
9. Read-only preflight before touching recovery state
Capture the recovery contract before taking or restoring a backup. These are inspection commands; they do not mutate SonarQube:
date -u +%FT%TZ
curl -fsS http://localhost:9000/api/system/status
sonar-scanner --version
git rev-parse HEAD
# PostgreSQL server identity (example local lab container):
docker exec sq28-pg-source psql -U sonar -d sonar -Atc 'select version();'
# Container/image identity and mounts:
docker inspect sq28-sonar-source --format '{{.Config.Image}}'
docker inspect sq28-sonar-source --format '{{json .Mounts}}'
# Record, do not print, secret ownership:
printf '%s\n' 'SONAR_TOKEN: environment-only, value omitted'
printf '%s\n' 'DB password: compose secret/env, value omitted'
The manifest should also capture active plugins, base URL/proxy assumptions, external IdP ownership, database host, backup destination, retention policy, encryption/access controls, and who is authorized to invoke destructive restore or cleanup steps.
Knowledge check
Why is a persisted sonarqube_data volume not
automatically a complete disaster-recovery backup?
It primarily contains search/index runtime data. The supported recovery authority is the external database backup plus separately managed configuration/plugin/secret state; indexes are rebuilt after restore.
What does RPO measure?
The maximum acceptable data loss measured in time between the recoverable backup point and the failure.
Why should a DR drill normally restore the same SonarQube release before testing an upgrade?
It isolates recovery correctness from schema/plugin/runtime changes introduced by an upgrade, making failures attributable and rollback clearer.
After restoring the database, what is the supported role of Elasticsearch indexes?
They are rebuilt from restored SonarQube state. Current guidance
clears data/es8 before restart rather than treating
a stale index copy as authoritative.
Your database dump restored, but LDAP login fails. Which layer should you investigate?
External authentication and startup configuration. Restoring the database does not guarantee the reverse proxy, IdP metadata, environment variables, trust material, or secret references were reconstructed.
Official references and version notes
- SonarQube Community Build — Backup and restore — use the database vendor's backup tooling; hot database backups are supported; restore the database and rebuild Elasticsearch indexes.
- SonarQube Community Build — Reindexing — startup after restore rebuilds Elasticsearch indexes; project availability and background reindex behavior are documented.
- SonarQube Community Build — Installing database — current supported database engines and versions; PostgreSQL 14–18 is supported starting with Community Build 26.2.
-
Configuration methods
— startup/system properties can live in environment variables,
command line, Helm configuration, or
sonar.properties, and are not all database state. - Performing an update — use a fresh distribution/config review and install compatible third-party plugins rather than blindly copying an old installation.
- Sensitive settings — secret-key material and encrypted configuration require separate protected handling.
- PostgreSQL — pg_dump — produces a consistent logical backup while the database can remain in use.
-
PostgreSQL — pg_restore
— restores non-plain-text archives created by
pg_dump. - Official SonarQube Docker tags — exact Community Build image identity used in the lab.
- Official PostgreSQL Docker tags — exact PostgreSQL image identity used in the lab.
Rechecked 2026-09-08. Mandatory examples target
SonarQube Community Build 26.9.0.129388, the
official image sonarqube:26.9.0.129388-community,
PostgreSQL 17.11 via
postgres:17.11-alpine, and
SonarScanner CLI 8.1.0.6389. Community Build 26.9
supports PostgreSQL 14–18. The lab restores a 26.9 database into a
fresh 26.9 instance first; it does not combine disaster recovery
with a SonarQube upgrade or database-vendor migration. Search
indexes are treated as rebuildable Sonar-owned state, not the
backup authority. Recheck current database requirements, update
path, plugins, Java/runtime requirements, and restore instructions
before applying these procedures to another release.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.