Database Backup, Restore, Disaster Recovery, and Data Protection: Configuration, Design Patterns, and Trade-Offs
Choose logical or physical database protection, configuration-as-code, search-index rebuilding, restore topology, and backup cadence from explicit RPO/RTO and security requirements.
Learning objectives
- Choose logical or physical database protection based on database-vendor support and recovery objectives rather than SonarQube folklore.
- Separate configuration-as-code and secret recovery from database backup and search-index rebuilding.
- Design cold, warm, and staged recovery patterns with explicit RPO/RTO consequences.
- Set backup cadence, retention, encryption, access, and verification controls from business objectives.
- Keep Community Build, commercial Server, Data Center, database, container, Kubernetes, and cloud-provider ownership boundaries distinct.
- Use a decision record that makes restore-version, plugin, configuration, and validation assumptions auditable.
1. Recovery design begins with the failure you intend to survive
A database backup protects against some failures, but not every failure. Losing a single SonarQube container is different from losing the database host, a cloud region, the secret store, the deployment repository, or an administrator account. A design that survives “container deleted” may still fail completely during “database and secret store unavailable.”
Define the failure domain first, then decide which independent copies and reconstruction inputs must exist outside it.
2. Logical versus physical database backup
| Choice | Strength | Trade-off | SonarQube rule |
|---|---|---|---|
Logical dump, e.g. PostgreSQL pg_dump |
Portable, inspectable, easy isolated restore; suitable for the teaching lab. | Restore can be slower on very large databases; database roles/global objects may require separate treatment. | Use the database vendor's supported tools and validate the restored SonarQube state. |
| Physical/base backup + WAL/PITR | Fast recovery and fine-grained RPO when engineered correctly. | More database-specific; requires consistent snapshot/WAL retention and recovery practice. | Database platform owns correctness; SonarQube does not bless arbitrary file copies. |
| Managed database snapshot | Operationally convenient and often policy-integrated. | Provider-specific consistency, retention, encryption, region/account dependencies. | Document provider guarantees and prove an isolated restore, not just snapshot creation. |
3. Configuration-as-code versus copying installation files
System properties can come from environment variables, Helm values,
Secrets/ConfigMaps, command-line options, or
sonar.properties. A reproducible recovery prefers a
reviewed source-of-truth that can recreate those settings rather
than a blind archive of a mutable installation directory.
Good reconstruction evidence
Versioned Compose/Helm/IaC, configuration manifest, documented secret references, checksumed plugin artifacts, and a known compatible SonarQube image.
Weak reconstruction evidence
“Copy whatever is under /opt/sonarqube” with no
version, source, owner, or secret provenance.
During an update Sonar recommends a fresh distribution and reviewed settings, not blindly copying old configuration files and plugins. The same mindset improves DR: restore intentional configuration, not accidental filesystem history.
4. Database backup versus search-index copies
Database state and Elasticsearch indexes solve different problems. The database is the recovery authority for SonarQube application state. Search indexes make that state queryable. If the restored database and indexes disagree, the supported answer is to rebuild indexes—not edit the database or Elasticsearch documents until they happen to match.
restored DB + compatible server → empty/cleared data/es8 →
SonarQube reindex → search/UI consistency
This is also why RTO includes reindex time. Large instances may be technically restored at the database layer while global issue search is still becoming fully available.
5. Cold, warm, and staged recovery patterns
| Pattern | Typical state before incident | RTO tendency | Cost / risk |
|---|---|---|---|
| Cold | Backups + IaC only; replacement compute is created after incident. | Longest. | Low standby cost; requires strong automation and practiced restore. |
| Warm | Preprovisioned network/host/database target, but no active SonarQube writer against the restored DB. | Shorter. | More cost; configuration drift must be controlled. |
| Staged drill environment | Regular isolated restore target used only for verification. | Does not itself provide failover, but measures achievable RTO. | Excellent confidence; must protect copied credentials/data. |
6. Backup frequency, retention, and recovery objectives
If the RPO is 15 minutes, a nightly dump cannot satisfy it. The database platform may need continuous WAL/log backups or managed point-in-time recovery. If the RTO is one hour, a multi-hour database restore plus multi-hour reindex cannot satisfy it without a different architecture or more capacity.
Track observed durations separately:
backup_start → backup_complete
restore_start → database_restored
sonarqube_start → HTTP operational
reindex_start → search fully available
validation_scan_start → CE SUCCESS
recovery_declared
Use drill measurements, not estimates, to determine whether RPO/RTO commitments are realistic.
7. Backup confidentiality and integrity are security controls
A SonarQube database may contain user identities, issue comments, project metadata, token records, integration settings, and analysis history. A copied recovery database can be sensitive even if source code itself is not stored as a full repository checkout.
- Encrypt backups at rest and in transit according to organizational policy.
- Restrict restore/read permissions independently from routine SonarQube administration.
- Keep integrity hashes or provider-native integrity evidence.
- Define retention and deletion rules; “keep forever” is not automatically safer.
- Use isolated networks for restore drills because restored credentials/identities may still be valid.
- Never include plaintext passwords/tokens in the recovery manifest or CI logs.
8. Restore version and upgrade path are separate decision records
Record the source SonarQube version, edition, Java/runtime, database engine/version, third-party plugins, and deployment configuration at the backup point. For DR, first prove that the backup can reconstruct that known state. If the original release is no longer acceptable for long-term operation, create a second controlled step using the documented update path.
Database-vendor migration is also separate. SonarQube provides a DB Copy Tool for vendor-to-vendor migration with SonarQube-specific checks; a disaster-recovery backup is not the place to improvise an Oracle→PostgreSQL conversion.
Recovery acceptance also needs fresh analysis provenance: preserve
the validation report-task.txt and its
ceTaskId so the restored server’s Compute Engine result
can be tied to the exact post-restore source revision rather than
inferred from a green UI.
9. Edition and deployment boundaries
| Environment | Mandatory recovery principle | Extra concern |
|---|---|---|
| Community Build / single-node Server | Protect database; preserve configuration/plugin/secret ownership; rebuild search indexes. | Lab path in this chapter. |
| Docker | Container/volume persistence does not replace database backup. | Image tag/digest, Compose/secret configuration, volumes, external DB reachability. |
| Kubernetes | External database remains primary durable application state; redeploy chart/config and rebuild search as documented. | Helm values, Secrets, ConfigMaps, PVC policy, ingress and secret-manager recovery. |
| Data Center Edition | Same durable database principle. | Application/search-node topology, licenses, cluster configuration, coordinated recovery; do not invent a one-node shortcut. |
10. Worked decision table
| Requirement | Choice | Observable evidence |
|---|---|---|
| RPO 24h, small lab instance |
Daily encrypted pg_dump + off-host retention
|
Timestamped dump, hash, restore drill, recovery-point comparison. |
| RPO 10m, production PostgreSQL | Database-native PITR/WAL or managed equivalent | WAL continuity, recovery-point test, provider/vendor restore logs. |
| Fast rebuild after host loss | Versioned Compose/Helm/IaC + pinned image/plugin artifacts | Fresh host can reconstruct configuration without copying an old mutable root filesystem. |
| Search data volume lost | Reindex from intact database | Startup/reindex logs and restored project/search behavior. |
| Database and backup account compromised | Independent immutable/offline copy and access separation | Separate credentials, retention lock, restore audit trail. |
Knowledge check
When is a physical PostgreSQL copy a valid backup?
Only when produced by a database-supported physical backup/snapshot/PITR procedure that guarantees consistency. An arbitrary live data-directory copy is not made valid by SonarQube.
Why prefer configuration-as-code over archiving an old SonarQube installation directory?
It records intentional settings and dependencies, makes review/rollback possible, and avoids carrying obsolete binaries/plugins/files into a fresh recovery.
Does deleting Elasticsearch indexes after DB restore destroy authoritative SonarQube history?
No. Current restore guidance expects the indexes to be rebuilt from the restored application state.
A 15-minute RPO is required. What is wrong with a nightly-only dump?
Its backup interval cannot satisfy the maximum acceptable data-loss window. Use a database-native strategy with a sufficiently fine recovery point.
Why is an isolated restore environment a security boundary?
The restored database can contain identities, token records, integration settings, and sensitive metadata. It should not automatically connect to production IdPs, providers, webhooks, or users.
Official references and version notes
- SonarQube Community Build — Backup and restore — use the database vendor's backup tooling; hot database backups are supported; restore the database and rebuild Elasticsearch indexes.
- SonarQube Community Build — Reindexing — startup after restore rebuilds Elasticsearch indexes; project availability and background reindex behavior are documented.
- SonarQube Community Build — Installing database — current supported database engines and versions; PostgreSQL 14–18 is supported starting with Community Build 26.2.
-
Configuration methods
— startup/system properties can live in environment variables,
command line, Helm configuration, or
sonar.properties, and are not all database state. - Performing an update — use a fresh distribution/config review and install compatible third-party plugins rather than blindly copying an old installation.
- Sensitive settings — secret-key material and encrypted configuration require separate protected handling.
- PostgreSQL — pg_dump — produces a consistent logical backup while the database can remain in use.
-
PostgreSQL — pg_restore
— restores non-plain-text archives created by
pg_dump. - Official SonarQube Docker tags — exact Community Build image identity used in the lab.
- Official PostgreSQL Docker tags — exact PostgreSQL image identity used in the lab.
Rechecked 2026-09-08. Mandatory examples target
SonarQube Community Build 26.9.0.129388, the
official image sonarqube:26.9.0.129388-community,
PostgreSQL 17.11 via
postgres:17.11-alpine, and
SonarScanner CLI 8.1.0.6389. Community Build 26.9
supports PostgreSQL 14–18. The lab restores a 26.9 database into a
fresh 26.9 instance first; it does not combine disaster recovery
with a SonarQube upgrade or database-vendor migration. Search
indexes are treated as rebuildable Sonar-owned state, not the
backup authority. Recheck current database requirements, update
path, plugins, Java/runtime requirements, and restore instructions
before applying these procedures to another release.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.