Chapter 28Lesson 04~125 minutes

Database Backup, Restore, Disaster Recovery, and Data Protection: Diagnostics, Failure Modes, and Production Practices

Diagnose failed restores without deleting first-failure evidence, treating Elasticsearch as authoritative, bypassing version compatibility, or destroying the surviving source before validation.

Backup & RestorePostgreSQLDisaster RecoveryRPO / RTOData Protection

Learning objectives

  • Diagnose recovery failures from preserved database, SonarQube startup, search, and scanner evidence.
  • Reject Docker-volume and Elasticsearch-copy shortcuts that do not satisfy the supported durable-state model.
  • Identify version/configuration incompatibility before allowing a restored database to be migrated accidentally.
  • Distinguish backup corruption from JDBC, credentials, plugin, search, resource, or network failures.
  • Protect backups as sensitive assets and retain the surviving source until the restored target is independently validated.

1. Evidence-first restore diagnostic sequence

  1. Preserve the backup archive, checksum, creation timestamp, database-tool version, restore command, and original error output.
  2. Confirm exact SonarQube version/edition, Java/runtime, database engine/version, plugin inventory, and configuration source.
  3. Confirm the restore target is isolated and the source database remains untouched.
  4. Verify archive integrity and database-level restore success before starting SonarQube.
  5. Inspect SonarQube sonar.log/web.log for JDBC, schema, migration, plugin, or startup errors.
  6. Inspect es.log and reindex progress only after database/startup state is sound.
  7. Inspect project history, profiles/gates/users/integrations, then run the smallest validation scan.
  8. Follow report-task.txt and ceTaskId to terminal Compute Engine state.
  9. Apply the least destructive correction to the owning layer and rerun only that recovery step.

2. Failure mode: “we copied the Docker volumes, so we are backed up”

Docker volumes are useful persistence mechanisms. They reduce data loss during container replacement and can carry logs/extensions/search data. They do not remove the requirement for a database backup. A volume snapshot that omits the external PostgreSQL database cannot reconstruct SonarQube application state.

Repair: inventory what each volume contains, add a database-vendor backup, protect deployment configuration and plugin artifacts separately, and prove an isolated reconstruction.

3. Failure mode: treating Elasticsearch as the authoritative restore source

An operator finds an old data/es8 copy and considers it “newer” than the database backup. Merging or editing Elasticsearch documents to manufacture newer state is unsupported and destroys causal evidence.

Repair: restore the chosen database recovery point, clear/freshen the SonarQube search index state according to current documentation, restart, and let SonarQube rebuild indexes. If the RPO is unacceptable, the solution is a better database backup strategy—not hand-editing search state.

4. Failure mode: restoring blindly into an incompatible server

Starting an arbitrary newer SonarQube against an older restored DB can turn a DR drill into an upgrade/migration attempt. Plugin APIs, database requirements, Java requirements, schema migration, and deprecated configuration can all change.

Repair: preserve the restored database, stop the server, identify the exact backup source version/edition, and reconstruct with a compatible release first. Plan the documented update as a separate controlled operation with its own backup and rollback.

5. Failure mode: a successful backup job that has never been restored

A scheduler exit code, object-storage upload, or snapshot “available” status proves creation—not recoverability. The only meaningful evidence is an isolated restore that reaches coherent SonarQube state and a successful fresh analysis.

Track restore drill date, archive hash, database restore result, SonarQube startup, reindex progress, project/history checks, scanner ceTaskId, and observed RTO. Alert on the age of the last successful restore test, not only the age of the newest backup file.

6. Failure mode: backups are readable by everyone who can run SonarQube

Operational access and backup access are different privileges. The backup may outlive source tokens, contain historical user identities and integration settings, and be portable to another host. Store it under separate IAM, encryption, retention, and audit controls.

Never solve a restore TLS problem by disabling certificate verification globally. Fix the trust chain, proxy, database certificate, or secret distribution at the owning layer.

7. Failure mode: deleting the surviving source before restore verification

A restore target that merely reaches the login page is not yet a validated replacement. Do not delete the source database, backup set, container volumes, or original logs until the restored target passes the recovery checklist and stakeholders explicitly accept the recovery point.

This is the rollback boundary: preserve known-good state until the replacement is proven.

8. Intentionally broken example: valid DB restore, wrong JDBC target

The archive checksum verifies. pg_restore finishes successfully into sonar_restore. Then the restore server is intentionally started with:

environment:
  SONAR_JDBC_URL: jdbc:postgresql://pg-restore:5432/sonar_restore_typo

Expected startup evidence is a PostgreSQL/JDBC error indicating that the requested database does not exist. Preserve it:

docker compose -p sq28dr up -d pg-restore
docker compose -p sq28dr up sonar-restore 2>&1 \
  | tee evidence/broken/wrong-jdbc-startup.log

# Prove the database that actually exists without mutating it:
docker compose -p sq28dr exec -T pg-restore \
  psql -U sonar -d postgres -Atc \
  "select datname from pg_database order by datname;" \
  | tee evidence/broken/database-list.txt

Diagnosis: database restore evidence is good; SonarQube cannot connect because deployment configuration points to the wrong target. Correct only the JDBC URL, preserve the broken log, and start the same SonarQube image again. Do not retake the backup, modify tables, or delete Elasticsearch internals to hide the error.

9. Causal failure map

Evidence Owning layer Least-destructive next step
SHA mismatch before restore Backup transport/storage Quarantine copy, fetch another protected copy, investigate corruption; do not restore.
pg_restore parse/error Database backup/restore Preserve stderr and tool versions; inspect archive compatibility and backup creation.
DB restored, JDBC “database does not exist” SonarQube deployment config Correct JDBC database/host/user; do not change restored data.
Schema migration required SonarQube version/update path Stop and verify intended version path before proceeding.
Server UP, global Issues unavailable Elasticsearch reindexing Inspect reindex progress/logs; do not edit indexes directly.
Server/project visible, validation scan CE fails Compute Engine/project analysis Follow ceTaskId, preserve CE task/error, fix that analysis path.

10. Recovery shortcuts to reject

  • Do not call docker cp of a running database data directory a backup.
  • Do not restore a stale Elasticsearch directory over a newer restored database.
  • Do not directly update SonarQube database tables or Elasticsearch documents to “repair” a restore.
  • Do not start repeated random SonarQube versions against the same restore target.
  • Do not delete first-failure logs, backup hashes, or restore stderr before diagnosis.
  • Do not expose restored credentials to production provider endpoints during a drill.
  • Do not destroy the surviving source or only good backup before validation completes.
  • Do not weaken Quality Gates, rules, permissions, or TLS to make the recovery look green.

Knowledge check

A restore target shows the login page but global issue search is incomplete. Is recovery complete?

The dump hash fails verification. Should you try pg_restore anyway?

The database restore succeeds but SonarQube reports “database does not exist.” What layer is most likely wrong?

Why is repeatedly restarting SonarQube a poor restore diagnostic?

Why should restore-test data remain isolated from production IdPs/providers?

Next lesson

Prove the complete disaster-recovery checkpoint

Lesson 5 packages the backup, recovery point, restore, reindex, analysis, and rollback evidence into an auditable checkpoint.

Official references and version notes

Version and compatibility note

Rechecked 2026-09-08. Mandatory examples target SonarQube Community Build 26.9.0.129388, the official image sonarqube:26.9.0.129388-community, PostgreSQL 17.11 via postgres:17.11-alpine, and SonarScanner CLI 8.1.0.6389. Community Build 26.9 supports PostgreSQL 14–18. The lab restores a 26.9 database into a fresh 26.9 instance first; it does not combine disaster recovery with a SonarQube upgrade or database-vendor migration. Search indexes are treated as rebuildable Sonar-owned state, not the backup authority. Recheck current database requirements, update path, plugins, Java/runtime requirements, and restore instructions before applying these procedures to another release.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.