Chapter 30Lesson 04~340 minutes

Self-Managed Administration: Configuration, Email, Object Storage, Backups, Restore, and Maintenance: Diagnostics, Failure Modes, Security, and Performance

Diagnose incomplete backups, missing encryption secrets, version/topology mismatches, object-storage mistakes, SMTP/TLS faults, and unsafe maintenance actions from preserved evidence.

DiagnosticsSecretsVersioningSMTPMaintenance

Learning objectives

  • Use an evidence-first diagnostic sequence for Self-Managed failures.
  • Diagnose missing encryption secrets separately from missing application data.
  • Interpret exact-version/type restore rejection rather than bypassing it.
  • Trace wrong object-storage environment and SMTP/TLS faults without exposing credentials.
  • Prevent maintenance/restore actions from becoming the incident.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. Backup/restore, health checks, Rake administration, SMTP configuration, and object-storage administration are Self-Managed concerns and have Free-compatible paths where the installation method supports them. Maintenance Mode is Premium/Ultimate on Self-Managed. The mandatory chapter path uses local fixtures and a separate synthetic restore target, so it requires no paid tier, cloud account, production administrator access, or full GitLab installation. Any live commands are explicitly optional and only for an isolated disposable Self-Managed instance you own.

1. Diagnostic sequence for administrator-level incidents

Administrator failures have large blast radii, so diagnosis begins by preserving evidence, not “trying commands.” Identify the exact offering, GitLab version/type, installation method, topology, repository storage names, object-store buckets, backup ID/hash, encryption-secret generation, affected object class, and maintenance state. Then inspect service health/configuration/logs before the least destructive correction.

preserve evidence
  -> identify version/type/install/topology/storage/object class
  -> inspect health, config inventory, backup manifest, logs, API/UI symptoms
  -> classify data loss vs configuration vs secret vs dependency failure
  -> choose least destructive correction
  -> verify representative identities and residual risk
  -> record sanitized incident evidence

2. Failure: backup exists, encryption secrets do not

The database may restore while encrypted values cannot be decrypted. This is not fixed by rerunning repository restore. The correct response is to locate the protected secret backup for the matching source generation and restore it according to the installation method. Current GitLab provides gitlab:doctor:secrets to verify decryptability.

Do not fabricate or regenerate replacement secrets in the hope that old encrypted values will decrypt. Regenerated keys are different keys. Follow GitLab’s lost-secrets recovery documentation and treat unavailable encrypted values as a data-recovery incident.

3. Intentionally broken example: exact-version mismatch

Suppose the evidence manifest says 19.3.0-ee but the target reports 19.3.1-ee. GitLab restore aborts with a version mismatch. That rejection is a safety signal.

EXPECTED_VERSION=19.3.0-ee
TARGET_VERSION=19.3.1-ee
if [ "$EXPECTED_VERSION" != "$TARGET_VERSION" ]; then
  echo "STOP: install exact backup version/type before restore" >&2
  exit 42
fi

The repair is to provision the exact matching GitLab build/type as a fresh working target and restore there. Do not edit backup metadata or suppress the mismatch check.

4. Failure: repository storage or target topology does not match

Database records can reference named repository storages. The target must define at least those names. Likewise, restoring local-storage data does not magically migrate it to object storage. A topology migration should be a separate documented operation before or after restore, not hidden inside disaster recovery.

Evidence Interpretation Correction
repo_storage=storage-a in source but target only has default Configuration mismatch. Define compatible storage names before restore; verify Gitaly health.
Backup from local uploads; target configured for object uploads Storage-mode mismatch. Restore as documented, then perform supported migration separately.
Existing same-name repository on target Target not clean enough. Use a fresh restore target; do not delete valuable repositories merely to make restore continue.

5. Failure: object-storage credentials work against the wrong environment

A 200 response from S3-compatible storage proves authentication, not correctness. Compare endpoint, account/project, bucket names, object prefixes, object counts, representative hashes, and backup timestamps against the source inventory. Use read-only list/head operations first. Wrong-environment writes can corrupt the recovery set.

# Fixture comparison only.
expected_bucket=gitlab-prod-artifacts
observed_bucket=gitlab-staging-artifacts
[ "$expected_bucket" = "$observed_bucket" ] || {
  echo 'STOP: object-storage environment mismatch' >&2
  exit 43
}

6. Failure: SMTP/TLS is misconfigured or credentials leak

GitLab documents SMTPS/TLS and STARTTLS as mutually exclusive choices. Enabling both produces a specific configuration error. Preserve the error, inspect effective SMTP settings without printing the password, correct the TLS mode, reconfigure the disposable instance, and retest against a controlled receiver.

# Broken fixture: do NOT copy to production.
gitlab_rails['smtp_tls'] = true
gitlab_rails['smtp_enable_starttls_auto'] = true
# Expected diagnosis: mutually exclusive TLS modes.

If credentials appeared in a log or ticket, rotate/revoke them first, then sanitize retained evidence. Deleting the log before credential rotation is not incident response.

7. Failure: “backup succeeded” but external data was never protected

On Linux package/Docker/self-compiled installations, ordinary application backup does not capture external object-storage contents. Compare the backup manifest to the object-storage map. If external artifacts/packages/registry objects were never backed up, the correct conclusion may be that the recovery point is incomplete—not that restore tooling failed.

8. Failure: maintenance command becomes the outage

Never run restore, migration, destructive cleanup, or storage-moving tasks against production because “a backup probably exists.” Preflight must prove current backup ID/hash, config/secrets availability, target/version compatibility, available disk/object-store capacity, expected downtime, rollback decision point, and owners. If any prerequisite is unknown, stop.

Production rule: the existence of a command in documentation is not authorization to execute it. Use a change record, peer review, current backup evidence, and an explicit rollback/recovery plan.

9. Read health signals causally

A readiness failure for Redis suggests a dependency path; an SMTP failure does not. A healthy Rails process does not prove Gitaly or object storage. A Sidekiq backlog can delay background effects without breaking simple liveness. Choose the probe/log/task that matches the failing resource instead of collecting every log or dumping every environment variable.

10. Incident runbook checkpoint

  1. Freeze destructive changes and preserve backup/config/log evidence.
  2. Record exact GitLab version/type and installation method.
  3. Classify missing data versus missing configuration/secrets versus dependency failure.
  4. Check health of only relevant services and storage.
  5. Use a fresh separate target for destructive recovery experiments.
  6. Verify repository SHA, database sample, upload/artifact hash, and registry/package identity as applicable.
  7. Rotate any leaked credential before log cleanup.
  8. Document the gap that allowed the incident and update the recovery drill.

Knowledge check

A backup restore aborts on version mismatch. What is the correct repair?

What does gitlab:doctor:secrets test?

Why can valid object-storage credentials still produce a recovery incident?

A real SMTP password appeared in logs. What comes first?

What is the least safe way to validate a restore?

11. Lesson summary and bridge

Failure analysis has converted recovery assumptions into explicit invariants. Lesson 5 combines them into one checkpoint: define objectives, create a complete backup set, restore separately, verify identities, and destroy the disposable environment without leaving secrets behind.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Self-Managed commands and file locations depend on installation method. Re-check the documentation for the exact version, topology, package/chart, and storage architecture before production administration or recovery.

Next lesson

Checkpoint Lab

Prove a complete recovery workflow with preflight, failure injection, separate restore target, integrity evidence, and cleanup.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.