Self-Managed Administration: Configuration, Email, Object Storage, Backups, Restore, and Maintenance: Diagnostics, Failure Modes, Security, and Performance
Diagnose incomplete backups, missing encryption secrets, version/topology mismatches, object-storage mistakes, SMTP/TLS faults, and unsafe maintenance actions from preserved evidence.
Learning objectives
- Use an evidence-first diagnostic sequence for Self-Managed failures.
- Diagnose missing encryption secrets separately from missing application data.
- Interpret exact-version/type restore rejection rather than bypassing it.
- Trace wrong object-storage environment and SMTP/TLS faults without exposing credentials.
- Prevent maintenance/restore actions from becoming the incident.
1. Diagnostic sequence for administrator-level incidents
Administrator failures have large blast radii, so diagnosis begins by preserving evidence, not “trying commands.” Identify the exact offering, GitLab version/type, installation method, topology, repository storage names, object-store buckets, backup ID/hash, encryption-secret generation, affected object class, and maintenance state. Then inspect service health/configuration/logs before the least destructive correction.
preserve evidence
-> identify version/type/install/topology/storage/object class
-> inspect health, config inventory, backup manifest, logs, API/UI symptoms
-> classify data loss vs configuration vs secret vs dependency failure
-> choose least destructive correction
-> verify representative identities and residual risk
-> record sanitized incident evidence
2. Failure: backup exists, encryption secrets do not
The database may restore while encrypted values cannot be decrypted.
This is not fixed by rerunning repository restore. The correct
response is to locate the protected secret backup for the matching
source generation and restore it according to the installation
method. Current GitLab provides
gitlab:doctor:secrets to verify decryptability.
3. Intentionally broken example: exact-version mismatch
Suppose the evidence manifest says 19.3.0-ee but the
target reports 19.3.1-ee. GitLab restore aborts with a
version mismatch. That rejection is a safety signal.
EXPECTED_VERSION=19.3.0-ee
TARGET_VERSION=19.3.1-ee
if [ "$EXPECTED_VERSION" != "$TARGET_VERSION" ]; then
echo "STOP: install exact backup version/type before restore" >&2
exit 42
fi
The repair is to provision the exact matching GitLab build/type as a fresh working target and restore there. Do not edit backup metadata or suppress the mismatch check.
4. Failure: repository storage or target topology does not match
Database records can reference named repository storages. The target must define at least those names. Likewise, restoring local-storage data does not magically migrate it to object storage. A topology migration should be a separate documented operation before or after restore, not hidden inside disaster recovery.
| Evidence | Interpretation | Correction |
|---|---|---|
repo_storage=storage-a in source but target
only has default
|
Configuration mismatch. | Define compatible storage names before restore; verify Gitaly health. |
| Backup from local uploads; target configured for object uploads | Storage-mode mismatch. | Restore as documented, then perform supported migration separately. |
| Existing same-name repository on target | Target not clean enough. | Use a fresh restore target; do not delete valuable repositories merely to make restore continue. |
5. Failure: object-storage credentials work against the wrong environment
A 200 response from S3-compatible storage proves authentication, not correctness. Compare endpoint, account/project, bucket names, object prefixes, object counts, representative hashes, and backup timestamps against the source inventory. Use read-only list/head operations first. Wrong-environment writes can corrupt the recovery set.
# Fixture comparison only.
expected_bucket=gitlab-prod-artifacts
observed_bucket=gitlab-staging-artifacts
[ "$expected_bucket" = "$observed_bucket" ] || {
echo 'STOP: object-storage environment mismatch' >&2
exit 43
}
6. Failure: SMTP/TLS is misconfigured or credentials leak
GitLab documents SMTPS/TLS and STARTTLS as mutually exclusive choices. Enabling both produces a specific configuration error. Preserve the error, inspect effective SMTP settings without printing the password, correct the TLS mode, reconfigure the disposable instance, and retest against a controlled receiver.
# Broken fixture: do NOT copy to production.
gitlab_rails['smtp_tls'] = true
gitlab_rails['smtp_enable_starttls_auto'] = true
# Expected diagnosis: mutually exclusive TLS modes.
If credentials appeared in a log or ticket, rotate/revoke them first, then sanitize retained evidence. Deleting the log before credential rotation is not incident response.
7. Failure: “backup succeeded” but external data was never protected
On Linux package/Docker/self-compiled installations, ordinary application backup does not capture external object-storage contents. Compare the backup manifest to the object-storage map. If external artifacts/packages/registry objects were never backed up, the correct conclusion may be that the recovery point is incomplete—not that restore tooling failed.
8. Failure: maintenance command becomes the outage
Never run restore, migration, destructive cleanup, or storage-moving tasks against production because “a backup probably exists.” Preflight must prove current backup ID/hash, config/secrets availability, target/version compatibility, available disk/object-store capacity, expected downtime, rollback decision point, and owners. If any prerequisite is unknown, stop.
9. Read health signals causally
A readiness failure for Redis suggests a dependency path; an SMTP failure does not. A healthy Rails process does not prove Gitaly or object storage. A Sidekiq backlog can delay background effects without breaking simple liveness. Choose the probe/log/task that matches the failing resource instead of collecting every log or dumping every environment variable.
10. Incident runbook checkpoint
- Freeze destructive changes and preserve backup/config/log evidence.
- Record exact GitLab version/type and installation method.
- Classify missing data versus missing configuration/secrets versus dependency failure.
- Check health of only relevant services and storage.
- Use a fresh separate target for destructive recovery experiments.
- Verify repository SHA, database sample, upload/artifact hash, and registry/package identity as applicable.
- Rotate any leaked credential before log cleanup.
- Document the gap that allowed the incident and update the recovery drill.
Knowledge check
A backup restore aborts on version mismatch. What is the correct repair?
Provision the exact same GitLab version and CE/EE type that created the backup, then restore there. Do not bypass the check.
What does gitlab:doctor:secrets test?
Whether encrypted database values can be decrypted using the current restored secrets.
Why can valid object-storage credentials still produce a recovery incident?
They may point to the wrong account, bucket, prefix, or environment; authentication does not prove object identity.
A real SMTP password appeared in logs. What comes first?
Rotate/revoke the credential, then sanitize retained logs and fix the configuration/logging cause.
What is the least safe way to validate a restore?
Overwriting the only source/production instance before proving prerequisites on a separate target.
11. Lesson summary and bridge
Failure analysis has converted recovery assumptions into explicit invariants. Lesson 5 combines them into one checkpoint: define objectives, create a complete backup set, restore separately, verify identities, and destroy the disposable environment without leaving secrets behind.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Self-Managed commands and file locations depend on installation method. Re-check the documentation for the exact version, topology, package/chart, and storage architecture before production administration or recovery.
- GitLab 19.3 release
- Administer GitLab
- Configure GitLab
- Back up GitLab
- Restore GitLab
- Linux package backup configuration
- Docker backup
- Helm chart backup and restore
- Object storage
- SMTP settings
- Encrypted configuration
- Health check
- Maintenance Mode
- Maintenance Rake tasks
- Integrity check Rake tasks
- Repository checks
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.