Chapter 30Lesson 03~310 minutes

Self-Managed Administration: Configuration, Email, Object Storage, Backups, Restore, and Maintenance: Configuration, Design Choices, and Tradeoffs

Choose deliberately between hosted and Self-Managed responsibility, local and object storage, backup/RPO strategies, and simple single-node versus externalized service architectures.

ArchitectureObject storageRPO/RTOTradeoffsOperations

Learning objectives

  • Choose between GitLab.com and Self-Managed based on operational responsibility rather than feature mythology.
  • Evaluate local versus external object storage using recovery, reliability, security, and cost dimensions.
  • Translate business RPO/RTO into backup frequency, retention, off-site copies, and restore drills.
  • Explain how externalized PostgreSQL/Redis/Gitaly/object storage change failure and recovery coordination.
  • Use a decision table to select a maintainable architecture for a realistic team.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. Backup/restore, health checks, Rake administration, SMTP configuration, and object-storage administration are Self-Managed concerns and have Free-compatible paths where the installation method supports them. Maintenance Mode is Premium/Ultimate on Self-Managed. The mandatory chapter path uses local fixtures and a separate synthetic restore target, so it requires no paid tier, cloud account, production administrator access, or full GitLab installation. Any live commands are explicitly optional and only for an isolated disposable Self-Managed instance you own.

1. From commands to design

Lesson 2 showed that recovery is a dependency graph. Design determines how large that graph becomes. A single Linux-package node is operationally simple but concentrates failure; external object storage can improve scalability but adds bucket policy, network, consistency, and separate-backup obligations. Self-Managed freedom is therefore coupled to platform-engineering ownership.

2. GitLab.com versus Self-Managed operational responsibility

Decision factor GitLab.com Self-Managed
Server patching / GitLab installation Provider responsibility. Your responsibility, including supported upgrade paths.
PostgreSQL/Redis/Gitaly operations Provider responsibility. Your topology, monitoring, backup, capacity, failover.
Application backup/restore Provider platform responsibility; project-level export is a different feature. You must design/test application + config/secrets + storage recovery.
Infrastructure customization Lower host-level control. High control but higher blast radius.
Compliance/data placement Use SaaS capabilities/contract. You choose infrastructure and inherit its controls/obligations.

3. Local storage versus external object storage

Local disk reduces moving parts and is suitable for small disposable/single-node environments, but capacity and durability are tied closely to the node/storage system. External object storage decouples large blobs and is generally preferred for larger setups, but each bucket and credential becomes part of recovery design. The consolidated GitLab object-store configuration reduces duplicated connection configuration across supported object types.

Choice Advantages Recovery/operations cost
Local storage Simple path model; easy lab reasoning. Backups may become large/slow; node storage failure affects more data.
External object storage Scalability/durability options; reduces local disk pressure. Separate provider backup/versioning, IAM, bucket mapping, egress/network, restore sequencing.
Mixed Migrate selected heavy object types first. More state-placement complexity; inventory must say where each object class lives.

4. Backup frequency is a business decision expressed as RPO/RTO

RPO is the maximum acceptable data loss window. RTO is the target time to restore service. A daily backup cannot honestly satisfy a one-hour RPO. A fast backup tool does not prove a two-hour RTO if operators first need to rebuild the exact GitLab version, recover secrets, recreate bucket IAM, and discover storage names.

# Example recovery objective fixture (not a universal recommendation)
rpo_minutes=60
rto_minutes=180
application_backup_frequency=hourly
config_backup_on_change=true
secrets_backup_on_change=true
object_storage_versioning_or_backup=true
restore_drill_frequency=quarterly

5. Single-node simplicity versus external services

External PostgreSQL, Redis, Gitaly nodes, registry metadata database, object stores, and load balancers can improve capacity and availability, but they split recovery across systems. A backup tool may require PgBouncer-specific parameters; repository storage names must match; object-store data and registry metadata should be captured close enough in time to reconstruct a coherent state. Chapter 31 will expand this into HA and disaster recovery.

6. Mail design: delivery is security-sensitive infrastructure

SMTP is not just “notifications.” It participates in account recovery and user workflows. Use authenticated TLS correctly, validate certificates, configure SPF/DKIM/DMARC for production domains, and keep credentials in encrypted configuration or an appropriate secret store. GitLab documents STARTTLS and SMTPS as mutually exclusive modes; enabling both can create a clear diagnostic failure rather than “silent mail instability.”

7. Backup policy: scope, separation, retention, and proof

A production policy should identify application data, config/secrets, certificates/SSH host keys, external object storage, and any external database/registry-specific backup method. Retain multiple generations, keep at least one failure-domain-separated copy, encrypt backups at rest/in transit, restrict restore credentials more strongly than read-only backup credentials where possible, and record deletion/retention ownership.

8. Maintenance Mode versus explicit downtime

Maintenance Mode is useful when Premium/Ultimate is available because it reduces external writes while allowing many reads. It is not a substitute for understanding what internal jobs still change state. Free operators can schedule explicit downtime and stop/write-quiesce services according to the backup method. Choose based on consistency requirement and availability objective, not on convenience.

9. Worked decision table

Scenario Recommended starting point Why / guardrail
10-person internal lab, modest data Single disposable Linux-package node + local storage + off-host backups. Minimize moving parts; still back up config/secrets separately and test restore.
Growing engineering org with heavy artifacts/packages Self-Managed with consolidated object storage and explicit bucket backup/versioning. Reduce local disk pressure; document object classes and IAM.
No platform team, low need for host customization Prefer GitLab.com evaluation. Avoid inheriting server/database/storage operational ownership.
Strict 15-minute RPO and multi-hour restore currently untested Do not claim objective met; redesign backup/replication/drill process. RPO/RTO are evidence-based SLOs, not policy text.
Future HA required Design storage/database/repository boundaries with Chapter 31 reference architecture constraints. Avoid accidental architecture that cannot scale without risky migration.

10. Anti-patterns to reject

  • “The VM snapshot is our backup” without application consistency or restore testing.
  • Keeping gitlab-secrets.json beside an unencrypted database backup with the same access control.
  • Moving to object storage but leaving backup ownership undefined.
  • Using a single permanent root/cloud key for routine backup and restore.
  • Measuring RTO from the moment the restore command starts rather than incident declaration to verified service.
  • Adding external services before there is monitoring, ownership, and a recovery runbook.

Knowledge check

Why can external object storage improve scalability but worsen recovery complexity?

What does a one-hour RPO require?

Does externalizing PostgreSQL automatically create HA?

Why might GitLab.com be a better design choice for a small team?

What should happen when storage design changes?

11. Lesson summary and bridge

Architecture is recovery policy made physical. Lesson 4 now stress-tests the design against the failures that most often expose false confidence: missing secrets, incompatible targets, wrong buckets, SMTP/TLS mistakes, and unsafe maintenance.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Self-Managed commands and file locations depend on installation method. Re-check the documentation for the exact version, topology, package/chart, and storage architecture before production administration or recovery.

Next lesson

Diagnostics, Failure Modes, Security, and Performance

Diagnose missing secrets, incompatible targets, wrong buckets, SMTP/TLS faults, and unsafe maintenance.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.