Upgrades, Zero-Downtime Planning, High Availability, Geo, Disaster Recovery, and Capacity: Concepts, Architecture, and Mental Model
Build a version-aware GitLab Self-Managed resilience mental model: upgrade paths and stops, background migrations, zero-downtime prerequisites, HA, Geo, DR objectives, rollback, and capacity evidence.
Learning objectives
- Explain why GitLab upgrades are data migrations and change-management events rather than package-install operations.
- Calculate the role of required upgrade stops and background migrations in a safe path.
- Distinguish ordinary downtime upgrades from the stricter zero-downtime procedure.
- Separate HA, backups, and Geo as different resilience mechanisms with different failure domains.
- Use RPO, RTO, RPS, latency, storage growth, and queue pressure as operational evidence instead of user count alone.
1. The practical problem: “the service started” is not an upgrade result
A GitLab upgrade changes application code, database schema, background data migrations, bundled services, configuration expectations, and sometimes operational sequencing. A node can restart successfully while background migrations are incomplete, a repository backend is unhealthy, encrypted values cannot be decrypted, or a later upgrade stop has already been invalidated. The operator therefore needs a path, not merely a target version.
2. Version cadence, upgrade notes, and required stops
GitLab publishes monthly major/minor releases and patch releases.
Current GitLab 19 upgrade stops follow the predictable pattern
19.2, 19.5, 19.8, and
19.11. GitLab 18 similarly uses 18.2,
18.5, 18.8, and 18.11. At
each required stop, use the latest available patch of that minor
release, read the upgrade notes for every intervening version
relevant to your installation method, and let required background
migrations finish before the next stop.
| Term | Meaning | Operational consequence |
|---|---|---|
| Target version | The version you eventually want to run. | It does not define the legal path from the current version. |
| Required upgrade stop | A minor release GitLab requires before later releases. | Install it, verify it, and finish background migrations before continuing. |
| Patch release | Bug/security fixes within a minor line. | Prefer the latest available patch at a stop instead of the first patch. |
| Upgrade note | Version-specific prerequisites, breaking changes, long migrations, or installation-specific actions. | Read notes for all versions crossed; do not infer from generic instructions. |
3. Database and background migrations are hard gates
GitLab uses ordinary migrations plus batched background migrations. Sidekiq runs batched migrations while the application remains available, but GitLab explicitly requires all relevant background migrations to be finished before the next upgrade. On GitLab 18.9 and later, the current Rake interface can list them directly.
# OPTIONAL read-only commands on an isolated/current Self-Managed instance.
sudo gitlab-rake gitlab:background_migrations:list
sudo gitlab-rake gitlab:check
sudo gitlab-rake gitlab:doctor:secrets
curl -fsS https://gitlab.example.invalid/-/readiness?all=1
An active, failed, or otherwise unfinished
migration is evidence to stop. Forcing a later upgrade can produce
migration errors or schema/data inconsistency. The right response is
to diagnose the migration and complete or repair it according to the
version-specific documentation—not to edit the database state to
make the warning disappear.
4. Downtime upgrade versus zero-downtime upgrade
A normal single-node or planned-downtime upgrade accepts an outage to make sequencing simpler. Zero-downtime is a distinct multi-node Linux-package procedure. It assumes external load balancing, readiness checks, and HA mechanisms for the components that must remain available. GitLab requires upgrading one minor release at a time for zero-downtime; skipping minors can apply database changes in the wrong sequence.
| Dimension | Planned downtime | Zero-downtime |
|---|---|---|
| Topology | Single-node or multi-node. | Multi-node Linux-package environment with load balancers and HA mechanisms. |
| Version stepping | Follow required stops and upgrade notes. | One minor release at a time, plus required stops/notes. |
| Traffic | Writes/app may be stopped during critical work. | Nodes are drained/rotated while other nodes serve traffic. |
| Database work | Simpler coordinated maintenance window. | Deploy node, post-deployment migration discipline, direct DB-leader access where required. |
| Risk | Business outage is explicit. | Operational complexity is higher; “zero” downtime is an engineering objective, not a magical flag. |
5. High availability removes some node failures, not all failure modes
HA means redundant components and failover mechanisms are designed so a single component failure does not necessarily stop service. It does not mean the whole platform is immune to correlated failures, bad migrations, operator error, storage corruption, network partitions, or a region loss. GitLab reference architectures use workload evidence such as requests per second (RPS) and then map that load to component sizing.
flowchart TD
CUR[Current GitLab version] --> PATH[Upgrade path + required stops]
PATH --> PRE[Backup/rollback + health preflight]
PRE --> STOP1[Upgrade stop]
STOP1 --> MIG[Background migrations finish]
MIG --> VERIFY[Health + data verification]
VERIFY --> NEXT{More stops?}
NEXT -->|Yes| STOP1
NEXT -->|No| TARGET[Target version]
TARGET --> POST[Post-upgrade evidence]
POST --> DR[Restore/failover readiness]
The current reference-architecture guidance recommends standalone designs for many environments below roughly the 3K-user / 60-RPS class unless HA is a business requirement; at 3K and above, GitLab publishes HA reference architectures. Treat these as current design guidance, not universal capacity limits. Automation-heavy installations can exceed a “user-count” model quickly.
6. Geo is replication and disaster recovery—not backup and not automatic zero-RPO
GitLab Geo uses a primary site and one or more secondary sites, and it is currently Premium/Ultimate for Self-Managed. Geo can replicate PostgreSQL data, repositories, and other tracked data to another site and can support planned or disaster failover. But Geo is explicitly not an out-of-the-box HA solution, and replication lag means the actual recoverable point depends on what reached the secondary before the incident.
7. RPO and RTO make resilience testable
Recovery point objective (RPO) is the maximum tolerable data loss measured in time. Recovery time objective (RTO) is the target time to restore a verified service after a disruption. Backups, Geo lag, DNS/load-balancer cutover, secret/config restoration, dependency health, and human decision time all affect whether those objectives are realistic.
| Evidence | RPO relevance | RTO relevance |
|---|---|---|
| Backup timestamp + successful restore test | Defines a known recoverable point. | Shows how long restore/verification actually takes. |
| Geo replication lag | Shows possible latest replicated point. | Affects confidence before promotion. |
| Runbook ownership | Little direct effect on data point. | Strong effect on decision and execution time. |
| Capacity headroom | Helps secondary/target absorb current workload. | Insufficient headroom can turn failover into a performance incident. |
8. Capacity is measured demand plus headroom
Current GitLab reference-architecture guidance uses peak RPS as the primary sizing signal, with user count as a fallback. Measure API, web, Git pull/push demand, Sidekiq queues, database/Redis load, Gitaly I/O, object-storage throughput, network latency, and growth. For HA, synchronous components need low latency; GitLab guidance generally targets less than about 5 ms within the supported regional topology. A single GitLab environment across multiple regions is not supported; use Geo for cross-region distribution.
9. Read-only inspection before any upgrade design
# OPTIONAL: isolated Linux-package Self-Managed instance only.
sudo gitlab-rake gitlab:env:info
sudo gitlab-rake gitlab:check SANITIZE=true
sudo gitlab-rake gitlab:doctor:secrets
sudo gitlab-rake gitlab:background_migrations:list
curl -fsS https://gitlab.example.invalid/-/health
curl -fsS https://gitlab.example.invalid/-/readiness?all=1
# Record versions/topology, not secret values.
Also inventory installation method, edition, OS, PostgreSQL/Redis/Gitaly topology, external object stores, registry, advanced search, Geo sites, runners, and current backup/restore evidence. Upgrade notes are topology-aware.
Knowledge check
Why can an upgrade path require intermediate versions?
Required stops carry migration and upgrade-process assumptions that later versions depend on; skipping them can leave schema/data changes out of sequence.
What must happen before moving from one required stop to the next?
Finish required background migrations, run health/integrity checks, and verify the current step before continuing.
Can a zero-downtime upgrade skip from 19.1 directly to 19.3 if 19.2 is healthy elsewhere?
No. The documented zero-downtime procedure requires one minor release at a time.
Does Geo replace backups?
No. Geo replicates state and can replicate corruption or unwanted changes; independent backups provide a different recovery mechanism and historical recovery point.
Why is RPS better than user count for capacity planning?
RPS measures actual peak workload, including automation, while equal user counts can generate very different request volumes.
10. Lesson summary and bridge
You now have the control model: version path → migration gate → health/data verification → rollback decision, surrounded by topology, HA/Geo/backup boundaries, and capacity evidence. Lesson 2 turns that model into a reproducible upgrade and DR planning lab.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Upgrade paths, required stops, supported infrastructure, migration behavior, Geo procedures, and reference architectures change over time. Recalculate the path and re-read every version-specific note immediately before a real upgrade.
- GitLab 19.3 release
- Upgrade GitLab
- Before you upgrade
- Plan your upgrade path
- GitLab 19 upgrade notes
- Background migrations
- Upgrade multi-node with downtime
- Upgrade multi-node with zero downtime
- Roll back earlier GitLab versions
- Reference architectures
- 1K / 20 RPS reference architecture
- 2K / 40 RPS reference architecture
- 3K / 60 RPS HA reference architecture
- GitLab Geo
- Set up Geo
- Geo disaster recovery
- Restore GitLab
- Health checks
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.