Upgrades, Zero-Downtime Planning, High Availability, Geo, Disaster Recovery, and Capacity: Configuration, Design Choices, and Tradeoffs
Choose deliberately among in-place and staged upgrades, downtime and zero-downtime, HA and Geo/backup, and capacity headroom and cost with operational consequences explicit.
Learning objectives
- Choose between in-place and staged-replacement upgrade strategies.
- Decide when downtime is safer than zero-downtime complexity.
- Separate availability engineering from disaster-recovery engineering.
- Design capacity headroom from measured workload rather than arbitrary node counts.
- Document rollback criteria and exception ownership before execution.
1. Start with constraints, not fashionable topology
There is no universally “best” GitLab upgrade or resilience architecture. The right choice depends on tolerated downtime, data-loss tolerance, workload, operator maturity, installation method, service dependencies, budget, and evidence from restore/failover tests. A simpler topology with a proven three-hour restore can be safer than an elaborate HA design nobody can operate under pressure.
2. In-place upgrade versus staged replacement
| Choice | Strengths | Risks / costs | Prefer when |
|---|---|---|---|
| In-place | Less duplicate infrastructure; preserves host/storage identity. | Rollback may require restoring the previous-version backup; configuration drift can accumulate. | Topology is well understood and rollback has been rehearsed. |
| Staged replacement / clone | Can validate target separately; easier before/after comparison. | Data synchronization/cutover is more complex; extra capacity/storage cost. | You need stronger isolation or major infrastructure change. |
| Disposable rehearsal clone | High learning value with low production risk. | Must represent production closely enough to expose real migration/runtime cost. | Always when practical before a significant production upgrade. |
3. Downtime versus zero-downtime
Zero-downtime is justified when outage cost is high enough to pay for topology and operational complexity. Current GitLab documentation requires a multi-node Linux-package environment, load balancing, readiness checks, HA backends, one-minor-at-a-time upgrades, and careful migration orchestration. If one stateful dependency is not HA, that dependency can still require downtime.
4. HA versus Geo versus backup
| Mechanism | Primary question answered | Does not guarantee |
|---|---|---|
| HA | Can this site continue through a component/node failure? | Region survival, historical recovery, protection from logical corruption. |
| Geo | Can another site hold replicated GitLab state and be promoted for DR? | Zero-RPO, instant failover, or recovery from replicated bad state. |
| Backup/restore | Can we reconstruct a known historical recovery point? | Low RTO unless restore is engineered/tested; continuous availability. |
For a production operating model, write down which failure scenario each control covers. If two controls claim the same scenario, test whether they actually fail independently.
5. Turn RPO/RTO into retention and test policy
An RPO of one hour implies backup/replication processes and monitoring must make a recoverable point at least that recent under the expected failure scenario. An RTO of three hours implies the full chain—incident declaration, access, infrastructure, restore/promotion, secrets/configuration, DNS, validation, and stakeholder decision—must fit within three hours. “Restore took 40 minutes” is not the same as an end-to-end RTO.
6. Capacity headroom versus cost
Reference architectures are starting points. Measure peak RPS, background queues, database/Redis saturation, repository I/O, network, storage growth, and burst patterns. Capacity headroom buys time during traffic growth, node loss, migrations, and failover, but unused capacity costs money. The correct headroom is a risk decision supported by load evidence and scaling lead time.
| Signal | Under-capacity symptom | Design response |
|---|---|---|
| API/web RPS + latency | Latency/error rate rises near peak. | Scale application/load-balancing layer and inspect downstream bottlenecks. |
| Sidekiq queues | Jobs age/grow faster than they drain. | Add/tune workers after confirming DB/Redis capacity. |
| Gitaly I/O | Clone/push latency, disk saturation. | Scale storage/shards according to supported architecture; avoid burstable disks. |
| PostgreSQL | CPU, I/O, connections, lock pressure. | Tune/scale the database architecture; application nodes alone will not help. |
| Geo lag | Secondary falls farther behind under load. | Increase replication/network/storage capacity before relying on DR objectives. |
7. Worked scenario: 45 RPS today, 70 RPS projected
A fictional environment currently peaks at 45 RPS and must survive one app-node failure. Forecast demand is 70 RPS within six months. Rather than select a reference architecture solely by user count, the team records actual RPS composition, database headroom, Sidekiq queue age, Gitaly throughput, and failover capacity. It then chooses a staged capacity increase before the forecasted peak and reruns load/failover tests. If cross-region DR is required, Geo is evaluated separately; it is not added merely because the local HA architecture grows.
8. Rollback is a data decision, not a package downgrade shortcut
Current GitLab rollback guidance requires a database backup from the exact version/edition you are rolling back to because schema migrations have changed the database. A safe runbook defines the point after which a roll-forward is preferable, the backup ID, restore target, exact source version, and data-loss consequence of reverting to the pre-upgrade point.
9. Architecture decision table
| Constraint | Recommended direction | Why |
|---|---|---|
| Small environment; 2-hour planned outage acceptable | Simpler standalone + tested backup/restore. | Lower operating complexity may beat HA cost. |
| High outage cost; mature multi-node Linux-package platform | Evaluate documented zero-downtime sequence. | HA/load-balancing can keep traffic flowing while nodes rotate. |
| Regional disaster is in threat model | Add Geo evaluation plus independent backup. | Local HA does not cover region loss; Geo still does not replace historical backups. |
| Fast automation growth | Size by measured RPS/queues, not seats. | Bots and CI can dominate workload. |
| Unproven recovery process | Prioritize restore rehearsal before topology expansion. | Complexity without recovery evidence increases operational risk. |
Knowledge check
Why might planned downtime be safer than zero-downtime?
If HA/load-balancer/stateful-service prerequisites or operator rehearsal are missing, the simpler coordinated outage reduces moving parts and hidden partial-upgrade states.
What failure does HA not solve by itself?
Examples include region loss, logical corruption, bad migrations, operator deletion, or shared-database/storage failure.
Why can a Geo secondary still fail RTO after promotion is technically possible?
It may lack capacity, current configuration/secrets, network/DNS readiness, operational ownership, or dependency validation.
What does rollback require after database migrations have run?
A compatible previous code version and typically a backup/database state from the exact version/edition being restored to.
Why should capacity planning include failure-mode headroom?
During node loss, upgrades, migrations, or DR, remaining components must absorb extra load; normal-state utilization alone can hide insufficient failover capacity.
10. Lesson summary and bridge
Architecture choices are risk allocations: downtime versus complexity, local availability versus regional recovery, and headroom versus cost. Lesson 4 applies this model to real failure signatures and diagnostic gates.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Upgrade paths, required stops, supported infrastructure, migration behavior, Geo procedures, and reference architectures change over time. Recalculate the path and re-read every version-specific note immediately before a real upgrade.
- GitLab 19.3 release
- Upgrade GitLab
- Before you upgrade
- Plan your upgrade path
- GitLab 19 upgrade notes
- Background migrations
- Upgrade multi-node with downtime
- Upgrade multi-node with zero downtime
- Roll back earlier GitLab versions
- Reference architectures
- 1K / 20 RPS reference architecture
- 2K / 40 RPS reference architecture
- 3K / 60 RPS HA reference architecture
- GitLab Geo
- Set up Geo
- Geo disaster recovery
- Restore GitLab
- Health checks
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.