Chapter 31Lesson 03~320 minutes

Upgrades, Zero-Downtime Planning, High Availability, Geo, Disaster Recovery, and Capacity: Configuration, Design Choices, and Tradeoffs

Choose deliberately among in-place and staged upgrades, downtime and zero-downtime, HA and Geo/backup, and capacity headroom and cost with operational consequences explicit.

ArchitectureTradeoffsHACapacityRollback

Learning objectives

  • Choose between in-place and staged-replacement upgrade strategies.
  • Decide when downtime is safer than zero-downtime complexity.
  • Separate availability engineering from disaster-recovery engineering.
  • Design capacity headroom from measured workload rather than arbitrary node counts.
  • Document rollback criteria and exception ownership before execution.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. Upgrade planning, required upgrade stops, background-migration checks, rollback planning, and the general reference-architecture guidance apply to GitLab Self-Managed and have Free-compatible learning paths. The documented zero-downtime procedure is available for Self-Managed but requires a properly designed multi-node Linux-package environment with load balancing and HA mechanisms. Geo is Premium/Ultimate on GitLab Self-Managed. The mandatory chapter path uses static fixtures and tabletop exercises; it requires no paid tier, enterprise topology, cloud spend, or real production upgrade.

1. Start with constraints, not fashionable topology

There is no universally “best” GitLab upgrade or resilience architecture. The right choice depends on tolerated downtime, data-loss tolerance, workload, operator maturity, installation method, service dependencies, budget, and evidence from restore/failover tests. A simpler topology with a proven three-hour restore can be safer than an elaborate HA design nobody can operate under pressure.

2. In-place upgrade versus staged replacement

Choice Strengths Risks / costs Prefer when
In-place Less duplicate infrastructure; preserves host/storage identity. Rollback may require restoring the previous-version backup; configuration drift can accumulate. Topology is well understood and rollback has been rehearsed.
Staged replacement / clone Can validate target separately; easier before/after comparison. Data synchronization/cutover is more complex; extra capacity/storage cost. You need stronger isolation or major infrastructure change.
Disposable rehearsal clone High learning value with low production risk. Must represent production closely enough to expose real migration/runtime cost. Always when practical before a significant production upgrade.

3. Downtime versus zero-downtime

Zero-downtime is justified when outage cost is high enough to pay for topology and operational complexity. Current GitLab documentation requires a multi-node Linux-package environment, load balancing, readiness checks, HA backends, one-minor-at-a-time upgrades, and careful migration orchestration. If one stateful dependency is not HA, that dependency can still require downtime.

Production pattern: choose an explicit, tested maintenance window when zero-downtime prerequisites are not genuinely satisfied. Claiming “zero downtime” while a non-HA database or storage layer remains is a documentation failure before it is a technical failure.

4. HA versus Geo versus backup

Mechanism Primary question answered Does not guarantee
HA Can this site continue through a component/node failure? Region survival, historical recovery, protection from logical corruption.
Geo Can another site hold replicated GitLab state and be promoted for DR? Zero-RPO, instant failover, or recovery from replicated bad state.
Backup/restore Can we reconstruct a known historical recovery point? Low RTO unless restore is engineered/tested; continuous availability.

For a production operating model, write down which failure scenario each control covers. If two controls claim the same scenario, test whether they actually fail independently.

5. Turn RPO/RTO into retention and test policy

An RPO of one hour implies backup/replication processes and monitoring must make a recoverable point at least that recent under the expected failure scenario. An RTO of three hours implies the full chain—incident declaration, access, infrastructure, restore/promotion, secrets/configuration, DNS, validation, and stakeholder decision—must fit within three hours. “Restore took 40 minutes” is not the same as an end-to-end RTO.

6. Capacity headroom versus cost

Reference architectures are starting points. Measure peak RPS, background queues, database/Redis saturation, repository I/O, network, storage growth, and burst patterns. Capacity headroom buys time during traffic growth, node loss, migrations, and failover, but unused capacity costs money. The correct headroom is a risk decision supported by load evidence and scaling lead time.

Signal Under-capacity symptom Design response
API/web RPS + latency Latency/error rate rises near peak. Scale application/load-balancing layer and inspect downstream bottlenecks.
Sidekiq queues Jobs age/grow faster than they drain. Add/tune workers after confirming DB/Redis capacity.
Gitaly I/O Clone/push latency, disk saturation. Scale storage/shards according to supported architecture; avoid burstable disks.
PostgreSQL CPU, I/O, connections, lock pressure. Tune/scale the database architecture; application nodes alone will not help.
Geo lag Secondary falls farther behind under load. Increase replication/network/storage capacity before relying on DR objectives.

7. Worked scenario: 45 RPS today, 70 RPS projected

A fictional environment currently peaks at 45 RPS and must survive one app-node failure. Forecast demand is 70 RPS within six months. Rather than select a reference architecture solely by user count, the team records actual RPS composition, database headroom, Sidekiq queue age, Gitaly throughput, and failover capacity. It then chooses a staged capacity increase before the forecasted peak and reruns load/failover tests. If cross-region DR is required, Geo is evaluated separately; it is not added merely because the local HA architecture grows.

8. Rollback is a data decision, not a package downgrade shortcut

Current GitLab rollback guidance requires a database backup from the exact version/edition you are rolling back to because schema migrations have changed the database. A safe runbook defines the point after which a roll-forward is preferable, the backup ID, restore target, exact source version, and data-loss consequence of reverting to the pre-upgrade point.

Never promise “we can always downgrade the package.” Code and database schema must be compatible. After migrations run, restoring the pre-upgrade database may be required.

9. Architecture decision table

Constraint Recommended direction Why
Small environment; 2-hour planned outage acceptable Simpler standalone + tested backup/restore. Lower operating complexity may beat HA cost.
High outage cost; mature multi-node Linux-package platform Evaluate documented zero-downtime sequence. HA/load-balancing can keep traffic flowing while nodes rotate.
Regional disaster is in threat model Add Geo evaluation plus independent backup. Local HA does not cover region loss; Geo still does not replace historical backups.
Fast automation growth Size by measured RPS/queues, not seats. Bots and CI can dominate workload.
Unproven recovery process Prioritize restore rehearsal before topology expansion. Complexity without recovery evidence increases operational risk.

Knowledge check

Why might planned downtime be safer than zero-downtime?

What failure does HA not solve by itself?

Why can a Geo secondary still fail RTO after promotion is technically possible?

What does rollback require after database migrations have run?

Why should capacity planning include failure-mode headroom?

10. Lesson summary and bridge

Architecture choices are risk allocations: downtime versus complexity, local availability versus regional recovery, and headroom versus cost. Lesson 4 applies this model to real failure signatures and diagnostic gates.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Upgrade paths, required stops, supported infrastructure, migration behavior, Geo procedures, and reference architectures change over time. Recalculate the path and re-read every version-specific note immediately before a real upgrade.

Next lesson

Diagnostics, Failure Modes, Security, and Performance

Diagnose skipped stops, unfinished migrations, fragile HA, Geo lag, and capacity failures from evidence.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.