Chapter 31Lesson 02~400 minutes

Upgrades, Zero-Downtime Planning, High Availability, Geo, Disaster Recovery, and Capacity: Guided Hands-On Workflow and Core Operations

Use Free-compatible fixtures to calculate a safe upgrade path, gate each stop on migrations and health, model HA/Geo topology, and run a tabletop failover and recovery exercise.

Hands-onUpgrade pathDRRPO/RTOEvidence

Learning objectives

  • Calculate a current required-stop path without hard-coding future patch versions.
  • Gate each simulated upgrade step on migration and health evidence.
  • Contrast a normal downtime path with the stricter zero-downtime minor-by-minor path.
  • Model HA and Geo topology without requiring enterprise infrastructure.
  • Run a tabletop failover and record RPO/RTO and capacity evidence.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. Upgrade planning, required upgrade stops, background-migration checks, rollback planning, and the general reference-architecture guidance apply to GitLab Self-Managed and have Free-compatible learning paths. The documented zero-downtime procedure is available for Self-Managed but requires a properly designed multi-node Linux-package environment with load balancing and HA mechanisms. Geo is Premium/Ultimate on GitLab Self-Managed. The mandatory chapter path uses static fixtures and tabletop exercises; it requires no paid tier, enterprise topology, cloud spend, or real production upgrade.

1. Disposable scenario and preflight

The fictional source is a Self-Managed 18.10.x-ee instance whose target is 19.3.x-ee. The mandatory lab never installs GitLab. Instead, it creates a small evidence bundle that validates the decision logic: stops, migrations, health, rollback readiness, HA failure domains, Geo lag, and recovery objectives.

Production warning: do not reuse the fixture’s version path blindly. Immediately before a real upgrade, calculate the path from the exact source version, consult every relevant upgrade note, and choose the latest available patch for each stop.

2. Build the upgrade evidence fixture

set -eu
rm -rf ch31-upgrade-lab
mkdir -p ch31-upgrade-lab/{evidence,backup,topology}
cd ch31-upgrade-lab
cat > evidence/instance.json <<'JSON'
{
  "offering": "Self-Managed",
  "edition": "EE",
  "installation": "Linux package fixture",
  "current": "18.10.x",
  "target": "19.3.x",
  "rpo_minutes": 60,
  "rto_minutes": 180,
  "peak_rps": 45
}
JSON
printf 'backup-id=fixture-20260822
restore-test=PASS
secrets-copy=PASS
' > backup/rollback-evidence.txt
sha256sum evidence/instance.json backup/rollback-evidence.txt > evidence/preflight.sha256
sha256sum -c evidence/preflight.sha256

3. Calculate two different paths: downtime and zero-downtime

For the current documentation snapshot, a normal path from 18.10 to 19.3 must include the 18.11 and 19.2 required stops. By contrast, zero-downtime cannot skip minors, so it must traverse each minor release sequentially.

cat > evidence/upgrade-plan.json <<'JSON'
{
  "normal_downtime": ["18.11.latest", "19.2.latest", "19.3.latest"],
  "zero_downtime": ["18.11.latest", "19.0.latest", "19.1.latest", "19.2.latest", "19.3.latest"],
  "gate_after_each_step": ["background_migrations_finished", "health_pass", "sample_data_pass"],
  "patch_rule": "resolve latest available patch immediately before execution"
}
JSON
jq . evidence/upgrade-plan.json

Why does the normal path not list every minor? Required-stop planning and zero-downtime sequencing are different rules. The standard path must include required stops; the zero-downtime procedure additionally requires one minor at a time.

4. Model the migration gate and halt behavior

cat > evidence/19.2-gate.json <<'JSON'
{"version":"19.2.latest","background_migrations":"active","readiness":"pass","sample_data":"pass"}
JSON
python - <<'PY'
import json, sys
s=json.load(open('evidence/19.2-gate.json'))
if s['background_migrations'] != 'finished':
    print('HALT: background migrations are not finished')
    sys.exit(42)
PY

The expected exit code is 42. Preserve it as evidence. Then repair the fixture rather than bypassing the gate:

python - <<'PY'
import json
p='evidence/19.2-gate.json'
s=json.load(open(p)); s['background_migrations']='finished'
open(p,'w').write(json.dumps(s,indent=2)+'\n')
PY
python - <<'PY'
import json
s=json.load(open('evidence/19.2-gate.json'))
assert s['background_migrations']=='finished'
assert s['readiness']=='pass'
assert s['sample_data']=='pass'
print('19.2 gate=PASS; continuation allowed')
PY

5. Optional read-only live preflight

If you own an isolated disposable Self-Managed instance, inspect rather than mutate first:

# OPTIONAL: current Self-Managed instance only.
sudo gitlab-rake gitlab:env:info
sudo gitlab-rake gitlab:check SANITIZE=true
sudo gitlab-rake gitlab:doctor:secrets
sudo gitlab-rake gitlab:background_migrations:list
curl -fsS https://gitlab.example.invalid/-/readiness?all=1
# Compare exact current/target versions with the current upgrade path and upgrade notes.

An actual upgrade is optional and deliberately omitted from the mandatory path because package/chart commands, OS support, exact patch versions, optional services, and migration duration are environment-specific. If you perform one, do it only on a disposable clone and preserve pre/post evidence.

6. Model HA and Geo as separate failure domains

cat > topology/resilience.json <<'JSON'
{
  "site_a": {
    "app_nodes": 2,
    "database": "HA conceptual",
    "redis": "HA conceptual",
    "gitaly": "redundant conceptual",
    "peak_rps": 45
  },
  "site_b_geo_secondary": {
    "tier": "Premium/Ultimate optional",
    "replication_lag_seconds": 75,
    "tested_capacity_rps": 55
  },
  "independent_backup": {
    "age_minutes": 35,
    "restore_test": "PASS"
  }
}
JSON
jq . topology/resilience.json

The HA model handles an app-node failure inside Site A. The Geo secondary models regional DR. The independent backup models recovery from destructive/corrupt state. These are deliberately three separate controls.

HA, Geo, backup, and capacity are different controls
flowchart LR
  USERS[Users / CI / API] --> LB[Load balancer]
  LB --> APP1[GitLab app node A]
  LB --> APP2[GitLab app node B]
  APP1 --> PG[(HA PostgreSQL)]
  APP2 --> PG
  APP1 --> REDIS[(HA Redis)]
  APP2 --> REDIS
  APP1 --> GIT[Gitaly / Praefect]
  APP2 --> GIT
  PRIMARY[Primary site] -. replication .-> GEO[Geo secondary site]
  BACKUP[Independent backups] --> VAULT[Separate recovery copy]
  METRICS[Capacity / RPS / latency] --> LB

7. Tabletop DR: calculate evidence before promotion

python - <<'PY'
import json
x=json.load(open('topology/resilience.json'))
lag=x['site_b_geo_secondary']['replication_lag_seconds']
rpo=x.get('rpo_seconds', 60*60)
capacity_ok=x['site_b_geo_secondary']['tested_capacity_rps'] >= x['site_a']['peak_rps']
print('geo_lag_seconds=',lag)
print('within_rpo=',lag <= rpo)
print('secondary_capacity_ok=',capacity_ok)
print('restore_test=',x['independent_backup']['restore_test'])
assert capacity_ok
assert x['independent_backup']['restore_test']=='PASS'
PY

A real Geo promotion also requires current Geo health/replication checks, configuration and secret parity, DNS/load-balancer/certificate planning, application dependency checks, and awareness that unreplicated data can be lost. “Secondary is online” is not enough.

8. Small challenge: choose the control

For each situation, choose required-stop planning, HA, Geo, backup/restore, or capacity scaling: (1) a Rails node dies; (2) the region is unavailable; (3) an operator deletes important data and the deletion replicates; (4) p95 latency rises as API RPS doubles; (5) an upgrade from 18.10 to 19.3 is planned. Explain why no single control solves all five.

9. Cleanup and retained evidence

# Keep only sanitized evidence if desired; remove the disposable lab otherwise.
cd ..
tar -czf ch31-upgrade-evidence.tar.gz ch31-upgrade-lab/evidence ch31-upgrade-lab/topology
sha256sum ch31-upgrade-evidence.tar.gz
rm -rf ch31-upgrade-lab

Knowledge check

Why does the normal fixture path include 18.11 and 19.2?

Why does the zero-downtime fixture also include 19.0 and 19.1?

What should happen when the migration gate reports active?

If Geo lag is inside RPO, is promotion automatically safe?

What is the purpose of the independent backup when Geo exists?

10. Lesson summary and bridge

You can now build an upgrade path, stop on migration evidence, and reason about HA, Geo, backup, and capacity separately. Lesson 3 turns these mechanics into architecture decisions and tradeoffs.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Upgrade paths, required stops, supported infrastructure, migration behavior, Geo procedures, and reference architectures change over time. Recalculate the path and re-read every version-specific note immediately before a real upgrade.

Next lesson

Configuration, Design Choices, and Tradeoffs

Choose upgrade, availability, DR, and capacity strategies with rollback and cost implications explicit.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.