Upgrades, Zero-Downtime Planning, High Availability, Geo, Disaster Recovery, and Capacity: Guided Hands-On Workflow and Core Operations
Use Free-compatible fixtures to calculate a safe upgrade path, gate each stop on migrations and health, model HA/Geo topology, and run a tabletop failover and recovery exercise.
Learning objectives
- Calculate a current required-stop path without hard-coding future patch versions.
- Gate each simulated upgrade step on migration and health evidence.
- Contrast a normal downtime path with the stricter zero-downtime minor-by-minor path.
- Model HA and Geo topology without requiring enterprise infrastructure.
- Run a tabletop failover and record RPO/RTO and capacity evidence.
1. Disposable scenario and preflight
The fictional source is a Self-Managed
18.10.x-ee instance whose target is
19.3.x-ee. The mandatory lab never installs GitLab.
Instead, it creates a small evidence bundle that validates the
decision logic: stops, migrations, health, rollback
readiness, HA failure domains, Geo lag, and recovery objectives.
2. Build the upgrade evidence fixture
set -eu
rm -rf ch31-upgrade-lab
mkdir -p ch31-upgrade-lab/{evidence,backup,topology}
cd ch31-upgrade-lab
cat > evidence/instance.json <<'JSON'
{
"offering": "Self-Managed",
"edition": "EE",
"installation": "Linux package fixture",
"current": "18.10.x",
"target": "19.3.x",
"rpo_minutes": 60,
"rto_minutes": 180,
"peak_rps": 45
}
JSON
printf 'backup-id=fixture-20260822
restore-test=PASS
secrets-copy=PASS
' > backup/rollback-evidence.txt
sha256sum evidence/instance.json backup/rollback-evidence.txt > evidence/preflight.sha256
sha256sum -c evidence/preflight.sha256
3. Calculate two different paths: downtime and zero-downtime
For the current documentation snapshot, a normal path from 18.10 to 19.3 must include the 18.11 and 19.2 required stops. By contrast, zero-downtime cannot skip minors, so it must traverse each minor release sequentially.
cat > evidence/upgrade-plan.json <<'JSON'
{
"normal_downtime": ["18.11.latest", "19.2.latest", "19.3.latest"],
"zero_downtime": ["18.11.latest", "19.0.latest", "19.1.latest", "19.2.latest", "19.3.latest"],
"gate_after_each_step": ["background_migrations_finished", "health_pass", "sample_data_pass"],
"patch_rule": "resolve latest available patch immediately before execution"
}
JSON
jq . evidence/upgrade-plan.json
Why does the normal path not list every minor? Required-stop planning and zero-downtime sequencing are different rules. The standard path must include required stops; the zero-downtime procedure additionally requires one minor at a time.
4. Model the migration gate and halt behavior
cat > evidence/19.2-gate.json <<'JSON'
{"version":"19.2.latest","background_migrations":"active","readiness":"pass","sample_data":"pass"}
JSON
python - <<'PY'
import json, sys
s=json.load(open('evidence/19.2-gate.json'))
if s['background_migrations'] != 'finished':
print('HALT: background migrations are not finished')
sys.exit(42)
PY
The expected exit code is 42. Preserve it as evidence.
Then repair the fixture rather than bypassing the gate:
python - <<'PY'
import json
p='evidence/19.2-gate.json'
s=json.load(open(p)); s['background_migrations']='finished'
open(p,'w').write(json.dumps(s,indent=2)+'\n')
PY
python - <<'PY'
import json
s=json.load(open('evidence/19.2-gate.json'))
assert s['background_migrations']=='finished'
assert s['readiness']=='pass'
assert s['sample_data']=='pass'
print('19.2 gate=PASS; continuation allowed')
PY
5. Optional read-only live preflight
If you own an isolated disposable Self-Managed instance, inspect rather than mutate first:
# OPTIONAL: current Self-Managed instance only.
sudo gitlab-rake gitlab:env:info
sudo gitlab-rake gitlab:check SANITIZE=true
sudo gitlab-rake gitlab:doctor:secrets
sudo gitlab-rake gitlab:background_migrations:list
curl -fsS https://gitlab.example.invalid/-/readiness?all=1
# Compare exact current/target versions with the current upgrade path and upgrade notes.
An actual upgrade is optional and deliberately omitted from the mandatory path because package/chart commands, OS support, exact patch versions, optional services, and migration duration are environment-specific. If you perform one, do it only on a disposable clone and preserve pre/post evidence.
6. Model HA and Geo as separate failure domains
cat > topology/resilience.json <<'JSON'
{
"site_a": {
"app_nodes": 2,
"database": "HA conceptual",
"redis": "HA conceptual",
"gitaly": "redundant conceptual",
"peak_rps": 45
},
"site_b_geo_secondary": {
"tier": "Premium/Ultimate optional",
"replication_lag_seconds": 75,
"tested_capacity_rps": 55
},
"independent_backup": {
"age_minutes": 35,
"restore_test": "PASS"
}
}
JSON
jq . topology/resilience.json
The HA model handles an app-node failure inside Site A. The Geo secondary models regional DR. The independent backup models recovery from destructive/corrupt state. These are deliberately three separate controls.
flowchart LR USERS[Users / CI / API] --> LB[Load balancer] LB --> APP1[GitLab app node A] LB --> APP2[GitLab app node B] APP1 --> PG[(HA PostgreSQL)] APP2 --> PG APP1 --> REDIS[(HA Redis)] APP2 --> REDIS APP1 --> GIT[Gitaly / Praefect] APP2 --> GIT PRIMARY[Primary site] -. replication .-> GEO[Geo secondary site] BACKUP[Independent backups] --> VAULT[Separate recovery copy] METRICS[Capacity / RPS / latency] --> LB
7. Tabletop DR: calculate evidence before promotion
python - <<'PY'
import json
x=json.load(open('topology/resilience.json'))
lag=x['site_b_geo_secondary']['replication_lag_seconds']
rpo=x.get('rpo_seconds', 60*60)
capacity_ok=x['site_b_geo_secondary']['tested_capacity_rps'] >= x['site_a']['peak_rps']
print('geo_lag_seconds=',lag)
print('within_rpo=',lag <= rpo)
print('secondary_capacity_ok=',capacity_ok)
print('restore_test=',x['independent_backup']['restore_test'])
assert capacity_ok
assert x['independent_backup']['restore_test']=='PASS'
PY
A real Geo promotion also requires current Geo health/replication checks, configuration and secret parity, DNS/load-balancer/certificate planning, application dependency checks, and awareness that unreplicated data can be lost. “Secondary is online” is not enough.
8. Small challenge: choose the control
For each situation, choose required-stop planning, HA, Geo, backup/restore, or capacity scaling: (1) a Rails node dies; (2) the region is unavailable; (3) an operator deletes important data and the deletion replicates; (4) p95 latency rises as API RPS doubles; (5) an upgrade from 18.10 to 19.3 is planned. Explain why no single control solves all five.
9. Cleanup and retained evidence
# Keep only sanitized evidence if desired; remove the disposable lab otherwise.
cd ..
tar -czf ch31-upgrade-evidence.tar.gz ch31-upgrade-lab/evidence ch31-upgrade-lab/topology
sha256sum ch31-upgrade-evidence.tar.gz
rm -rf ch31-upgrade-lab
Knowledge check
Why does the normal fixture path include 18.11 and 19.2?
Those are required upgrade stops between the fictional source 18.10 and target 19.3 in the current documentation.
Why does the zero-downtime fixture also include 19.0 and 19.1?
Zero-downtime requires upgrading one minor release at a time, even when those minors are not general required stops.
What should happen when the migration gate reports active?
Halt. Diagnose and let/fix the migrations until finished, then re-run the gate before continuing.
If Geo lag is inside RPO, is promotion automatically safe?
No. Also verify Geo health, capacity, configuration/secrets, dependencies, network/DNS cutover, and the current failover procedure.
What is the purpose of the independent backup when Geo exists?
It gives historical recovery from deletion/corruption or other state that replication may faithfully copy to the secondary.
10. Lesson summary and bridge
You can now build an upgrade path, stop on migration evidence, and reason about HA, Geo, backup, and capacity separately. Lesson 3 turns these mechanics into architecture decisions and tradeoffs.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Upgrade paths, required stops, supported infrastructure, migration behavior, Geo procedures, and reference architectures change over time. Recalculate the path and re-read every version-specific note immediately before a real upgrade.
- GitLab 19.3 release
- Upgrade GitLab
- Before you upgrade
- Plan your upgrade path
- GitLab 19 upgrade notes
- Background migrations
- Upgrade multi-node with downtime
- Upgrade multi-node with zero downtime
- Roll back earlier GitLab versions
- Reference architectures
- 1K / 20 RPS reference architecture
- 2K / 40 RPS reference architecture
- 3K / 60 RPS HA reference architecture
- GitLab Geo
- Set up Geo
- Geo disaster recovery
- Restore GitLab
- Health checks
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.