Chapter 31Lesson 04~350 minutes

Upgrades, Zero-Downtime Planning, High Availability, Geo, Disaster Recovery, and Capacity: Diagnostics, Failure Modes, Security, and Performance

Diagnose skipped upgrade stops, unfinished migrations, untested rollback, hidden HA single points, Geo lag, and capacity exhaustion from preserved evidence before taking corrective action.

DiagnosticsMigrationsGeoFailure modesRecovery

Learning objectives

  • Diagnose an unsupported upgrade jump before changing a real instance.
  • Interpret unfinished background migrations as a hard stop rather than a nuisance warning.
  • Recognize when HA still contains a shared single point of failure.
  • Treat Geo lag and selective replication as data-loss evidence, not cosmetic metrics.
  • Use preserved pre/post evidence to choose the least destructive correction.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. Upgrade planning, required upgrade stops, background-migration checks, rollback planning, and the general reference-architecture guidance apply to GitLab Self-Managed and have Free-compatible learning paths. The documented zero-downtime procedure is available for Self-Managed but requires a properly designed multi-node Linux-package environment with load balancing and HA mechanisms. Geo is Premium/Ultimate on GitLab Self-Managed. The mandatory chapter path uses static fixtures and tabletop exercises; it requires no paid tier, enterprise topology, cloud spend, or real production upgrade.

1. Evidence-first diagnostic sequence

Use the same sequence for upgrade and resilience incidents: preserve evidence → identify exact source/target version and installation/topology → inspect required stops and version notes → inspect migration/health/backup/Geo/capacity state → choose the least destructive correction → verify independently → only then continue. Do not “fix” the symptom by forcing migrations, disabling checks, or promoting a secondary before understanding the state.

2. Failure: required upgrade stop was skipped

Symptom: an operator plans 18.10 → 19.3 as one package action. The plan is invalid before execution because current upgrade-path documentation requires intervening stops. For a normal path, 18.11 and 19.2 are relevant required stops; for zero-downtime, every minor must be traversed.

# Broken planning fixture — no real package command.
current=18.10.x
target=19.3.x
proposed_path=[19.3.x]
result=REJECTED
reason="required stops and migration gates missing"

Repair the runbook, not the server. Resolve latest patches for each stop on execution day and preserve the calculated path.

3. Failure: background migrations are still active

Symptom: the next upgrade attempts to run while an earlier batched migration is active, and GitLab reports that the expected migration has not finished. This is a data-state gate.

# OPTIONAL read-only diagnostic on a current isolated instance.
sudo gitlab-rake gitlab:background_migrations:list
# Expected safe continuation: every migration required by this step is finished/finalized.
# If active/failed: STOP and diagnose. Do not edit status rows to silence the check.

On large installations, background migrations can take substantial time and can compete for database resources. Capacity planning therefore affects upgrade duration. If a migration has genuinely failed, inspect its details/logs and follow the documented retry/finalize path for that specific version.

4. Failure: the rollback path was never tested

A pre-upgrade backup exists, but nobody has restored it. During rollback, the team discovers missing secrets, version mismatch, or unavailable object-storage data. This is exactly why Chapter 30 treated recovery as a test, not an archive. Before the upgrade window, restore representative data into a separate target and record the actual recovery duration.

Destructive boundary: a real rollback can overwrite newer database state with the older backup. Record the data-loss consequence explicitly before choosing rollback.

5. Failure: “HA” still has a shared bottleneck

Two Rails nodes sit behind a load balancer, but both depend on one non-HA PostgreSQL instance. Losing one Rails node is survivable; losing PostgreSQL is not. Adding more front ends does not remove the shared stateful failure. Inventory every dependency and ask whether its failure is independent and whether failover has been tested.

Redundant frontends can still share fragile state
flowchart LR
  USERS[Users / CI / API] --> LB[Load balancer]
  LB --> APP1[GitLab app node A]
  LB --> APP2[GitLab app node B]
  APP1 --> PG[(HA PostgreSQL)]
  APP2 --> PG
  APP1 --> REDIS[(HA Redis)]
  APP2 --> REDIS
  APP1 --> GIT[Gitaly / Praefect]
  APP2 --> GIT
  PRIMARY[Primary site] -. replication .-> GEO[Geo secondary site]
  BACKUP[Independent backups] --> VAULT[Separate recovery copy]
  METRICS[Capacity / RPS / latency] --> LB

6. Failure: Geo is mistaken for a backup or zero-RPO guarantee

A destructive database change is replicated to the secondary, so Geo faithfully preserves the wrong state. In another event, the primary fails while the secondary is 90 seconds behind, so the newest changes are absent after promotion. Both scenarios are consistent with replication semantics. The mitigation is independent backup/history plus monitored replication and an RPO that acknowledges actual lag.

Selective synchronization warning: promoting a Geo secondary that intentionally did not replicate some data can cause permanent loss of that unreplicated data. Promotion decisions must include replication-scope evidence.

7. Failure: capacity headroom disappears during maintenance

Normal traffic uses 70% of available app capacity. During a rolling upgrade one node is drained, pushing the remainder beyond saturation. The upgrade is technically correct but causes latency/errors. Test N-1 capacity and maintenance load before claiming zero-downtime readiness.

8. Intentionally broken gate: preserve the cause, then repair it

set -eu
rm -rf ch31-diagnostic-fixture
mkdir -p ch31-diagnostic-fixture
cd ch31-diagnostic-fixture
cat > state.json <<'JSON'
{
  "required_stop_present": true,
  "background_migrations": "finished",
  "backup_restore_test": "PASS",
  "readiness": "FAIL",
  "sample_repository": "PASS",
  "geo_lag_seconds": 30,
  "rpo_seconds": 3600
}
JSON
python - <<'PY'
import json, sys
s=json.load(open('state.json'))
checks=[
 ('required stop',s['required_stop_present'] is True),
 ('migrations',s['background_migrations']=='finished'),
 ('restore test',s['backup_restore_test']=='PASS'),
 ('readiness',s['readiness']=='PASS'),
 ('repository',s['sample_repository']=='PASS'),
 ('geo inside RPO',s['geo_lag_seconds'] <= s['rpo_seconds']),
]
for name,ok in checks: print(name, 'PASS' if ok else 'FAIL')
if not all(ok for _,ok in checks): sys.exit(23)
PY

The expected exit is 23, and readiness is the preserved cause. Do not change the script to ignore readiness. Diagnose the failed dependency, then update the fixture only after the dependency is actually healthy:

python - <<'PY'
import json
p='state.json'; s=json.load(open(p)); s['readiness']='PASS'
open(p,'w').write(json.dumps(s,indent=2)+'\n')
PY
python - <<'PY'
import json
s=json.load(open('state.json'))
assert s['readiness']=='PASS'
print('gate repaired=PASS')
PY

9. Security, reliability, and performance are causal

Upgrade windows stress databases, Sidekiq, storage, and network. Geo synchronization can compete for bandwidth and storage I/O. Capacity constraints therefore change migration duration and RTO. Security also matters: rollback and Geo require configuration/secrets parity, but evidence bundles must not contain actual secret values. Preserve hashes, backup IDs, versions, timestamps, topology, lag, and sanitized logs instead.

Knowledge check

What is the safest response to an active background migration before the next stop?

Why do two Rails nodes not automatically mean HA?

What does Geo lag tell you?

Why is a passing process-status check insufficient after upgrade?

Why should an intentionally broken gate preserve its original failure?

10. Lesson summary and bridge

You can now diagnose upgrade and resilience failures without bypassing the control that exposed them. Lesson 5 integrates the whole chapter into a version-aware upgrade/DR runbook.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Upgrade paths, required stops, supported infrastructure, migration behavior, Geo procedures, and reference architectures change over time. Recalculate the path and re-read every version-specific note immediately before a real upgrade.

Next lesson

Checkpoint Lab

Produce and validate a version-aware upgrade/DR runbook with injected failures, RPO/RTO, rollback, and capacity evidence.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.