Chapter 31Lesson 05~420 minutes

Checkpoint Lab — Upgrades, Zero-Downtime Planning, High Availability, Geo, Disaster Recovery, and Capacity

Produce and validate a complete upgrade/DR runbook with current version checkpoints, injected gate failures, RPO/RTO, capacity signals, recovery evidence, and safe cleanup.

CheckpointRunbookDR drillCapacityVerification

Learning objectives

  • Produce a version-aware upgrade and DR runbook from explicit assumptions.
  • Prove the runbook halts when a migration or health gate fails.
  • Define and verify RPO/RTO with backup, Geo, and capacity evidence.
  • Separate HA failover, Geo promotion, and restore as different recovery paths.
  • Retain sanitized operational evidence and remove disposable lab state.
Availability and safety baseline — verified 2026-08-22 against GitLab 19.3. Upgrade planning, required upgrade stops, background-migration checks, rollback planning, and the general reference-architecture guidance apply to GitLab Self-Managed and have Free-compatible learning paths. The documented zero-downtime procedure is available for Self-Managed but requires a properly designed multi-node Linux-package environment with load balancing and HA mechanisms. Geo is Premium/Ultimate on GitLab Self-Managed. The mandatory chapter path uses static fixtures and tabletop exercises; it requires no paid tier, enterprise topology, cloud spend, or real production upgrade.

1. Checkpoint scenario and acceptance criteria

You are preparing a fictional Self-Managed instance for an upgrade from 18.10.x-ee to 19.3.x-ee. Production needs are: a one-hour RPO, a three-hour RTO, measured peak load of 45 RPS, and a regional DR requirement. The mandatory exercise is a complete fixture/tabletop runbook. An actual Self-Managed upgrade or Geo deployment remains optional.

Acceptance rule: the runbook must refuse to continue when a required stop, migration, readiness, restore, or capacity gate fails. A “successful” script that ignores failed gates does not pass this checkpoint.

2. Setup and preflight evidence

set -eu
rm -rf ch31-checkpoint
mkdir -p ch31-checkpoint/{evidence,restore,dr}
cd ch31-checkpoint
cat > evidence/baseline.json <<'JSON'
{
  "source":"18.10.x-ee",
  "target":"19.3.x-ee",
  "installation":"Self-Managed Linux-package fixture",
  "normal_required_stops":["18.11.latest","19.2.latest","19.3.latest"],
  "rpo_seconds":3600,
  "rto_seconds":10800,
  "peak_rps":45,
  "backup_restore_test":"PASS",
  "secrets_reference":"vault://fixture-not-a-secret"
}
JSON
printf 'project=platform/sample\ncommit=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa\n' > restore/repository.txt
printf 'issue_count=4\nnamespace=platform\n' > restore/database.txt
sha256sum evidence/baseline.json restore/repository.txt restore/database.txt > evidence/baseline.sha256
sha256sum -c evidence/baseline.sha256

3. Predict before action

Write the predictions before running the gates: (1) the first planned normal-path stop after 18.10 is 18.11; (2) the path must reach 19.2 before 19.3; (3) an unfinished migration must halt; (4) a DR secondary with less than 45-RPS tested capacity fails the capacity gate even if replication is healthy.

cat > evidence/predictions.txt <<'EOF'
P1 first_stop=18.11.latest
P2 required_before_target=19.2.latest
P3 active_migration=HALT
P4 secondary_capacity_below_45=HALT
EOF
cat evidence/predictions.txt

4. Validate the current documented path fixture

python - <<'PY'
import json
b=json.load(open('evidence/baseline.json'))
p=b['normal_required_stops']
assert p[0]=='18.11.latest'
assert '19.2.latest' in p
assert p[-1]=='19.3.latest'
print('path_fixture=PASS')
print('IMPORTANT: resolve each .latest to the current latest patch on execution day')
PY

5. Inject a failed migration/health gate and halt safely

cat > evidence/stop-19.2.json <<'JSON'
{
  "version":"19.2.latest",
  "background_migrations":"active",
  "gitlab_check":"PASS",
  "doctor_secrets":"PASS",
  "readiness":"PASS",
  "sample_repository":"PASS"
}
JSON
set +e
python - <<'PY'
import json,sys
s=json.load(open('evidence/stop-19.2.json'))
required={
 'background_migrations': 'finished',
 'gitlab_check':'PASS',
 'doctor_secrets':'PASS',
 'readiness':'PASS',
 'sample_repository':'PASS'
}
failed=[k for k,v in required.items() if s.get(k)!=v]
print('failed_gates=',','.join(failed) if failed else 'none')
sys.exit(31 if failed else 0)
PY
rc=$?
set -e
printf 'first_gate_exit=%s\n' "$rc" | tee evidence/failed-gate.txt
test "$rc" -eq 31

Expected result: the runbook halts and preserves background_migrations as the cause. Now simulate the actual remediation outcome—migration finished—then rerun:

python - <<'PY'
import json
p='evidence/stop-19.2.json'; s=json.load(open(p)); s['background_migrations']='finished'
open(p,'w').write(json.dumps(s,indent=2)+'\n')
PY
python - <<'PY'
import json
s=json.load(open('evidence/stop-19.2.json'))
assert s['background_migrations']=='finished'
for k in ['gitlab_check','doctor_secrets','readiness','sample_repository']:
    assert s[k]=='PASS'
print('19.2 continuation gate=PASS')
PY

6. Verify rollback readiness before target progression

The checkpoint does not execute a rollback, but it verifies the evidence required to make one credible: exact previous version/edition, pre-upgrade backup ID, configuration/secrets backup reference, tested restore, and documented data-loss point. Current GitLab guidance requires restoring a compatible previous-version database after migrations if rollback is needed.

cat > evidence/rollback.json <<'JSON'
{
  "previous_version":"18.11.latest-ee",
  "backup_id":"fixture-pre-upgrade",
  "configuration_backup":"present",
  "secrets_backup":"present",
  "restore_rehearsal":"PASS",
  "decision":"rollback only if post-upgrade defects cannot be safely rolled forward"
}
JSON
jq . evidence/rollback.json

7. DR tabletop: RPO, capacity, and independent backup

cat > dr/state.json <<'JSON'
{
  "geo_tier":"Premium/Ultimate optional",
  "replication_lag_seconds":90,
  "secondary_tested_rps":40,
  "primary_peak_rps":45,
  "backup_age_seconds":2100,
  "backup_restore_test":"PASS"
}
JSON
set +e
python - <<'PY'
import json,sys
b=json.load(open('evidence/baseline.json')); d=json.load(open('dr/state.json'))
checks={
 'geo_inside_rpo': d['replication_lag_seconds'] <= b['rpo_seconds'],
 'backup_inside_rpo': d['backup_age_seconds'] <= b['rpo_seconds'],
 'backup_restorable': d['backup_restore_test']=='PASS',
 'secondary_capacity': d['secondary_tested_rps'] >= d['primary_peak_rps']
}
for k,v in checks.items(): print(k,'PASS' if v else 'FAIL')
sys.exit(32 if not all(checks.values()) else 0)
PY
rc=$?
set -e
printf 'dr_gate_exit=%s\n' "$rc" | tee evidence/dr-failure.txt
test "$rc" -eq 32

The expected failure is secondary capacity. This is important: replication health alone is not sufficient for a viable failover target. Repair the design by proving more capacity, not by weakening the assertion:

python - <<'PY'
import json
p='dr/state.json'; d=json.load(open(p)); d['secondary_tested_rps']=55
open(p,'w').write(json.dumps(d,indent=2)+'\n')
PY
python - <<'PY'
import json
b=json.load(open('evidence/baseline.json')); d=json.load(open('dr/state.json'))
assert d['replication_lag_seconds'] <= b['rpo_seconds']
assert d['backup_age_seconds'] <= b['rpo_seconds']
assert d['backup_restore_test']=='PASS'
assert d['secondary_tested_rps'] >= d['primary_peak_rps']
print('DR readiness fixture=PASS')
PY

8. RTO is the entire recovery timeline

cat > evidence/rto-seconds.json <<'JSON'
{
  "detect_and_decide":900,
  "promotion_or_restore":3600,
  "dns_and_dependency_cutover":900,
  "verification":1200
}
JSON
python - <<'PY'
import json
b=json.load(open('evidence/baseline.json')); x=json.load(open('evidence/rto-seconds.json'))
total=sum(x.values())
print('modeled_rto_seconds=',total)
print('objective_seconds=',b['rto_seconds'])
assert total <= b['rto_seconds']
print('modeled_rto=PASS')
PY

For a real exercise, replace modeled times with measured timestamps from incident declaration through verified service and stakeholder sign-off.

9. Optional live runbook checks

# OPTIONAL only on a disposable Self-Managed instance you own.
sudo gitlab-rake gitlab:env:info
sudo gitlab-rake gitlab:check SANITIZE=true
sudo gitlab-rake gitlab:doctor:secrets
sudo gitlab-rake gitlab:background_migrations:list
curl -fsS https://gitlab.example.invalid/-/readiness?all=1
# If Geo is configured (Premium/Ultimate), use the current Geo check/status procedures.
# Do not promote a real secondary merely to complete this course.

10. Final verification checklist

  • Current and target version/edition are explicit.
  • Required stops were recalculated from current docs.
  • Every stop has migration, health, secrets, and sample-data gates.
  • Rollback names the exact prior version/backup and has a proven restore.
  • HA failure domains are enumerated, including shared stateful dependencies.
  • Geo is treated as optional Premium/Ultimate DR, not backup.
  • RPO uses actual backup age/replication lag.
  • RTO includes decision, restore/promotion, cutover, and verification.
  • Secondary/failover capacity has been tested against peak demand.
  • No evidence file contains real tokens, passwords, secrets, private keys, or production host data.

11. Cleanup and sanitized handoff

sha256sum evidence/*.json evidence/*.txt dr/state.json > evidence/final-manifest.sha256
# Preserve sanitized runbook evidence if desired.
cd ..
tar -czf ch31-checkpoint-evidence.tar.gz ch31-checkpoint/evidence ch31-checkpoint/dr
sha256sum ch31-checkpoint-evidence.tar.gz
rm -rf ch31-checkpoint

12. What the production handoff must contain

Retain the exact path calculation date, source/target versions, resolved patch versions, upgrade-note links, installation/topology inventory, backup/restore proof, configuration/secrets references, migration status, pre/post health evidence, sample repository/project checks, rollback criteria, actual maintenance timestamps, Geo lag/scope if used, RPO/RTO measurements, capacity signals, N-1/failover load evidence, failures encountered, corrective actions, and final go/no-go decision. Keep sensitive secret material in its protected store, not in the runbook bundle.

Knowledge check

What made the first 19.2 gate fail?

Why did the first DR gate fail even though Geo lag was inside RPO?

What is the difference between a rollback plan and a package downgrade command?

Why does RTO include decision and verification time?

What does Chapter 31 add to the production operating model?

What comes next?

13. Chapter close

You can now treat upgrades and resilience as evidence-driven platform operations rather than infrastructure labels. Chapter 32 builds the observability and diagnostic evidence needed to operate these controls continuously.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Upgrade paths, required stops, supported infrastructure, migration behavior, Geo procedures, and reference architectures change over time. Recalculate the path and re-read every version-specific note immediately before a real upgrade.

Next chapter

Audit Events, Logs, Metrics, Prometheus Integration, Troubleshooting, and Operational Diagnostics

Use observability and audit evidence to operate the resilient platform continuously.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.