Checkpoint Lab — Upgrades, Zero-Downtime Planning, High Availability, Geo, Disaster Recovery, and Capacity
Produce and validate a complete upgrade/DR runbook with current version checkpoints, injected gate failures, RPO/RTO, capacity signals, recovery evidence, and safe cleanup.
Learning objectives
- Produce a version-aware upgrade and DR runbook from explicit assumptions.
- Prove the runbook halts when a migration or health gate fails.
- Define and verify RPO/RTO with backup, Geo, and capacity evidence.
- Separate HA failover, Geo promotion, and restore as different recovery paths.
- Retain sanitized operational evidence and remove disposable lab state.
1. Checkpoint scenario and acceptance criteria
You are preparing a fictional Self-Managed instance for an upgrade
from 18.10.x-ee to 19.3.x-ee. Production
needs are: a one-hour RPO, a three-hour RTO, measured peak load of
45 RPS, and a regional DR requirement. The mandatory exercise is a
complete fixture/tabletop runbook. An actual Self-Managed upgrade or
Geo deployment remains optional.
2. Setup and preflight evidence
set -eu
rm -rf ch31-checkpoint
mkdir -p ch31-checkpoint/{evidence,restore,dr}
cd ch31-checkpoint
cat > evidence/baseline.json <<'JSON'
{
"source":"18.10.x-ee",
"target":"19.3.x-ee",
"installation":"Self-Managed Linux-package fixture",
"normal_required_stops":["18.11.latest","19.2.latest","19.3.latest"],
"rpo_seconds":3600,
"rto_seconds":10800,
"peak_rps":45,
"backup_restore_test":"PASS",
"secrets_reference":"vault://fixture-not-a-secret"
}
JSON
printf 'project=platform/sample\ncommit=aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa\n' > restore/repository.txt
printf 'issue_count=4\nnamespace=platform\n' > restore/database.txt
sha256sum evidence/baseline.json restore/repository.txt restore/database.txt > evidence/baseline.sha256
sha256sum -c evidence/baseline.sha256
3. Predict before action
Write the predictions before running the gates: (1) the first planned normal-path stop after 18.10 is 18.11; (2) the path must reach 19.2 before 19.3; (3) an unfinished migration must halt; (4) a DR secondary with less than 45-RPS tested capacity fails the capacity gate even if replication is healthy.
cat > evidence/predictions.txt <<'EOF'
P1 first_stop=18.11.latest
P2 required_before_target=19.2.latest
P3 active_migration=HALT
P4 secondary_capacity_below_45=HALT
EOF
cat evidence/predictions.txt
4. Validate the current documented path fixture
python - <<'PY'
import json
b=json.load(open('evidence/baseline.json'))
p=b['normal_required_stops']
assert p[0]=='18.11.latest'
assert '19.2.latest' in p
assert p[-1]=='19.3.latest'
print('path_fixture=PASS')
print('IMPORTANT: resolve each .latest to the current latest patch on execution day')
PY
5. Inject a failed migration/health gate and halt safely
cat > evidence/stop-19.2.json <<'JSON'
{
"version":"19.2.latest",
"background_migrations":"active",
"gitlab_check":"PASS",
"doctor_secrets":"PASS",
"readiness":"PASS",
"sample_repository":"PASS"
}
JSON
set +e
python - <<'PY'
import json,sys
s=json.load(open('evidence/stop-19.2.json'))
required={
'background_migrations': 'finished',
'gitlab_check':'PASS',
'doctor_secrets':'PASS',
'readiness':'PASS',
'sample_repository':'PASS'
}
failed=[k for k,v in required.items() if s.get(k)!=v]
print('failed_gates=',','.join(failed) if failed else 'none')
sys.exit(31 if failed else 0)
PY
rc=$?
set -e
printf 'first_gate_exit=%s\n' "$rc" | tee evidence/failed-gate.txt
test "$rc" -eq 31
Expected result: the runbook halts and preserves
background_migrations as the cause. Now simulate the
actual remediation outcome—migration finished—then rerun:
python - <<'PY'
import json
p='evidence/stop-19.2.json'; s=json.load(open(p)); s['background_migrations']='finished'
open(p,'w').write(json.dumps(s,indent=2)+'\n')
PY
python - <<'PY'
import json
s=json.load(open('evidence/stop-19.2.json'))
assert s['background_migrations']=='finished'
for k in ['gitlab_check','doctor_secrets','readiness','sample_repository']:
assert s[k]=='PASS'
print('19.2 continuation gate=PASS')
PY
6. Verify rollback readiness before target progression
The checkpoint does not execute a rollback, but it verifies the evidence required to make one credible: exact previous version/edition, pre-upgrade backup ID, configuration/secrets backup reference, tested restore, and documented data-loss point. Current GitLab guidance requires restoring a compatible previous-version database after migrations if rollback is needed.
cat > evidence/rollback.json <<'JSON'
{
"previous_version":"18.11.latest-ee",
"backup_id":"fixture-pre-upgrade",
"configuration_backup":"present",
"secrets_backup":"present",
"restore_rehearsal":"PASS",
"decision":"rollback only if post-upgrade defects cannot be safely rolled forward"
}
JSON
jq . evidence/rollback.json
7. DR tabletop: RPO, capacity, and independent backup
cat > dr/state.json <<'JSON'
{
"geo_tier":"Premium/Ultimate optional",
"replication_lag_seconds":90,
"secondary_tested_rps":40,
"primary_peak_rps":45,
"backup_age_seconds":2100,
"backup_restore_test":"PASS"
}
JSON
set +e
python - <<'PY'
import json,sys
b=json.load(open('evidence/baseline.json')); d=json.load(open('dr/state.json'))
checks={
'geo_inside_rpo': d['replication_lag_seconds'] <= b['rpo_seconds'],
'backup_inside_rpo': d['backup_age_seconds'] <= b['rpo_seconds'],
'backup_restorable': d['backup_restore_test']=='PASS',
'secondary_capacity': d['secondary_tested_rps'] >= d['primary_peak_rps']
}
for k,v in checks.items(): print(k,'PASS' if v else 'FAIL')
sys.exit(32 if not all(checks.values()) else 0)
PY
rc=$?
set -e
printf 'dr_gate_exit=%s\n' "$rc" | tee evidence/dr-failure.txt
test "$rc" -eq 32
The expected failure is secondary capacity. This is important: replication health alone is not sufficient for a viable failover target. Repair the design by proving more capacity, not by weakening the assertion:
python - <<'PY'
import json
p='dr/state.json'; d=json.load(open(p)); d['secondary_tested_rps']=55
open(p,'w').write(json.dumps(d,indent=2)+'\n')
PY
python - <<'PY'
import json
b=json.load(open('evidence/baseline.json')); d=json.load(open('dr/state.json'))
assert d['replication_lag_seconds'] <= b['rpo_seconds']
assert d['backup_age_seconds'] <= b['rpo_seconds']
assert d['backup_restore_test']=='PASS'
assert d['secondary_tested_rps'] >= d['primary_peak_rps']
print('DR readiness fixture=PASS')
PY
8. RTO is the entire recovery timeline
cat > evidence/rto-seconds.json <<'JSON'
{
"detect_and_decide":900,
"promotion_or_restore":3600,
"dns_and_dependency_cutover":900,
"verification":1200
}
JSON
python - <<'PY'
import json
b=json.load(open('evidence/baseline.json')); x=json.load(open('evidence/rto-seconds.json'))
total=sum(x.values())
print('modeled_rto_seconds=',total)
print('objective_seconds=',b['rto_seconds'])
assert total <= b['rto_seconds']
print('modeled_rto=PASS')
PY
For a real exercise, replace modeled times with measured timestamps from incident declaration through verified service and stakeholder sign-off.
9. Optional live runbook checks
# OPTIONAL only on a disposable Self-Managed instance you own.
sudo gitlab-rake gitlab:env:info
sudo gitlab-rake gitlab:check SANITIZE=true
sudo gitlab-rake gitlab:doctor:secrets
sudo gitlab-rake gitlab:background_migrations:list
curl -fsS https://gitlab.example.invalid/-/readiness?all=1
# If Geo is configured (Premium/Ultimate), use the current Geo check/status procedures.
# Do not promote a real secondary merely to complete this course.
10. Final verification checklist
- Current and target version/edition are explicit.
- Required stops were recalculated from current docs.
- Every stop has migration, health, secrets, and sample-data gates.
- Rollback names the exact prior version/backup and has a proven restore.
- HA failure domains are enumerated, including shared stateful dependencies.
- Geo is treated as optional Premium/Ultimate DR, not backup.
- RPO uses actual backup age/replication lag.
- RTO includes decision, restore/promotion, cutover, and verification.
- Secondary/failover capacity has been tested against peak demand.
- No evidence file contains real tokens, passwords, secrets, private keys, or production host data.
11. Cleanup and sanitized handoff
sha256sum evidence/*.json evidence/*.txt dr/state.json > evidence/final-manifest.sha256
# Preserve sanitized runbook evidence if desired.
cd ..
tar -czf ch31-checkpoint-evidence.tar.gz ch31-checkpoint/evidence ch31-checkpoint/dr
sha256sum ch31-checkpoint-evidence.tar.gz
rm -rf ch31-checkpoint
12. What the production handoff must contain
Retain the exact path calculation date, source/target versions, resolved patch versions, upgrade-note links, installation/topology inventory, backup/restore proof, configuration/secrets references, migration status, pre/post health evidence, sample repository/project checks, rollback criteria, actual maintenance timestamps, Geo lag/scope if used, RPO/RTO measurements, capacity signals, N-1/failover load evidence, failures encountered, corrective actions, and final go/no-go decision. Keep sensitive secret material in its protected store, not in the runbook bundle.
Knowledge check
What made the first 19.2 gate fail?
The background migration state was active. The runbook correctly halted before continuing.
Why did the first DR gate fail even though Geo lag was inside RPO?
The secondary was tested below the primary peak RPS, so failover capacity was not proven.
What is the difference between a rollback plan and a package downgrade command?
A rollback plan includes compatible previous code, exact-version database backup/restore, config/secrets, decision criteria, data-loss impact, and verification.
Why does RTO include decision and verification time?
Users do not have a recovered service until the incident is declared, recovery executed, dependencies/cutover completed, and the result independently verified.
What does Chapter 31 add to the production operating model?
Version-aware change control, migration/health gates, rollback evidence, HA/Geo/backup separation, DR objectives, and measured capacity/failover readiness.
What comes next?
Chapter 32 moves from resilience design to audit events, logs, metrics, Prometheus integration, troubleshooting, and operational diagnostics.
13. Chapter close
You can now treat upgrades and resilience as evidence-driven platform operations rather than infrastructure labels. Chapter 32 builds the observability and diagnostic evidence needed to operate these controls continuously.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Upgrade paths, required stops, supported infrastructure, migration behavior, Geo procedures, and reference architectures change over time. Recalculate the path and re-read every version-specific note immediately before a real upgrade.
- GitLab 19.3 release
- Upgrade GitLab
- Before you upgrade
- Plan your upgrade path
- GitLab 19 upgrade notes
- Background migrations
- Upgrade multi-node with downtime
- Upgrade multi-node with zero downtime
- Roll back earlier GitLab versions
- Reference architectures
- 1K / 20 RPS reference architecture
- 2K / 40 RPS reference architecture
- 3K / 60 RPS HA reference architecture
- GitLab Geo
- Set up Geo
- Geo disaster recovery
- Restore GitLab
- Health checks
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.