Checkpoint Lab — Self-Managed Administration: Configuration, Email, Object Storage, Backups, Restore, and Maintenance
Prove a disposable recovery workflow from inventory and RPO/RTO through backup, separate-target restore verification, integrity checks, evidence retention, and complete cleanup.
Learning objectives
- Define RPO/RTO and recovery scope before execution.
- Predict source, backup, target, and secret/storage state transitions before changing them.
- Produce and verify a complete synthetic backup set with separate prerequisites.
- Restore into a separate disposable target and validate representative identities independently.
- Clean up lab data/secrets and retain only sanitized evidence suitable for an operational review.
1. Checkpoint scenario and acceptance criteria
You are proving that the fictional
platform-lab Self-Managed service can be recovered. The
mandatory path is entirely local and synthetic. Passing means you
can reconstruct representative application state from a documented
backup set while proving exact version/type, separate-secret
custody, storage scope, and verification. Simply producing a tar
file is a failed checkpoint.
2. Preflight: write the objective before the backup
mkdir -p ch30-checkpoint/{source,control,evidence,target}
cat > ch30-checkpoint/control/objective.json <<'JSON'
{
"source_instance": "gitlab.example.invalid",
"gitlab_version": "19.3.x-ee",
"rpo_minutes": 60,
"rto_minutes": 180,
"restore_target": "separate-disposable-target",
"required_objects": ["repository", "database-record", "upload", "artifact", "registry-digest"],
"separate_prerequisites": ["configuration", "encryption-secrets", "object-storage-backup"]
}
JSON
jq . ch30-checkpoint/control/objective.json
Prediction 1: creating the application backup must
not modify source data. Prediction 2: restoring
into target/ must reproduce the source identities while
leaving source/ untouched.
Prediction 3: deleting the target after evidence
capture must not delete retained backup/control evidence.
3. Create representative source state
set -eu
cd ch30-checkpoint
mkdir -p source/{repository,database,uploads,artifacts,registry}
printf 'HEAD=cccccccccccccccccccccccccccccccccccccccc\n' > source/repository/identity.txt
printf '{"project":"platform-lab/recovery","members":3,"issues":1}\n' > source/database/project.json
printf 'upload-30\n' > source/uploads/readme.txt
printf 'artifact-30\n' > source/artifacts/report.txt
printf 'sha256:dddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddddd\n' > source/registry/digest.txt
find source -type f -print0 | sort -z | xargs -0 sha256sum > evidence/source.sha256
cat evidence/source.sha256
4. Document backup scope and separate prerequisites
cat > control/backup-scope.json <<'JSON'
{
"application_backup": ["repository", "database", "uploads", "artifacts", "registry-marker"],
"not_inside_application_archive": ["gitlab configuration", "encryption secrets", "external object storage"],
"repository_storages": ["default"],
"secret_values_recorded_here": false
}
JSON
printf 'config_backup_id=config-20260822-01\nsecrets_backup_id=secrets-20260822-01\nobject_store_backup_id=objects-20260822-01\n' > control/prerequisite-ids.txt
jq . control/backup-scope.json
cat control/prerequisite-ids.txt
5. Create and verify the backup set
tar -czf evidence/application.tar.gz source
sha256sum evidence/application.tar.gz > evidence/application.tar.gz.sha256
sha256sum control/objective.json control/backup-scope.json control/prerequisite-ids.txt > evidence/control.sha256
sha256sum -c evidence/application.tar.gz.sha256
sha256sum -c evidence/control.sha256
Record backup start/end timestamps in a real drill and compare them to the RPO. For a live Linux-package exercise, create the application backup with the official tool and configuration backup separately; for external object storage, record the provider-native backup/version/snapshot ID without copying access keys into evidence.
6. Inject one broken prerequisite and diagnose it
Before restore, simulate a target version mismatch. The restore gate must fail before extraction.
SOURCE_VERSION=19.3.x-ee
TARGET_VERSION=19.2.x-ee
if [ "$SOURCE_VERSION" != "$TARGET_VERSION" ]; then
printf 'version_gate=DENIED source=%s target=%s\n' "$SOURCE_VERSION" "$TARGET_VERSION" | tee evidence/version-gate.txt
else
echo 'unexpected pass' >&2; exit 1
fi
# Repair the fixture target metadata only after preserving the denial.
TARGET_VERSION="$SOURCE_VERSION"
[ "$SOURCE_VERSION" = "$TARGET_VERSION" ]
7. Restore to the separate target
sha256sum -c evidence/application.tar.gz.sha256
rm -rf target/*
tar -xzf evidence/application.tar.gz -C target
# The archive contains a source/ top-level directory.
test -f target/source/repository/identity.txt
test -f target/source/database/project.json
test -f target/source/uploads/readme.txt
test -f target/source/artifacts/report.txt
test -f target/source/registry/digest.txt
8. Independent recovery verification
Do not verify only “files exist.” Verify exact representative identity: commit SHA marker, database record, upload/artifact content hash, and image digest marker.
find target/source -type f -print0 | sort -z | xargs -0 sha256sum > evidence/target.sha256
# Normalize root prefix and compare to source hashes.
sed 's# target/source/# source/#' evidence/target.sha256 > evidence/target.normalized.sha256
diff -u evidence/source.sha256 evidence/target.normalized.sha256
jq -e '.project == "platform-lab/recovery" and .members == 3' target/source/database/project.json >/dev/null
grep -Eq '^HEAD=[0-9a-f]{40}$' target/source/repository/identity.txt
grep -Eq '^sha256:[0-9a-f]{64}$' target/source/registry/digest.txt
echo PASS > evidence/recovery-result.txt
9. Optional Self-Managed extension
If you have an already-provisioned disposable Linux-package source
and fresh target, repeat the same logic with real GitLab tools. The
checkpoint still requires exact version/type, separately restored
secrets/configuration, external object-store coverage, a fresh
target, post-restore gitlab:check SANITIZE=true,
gitlab:doctor:secrets, and representative
project/repository/upload/artifact checks. Do not use production.
10. Rollback and deletion consequences
This checkpoint does not “roll back” the source: the source was never overwritten. In production, document the decision point at which a restore becomes authoritative, what data written after the chosen RPO is intentionally lost, how DNS/load balancer traffic changes, and how you prevent clients from writing to both old and restored instances.
11. Cleanup and residual-access verification
# Remove only disposable reconstructed state.
rm -rf target
# No secrets should exist in retained evidence/control.
if grep -RniE '(BEGIN .*PRIVATE KEY|smtp_password|aws_secret_access_key|PRIVATE-TOKEN|glpat-[A-Za-z0-9_-]+)' control evidence; then
echo 'FAIL: sensitive-looking value retained' >&2; exit 1
fi
# Prove source fixture still exists and target is gone.
test -f source/repository/identity.txt
test ! -e target
printf 'cleanup=PASS\n' >> evidence/recovery-result.txt
cat evidence/recovery-result.txt
12. Operational handoff package
A production-grade recovery record should retain: exact GitLab version/type, installation/topology inventory, RPO/RTO, backup IDs/hashes/timestamps, config/secrets/object-storage backup references, repository storage names, restore target identity, health/integrity test results, representative resource hashes/SHAs/digests, failures encountered, corrective actions, actual recovery duration, and cleanup/traffic-cutover decision. Store sensitive secret values separately.
Knowledge check
What makes this checkpoint a restore test rather than a backup test?
It reconstructs state on a separate target and independently compares representative identities to the source.
Why is the version-gate failure preserved in evidence?
It proves the runbook stops on an incompatible target instead of bypassing a safety requirement.
What is the difference between RPO and RTO?
RPO limits acceptable data loss; RTO targets elapsed time to restore verified service.
Should a restore drill retain the actual encryption secret in its evidence bundle?
No. Retain only a reference/backup ID and verification result; secret material stays in its protected secret backup location.
What does Chapter 30 add to the production operating model?
Explicit platform ownership, configuration/secrets inventory, storage boundaries, recovery objectives, tested backup/restore, health/integrity verification, and safe maintenance discipline.
What comes next?
Chapter 31 expands from single-instance recovery into upgrades, zero-downtime planning, high availability, Geo, disaster recovery, and capacity.
13. Chapter close
You can now distinguish data protection from recoverability and run a safe evidence-driven restore drill. Carry that discipline into Chapter 31, where topology, upgrades, HA, Geo, DR, and capacity make recovery coordination even more demanding.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22 and GitLab 19.3. Self-Managed commands and file locations depend on installation method. Re-check the documentation for the exact version, topology, package/chart, and storage architecture before production administration or recovery.
- GitLab 19.3 release
- Administer GitLab
- Configure GitLab
- Back up GitLab
- Restore GitLab
- Linux package backup configuration
- Docker backup
- Helm chart backup and restore
- Object storage
- SMTP settings
- Encrypted configuration
- Health check
- Maintenance Mode
- Maintenance Rake tasks
- Integrity check Rake tasks
- Repository checks
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.