Restore into isolation, verify counts/checksums/invariants, detect a point-in-time gap, reconcile against retained evidence, and prove the corrected state before reconnecting traffic.
Test Restores, Integrity Checks, Reconciliation, and Repair After Partial Recovery
The chapter moves from recovery intent to proof: AtlasMart restores into isolation, injects defects, detects them with checksums/counts/versions, and repairs from retained evidence.
Restore a backup into isolation before considering it production-ready.
Verify counts, hashes/checksums, versions, and domain invariants rather than trusting a successful restore command.
Detect and repair missing/corrupt state using surviving authoritative logs/streams or replicas under an explicit reconciliation policy.
Treat restore testing as recurring evidence of recoverability.
1. A backup that has never been restored is an unproven hypothesis
A successful backup job proves that some artifacts were written; it does not prove that they are complete, decryptable, compatible, internally consistent, or fast enough to restore. Recovery engineering therefore includes scheduled restores into an isolated environment where validation cannot accidentally overwrite surviving production state.
2. Restore validation layers
Start with artifact integrity (manifest hashes, archive continuity, expected object counts), then database/storage integrity, then logical invariants. Useful checks include row/document counts by partition/time range, checksums over canonicalized records, primary-key uniqueness, version monotonicity, foreign/business relationships where applicable, and known sentinel records. A clean checksum proves equality to a known manifest; it does not prove the manifest represented a correct business state.
| Layer | Evidence | Detects |
|---|---|---|
| Artifact | hash/size/archive chain | missing or altered backup files |
| Engine | database-native integrity checks | structural corruption detectable by engine |
| Logical | counts/versions/invariants | missing, stale, semantically invalid state |
| Cross-system | stream/reconciliation comparison | partial PITR gaps and projection divergence |
3. Reconciliation and repair after partial recovery
If the restore target intentionally predates some surviving change stream, replay can move it forward. If one replica has newer data, do not blindly copy “newest” bytes without understanding versions/conflicts. Wide-column systems may need repair to reconcile replicas; relational systems may use WAL/PITR or logical audit/event sources; derived stores may be dropped and rebuilt entirely. Repair authority must be explicit.
Reconstruct and validate in isolation, record the source of each repaired record/range, then use a controlled cutover. Otherwise the repair process itself can overwrite the best surviving evidence.
4. AtlasMart lab: inject a broken restore, detect it, repair it
import hashlib, json
from copy import deepcopy
source = {
"o-1":{"status":"paid","total":80,"version":5},
"o-2":{"status":"shipped","total":45,"version":3},
"o-3":{"status":"created","total":30,"version":2},
}
backup = deepcopy(source)
def digest(rows):
raw=json.dumps(rows, sort_keys=True, separators=(",",":")).encode()
return hashlib.sha256(raw).hexdigest()
manifest={"count":len(backup), "sha256":digest(backup), "max_versions":{k:v["version"] for k,v in backup.items()}}
print("BACKUP MANIFEST", manifest)
print("\nISOLATED RESTORE WITH INJECTED DEFECT")
restored=deepcopy(backup)
del restored["o-3"]
restored["o-2"]["total"] = 4500
print("count ok:", len(restored)==manifest["count"])
print("checksum ok:", digest(restored)==manifest["sha256"])
print("\nRETAINED CHANGE/EVENT EVIDENCE")
events=[
{"id":"e1","order":"o-2","version":3,"state":source["o-2"]},
{"id":"e2","order":"o-3","version":2,"state":source["o-3"]},
]
for e in events:
oid=e["order"]
cur=restored.get(oid)
if cur is None or cur.get("version",-1) <= e["version"]:
restored[oid]=deepcopy(e["state"])
print("repair", oid, "from", e["id"])
print("\nPOST-REPAIR VERIFICATION")
print("count ok:", len(restored)==manifest["count"])
print("checksum ok:", digest(restored)==manifest["sha256"])
print("versions:", {k:v["version"] for k,v in restored.items()})
print("invariant totals_nonnegative:", all(v["total"] >= 0 for v in restored.values()))
print("only now is the restored copy evidence-backed; a backup file existing was not proof of recoverability")
The restored copy initially fails both count and checksum
checks: order o-3 is missing and
o-2 is corrupted. Retained versioned event
evidence repairs both records, after which
count/checksum/version/invariant checks pass.
5. Product-specific integrity/repair examples remain implementation details
PostgreSQL provides native backup/PITR and checksum mechanisms under defined configurations. Cassandra repair compares common token ranges and streams differences; current Cassandra documentation explicitly notes repair is operationally expensive and that incremental repair is not a substitute for every corruption/operator-error scenario. Use each engine’s supported tooling and version-specific recovery compatibility rules.
6. Production judgment
Track last successful restore test, restore duration, checksum coverage, sampled invariant coverage, mismatch counts, repair volume, replay lag, key retrieval, and operator interventions. A restore exercise should fail loudly on bad evidence rather than “complete” with silent drift. The next lesson turns these checks into a timed disaster-recovery game day with ownership and communication.
Check your understanding
- Why is a successful backup job insufficient evidence of recoverability?
- Why restore into an isolated environment first?
- What is the difference between a checksum and a business invariant?
- When can change streams help a partial restore?
- Why should repair authority be explicit?
Review the answers
1. It proves artifacts were produced, not that they are complete, decryptable, compatible, logically correct, or restorable within the RTO.
2. To validate and repair without overwriting surviving production evidence or accidentally serving unverified state.
3. A checksum can prove byte/logical equality to a reference; an invariant checks whether the state obeys domain correctness rules.
4. When retained ordered/versioned changes after the restore point can be safely replayed to close a known gap.
5. Blindly choosing a replica/event as authoritative can propagate stale or corrupted state; operators need defined version/conflict rules.
References
Foundational claims use standards, specifications, primary research, or current official documentation where practical. Product references are optional implementation anchors; the mandatory labs are vendor-neutral.
- PostgreSQL 18 — Backup and Restore — Current restore mechanisms and assumptions for PostgreSQL.
- PostgreSQL 18 — Data Checksums — Official integrity-check context for PostgreSQL data pages.
- Apache Cassandra — Repair — Current repair mechanics, Merkle-tree comparison, streaming, and operational cautions.
- SQLite PRAGMA integrity_check — Example of an engine-native structural integrity check for local SQLite databases.