Restore into isolation, verify counts/checksums/invariants, detect a point-in-time gap, reconcile against retained evidence, and prove the corrected state before reconnecting traffic.

Test Restores, Integrity Checks, Reconciliation, and Repair After Partial Recovery

The chapter moves from recovery intent to proof: AtlasMart restores into isolation, injects defects, detects them with checksums/counts/versions, and repairs from retained evidence.

Advanced130–170 minutesRestore-validation and repair labPython 3.13+ · standard libraryVendor-neutral · free/local mandatory pathLast reviewed: August 2026
01

Restore a backup into isolation before considering it production-ready.

02

Verify counts, hashes/checksums, versions, and domain invariants rather than trusting a successful restore command.

03

Detect and repair missing/corrupt state using surviving authoritative logs/streams or replicas under an explicit reconciliation policy.

04

Treat restore testing as recurring evidence of recoverability.

1. A backup that has never been restored is an unproven hypothesis

A successful backup job proves that some artifacts were written; it does not prove that they are complete, decryptable, compatible, internally consistent, or fast enough to restore. Recovery engineering therefore includes scheduled restores into an isolated environment where validation cannot accidentally overwrite surviving production state.

2. Restore validation layers

Start with artifact integrity (manifest hashes, archive continuity, expected object counts), then database/storage integrity, then logical invariants. Useful checks include row/document counts by partition/time range, checksums over canonicalized records, primary-key uniqueness, version monotonicity, foreign/business relationships where applicable, and known sentinel records. A clean checksum proves equality to a known manifest; it does not prove the manifest represented a correct business state.

Layer Evidence Detects
Artifact hash/size/archive chain missing or altered backup files
Engine database-native integrity checks structural corruption detectable by engine
Logical counts/versions/invariants missing, stale, semantically invalid state
Cross-system stream/reconciliation comparison partial PITR gaps and projection divergence

3. Reconciliation and repair after partial recovery

If the restore target intentionally predates some surviving change stream, replay can move it forward. If one replica has newer data, do not blindly copy “newest” bytes without understanding versions/conflicts. Wide-column systems may need repair to reconcile replicas; relational systems may use WAL/PITR or logical audit/event sources; derived stores may be dropped and rebuilt entirely. Repair authority must be explicit.

Do not repair into production first

Reconstruct and validate in isolation, record the source of each repaired record/range, then use a controlled cutover. Otherwise the repair process itself can overwrite the best surviving evidence.

4. AtlasMart lab: inject a broken restore, detect it, repair it

python · AtlasMart deterministic simulation
import hashlib, json
from copy import deepcopy

source = {
    "o-1":{"status":"paid","total":80,"version":5},
    "o-2":{"status":"shipped","total":45,"version":3},
    "o-3":{"status":"created","total":30,"version":2},
}
backup = deepcopy(source)

def digest(rows):
    raw=json.dumps(rows, sort_keys=True, separators=(",",":")).encode()
    return hashlib.sha256(raw).hexdigest()

manifest={"count":len(backup), "sha256":digest(backup), "max_versions":{k:v["version"] for k,v in backup.items()}}
print("BACKUP MANIFEST", manifest)

print("\nISOLATED RESTORE WITH INJECTED DEFECT")
restored=deepcopy(backup)
del restored["o-3"]
restored["o-2"]["total"] = 4500
print("count ok:", len(restored)==manifest["count"])
print("checksum ok:", digest(restored)==manifest["sha256"])

print("\nRETAINED CHANGE/EVENT EVIDENCE")
events=[
    {"id":"e1","order":"o-2","version":3,"state":source["o-2"]},
    {"id":"e2","order":"o-3","version":2,"state":source["o-3"]},
]
for e in events:
    oid=e["order"]
    cur=restored.get(oid)
    if cur is None or cur.get("version",-1) <= e["version"]:
        restored[oid]=deepcopy(e["state"])
        print("repair", oid, "from", e["id"])

print("\nPOST-REPAIR VERIFICATION")
print("count ok:", len(restored)==manifest["count"])
print("checksum ok:", digest(restored)==manifest["sha256"])
print("versions:", {k:v["version"] for k,v in restored.items()})
print("invariant totals_nonnegative:", all(v["total"] >= 0 for v in restored.values()))
print("only now is the restored copy evidence-backed; a backup file existing was not proof of recoverability")
Expected evidence

The restored copy initially fails both count and checksum checks: order o-3 is missing and o-2 is corrupted. Retained versioned event evidence repairs both records, after which count/checksum/version/invariant checks pass.

5. Product-specific integrity/repair examples remain implementation details

PostgreSQL provides native backup/PITR and checksum mechanisms under defined configurations. Cassandra repair compares common token ranges and streams differences; current Cassandra documentation explicitly notes repair is operationally expensive and that incremental repair is not a substitute for every corruption/operator-error scenario. Use each engine’s supported tooling and version-specific recovery compatibility rules.

6. Production judgment

Track last successful restore test, restore duration, checksum coverage, sampled invariant coverage, mismatch counts, repair volume, replay lag, key retrieval, and operator interventions. A restore exercise should fail loudly on bad evidence rather than “complete” with silent drift. The next lesson turns these checks into a timed disaster-recovery game day with ownership and communication.

Check your understanding

  1. Why is a successful backup job insufficient evidence of recoverability?
  2. Why restore into an isolated environment first?
  3. What is the difference between a checksum and a business invariant?
  4. When can change streams help a partial restore?
  5. Why should repair authority be explicit?
Review the answers

1. It proves artifacts were produced, not that they are complete, decryptable, compatible, logically correct, or restorable within the RTO.

2. To validate and repair without overwriting surviving production evidence or accidentally serving unverified state.

3. A checksum can prove byte/logical equality to a reference; an invariant checks whether the state obeys domain correctness rules.

4. When retained ordered/versioned changes after the restore point can be safely replayed to close a known gap.

5. Blindly choosing a replica/event as authoritative can propagate stale or corrupted state; operators need defined version/conflict rules.

References

Foundational claims use standards, specifications, primary research, or current official documentation where practical. Product references are optional implementation anchors; the mandatory labs are vendor-neutral.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.