Chapter 16 · Orchestration, Dependencies, Scheduling, Backfills, SLAs, and Failure Recovery

Partition-Aware Backfills, Historical Reprocessing, Resource Isolation, and Protecting Current Loads

Reprocess only the historical partitions that need repair, isolate backfill resources from current loads, preserve immutable inputs and partition manifests, and prove that unrelated current output remains unchanged.

Intermediate → Advanced130–160 minutesPartition backfill/isolation labPython 3 stdlib · local/syntheticLast reviewed: September 2026

Learning outcomes

AtlasMart discovers that the certified 2026-09-20 historical publication has the correct 690 USD GMV but an incorrect 435 USD cost total because one line carried a legacy cost of 40 instead of 30. The immutable raw evidence for that partition still produces 425 USD cost and 265 USD gross profit. The operational question is not “can we rebuild history?” It is “how narrowly can we rebuild the affected history while protecting today’s load?”

01

Define a backfill as governed historical reprocessing over an explicit partition/range rather than as “rerun everything.”

02

Use immutable input identity and partition manifests to make historical replay reproducible and reviewable.

03

Separate logical resource pools/concurrency policies so backfills cannot silently starve current freshness work.

04

Prove a repaired historical partition changes while an unrelated current partition checksum remains unchanged.

05

Choose backfill scope from the semantic blast radius of the correction instead of from convenient storage boundaries.

Chapter 16 continuity contract

Chapter 16 does not change the accepted Chapter 15 analytical state. The current certified AtlasMart sales target remains nine paid order-line facts, seven paid orders, eleven units, 740 USD paid GMV, 450 USD cost-at-sale, and 290 USD gross profit at committed source sequence 206. Orchestration adds run/task state, dependency evidence, partition manifests, SLO measurements, resource-pool labels, and operator notes around that data state. The historical backfill fixture deliberately starts from one corrupted 2026-09-20 publication (cost 435 USD instead of 425 USD) and repairs only that historical partition from immutable raw evidence; the current 2026-09-21 checksum must remain unchanged.

Execution and guarantee boundary

The mandatory labs are synthetic, local, and free. They use Python 3 standard-library modules and local JSON/filesystem state; generation-time validation ran with Python 3.13.5. Logical timestamps are simulated—no real waiting, cluster scheduler, queue, cloud warehouse, distributed lock, or resource manager is involved. A labeled current versus backfill resource pool proves orchestration intent in the fixture, not operating-system or cloud compute isolation. Apache Airflow is referenced only as an optional later-course implementation example and is not a prerequisite.

1. Backfill is scoped historical execution

A backfill re-executes pipeline logic for one or more historical processing partitions, event-time intervals, or version ranges. A partition-aware backfill carries that scope explicitly through readiness, transform, quality, publication, and certification so unrelated partitions remain untouched. Historical reprocessing is the broader category: it may be a backfill, a restatement, a rebuild from raw, or replay from a change log.

The orchestrator should know both the requested scope and the true semantic dependency radius. If a dimension correction affects all facts after a date, a one-day storage partition may be too narrow. Conversely, a single corrupted partition checksum does not justify rebuilding five years of history.

Surface Before backfill After BF-20260920-0900-R1
2026-09-20 rows/orders/units 8 / 5 / 10 8 / 5 / 10
2026-09-20 GMV 690 USD 690 USD
2026-09-20 cost 435 USD (wrong) 425 USD (reconciled)
2026-09-20 gross profit 255 USD (wrong) 265 USD
2026-09-21 current controls 9 / 7 / 11 / 740 / 450 / 290 unchanged
2026-09-21 current SHA-256 02715e… same 02715e…

2. Partition identity needs manifests and immutable inputs

The local raw files are named by logical partition and carry the source high sequence that produced them. A production manifest should normally identify source version/range, schema/contract version, file/object set, row/byte counts, checksums, extraction batch, and code/model version. Without those identities, “rerun 2026-09-20” can silently mean a different input than the original run.

backfill_contract.json
{  "run_id": "BF-20260920-0900-R1",  "partition": "2026-09-20",  "input": "raw/partition=2026-09-20.json",  "resource_pool": "backfill",  "max_concurrency": 1,  "expected_controls": [8, 5, 10, 690, 425, 265],  "protect_partitions": ["2026-09-21"]}

In a real platform, a backfill may reference object versions, table snapshots, transaction offsets, or immutable landing batches. The important property is reproducibility, not this file naming convention.

3. Resource isolation protects current freshness

Historical work competes for CPU, memory, I/O, warehouse slots, locks, object-store bandwidth, and source extraction capacity. The local fixture labels two pools: current with logical concurrency 2 and backfill with concurrency 1. That does not enforce OS-level isolation, but it makes the policy visible: backfill is lower-priority and separately bounded.

Real controls are engine/orchestrator-specific—separate queues, warehouses, resource groups, worker pools, priorities, quotas, or time windows. Measure whether current-run latency and queue time remain within SLOs while backfill is active. Do not declare isolation because a config file contains two names.

4. Controlled failure: rebuild all history in the production pool

A convenient for date in all_dates loop can saturate the same warehouse used by the current load. Current data then misses freshness targets precisely while the team is “fixing data quality.” Wider reprocessing also increases the number of partitions that could be changed by a new code bug.

The repair pattern is to calculate affected partitions, freeze/record current hashes, run the smallest backfill set in a bounded resource class, reconcile each output, and compare protected partitions before/after. If an upstream correction changes semantics across many dates, widen scope deliberately and communicate the blast radius.

5. Hands-on: prove current output survives historical repair

Run the integrated lab from Lesson 5, then compare the historical and current publications.

verify_backfill.py
import jsonfrom pathlib import Pathroot = Path("atlasmart_ch16_lab/published")hist = json.loads((root / "partition=2026-09-20.json").read_text())cur = json.loads((root / "partition=2026-09-21.json").read_text())print("historical:", hist["controls"])print("current:", cur["controls"])print("current_sha256:", cur["sha256"])# Expected historical controls: [8, 5, 10, 690, 425, 265]# Expected current controls:    [9, 7, 11, 740, 450, 290]

The acceptance summary additionally asserts current_hash_unchanged_after_backfill = true. That is the key isolation evidence for this fixture.

6. Production judgment and bridge

Backfills are change operations, not housekeeping. Define approval, scope, input identity, code version, resource budget, downstream consumers, reconciliation, rollback, and certification exactly as you would for a current production load. When a backfill changes historical metrics, notify consumers that old extracts/caches may no longer match.

Lesson 4 turns those operational outcomes into explicit measurements. A fast green run that publishes only 8/9 expected rows is not reliable; freshness, completeness, duration, and downstream availability need separate indicators and objectives.

Knowledge check

Check your understanding

  1. What is the historical defect in the lab?
  2. Why is the current partition hash captured before backfill?
  3. Does a partition always equal a physical database partition?
  4. What does resource_pool=backfill prove?
  5. When should backfill scope widen?
Review the answers

1. The 2026-09-20 published cost is 435 USD instead of the raw/reconciled 425 USD, making gross profit 255 instead of 265.

2. To prove the historical repair did not alter the protected current publication.

3. No. Here it is a logical processing/publication scope; physical partitioning is Chapter 17.

4. Only the intended orchestration policy in this local simulation; production isolation must be enforced/measured by the actual platform.

5. When the semantic correction affects more partitions than the initially detected storage/output defect.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.