Chapter 16 · Orchestration, Dependencies, Scheduling, Backfills, SLAs, and Failure Recovery
Partition-Aware Backfills, Historical Reprocessing, Resource Isolation, and Protecting Current Loads
Reprocess only the historical partitions that need repair, isolate backfill resources from current loads, preserve immutable inputs and partition manifests, and prove that unrelated current output remains unchanged.
Learning outcomes
AtlasMart discovers that the certified 2026-09-20 historical publication has the correct 690 USD GMV but an incorrect 435 USD cost total because one line carried a legacy cost of 40 instead of 30. The immutable raw evidence for that partition still produces 425 USD cost and 265 USD gross profit. The operational question is not “can we rebuild history?” It is “how narrowly can we rebuild the affected history while protecting today’s load?”
Define a backfill as governed historical reprocessing over an explicit partition/range rather than as “rerun everything.”
Use immutable input identity and partition manifests to make historical replay reproducible and reviewable.
Separate logical resource pools/concurrency policies so backfills cannot silently starve current freshness work.
Prove a repaired historical partition changes while an unrelated current partition checksum remains unchanged.
Choose backfill scope from the semantic blast radius of the correction instead of from convenient storage boundaries.
Chapter 16 does not change the accepted Chapter 15 analytical state. The current certified AtlasMart sales target remains nine paid order-line facts, seven paid orders, eleven units, 740 USD paid GMV, 450 USD cost-at-sale, and 290 USD gross profit at committed source sequence 206. Orchestration adds run/task state, dependency evidence, partition manifests, SLO measurements, resource-pool labels, and operator notes around that data state. The historical backfill fixture deliberately starts from one corrupted 2026-09-20 publication (cost 435 USD instead of 425 USD) and repairs only that historical partition from immutable raw evidence; the current 2026-09-21 checksum must remain unchanged.
The mandatory labs are synthetic, local, and free. They use
Python 3 standard-library modules and local JSON/filesystem
state; generation-time validation ran with Python 3.13.5.
Logical timestamps are simulated—no real waiting, cluster
scheduler, queue, cloud warehouse, distributed lock, or
resource manager is involved. A labeled
current versus backfill resource
pool proves orchestration intent in the fixture, not
operating-system or cloud compute isolation. Apache Airflow is
referenced only as an optional later-course implementation
example and is not a prerequisite.
1. Backfill is scoped historical execution
A backfill re-executes pipeline logic for one or more historical processing partitions, event-time intervals, or version ranges. A partition-aware backfill carries that scope explicitly through readiness, transform, quality, publication, and certification so unrelated partitions remain untouched. Historical reprocessing is the broader category: it may be a backfill, a restatement, a rebuild from raw, or replay from a change log.
The orchestrator should know both the requested scope and the true semantic dependency radius. If a dimension correction affects all facts after a date, a one-day storage partition may be too narrow. Conversely, a single corrupted partition checksum does not justify rebuilding five years of history.
| Surface | Before backfill | After BF-20260920-0900-R1 |
|---|---|---|
| 2026-09-20 rows/orders/units | 8 / 5 / 10 | 8 / 5 / 10 |
| 2026-09-20 GMV | 690 USD | 690 USD |
| 2026-09-20 cost | 435 USD (wrong) | 425 USD (reconciled) |
| 2026-09-20 gross profit | 255 USD (wrong) | 265 USD |
| 2026-09-21 current controls | 9 / 7 / 11 / 740 / 450 / 290 | unchanged |
| 2026-09-21 current SHA-256 | 02715e… | same 02715e… |
2. Partition identity needs manifests and immutable inputs
The local raw files are named by logical partition and carry the source high sequence that produced them. A production manifest should normally identify source version/range, schema/contract version, file/object set, row/byte counts, checksums, extraction batch, and code/model version. Without those identities, “rerun 2026-09-20” can silently mean a different input than the original run.
{ "run_id": "BF-20260920-0900-R1", "partition": "2026-09-20", "input": "raw/partition=2026-09-20.json", "resource_pool": "backfill", "max_concurrency": 1, "expected_controls": [8, 5, 10, 690, 425, 265], "protect_partitions": ["2026-09-21"]}
In a real platform, a backfill may reference object versions, table snapshots, transaction offsets, or immutable landing batches. The important property is reproducibility, not this file naming convention.
3. Resource isolation protects current freshness
Historical work competes for CPU, memory, I/O, warehouse slots,
locks, object-store bandwidth, and source extraction capacity.
The local fixture labels two pools: current with
logical concurrency 2 and backfill with concurrency
1. That does not enforce OS-level isolation, but it makes the
policy visible: backfill is lower-priority and separately
bounded.
Real controls are engine/orchestrator-specific—separate queues, warehouses, resource groups, worker pools, priorities, quotas, or time windows. Measure whether current-run latency and queue time remain within SLOs while backfill is active. Do not declare isolation because a config file contains two names.
4. Controlled failure: rebuild all history in the production pool
A convenient for date in all_dates loop can
saturate the same warehouse used by the current load. Current
data then misses freshness targets precisely while the team is
“fixing data quality.” Wider reprocessing also increases the
number of partitions that could be changed by a new code bug.
The repair pattern is to calculate affected partitions, freeze/record current hashes, run the smallest backfill set in a bounded resource class, reconcile each output, and compare protected partitions before/after. If an upstream correction changes semantics across many dates, widen scope deliberately and communicate the blast radius.
5. Hands-on: prove current output survives historical repair
Run the integrated lab from Lesson 5, then compare the historical and current publications.
import jsonfrom pathlib import Pathroot = Path("atlasmart_ch16_lab/published")hist = json.loads((root / "partition=2026-09-20.json").read_text())cur = json.loads((root / "partition=2026-09-21.json").read_text())print("historical:", hist["controls"])print("current:", cur["controls"])print("current_sha256:", cur["sha256"])# Expected historical controls: [8, 5, 10, 690, 425, 265]# Expected current controls: [9, 7, 11, 740, 450, 290]
The acceptance summary additionally asserts
current_hash_unchanged_after_backfill = true. That
is the key isolation evidence for this fixture.
6. Production judgment and bridge
Backfills are change operations, not housekeeping. Define approval, scope, input identity, code version, resource budget, downstream consumers, reconciliation, rollback, and certification exactly as you would for a current production load. When a backfill changes historical metrics, notify consumers that old extracts/caches may no longer match.
Lesson 4 turns those operational outcomes into explicit measurements. A fast green run that publishes only 8/9 expected rows is not reliable; freshness, completeness, duration, and downstream availability need separate indicators and objectives.
Knowledge check
Check your understanding
- What is the historical defect in the lab?
- Why is the current partition hash captured before backfill?
- Does a partition always equal a physical database partition?
- What does resource_pool=backfill prove?
- When should backfill scope widen?
Review the answers
1. The 2026-09-20 published cost is 435 USD instead of the raw/reconciled 425 USD, making gross profit 255 instead of 265.
2. To prove the historical repair did not alter the protected current publication.
3. No. Here it is a logical processing/publication scope; physical partitioning is Chapter 17.
4. Only the intended orchestration policy in this local simulation; production isolation must be enforced/measured by the actual platform.
5. When the semantic correction affects more partitions than the initially detected storage/output defect.
Authoritative references
- Python documentation — graphlibStandard-library topological ordering concepts used to explain dependency graphs without requiring an orchestrator product.
- Python documentation — os.replaceLocal atomic file-replacement primitive used by the fixture to demonstrate partition publication without partial output files.
- Python documentation — hashlibDeterministic SHA-256 evidence for replay/backfill comparisons in the local lab.
- Google SRE Book — Service Level ObjectivesFoundational distinction among service indicators/objectives and externally meaningful reliability goals.
- Apache Airflow documentation — DAGsOptional later-course example of a production orchestrator's DAG concept; no Airflow command or installation is required here.
- Apache Airflow documentation — BackfillOptional implementation reference for historical run concepts; Chapter 16 teaches the vendor-neutral semantics first.
- Kimball Group — Dimensional Modeling TechniquesBackground for the grains and dimensional facts whose correctness orchestration must preserve.