Chapter 16 · Orchestration, Dependencies, Scheduling, Backfills, SLAs, and Failure Recovery
SLAs/SLOs for Freshness, Completeness, Duration, and Downstream Availability
Measure freshness, completeness, duration, and downstream availability with explicit SLIs and fixture-specific SLOs; distinguish task success from data-product correctness and contractual SLAs from engineering objectives.
Learning outcomes
The successful AtlasMart run has every task in SUCCESS, but that alone does not answer the consumer questions: Is the data fresh? Is all expected data present? How long did recovery take? When could the certified product be used downstream? Reliability needs measurable indicators attached to consumer-relevant states.
Distinguish SLI, SLO, and SLA and tie each measurement to an explicit data-product state surface.
Calculate freshness, completeness, run duration, and downstream availability from stated timestamps/counts rather than from invented reliability numbers.
Explain why a successful task graph can violate a data SLO and why one SLO cannot stand in for all consumer expectations.
Use thresholds as fixture-specific policy choices, not universal warehouse tuning values.
Connect SLO breaches to alerting, certification, consumer communication, and recovery priorities.
Chapter 16 does not change the accepted Chapter 15 analytical state. The current certified AtlasMart sales target remains nine paid order-line facts, seven paid orders, eleven units, 740 USD paid GMV, 450 USD cost-at-sale, and 290 USD gross profit at committed source sequence 206. Orchestration adds run/task state, dependency evidence, partition manifests, SLO measurements, resource-pool labels, and operator notes around that data state. The historical backfill fixture deliberately starts from one corrupted 2026-09-20 publication (cost 435 USD instead of 425 USD) and repairs only that historical partition from immutable raw evidence; the current 2026-09-21 checksum must remain unchanged.
The mandatory labs are synthetic, local, and free. They use
Python 3 standard-library modules and local JSON/filesystem
state; generation-time validation ran with Python 3.13.5.
Logical timestamps are simulated—no real waiting, cluster
scheduler, queue, cloud warehouse, distributed lock, or
resource manager is involved. A labeled
current versus backfill resource
pool proves orchestration intent in the fixture, not
operating-system or cloud compute isolation. Apache Airflow is
referenced only as an optional later-course implementation
example and is not a prerequisite.
1. SLI, SLO, and SLA are different artifacts
A service level indicator (SLI) is a measured quantity such as freshness lag or completeness percentage. A service level objective (SLO) is a target/range for an SLI over a defined scope/window. A service level agreement (SLA) is an agreement—often external or contractual—that may include consequences, exclusions, and responsibilities. Teams should not call every dashboard threshold an SLA.
For data products, define the observation surface. “Pipeline duration” measures control-plane execution time. “Freshness” compares source availability/event state to certified data availability. “Downstream availability” asks when a consumer can access the certified product. These clocks can diverge.
| Indicator | Fixture formula | Observed | Fixture SLO | Result |
|---|---|---|---|---|
| Freshness | certified_at 08:18 − source_ready_at 08:12 | 6 min | ≤ 30 min | MET |
| Completeness | actual 9 rows / expected 9 rows | 100% | = 100% | MET |
| Duration | 08:18 − successful run start 08:12 | 6 min | ≤ 10 min | MET |
| Downstream availability | 08:18 − scheduled 08:00 | 18 min | ≤ 30 min | MET |
2. These thresholds are policy, not universal constants
The numbers above are deliberately simple fixture objectives chosen to make the mechanics visible and remain compatible with AtlasMart's earlier freshness expectations. Another warehouse may need five-minute finance availability, next-day regulatory certification, or different completeness tolerances by source. Use business decisions and source capabilities to set objectives, then measure actual distributions over time before committing to operational promises.
from datetime import datetimesource_ready = datetime.fromisoformat("2026-09-21T08:12:00+00:00")run_start = datetime.fromisoformat("2026-09-21T08:12:00+00:00")scheduled = datetime.fromisoformat("2026-09-21T08:00:00+00:00")certified = datetime.fromisoformat("2026-09-21T08:18:00+00:00")print("freshness_min:", int((certified-source_ready).total_seconds()/60))print("duration_min:", int((certified-run_start).total_seconds()/60))print("availability_min:", int((certified-scheduled).total_seconds()/60))print("completeness_pct:", 100 * 9 / 9)
3. Controlled failure: all tasks green, data incomplete
If the source manifest expects nine rows but the pipeline
publishes eight, the task graph may still be green. Completeness
is 8/9 = 88.89%. If certification trusts only
process state, consumers receive a formally “successful” but
incomplete product.
expected = 9actual = 8completeness = 100 * actual / expectedprint(round(completeness, 2))print("certify:", completeness == 100)# 88.89# False
Whether 100% is required depends on the data contract. Some telemetry products legitimately tolerate bounded lateness/partialness. The critical point is to make the tolerance explicit and observable rather than assuming task success equals data success.
4. Freshness and completeness can conflict
A system can publish quickly with missing data or wait for complete data and become stale. The right behavior depends on the product. AtlasMart chooses to block certification until expected current input is complete, then still measures how long that wait affected availability. For some products, the contract may publish a preliminary dataset with a visible completeness flag and later certify a final version. That is a semantic/product decision, not merely an orchestrator switch.
Measure retry counts, reject counts, source lag, queue time, current/backfill contention, and certification delay as diagnostic telemetry. They are not necessarily consumer-facing SLOs, but they help explain why a consumer-facing SLI moved.
5. SLO breaches should change operational behavior
An SLO that never gates publication, alerts an owner, changes incident priority, or informs capacity planning is merely a decorative number. Define who owns each SLI, what window is evaluated, how planned maintenance/backfills are classified, how late source data is attributed, and what consumer communication occurs when the product is outside objective.
Do not hide a source delay inside “pipeline duration.” RUN-A blocks because upstream data is unavailable; RUN-B then completes in six minutes. Reporting only RUN-B duration would make the end-to-end availability delay disappear.
6. Hands-on: inspect the acceptance summary
After running Lesson 5's script, open
atlasmart_ch16_lab/acceptance_summary.json. Confirm
current_run.slo reports 6-minute freshness, 100%
completeness, 6-minute duration, 18-minute downstream
availability, and all_met: true. Then change the
fixture's expected_rows to 10 and rerun: the
quality gate should fail before certification rather than
silently lowering the SLO.
7. Production judgment and bridge
Choose indicators that correspond to decisions consumers make, and keep measurement definitions versioned. If source availability semantics change, the freshness clock can change even if the pipeline code does not. If a metric changes from current-state to event-time completeness, historical SLO comparisons need interpretation.
Lesson 5 brings dependency, retry, backfill, and SLO evidence together into an operational runbook. The runbook exists so the response is not improvised while consumers are already affected.
Knowledge check
Check your understanding
- What is the difference between an SLI and an SLO?
- What is AtlasMart freshness in the fixture?
- Why is availability 18 minutes rather than 6?
- Can a green DAG violate completeness?
- Are the 30/10-minute thresholds universal?
Review the answers
1. An SLI is a measured quantity; an SLO is a target for that indicator over a defined scope/window.
2. 6 minutes from source readiness at 08:12 to certification at 08:18.
3. It is measured from the scheduled 08:00 consumer expectation to certification at 08:18, including upstream delay.
4. Yes. Process success can coexist with missing rows unless the data state is explicitly checked.
5. No. They are explicit fixture policy choices used to teach the mechanics.
Authoritative references
- Python documentation — graphlibStandard-library topological ordering concepts used to explain dependency graphs without requiring an orchestrator product.
- Python documentation — os.replaceLocal atomic file-replacement primitive used by the fixture to demonstrate partition publication without partial output files.
- Python documentation — hashlibDeterministic SHA-256 evidence for replay/backfill comparisons in the local lab.
- Google SRE Book — Service Level ObjectivesFoundational distinction among service indicators/objectives and externally meaningful reliability goals.
- Apache Airflow documentation — DAGsOptional later-course example of a production orchestrator's DAG concept; no Airflow command or installation is required here.
- Apache Airflow documentation — BackfillOptional implementation reference for historical run concepts; Chapter 16 teaches the vendor-neutral semantics first.
- Kimball Group — Dimensional Modeling TechniquesBackground for the grains and dimensional facts whose correctness orchestration must preserve.