Chapter 16 · Orchestration, Dependencies, Scheduling, Backfills, SLAs, and Failure Recovery

SLAs/SLOs for Freshness, Completeness, Duration, and Downstream Availability

Measure freshness, completeness, duration, and downstream availability with explicit SLIs and fixture-specific SLOs; distinguish task success from data-product correctness and contractual SLAs from engineering objectives.

Intermediate → Advanced120–145 minutesSLI/SLO measurement labPython 3 stdlib · local/syntheticLast reviewed: September 2026

Learning outcomes

The successful AtlasMart run has every task in SUCCESS, but that alone does not answer the consumer questions: Is the data fresh? Is all expected data present? How long did recovery take? When could the certified product be used downstream? Reliability needs measurable indicators attached to consumer-relevant states.

01

Distinguish SLI, SLO, and SLA and tie each measurement to an explicit data-product state surface.

02

Calculate freshness, completeness, run duration, and downstream availability from stated timestamps/counts rather than from invented reliability numbers.

03

Explain why a successful task graph can violate a data SLO and why one SLO cannot stand in for all consumer expectations.

04

Use thresholds as fixture-specific policy choices, not universal warehouse tuning values.

05

Connect SLO breaches to alerting, certification, consumer communication, and recovery priorities.

Chapter 16 continuity contract

Chapter 16 does not change the accepted Chapter 15 analytical state. The current certified AtlasMart sales target remains nine paid order-line facts, seven paid orders, eleven units, 740 USD paid GMV, 450 USD cost-at-sale, and 290 USD gross profit at committed source sequence 206. Orchestration adds run/task state, dependency evidence, partition manifests, SLO measurements, resource-pool labels, and operator notes around that data state. The historical backfill fixture deliberately starts from one corrupted 2026-09-20 publication (cost 435 USD instead of 425 USD) and repairs only that historical partition from immutable raw evidence; the current 2026-09-21 checksum must remain unchanged.

Execution and guarantee boundary

The mandatory labs are synthetic, local, and free. They use Python 3 standard-library modules and local JSON/filesystem state; generation-time validation ran with Python 3.13.5. Logical timestamps are simulated—no real waiting, cluster scheduler, queue, cloud warehouse, distributed lock, or resource manager is involved. A labeled current versus backfill resource pool proves orchestration intent in the fixture, not operating-system or cloud compute isolation. Apache Airflow is referenced only as an optional later-course implementation example and is not a prerequisite.

1. SLI, SLO, and SLA are different artifacts

A service level indicator (SLI) is a measured quantity such as freshness lag or completeness percentage. A service level objective (SLO) is a target/range for an SLI over a defined scope/window. A service level agreement (SLA) is an agreement—often external or contractual—that may include consequences, exclusions, and responsibilities. Teams should not call every dashboard threshold an SLA.

For data products, define the observation surface. “Pipeline duration” measures control-plane execution time. “Freshness” compares source availability/event state to certified data availability. “Downstream availability” asks when a consumer can access the certified product. These clocks can diverge.

Indicator Fixture formula Observed Fixture SLO Result
Freshness certified_at 08:18 − source_ready_at 08:12 6 min ≤ 30 min MET
Completeness actual 9 rows / expected 9 rows 100% = 100% MET
Duration 08:18 − successful run start 08:12 6 min ≤ 10 min MET
Downstream availability 08:18 − scheduled 08:00 18 min ≤ 30 min MET

2. These thresholds are policy, not universal constants

The numbers above are deliberately simple fixture objectives chosen to make the mechanics visible and remain compatible with AtlasMart's earlier freshness expectations. Another warehouse may need five-minute finance availability, next-day regulatory certification, or different completeness tolerances by source. Use business decisions and source capabilities to set objectives, then measure actual distributions over time before committing to operational promises.

slo_math.py
from datetime import datetimesource_ready = datetime.fromisoformat("2026-09-21T08:12:00+00:00")run_start = datetime.fromisoformat("2026-09-21T08:12:00+00:00")scheduled = datetime.fromisoformat("2026-09-21T08:00:00+00:00")certified = datetime.fromisoformat("2026-09-21T08:18:00+00:00")print("freshness_min:", int((certified-source_ready).total_seconds()/60))print("duration_min:", int((certified-run_start).total_seconds()/60))print("availability_min:", int((certified-scheduled).total_seconds()/60))print("completeness_pct:", 100 * 9 / 9)

3. Controlled failure: all tasks green, data incomplete

If the source manifest expects nine rows but the pipeline publishes eight, the task graph may still be green. Completeness is 8/9 = 88.89%. If certification trusts only process state, consumers receive a formally “successful” but incomplete product.

green_but_incomplete.py
expected = 9actual = 8completeness = 100 * actual / expectedprint(round(completeness, 2))print("certify:", completeness == 100)# 88.89# False

Whether 100% is required depends on the data contract. Some telemetry products legitimately tolerate bounded lateness/partialness. The critical point is to make the tolerance explicit and observable rather than assuming task success equals data success.

4. Freshness and completeness can conflict

A system can publish quickly with missing data or wait for complete data and become stale. The right behavior depends on the product. AtlasMart chooses to block certification until expected current input is complete, then still measures how long that wait affected availability. For some products, the contract may publish a preliminary dataset with a visible completeness flag and later certify a final version. That is a semantic/product decision, not merely an orchestrator switch.

Measure retry counts, reject counts, source lag, queue time, current/backfill contention, and certification delay as diagnostic telemetry. They are not necessarily consumer-facing SLOs, but they help explain why a consumer-facing SLI moved.

5. SLO breaches should change operational behavior

An SLO that never gates publication, alerts an owner, changes incident priority, or informs capacity planning is merely a decorative number. Define who owns each SLI, what window is evaluated, how planned maintenance/backfills are classified, how late source data is attributed, and what consumer communication occurs when the product is outside objective.

Do not hide a source delay inside “pipeline duration.” RUN-A blocks because upstream data is unavailable; RUN-B then completes in six minutes. Reporting only RUN-B duration would make the end-to-end availability delay disappear.

6. Hands-on: inspect the acceptance summary

After running Lesson 5's script, open atlasmart_ch16_lab/acceptance_summary.json. Confirm current_run.slo reports 6-minute freshness, 100% completeness, 6-minute duration, 18-minute downstream availability, and all_met: true. Then change the fixture's expected_rows to 10 and rerun: the quality gate should fail before certification rather than silently lowering the SLO.

7. Production judgment and bridge

Choose indicators that correspond to decisions consumers make, and keep measurement definitions versioned. If source availability semantics change, the freshness clock can change even if the pipeline code does not. If a metric changes from current-state to event-time completeness, historical SLO comparisons need interpretation.

Lesson 5 brings dependency, retry, backfill, and SLO evidence together into an operational runbook. The runbook exists so the response is not improvised while consumers are already affected.

Knowledge check

Check your understanding

  1. What is the difference between an SLI and an SLO?
  2. What is AtlasMart freshness in the fixture?
  3. Why is availability 18 minutes rather than 6?
  4. Can a green DAG violate completeness?
  5. Are the 30/10-minute thresholds universal?
Review the answers

1. An SLI is a measured quantity; an SLO is a target for that indicator over a defined scope/window.

2. 6 minutes from source readiness at 08:12 to certification at 08:18.

3. It is measured from the scheduled 08:00 consumer expectation to certification at 08:18, including upstream delay.

4. Yes. Process success can coexist with missing rows unless the data state is explicitly checked.

5. No. They are explicit fixture policy choices used to teach the mechanics.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.