Chapter 16 · Orchestration, Dependencies, Scheduling, Backfills, SLAs, and Failure Recovery
Retries, Timeouts, Idempotent Tasks, Partial Failure, Compensation, and Operator Intervention
Classify failures before retrying, bound attempts with timeouts, make task side effects idempotent or compensatable, and define when operator intervention is safer than automated repetition.
Learning outcomes
At 08:12 the AtlasMart input is ready. The transform begins and a synthetic parser exception occurs before publication. The orchestrator now has a dangerous choice: retry automatically, stop for an operator, or attempt compensation. “Retries are good” is not a policy. The correct action depends on what side effects occurred and whether repeating the task converges to the same state.
Classify retryable, non-retryable, and operator-decision failures from side effects and root cause rather than from exception type alone.
Define timeout, retry, idempotency, partial failure, compensation, and operator intervention in an analytical pipeline context.
Explain why retrying a non-idempotent append or external side effect can duplicate facts or notifications.
Use atomic partition replacement and stable run/partition identity to make the AtlasMart transform safe to retry.
Record retry traces and failure evidence so automation does not erase the cause of an incident.
Chapter 16 does not change the accepted Chapter 15 analytical state. The current certified AtlasMart sales target remains nine paid order-line facts, seven paid orders, eleven units, 740 USD paid GMV, 450 USD cost-at-sale, and 290 USD gross profit at committed source sequence 206. Orchestration adds run/task state, dependency evidence, partition manifests, SLO measurements, resource-pool labels, and operator notes around that data state. The historical backfill fixture deliberately starts from one corrupted 2026-09-20 publication (cost 435 USD instead of 425 USD) and repairs only that historical partition from immutable raw evidence; the current 2026-09-21 checksum must remain unchanged.
The mandatory labs are synthetic, local, and free. They use
Python 3 standard-library modules and local JSON/filesystem
state; generation-time validation ran with Python 3.13.5.
Logical timestamps are simulated—no real waiting, cluster
scheduler, queue, cloud warehouse, distributed lock, or
resource manager is involved. A labeled
current versus backfill resource
pool proves orchestration intent in the fixture, not
operating-system or cloud compute isolation. Apache Airflow is
referenced only as an optional later-course implementation
example and is not a prerequisite.
1. Timeout and retry solve different problems
A timeout bounds how long an attempt may remain active before the control plane treats it as failed or uncertain. A retry starts another attempt according to policy. A timeout does not undo work that already happened, and a retry does not become safe merely because the first attempt timed out.
An idempotent task is designed so repeating the same logical input identity converges to the same intended target state. Partial failure means some side effects completed while other required effects did not. Compensation is a governed action that neutralizes or supersedes an earlier side effect when true rollback is unavailable. Operator intervention is deliberate human control used when automated policy cannot safely decide.
| Failure | Automatic retry? | Reason / prerequisite |
|---|---|---|
| Transient read failure before any write | Usually reasonable | Same immutable input identity; no side effect yet. |
| Atomic partition transform failed before replace | Reasonable | Temporary output is disposable; published partition unchanged. |
| Append succeeded, acknowledgment lost | Dangerous without idempotency | Retry can append duplicate facts. |
| Quality rule fails deterministically | Usually stop | Repeating unchanged input/code reproduces the defect. |
| Schema contract changed | Stop / operator or migration workflow | Retry cannot make unsupported semantics compatible. |
| Downstream notification sent, publish later fails | Needs compensation/idempotent notification key | External side effect may already be visible. |
2. AtlasMart retry trace
The successful run ID is RUN-20260921-0812-B. Its
transform trace is deliberately:
| Attempt | State | Logical time | Evidence |
|---|---|---|---|
| 1 | FAILED | 08:14Z | Synthetic parser exception occurs before publication. |
| 1 | RETRYING | 08:15Z | Policy allows retry because input is immutable and target is not yet replaced. |
| 2 | SUCCESS | 08:16Z | Nine transformed rows produced; no duplicate publication exists. |
import json, osfrom pathlib import Pathdef atomic_publish(path: Path, payload: dict): tmp = path.with_suffix(path.suffix + ".tmp") tmp.write_text(json.dumps(payload, sort_keys=True), encoding="utf-8") os.replace(tmp, path) # replaces the destination atomically on supported local filesystems# A failed attempt before os.replace leaves the prior published file untouched.
The local pattern provides a narrow guarantee: the fixture never exposes a half-written JSON publication. Production object stores, warehouses, table formats, and distributed filesystems have different commit primitives; verify their exact atomicity semantics before translating this pattern.
3. Controlled failure: retry a non-idempotent append
Imagine transform_partition writes rows directly
with INSERT into a target lacking a unique grain
key, then crashes after eight rows. A blind retry inserts those
eight again plus the ninth. DISTINCT in a dashboard
does not safely repair the warehouse because measures, row
lineage, and downstream joins may already be duplicated.
-- Illustrative anti-pattern; do not run against production.INSERT INTO fact_salesSELECT * FROM transformed_partition;-- Crash after partial commit + retry => duplicate business rows unless the target contract prevents it.
Safer choices include transactionally replacing a partition, merging on a stable business/grain key with sequence/version guards, staging then swapping, or restoring before retry. The exact mechanism is engine-specific; the invariant is vendor-neutral: repeating the same logical attempt must not multiply business effects.
4. Compensation and operator boundaries
Compensation is not necessarily “run the inverse SQL.” A published analytical partition may have been consumed by dashboards, exports, alerts, or finance workflows. A safer compensation may be to mark the partition uncertified, withdraw a semantic-layer version, restore a previous snapshot, and run a controlled rebuild. The operator needs the run ID, partition, input hash, current publication hash, last successful task, failure trace, and consumer blast radius.
Automatic retry budgets should be bounded. Repeating a deterministic failure hundreds of times wastes resources and delays diagnosis. Conversely, a single transient network error should not always wake an operator. Define retry count/backoff/timeout from measured failure modes and source SLOs, not from folklore.
5. Hands-on: verify retry safety in the integrated fixture
Lesson 5 contains the complete
atlasmart_ch16.py script. When executed, inspect
task_events.jsonl and filter
RUN-20260921-0812-B /
transform_partition. The required state sequence is
FAILED → RETRYING → SUCCESS. The publication
appears only after the successful retry and quality gate.
import jsonfrom pathlib import Pathrows = [json.loads(x) for x in Path("atlasmart_ch16_lab/task_events.jsonl").read_text().splitlines()]trace = [r for r in rows if r["run_id"] == "RUN-20260921-0812-B" and r["task"] == "transform_partition"]print([r["state"] for r in trace])# Expected: ['FAILED', 'RETRYING', 'SUCCESS']
The retry trace plus unchanged pre-publication target state proves the local task was safe to repeat in this fixture. It does not prove arbitrary transforms, notifications, APIs, or warehouse MERGE statements are idempotent.
6. Production judgment and bridge
Retries should be attached to known failure classes and known side-effect contracts. Timeouts should surface uncertainty rather than pretending to cancel remote work. Operator intervention should be explicit when the system cannot tell whether a side effect happened or when a contract change requires a decision.
Lesson 3 widens the recovery scope. Instead of retrying one current task, AtlasMart must repair a historical 2026-09-20 partition without consuming the same resources or overwriting the current 2026-09-21 publication.
Knowledge check
Check your understanding
- Why is a timeout not a rollback?
- What makes the AtlasMart transform retry safe in the fixture?
- Why is a deterministic quality failure a poor retry candidate?
- What is compensation in an analytical pipeline?
- What retry trace is required in the lab?
Review the answers
1. Because the task or remote system may already have committed side effects even if the orchestrator stopped waiting.
2. Immutable input plus temporary output and atomic replacement; the failed attempt does not mutate the published partition.
3. The same input/code will usually produce the same failure until data or logic is repaired.
4. A governed action that neutralizes/supersedes prior side effects when true rollback is unavailable, such as withdrawing certification or restoring a prior partition.
5. FAILED → RETRYING → SUCCESS for the transform task.
Authoritative references
- Python documentation — graphlibStandard-library topological ordering concepts used to explain dependency graphs without requiring an orchestrator product.
- Python documentation — os.replaceLocal atomic file-replacement primitive used by the fixture to demonstrate partition publication without partial output files.
- Python documentation — hashlibDeterministic SHA-256 evidence for replay/backfill comparisons in the local lab.
- Google SRE Book — Service Level ObjectivesFoundational distinction among service indicators/objectives and externally meaningful reliability goals.
- Apache Airflow documentation — DAGsOptional later-course example of a production orchestrator's DAG concept; no Airflow command or installation is required here.
- Apache Airflow documentation — BackfillOptional implementation reference for historical run concepts; Chapter 16 teaches the vendor-neutral semantics first.
- Kimball Group — Dimensional Modeling TechniquesBackground for the grains and dimensional facts whose correctness orchestration must preserve.