Chapter 16 · Orchestration, Dependencies, Scheduling, Backfills, SLAs, and Failure Recovery

Retries, Timeouts, Idempotent Tasks, Partial Failure, Compensation, and Operator Intervention

Classify failures before retrying, bound attempts with timeouts, make task side effects idempotent or compensatable, and define when operator intervention is safer than automated repetition.

Intermediate → Advanced130–155 minutesRetry + idempotency failure labPython 3 stdlib · local/syntheticLast reviewed: September 2026

Learning outcomes

At 08:12 the AtlasMart input is ready. The transform begins and a synthetic parser exception occurs before publication. The orchestrator now has a dangerous choice: retry automatically, stop for an operator, or attempt compensation. “Retries are good” is not a policy. The correct action depends on what side effects occurred and whether repeating the task converges to the same state.

01

Classify retryable, non-retryable, and operator-decision failures from side effects and root cause rather than from exception type alone.

02

Define timeout, retry, idempotency, partial failure, compensation, and operator intervention in an analytical pipeline context.

03

Explain why retrying a non-idempotent append or external side effect can duplicate facts or notifications.

04

Use atomic partition replacement and stable run/partition identity to make the AtlasMart transform safe to retry.

05

Record retry traces and failure evidence so automation does not erase the cause of an incident.

Chapter 16 continuity contract

Chapter 16 does not change the accepted Chapter 15 analytical state. The current certified AtlasMart sales target remains nine paid order-line facts, seven paid orders, eleven units, 740 USD paid GMV, 450 USD cost-at-sale, and 290 USD gross profit at committed source sequence 206. Orchestration adds run/task state, dependency evidence, partition manifests, SLO measurements, resource-pool labels, and operator notes around that data state. The historical backfill fixture deliberately starts from one corrupted 2026-09-20 publication (cost 435 USD instead of 425 USD) and repairs only that historical partition from immutable raw evidence; the current 2026-09-21 checksum must remain unchanged.

Execution and guarantee boundary

The mandatory labs are synthetic, local, and free. They use Python 3 standard-library modules and local JSON/filesystem state; generation-time validation ran with Python 3.13.5. Logical timestamps are simulated—no real waiting, cluster scheduler, queue, cloud warehouse, distributed lock, or resource manager is involved. A labeled current versus backfill resource pool proves orchestration intent in the fixture, not operating-system or cloud compute isolation. Apache Airflow is referenced only as an optional later-course implementation example and is not a prerequisite.

1. Timeout and retry solve different problems

A timeout bounds how long an attempt may remain active before the control plane treats it as failed or uncertain. A retry starts another attempt according to policy. A timeout does not undo work that already happened, and a retry does not become safe merely because the first attempt timed out.

An idempotent task is designed so repeating the same logical input identity converges to the same intended target state. Partial failure means some side effects completed while other required effects did not. Compensation is a governed action that neutralizes or supersedes an earlier side effect when true rollback is unavailable. Operator intervention is deliberate human control used when automated policy cannot safely decide.

Failure Automatic retry? Reason / prerequisite
Transient read failure before any write Usually reasonable Same immutable input identity; no side effect yet.
Atomic partition transform failed before replace Reasonable Temporary output is disposable; published partition unchanged.
Append succeeded, acknowledgment lost Dangerous without idempotency Retry can append duplicate facts.
Quality rule fails deterministically Usually stop Repeating unchanged input/code reproduces the defect.
Schema contract changed Stop / operator or migration workflow Retry cannot make unsupported semantics compatible.
Downstream notification sent, publish later fails Needs compensation/idempotent notification key External side effect may already be visible.

2. AtlasMart retry trace

The successful run ID is RUN-20260921-0812-B. Its transform trace is deliberately:

Attempt State Logical time Evidence
1 FAILED 08:14Z Synthetic parser exception occurs before publication.
1 RETRYING 08:15Z Policy allows retry because input is immutable and target is not yet replaced.
2 SUCCESS 08:16Z Nine transformed rows produced; no duplicate publication exists.
safe_partition_publish.py
import json, osfrom pathlib import Pathdef atomic_publish(path: Path, payload: dict):    tmp = path.with_suffix(path.suffix + ".tmp")    tmp.write_text(json.dumps(payload, sort_keys=True), encoding="utf-8")    os.replace(tmp, path)  # replaces the destination atomically on supported local filesystems# A failed attempt before os.replace leaves the prior published file untouched.

The local pattern provides a narrow guarantee: the fixture never exposes a half-written JSON publication. Production object stores, warehouses, table formats, and distributed filesystems have different commit primitives; verify their exact atomicity semantics before translating this pattern.

3. Controlled failure: retry a non-idempotent append

Imagine transform_partition writes rows directly with INSERT into a target lacking a unique grain key, then crashes after eight rows. A blind retry inserts those eight again plus the ninth. DISTINCT in a dashboard does not safely repair the warehouse because measures, row lineage, and downstream joins may already be duplicated.

unsafe_retry.sql
-- Illustrative anti-pattern; do not run against production.INSERT INTO fact_salesSELECT * FROM transformed_partition;-- Crash after partial commit + retry => duplicate business rows unless the target contract prevents it.

Safer choices include transactionally replacing a partition, merging on a stable business/grain key with sequence/version guards, staging then swapping, or restoring before retry. The exact mechanism is engine-specific; the invariant is vendor-neutral: repeating the same logical attempt must not multiply business effects.

4. Compensation and operator boundaries

Compensation is not necessarily “run the inverse SQL.” A published analytical partition may have been consumed by dashboards, exports, alerts, or finance workflows. A safer compensation may be to mark the partition uncertified, withdraw a semantic-layer version, restore a previous snapshot, and run a controlled rebuild. The operator needs the run ID, partition, input hash, current publication hash, last successful task, failure trace, and consumer blast radius.

Automatic retry budgets should be bounded. Repeating a deterministic failure hundreds of times wastes resources and delays diagnosis. Conversely, a single transient network error should not always wake an operator. Define retry count/backoff/timeout from measured failure modes and source SLOs, not from folklore.

5. Hands-on: verify retry safety in the integrated fixture

Lesson 5 contains the complete atlasmart_ch16.py script. When executed, inspect task_events.jsonl and filter RUN-20260921-0812-B / transform_partition. The required state sequence is FAILED → RETRYING → SUCCESS. The publication appears only after the successful retry and quality gate.

inspect_retry.py
import jsonfrom pathlib import Pathrows = [json.loads(x) for x in Path("atlasmart_ch16_lab/task_events.jsonl").read_text().splitlines()]trace = [r for r in rows if r["run_id"] == "RUN-20260921-0812-B" and r["task"] == "transform_partition"]print([r["state"] for r in trace])# Expected: ['FAILED', 'RETRYING', 'SUCCESS']
What this proves

The retry trace plus unchanged pre-publication target state proves the local task was safe to repeat in this fixture. It does not prove arbitrary transforms, notifications, APIs, or warehouse MERGE statements are idempotent.

6. Production judgment and bridge

Retries should be attached to known failure classes and known side-effect contracts. Timeouts should surface uncertainty rather than pretending to cancel remote work. Operator intervention should be explicit when the system cannot tell whether a side effect happened or when a contract change requires a decision.

Lesson 3 widens the recovery scope. Instead of retrying one current task, AtlasMart must repair a historical 2026-09-20 partition without consuming the same resources or overwriting the current 2026-09-21 publication.

Knowledge check

Check your understanding

  1. Why is a timeout not a rollback?
  2. What makes the AtlasMart transform retry safe in the fixture?
  3. Why is a deterministic quality failure a poor retry candidate?
  4. What is compensation in an analytical pipeline?
  5. What retry trace is required in the lab?
Review the answers

1. Because the task or remote system may already have committed side effects even if the orchestrator stopped waiting.

2. Immutable input plus temporary output and atomic replacement; the failed attempt does not mutate the published partition.

3. The same input/code will usually produce the same failure until data or logic is repaired.

4. A governed action that neutralizes/supersedes prior side effects when true rollback is unavailable, such as withdrawing certification or restoring a prior partition.

5. FAILED → RETRYING → SUCCESS for the transform task.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.