Model a multi-service AtlasMart saga with durable intermediate states, idempotent steps, semantic compensation, and explicit orchestration/choreography tradeoffs.
Sagas: Orchestration vs Choreography, Compensating Actions, and Business-Level Atomicity
Sagas trade one long atomic transaction for a sequence of committed local transactions plus business-level compensation. This lesson shows why compensation is not rollback and why duplicate delivery must be expected.
Explain sagas as sequences of committed local transactions with compensating actions.
Compare orchestration and choreography without treating either as automatically superior.
Design idempotent forward and compensating steps for retries and duplicate delivery.
Recognize visible intermediate states and distinguish semantic compensation from rollback.
1. The problem: an order spans services that cannot share one lock
AtlasMart checkout may reserve inventory, authorize payment, arrange shipment, update loyalty points, and publish customer notifications. Holding one distributed transaction open across those services can be impractical or impossible. A saga decomposes a long-lived business transaction into a sequence of smaller committed transactions. If a later step fails, the system runs compensating actions that semantically amend earlier effects.
Compensation is not database rollback. After an authorization was visible to an external payment system, “undo” may mean issuing a void or refund. After an email was sent, it cannot be unsent; a later correction may be the only compensation. Therefore saga design is business semantics, not merely a technical transaction primitive.
2. Orchestration vs choreography
| Style | Coordination shape | Strength | Risk |
|---|---|---|---|
| Orchestration | One workflow component commands/records steps | Clear global state, explicit retries/compensations | Central workflow dependency; state machine complexity |
| Choreography | Services react to events and emit new events | Loose runtime coupling; local autonomy | Emergent control flow, harder end-to-end diagnosis |
| Hybrid | Critical workflow orchestrated, side effects event-driven | Balances explicit core with decoupled observers | Requires clear ownership and event contracts |
Neither style eliminates duplicate delivery, delayed messages, or partial failure. Each step needs an idempotency key, durable workflow/event state, observability, and a documented compensation or escalation path.
3. AtlasMart lab: shipment fails after earlier commits
Python 3.13+ standard library only. The generated lab was verified with Python 3.13.5. No database server, Docker, cloud account, paid feature, credential, firewall change, clock manipulation, or destructive failure injection is required. All failures are deterministic in-memory simulations.
state = {
"order": "CREATED",
"inventory_reserved": False,
"payment_authorized": False,
"shipment_created": False,
}
processed = set()
log = []
def once(step_id, action):
if step_id in processed:
log.append((step_id, "duplicate ignored"))
return
action()
processed.add(step_id)
log.append((step_id, "applied"))
print("FORWARD SAGA")
once("reserve:o-300", lambda: state.update(inventory_reserved=True))
once("authorize:o-300", lambda: state.update(payment_authorized=True))
print("after two steps:", state)
print("shipment step fails: carrier unavailable")
print("\nCOMPENSATION")
once("void-auth:o-300", lambda: state.update(payment_authorized=False))
once("release:o-300", lambda: state.update(inventory_reserved=False))
state["order"] = "CANCELLED"
# Retry the compensation to prove idempotency.
once("release:o-300", lambda: state.update(inventory_reserved=False))
print("final state:", state)
print("log:", log)
print("business restored:", state["order"] == "CANCELLED" and not state["inventory_reserved"] and not state["payment_authorized"])
print("note: compensation is a new business action, not time travel")
Inventory reservation and payment authorization commit before shipment fails. The saga then voids authorization and releases inventory, marks the order CANCELLED, and safely ignores a duplicate release compensation. The final business state is repaired, but the history still contains committed forward and compensating actions—evidence that this was not an ACID rollback.
4. Intermediate states are part of the contract
During a saga, an order may legitimately be
PAYMENT_AUTHORIZED but not yet
SHIPMENT_CREATED. User interfaces, support tools,
fraud processes, and retries must understand these states.
Hiding them behind a generic “processing” label without durable
workflow state makes incidents hard to reason about. Record step
IDs, attempt count, last error, next retry time, compensation
status, and correlation IDs.
Exactly-once delivery is usually the wrong assumption. Design for at-least-once execution plus idempotent effects, or use a transactional outbox/inbox pattern so local database changes and message publication have an auditable handoff.
5. Deliberately wrong approach: compensate by deleting history
Deleting the order row after a failed shipment may erase the very evidence needed for refund, fraud, audit, or customer support. Compensation should create a new business fact—released reservation, voided authorization, cancelled order—with a reason and correlation to the original step. Some effects are irreversible and require human escalation rather than automated “undo.”
6. Production judgment and bridge
Use sagas when the workflow spans autonomous services/resources and the business can tolerate visible intermediate states plus semantic compensation. Define timeouts as workflow policy, not proof that prior steps failed. Test duplicate messages, out-of-order events, compensation failure, poison messages, and operator replay. Monitor stuck sagas by age/state, compensation error rate, retry count, and end-to-end latency. The final lesson of this chapter asks the design question before choosing a protocol: which invariants can be re-shaped into smaller ownership boundaries, and which still require explicit strong coordination?
Check your understanding
- Why is compensation not the same as rollback?
- What must a saga expose to operators?
- What is a major choreography risk?
- Why are idempotency keys important?
- When is a saga a poor fit?
Review the answers
1. Earlier local transactions already committed; compensation is a new business action that amends their effects.
2. Durable workflow state, step IDs, attempts, errors, compensation status, and correlation identifiers.
3. The global workflow can become implicit across many event handlers, making diagnosis and evolution difficult.
4. Retries and duplicate delivery are normal; repeating the same step must not duplicate side effects.
5. When the invariant cannot tolerate intermediate states or compensation and requires synchronous all-or-nothing atomicity.
References
Foundational claims are vendor-neutral. Product documentation is used only as a current implementation example and is not required for the mandatory labs.
- Garcia-Molina & Salem — Sagas (SIGMOD 1987) — Original saga formulation for long-lived transactions and compensating transactions.
- Princeton Technical Report — SAGAS — Primary technical report version describing saga execution and compensation.
- Microsoft Azure Architecture Center — Saga pattern — Current implementation-oriented reference for orchestration/choreography tradeoffs and compensating actions.