Trace two-phase commit through prepare, durable decision, coordinator failure, blocking prepared state, and recovery without confusing atomic commit with consensus.
Two-Phase Commit: Coordinator State, Blocking Failure Modes, and Recovery Complexity
Two-phase commit can provide all-or-nothing outcome across transactional participants, but prepared state creates locks, uncertainty, and recovery obligations. AtlasMart makes the blocking failure mode visible.
Trace two-phase commit (2PC) through prepare and decision phases.
Identify prepared locks/state, coordinator logging, uncertain outcomes, and recovery obligations.
Explain why 2PC can block and why atomic commit is not the same problem as consensus.
Diagnose the unsafe practice of letting a prepared participant guess its own outcome.
1. The problem: inventory and payment must agree on one outcome
Suppose AtlasMart has two transactional participants: an inventory database and a payment ledger. For a specific operation, the business requires all-or-nothing commit across both. Two-phase commit (2PC) is an atomic-commit protocol with a coordinator and participants. Phase 1 asks participants to prepare: can each durably promise that it can later commit? Phase 2 distributes the coordinator’s final COMMIT or ABORT decision.
Prepared is not the same as committed. A prepared participant may hold locks or reserved resources and must retain enough recovery state to obey the final decision later. That intermediate state is exactly why 2PC can block when communication with the decision authority is lost.
2. Timeline and durable evidence
| Step | Coordinator | Inventory | Payment | Meaning |
|---|---|---|---|---|
| 1 | BEGIN tx-200 | ACTIVE | ACTIVE | Work is tentative |
| 2 | PREPARE? | PREPARED / YES | PREPARED / YES | Participants promise they can commit |
| 3 | Durably log COMMIT | PREPARED | PREPARED | Global decision exists |
| 4 | Send decision | COMMITTED | PREPARED if message lost | One participant still uncertain |
| 5 | Recovery/retry COMMIT | COMMITTED | COMMITTED | Both converge to durable decision |
The coordinator’s durable decision record is the crucial recovery evidence. If a participant is prepared and cannot learn the decision, it cannot safely invent one without risking atomicity. Timeouts help detect that progress is stalled; they do not magically reveal whether the other side committed.
3. AtlasMart lab: coordinator crashes after deciding
Python 3.13+ standard library only. The generated lab was verified with Python 3.13.5. No database server, Docker, cloud account, paid feature, credential, firewall change, clock manipulation, or destructive failure injection is required. All failures are deterministic in-memory simulations.
participants = {
"inventory": {"state": "ACTIVE", "locks": [], "decision": None},
"payment": {"state": "ACTIVE", "locks": [], "decision": None},
}
coordinator_log = []
def prepare(name):
p = participants[name]
p["state"] = "PREPARED"
p["locks"] = ["order:o-200"]
return True
def apply_decision(name, decision):
p = participants[name]
p["decision"] = decision
p["state"] = "COMMITTED" if decision == "COMMIT" else "ABORTED"
p["locks"] = []
print("PHASE 1: PREPARE")
for name in participants:
print(name, "vote YES:", prepare(name), participants[name])
print("\nCOORDINATOR DURABLY DECIDES COMMIT")
coordinator_log.append("COMMIT tx-200")
print("coordinator_log:", coordinator_log)
print("\nPARTIAL PHASE 2 + COORDINATOR CRASH")
apply_decision("inventory", "COMMIT")
print("inventory notified:", participants["inventory"])
print("payment still prepared:", participants["payment"])
print("payment can safely guess ABORT?", False)
print("\nRECOVERY")
recovered_decision = coordinator_log[-1].split()[0]
apply_decision("payment", recovered_decision)
print("recovered decision:", recovered_decision)
print("final states:", {k:v["state"] for k,v in participants.items()})
print("locks remaining:", {k:v["locks"] for k,v in participants.items()})
Both participants enter PREPARED and hold an order lock. The coordinator durably records COMMIT, informs inventory, then “crashes.” Payment remains PREPARED and cannot safely guess ABORT. Recovery replays the durable coordinator decision, commits payment, and releases all locks. The simulation proves the blocking/decision-recovery logic, not the fault tolerance of any real transaction manager.
4. 2PC solves atomic commit; it does not manufacture fault tolerance
2PC and consensus are related in distributed systems discussions but solve different core problems. 2PC coordinates one atomic commit across participants that can prepare; classic 2PC can block if the decision is unavailable. Consensus protocols are designed to get a replicated group to agree on values under specified failure assumptions. A production transaction manager may itself use replicated consensus, but saying “2PC is consensus” erases the distinction.
PostgreSQL 18 exposes prepared transactions through
PREPARE TRANSACTION, COMMIT PREPARED,
and ROLLBACK PREPARED; its documentation warns not
to leave prepared transactions open for long because they retain
locks/resources and can interfere with maintenance. That is an
implementation example of the general operational cost of
prepared state.
5. Deliberately wrong approach: timeout means abort locally
If inventory has already committed under a durable global COMMIT decision, payment independently aborting because its coordinator connection timed out creates a split outcome. The safe recovery procedure is to discover the authoritative decision and apply it idempotently. If no decision was durably recorded before the coordinator failed, the protocol may remain blocked until coordinator recovery or a higher-level replicated transaction manager resolves the uncertainty.
6. Production judgment and bridge
Use 2PC only when the invariant truly requires synchronous all-or-nothing commit across participating transactional resources and the latency/availability cost is acceptable. Monitor prepared-transaction age, lock wait time, coordinator health, recovery queue depth, in-doubt transaction count, and decision-log durability. Exercise coordinator and participant crashes in a disposable environment and document operator recovery. The next lesson considers a different business contract: instead of hiding intermediate states behind one atomic commit, a saga exposes a sequence of committed steps and uses semantic compensation if the workflow later fails.
Check your understanding
- What promise does a participant make when it votes YES to prepare?
- Why can 2PC block?
- What evidence lets recovery finish the lab transaction?
- Why is a timeout not proof of abort?
- Why is 2PC not identical to consensus?
Review the answers
1. It durably enters a state from which it can obey the later global commit decision and typically retains necessary locks/resources.
2. A prepared participant may be unable to learn the final decision and cannot safely guess without risking atomicity.
3. The coordinator durably logged COMMIT before crashing.
4. The coordinator or another participant may already have committed; timeout only proves the response did not arrive in time.
5. Atomic commit and replicated agreement are distinct problems; classic 2PC can rely on a single coordinator and block on its unavailability.
References
Foundational claims are vendor-neutral. Product documentation is used only as a current implementation example and is not required for the mandatory labs.
- PostgreSQL 18 — Two-Phase Transactions — Current implementation example of PREPARE/COMMIT PREPARED/ROLLBACK PREPARED and durable prepared state.
- PostgreSQL 18 — pg_prepared_xacts — Observable view of transactions currently prepared for two-phase commit.
- Gray & Reuter — Transaction Processing: Concepts and Techniques — Foundational transaction-processing reference for logging, atomic commit, and recovery concepts.