Chapter 11 · Transactions, Isolation, Locking, Deadlocks, Retries, and Consistency
Driver Transaction Functions, Retryable Failures, Idempotency, and Backoff
Design Neo4j driver retries so transient failures are recoverable without duplicating AtlasMart business state or external effects.
Learning outcomes
A deadlock is often safe to retry because Neo4j rolls one database transaction back. The dangerous part is application code around that retry: the Python driver may invoke a managed transaction callback more than once, so anything inside that callback must tolerate re-execution.
Explain the managed transaction retry contract in the official Python driver 6.3 line.
Distinguish retryable database failure from a blanket “retry every exception” policy.
Design idempotent graph writes using stable operation identifiers and constraints.
Keep irreversible external side effects outside a callback that may run more than once.
Instrument attempts, retry budget and final reconciliation without hiding failed attempts.
The mandatory lab continues Neo4j Community
2026.07.1, database neo4j, explicit
CYPHER 25 for version-sensitive examples,
container atlasmart-neo4j, authentication
enabled, loopback Bolt/HTTP endpoints, no mandatory APOC/GDS
plugin, and AtlasMart domain identifiers established in
Chapters 01–10. Neo4j 5.26.30 remains the LTS
comparison line. Optional application examples use the
official Python driver 6.3.x.
This generation environment does not run Docker/Neo4j, so no
deadlock frequency, wait duration, retry count, latency, lock
list or cluster routing output is invented. The labs define
deterministic invariants and controlled concurrency procedures
that you execute on the disposable local instance. Community
can use SHOW TRANSACTIONS for its own work;
dbms.listActiveLocks() is currently
Enterprise-only and is therefore optional evidence, not a
mandatory lab dependency.
1. Managed transaction functions may execute more than once
session.execute_write(fn) commits when the callback
returns normally and retries retryable failures according to
driver policy. Current Python driver configuration exposes
max_transaction_retry_time, defaulting to 30
seconds. That is a retry time budget, not a guarantee of a
particular attempt count.
from neo4j import GraphDatabasedriver = GraphDatabase.driver( 'bolt://127.0.0.1:7687', auth=('neo4j','atlasmart-course-2026'), max_transaction_retry_time=15.0,)
2. Idempotency begins with business identity
If every attempt creates a new anonymous reservation, retry can
duplicate state. Instead, use a stable
reservationId supplied by the caller, constrain it
to be unique, MERGE the reservation by that
identity, and reconcile the final graph by that key.
CYPHER 25MATCH (s:Stock {stockId:'CH11-ST-001'})MERGE (r:Reservation {reservationId:$reservationId})ON CREATE SET r.quantity=$qty,r.createdAt=datetime(),r.labTag='ch11'WITH s,rMERGE (r)-[rel:RESERVES]->(s)ON CREATE SET rel.createdAt=datetime()RETURN r.reservationId,r.quantity;
3. External side effects make callback retries dangerous
Do not send an email, charge a card, publish an irreversible message or call a non-idempotent third-party API inside a managed transaction callback unless that side effect has its own idempotency/reconciliation protocol. A callback can fail after the external call and then be invoked again.
| Location | Safe default | Reason |
|---|---|---|
| Inside managed callback | Database reads/writes that are retry-safe | Driver may re-invoke the callback |
| After successful callback | External effect keyed by stable operation ID | Database result is known to the application |
| Outbox/event node in same transaction | Persist intent atomically | Separate worker can publish and mark delivery idempotently |
| On ambiguous client failure | Reconcile by operation ID | Do not infer rollback from network symptoms |
4. Do not manually retry every exception
Authentication errors, syntax errors, constraint violations expressing a real business conflict, and programmer errors are not fixed by sleeping and trying again. Use the driver’s classification and managed transaction APIs for transient work, set a bounded retry budget, and surface exhausted attempts with enough context to reconcile.
def reserve_tx(tx, reservation_id, qty): result = tx.run(""" MATCH (s:Stock {stockId:'CH11-ST-001'}) SET s._ch11_lock=true REMOVE s._ch11_lock WITH s WHERE s.onHand - s.reserved >= $qty MERGE (r:Reservation {reservationId:$reservationId}) ON CREATE SET r.quantity=$qty,r.labTag='ch11' WITH s,r MERGE (r)-[x:RESERVES]->(s) ON CREATE SET x.counted=false WITH s,r,x FOREACH (_ IN CASE WHEN x.counted THEN [] ELSE [1] END | SET s.reserved=s.reserved+r.quantity, s.version=s.version+1, x.counted=true) RETURN s.reserved AS reserved,r.reservationId AS reservationId """, reservationId=reservation_id, qty=qty) row=result.single() if row is None: raise ValueError('insufficient capacity') return dict(row)with driver.session(database='neo4j') as session: outcome=session.execute_write(reserve_tx,'CH11-R-042',2)print(outcome)
5. Retry telemetry is part of correctness evidence
Record the operation ID, attempt/final outcome, Neo4j error code/GQLSTATUS when available, total elapsed retry budget and the final reconciliation query. Do not collapse “one request eventually succeeded” into a clean success metric if it required repeated deadlock retries; that contention is an SLO and capacity signal.
Check your understanding
- Can a managed transaction callback run more than once?
- What is the current Python driver default max_transaction_retry_time?
- Why is a stable reservationId important?
- Should a payment API call be placed inside a retryable callback by default?
- Why log retries even if the final request succeeds?
Review the answers
1. Yes. The official driver explicitly warns that execute_write/execute_read callbacks can be retried.
2. 30 seconds; applications can configure a different bounded value.
3. It gives repeated attempts the same business identity and enables uniqueness/reconciliation.
4. No. Put non-idempotent external side effects outside or give them a separate idempotency protocol.
5. Retry pressure reveals contention/transient-failure behavior that affects latency, capacity and correctness risk.
Summary and next step
Retries repair transient database failure only when the unit of work is safe to repeat. Next we examine bookmarks, which solve a different problem: causal read-after-write ordering across sessions and routed cluster work.
Authoritative references
- Current Neo4j versions — Release/LTS snapshot used for this chapter.
- Database internals and transactional behavior — ACID behavior, read-committed isolation and transaction internals.
- Database transactions — Transaction lifecycle, memory and completion behavior.
- Concurrent data access — Locks, lost updates, contention, deadlocks and lock timeout semantics.
- Show and terminate transactions — SHOW TRANSACTIONS visibility and transaction diagnostics.
- Neo4j Python Driver 6.3 API — Current official Python driver, Bolt compatibility and transaction APIs.
- Python driver managed transactions — Managed transaction functions, retries and explicit transaction patterns.
- Python driver bookmarks — Bookmark propagation and causal chaining across sessions.