Chapter 11 · Transactions, Isolation, Locking, Deadlocks, Retries, and Consistency
Design and Load-Test a High-Contention Write Workflow Without Hiding Correctness Failures
Load-test a deliberately contended AtlasMart reservation workflow and prove correctness through graph reconciliation, not optimistic request metrics.
Learning outcomes
The final lab deliberately creates contention on one AtlasMart stock record. The goal is not to report the highest requests per second; it is to prove that every accepted reservation is represented exactly once, capacity is never exceeded, retries are bounded and failed/ambiguous requests remain reconcilable.
Define correctness invariants before generating concurrent load.
Run a bounded concurrent reservation workload with stable operation IDs.
Separate successful requests, retry pressure, rejected capacity and unexpected errors.
Reconcile stock counters against reservation facts after the run.
Identify redesign options for a hot entity without hiding changed semantics or consistency cost.
The mandatory lab continues Neo4j Community
2026.07.1, database neo4j, explicit
CYPHER 25 for version-sensitive examples,
container atlasmart-neo4j, authentication
enabled, loopback Bolt/HTTP endpoints, no mandatory APOC/GDS
plugin, and AtlasMart domain identifiers established in
Chapters 01–10. Neo4j 5.26.30 remains the LTS
comparison line. Optional application examples use the
official Python driver 6.3.x.
This generation environment does not run Docker/Neo4j, so no
deadlock frequency, wait duration, retry count, latency, lock
list or cluster routing output is invented. The labs define
deterministic invariants and controlled concurrency procedures
that you execute on the disposable local instance. Community
can use SHOW TRANSACTIONS for its own work;
dbms.listActiveLocks() is currently
Enterprise-only and is therefore optional evidence, not a
mandatory lab dependency.
1. Acceptance criteria come before throughput
| Invariant / SLO | How to verify |
|---|---|
| No oversubscription |
stock.reserved <= stock.onHand after every
accepted commit and final reconciliation
|
| Exactly one reservation per operation ID | Uniqueness constraint + duplicate-ID query |
| Counter equals reservation facts |
stock.reserved = sum(r.quantity) for counted
reservations
|
| Retry budget bounded | Driver max retry time + per-operation elapsed/attempt telemetry |
| No hidden failures | Count rejected capacity, deadlock/transient retries, ambiguous errors and unexpected errors separately |
| Transaction duration bounded | Measure transaction/request distributions; no sleeps/user waits inside real workflow |
2. Reset the isolated fixture
CYPHER 25CREATE CONSTRAINT ch11_stock_id IF NOT EXISTSFOR (s:Stock) REQUIRE s.stockId IS UNIQUE;CREATE CONSTRAINT ch11_reservation_id IF NOT EXISTSFOR (r:Reservation) REQUIRE r.reservationId IS UNIQUE;MERGE (p:Product {productId:'CH11-P-001'})SET p.name='AtlasCam Pro', p.labTag='ch11'MERGE (store:Store {storeId:'CH11-S-001'})SET store.name='Central Store', store.labTag='ch11'MERGE (stock:Stock {stockId:'CH11-ST-001'})SET stock.onHand=20, stock.reserved=0, stock.version=0, stock.labTag='ch11'MERGE (store)-[:HOLDS_STOCK]->(stock)-[:FOR_PRODUCT]->(p);MERGE (a:Counter {counterId:'CH11-A'}) SET a.value=0,a.labTag='ch11';MERGE (b:Counter {counterId:'CH11-B'}) SET b.value=0,b.labTag='ch11';
CYPHER 25MATCH (s:Stock {stockId:'CH11-ST-001'})OPTIONAL MATCH (s)<-[:RESERVES]-(r:Reservation {labTag:'ch11'})RETURN s.onHand AS onHand, s.reserved AS reserved, coalesce(sum(r.quantity),0) AS reservationQuantity, count(r) AS reservationCount, s.reserved <= s.onHand AS capacityInvariant, s.reserved = coalesce(sum(r.quantity),0) AS reconciliationInvariant;
3. Concurrent load harness
from concurrent.futures import ThreadPoolExecutor, as_completedfrom time import perf_counterfrom neo4j import GraphDatabaseURI='bolt://127.0.0.1:7687'AUTH=('neo4j','atlasmart-course-2026')def reserve_tx(tx, rid): row=tx.run(""" MATCH (s:Stock {stockId:'CH11-ST-001'}) SET s._ch11_lock=true REMOVE s._ch11_lock WITH s WHERE s.onHand-s.reserved >= 1 MERGE (r:Reservation {reservationId:$rid}) ON CREATE SET r.quantity=1,r.labTag='ch11' MERGE (r)-[x:RESERVES]->(s) ON CREATE SET x.counted=false WITH s,r,x FOREACH (_ IN CASE WHEN x.counted THEN [] ELSE [1] END | SET s.reserved=s.reserved+1,s.version=s.version+1,x.counted=true) RETURN r.reservationId AS rid,s.reserved AS reserved """, rid=rid).single() return None if row is None else dict(row)def one(driver, i): rid=f'CH11-R-{i:03d}' t0=perf_counter() try: with driver.session(database='neo4j') as session: result=session.execute_write(reserve_tx,rid) return rid,'accepted' if result else 'capacity_rejected',(perf_counter()-t0)*1000,None except Exception as exc: return rid,'error',(perf_counter()-t0)*1000,repr(exc)with GraphDatabase.driver(URI,auth=AUTH,max_transaction_retry_time=10.0) as driver: with ThreadPoolExecutor(max_workers=8) as pool: futures=[pool.submit(one,driver,i) for i in range(1,41)] rows=[f.result() for f in as_completed(futures)]for row in sorted(rows): print(row)
Forty requests compete for twenty units. The expected business outcome is at most twenty accepted one-unit reservations, but exact scheduling, retries and latency are deliberately not pre-filled. Your evidence is the result classification plus the final graph invariants.
4. Reconcile rather than trust request logs
CYPHER 25MATCH (s:Stock {stockId:'CH11-ST-001'})OPTIONAL MATCH (s)<-[:RESERVES]-(r:Reservation {labTag:'ch11'})RETURN s.onHand AS onHand, s.reserved AS reserved, coalesce(sum(r.quantity),0) AS reservationQuantity, count(r) AS reservationCount, s.reserved <= s.onHand AS capacityInvariant, s.reserved = coalesce(sum(r.quantity),0) AS reconciliationInvariant;
CYPHER 25MATCH (r:Reservation {labTag:'ch11'})WITH r.reservationId AS id,count(*) AS copies,sum(r.quantity) AS qtyWHERE copies <> 1 OR qty <> 1RETURN id,copies,qty;MATCH (r:Reservation {labTag:'ch11'})WHERE NOT (r)-[:RESERVES]->(:Stock {stockId:'CH11-ST-001'})RETURN r.reservationId AS orphanReservation;
5. Hot-node mitigation is a design decision
| Option | Benefit | New cost/risk |
|---|---|---|
| Keep one stock node + short locked transaction | Strong, simple invariant | Serializes conflicting reservations |
| Partition stock by Store/SKU/location | Spreads writes along real business dimensions | Cross-partition availability queries/rebalancing |
| Reservation fact nodes + asynchronous aggregate | Append-oriented workflow | Aggregate is derived/eventually updated; reads need freshness semantics |
| Pre-allocated inventory buckets | Reduces central contention | Allocation/reconciliation complexity |
| Queue/serialize commands before Neo4j | Controls write concurrency | Extra infrastructure and queue availability/ordering semantics |
Do not label an eventually reconciled design “equivalent” to an immediately checked inventory invariant. Performance changes that weaken guarantees must be named as product/architecture changes.
6. Deliberately wrong: hide retries and report only successful throughput
A benchmark that discards deadlock retries, capacity rejections or error latency can look excellent while users experience poor tails and the database spends resources redoing work. Report latency distributions by outcome, retry/deadlock counts, concurrency, graph shape, resource limits and final invariant results.
Check your understanding
- What is the first success criterion of this load test?
- Why use stable reservation IDs under load?
- What should happen when forty requests compete for twenty units?
- Can moving to asynchronous aggregates be described as a pure performance optimization?
- What should a final report include beyond requests per second?
Review the answers
1. Business invariant preservation, not maximum throughput.
2. They make retries and reconciliation idempotent and auditable.
3. At most twenty should be accepted; the rest should be explicitly rejected or fail visibly, never silently oversubscribe capacity.
4. No. It changes freshness/consistency semantics and must be evaluated as an architecture decision.
5. Outcome counts, p50/p95/p99 by outcome, retry/deadlock pressure, graph/concurrency settings, resource state and reconciliation/invariant results.
7. Cleanup and Chapter 12 bridge
Remove only Chapter 11 disposable entities/constraints when finished. Chapter 12 moves into APOC/procedures/functions and the operational/security consequences of extending Cypher.
CYPHER 25MATCH (n) WHERE n.labTag='ch11' DETACH DELETE n;DROP CONSTRAINT ch11_stock_id IF EXISTS;DROP CONSTRAINT ch11_reservation_id IF EXISTS;
Summary
Correct concurrent systems make failure modes observable. Neo4j provides transactional atomicity, read-committed isolation, locks, deadlock detection, managed retry primitives and causal bookmarks; AtlasMart still has to define idempotency, contention ownership, retry budgets, external side-effect coordination and reconciliation.
Authoritative references
- Current Neo4j versions — Release/LTS snapshot used for this chapter.
- Database internals and transactional behavior — ACID behavior, read-committed isolation and transaction internals.
- Database transactions — Transaction lifecycle, memory and completion behavior.
- Concurrent data access — Locks, lost updates, contention, deadlocks and lock timeout semantics.
- Show and terminate transactions — SHOW TRANSACTIONS visibility and transaction diagnostics.
- Neo4j Python Driver 6.3 API — Current official Python driver, Bolt compatibility and transaction APIs.
- Python driver managed transactions — Managed transaction functions, retries and explicit transaction patterns.
- Python driver bookmarks — Bookmark propagation and causal chaining across sessions.