Chapter 13 · Transactions, WATCH, Optimistic Locking, and Atomicity
Build and Stress-Test a Conditional Update Workflow Under Concurrent Clients
Stress a bounded AtlasMart reservation workflow with concurrent redis-py clients, measure conflicts/retries/tail latency, and verify the final invariant instead of trusting happy-path tests.
Learning outcomes
A WATCH loop that works with two terminals may still collapse under a hot production key. The Chapter 13 capstone turns AtlasMart inventory reservation into a measured concurrency experiment: many clients compete for one bounded stock value, every successful reservation updates a second Redis key in the same transaction, and the final invariant is checked exactly.
Build a redis-py WATCH/MULTI/EXEC workflow that never allows synthetic stock to go below zero.
Collect watch conflicts, retries, exhausted attempts, committed/sold-out outcomes, and p50/p95/p99 latency.
Verify the final invariant from Redis state instead of trusting per-client success logs.
Vary concurrency deliberately and relate contention to retries/tail latency without inventing universal thresholds.
Identify when hot-key contention should trigger a primitive/data-model redesign or server-side logic.
All Chapter 13 mandatory labs reuse the disposable Chapter 01
environment: Redis Open Source 8.10.1 from Docker
Official Image redis:8.10.1, container
atlasmart-redis-ch01, standalone topology, host
publication 127.0.0.1:6379, TLS disabled only
because traffic stays on loopback, default ACL user disabled,
named users academy-admin and
atlasmart-app, logical database 0, AOF with
appendfsync everysec plus RDB snapshots,
persistent /data, and no explicit
maxmemory limit or eviction policy. Transaction
exercises use academy-admin because the Chapter
01 application ACL was intentionally not broadened with the
separate @transaction category; production
applications should grant only the commands and key patterns
they need. Fixtures stay under atlasmart:ch13:*.
Client-oriented exercises pin redis-py 8.1.0 on
Python 3.12 and use one pipeline/context per watched
transaction so connection-scoped WATCH state is not
accidentally lost.
1. Define the invariant before the load test
The synthetic SKU starts with 60 units. A successful logical
reservation must decrement stock by exactly one and increment
one worker field in the sold Hash. At the end,
final_stock >= 0 and
final_stock + exact_sold == initial_stock. These
exact assertions are more important than any throughput number.
| State | Role |
|---|---|
| atlasmart:ch13:l5:{sku-42}:stock | Authoritative synthetic remaining units |
| atlasmart:ch13:l5:{sku-42}:sold | Per-worker committed reservation counts |
| WATCH conflicts | Contention evidence |
| p50/p95/p99 operation latency | User-visible cost distribution |
2. Why the keys use one hash tag
The mandatory lab is standalone, but the keys deliberately use
{sku-42}. If the workflow later moves to Redis Open
Source Cluster, the stock and sold keys must share one hash slot
for the transaction. This naming choice is documented rather
than hidden.
UNLINK atlasmart:ch13:l5:{sku-42}:stock atlasmart:ch13:l5:{sku-42}:soldSET atlasmart:ch13:l5:{sku-42}:stock 60GET atlasmart:ch13:l5:{sku-42}:stock# Expected: "60"
3. The conditional reservation protocol
Each attempt watches the stock key, reads it, exits with UNWATCH
if sold out, or enters MULTI and queues DECR plus HINCRBY. If
another client changes stock before EXEC, redis-py raises
WatchError; the operation re-reads within a finite
retry budget. The sold Hash is written only in the same
successful transaction as the decrement.
| Step | Reason |
|---|---|
| WATCH stock | Detect any intervening stock change |
| GET stock | Decide whether reservation is still possible |
| UNWATCH on zero | Release connection state without pointless transaction |
| MULTI + DECR + HINCRBY | Queue coupled Redis changes |
| EXEC | Commit only if watched stock stayed unchanged |
| WatchError | Retry from a fresh read, within budget |
4. Pin the client and run the stress harness
The harness uses redis-py 8.1.0 and Python’s standard-library thread pool. It intentionally does not add random sleeps by default; you can vary concurrency in a controlled second run. Every thread uses its own Redis client/pool, while each watched operation uses a pipeline context that pins the transaction connection.
python -m pip install "redis==8.1.0"python -c "import redis; print(redis.__version__)"
from concurrent.futures import ThreadPoolExecutorfrom collections import Counterimport statisticsimport threadingimport timeimport redisHOST = "127.0.0.1"PORT = 6379USER = "academy-admin"PASSWORD = "AtlasMart-Admin-Lab-Only-2026"STOCK_KEY = "atlasmart:ch13:l5:{sku-42}:stock"SOLD_KEY = "atlasmart:ch13:l5:{sku-42}:sold"INITIAL_STOCK = 60WORKERS = 8ATTEMPTS_PER_WORKER = 20MAX_WATCH_ATTEMPTS = 10admin = redis.Redis(host=HOST, port=PORT, username=USER, password=PASSWORD, decode_responses=True)admin.unlink(STOCK_KEY, SOLD_KEY)admin.set(STOCK_KEY, INITIAL_STOCK)results = Counter()latencies_ms = []retries = []metrics_lock = threading.Lock()def reserve(worker_id: int): client = redis.Redis(host=HOST, port=PORT, username=USER, password=PASSWORD, decode_responses=True) local = Counter() local_latencies = [] local_retries = [] for _ in range(ATTEMPTS_PER_WORKER): start = time.perf_counter() for attempt in range(1, MAX_WATCH_ATTEMPTS + 1): try: with client.pipeline() as pipe: pipe.watch(STOCK_KEY) stock = int(pipe.get(STOCK_KEY) or 0) if stock <= 0: pipe.unwatch() local["sold_out"] += 1 local_retries.append(attempt - 1) break pipe.multi() pipe.decr(STOCK_KEY) pipe.hincrby(SOLD_KEY, f"worker:{worker_id}", 1) pipe.execute() local["committed"] += 1 local_retries.append(attempt - 1) break except redis.WatchError: local["watch_conflict"] += 1 if attempt == MAX_WATCH_ATTEMPTS: local["retry_exhausted"] += 1 local_retries.append(attempt) local_latencies.append((time.perf_counter() - start) * 1000) with metrics_lock: results.update(local) latencies_ms.extend(local_latencies) retries.extend(local_retries)with ThreadPoolExecutor(max_workers=WORKERS) as pool: list(pool.map(reserve, range(WORKERS)))final_stock = int(admin.get(STOCK_KEY))exact_sold = sum(int(v) for v in admin.hvals(SOLD_KEY))assert final_stock >= 0assert final_stock + exact_sold == INITIAL_STOCKlatencies_ms.sort()def pct(values, p): if not values: return 0.0 i = min(len(values) - 1, max(0, int(round((p / 100) * (len(values) - 1))))) return values[i]print(dict(results))print("final_stock", final_stock, "exact_sold", exact_sold)print("retry_mean", round(statistics.fmean(retries), 3))print("latency_ms", {"p50": round(pct(latencies_ms, 50), 3), "p95": round(pct(latencies_ms, 95), 3), "p99": round(pct(latencies_ms, 99), 3)})
5. Read the output as measurements, not promised numbers
The script prints actual conflict/retry counts, exact final Redis state, and observed latency percentiles from your machine. No particular p95 or conflict percentage is a universal pass/fail threshold because CPU scheduling, Docker networking, persistence, payloads, and machine load change the result. The invariant assertions are deterministic; performance is empirical.
| Output | Interpretation |
|---|---|
| committed | Number of successful reservations |
| sold_out | Attempts that observed zero stock |
| watch_conflict | Optimistic conflicts before successful/failed attempt |
| retry_exhausted | Operations that hit the explicit retry budget |
| retry_mean | Client work amplification |
| latency p50/p95/p99 | Distribution; tail growth is especially important |
6. Verify authoritative Redis state independently
Never trust only thread-local counters. Re-read Redis after the workload and reconcile the two keys.
GET atlasmart:ch13:l5:{sku-42}:stockHGETALL atlasmart:ch13:l5:{sku-42}:soldHLEN atlasmart:ch13:l5:{sku-42}:soldMEMORY USAGE atlasmart:ch13:l5:{sku-42}:stockMEMORY USAGE atlasmart:ch13:l5:{sku-42}:sold# Sum sold fields and verify final_stock + sold == 60 and final_stock >= 0.
7. Run a controlled concurrency sweep
Change only WORKERS—for example 1, 2, 4, then
8—reset the two keys before each run, and record conflict counts
and p50/p95/p99. Keep INITIAL_STOCK, attempts,
Redis configuration, persistence, payloads, and machine load as
constant as practical. This isolates how concurrent clients
affect the optimistic retry loop.
workers | committed | watch_conflicts | retry_exhausted | retry_mean | p50_ms | p95_ms | p99_ms1 | | | | | | |2 | | | | | | |4 | | | | | | |8 | | | | | | |
8. Deliberately wrong design: infinite retries on one hot key
Remove the attempt bound and the system can self-amplify: more conflicts create more retries, which create more competing commands. Tail latency grows and useful work may fall. Repair by restoring the budget, surfacing retry exhaustion, and setting a redesign trigger based on measured conflict/latency/SLO evidence.
9. Compare against a simpler or server-side primitive
The condition here is “decrement only when positive and update another key.” A single INCR/DECR cannot express the whole invariant; SET IFEQ can handle one exact String comparison but not the coupled Hash update. WATCH is therefore valid. If conflict is structurally high, Chapter 14 can move the bounded check/decrement/update logic into one script/function invocation and compare latency/operational tradeoffs.
10. External side effects remain outside Redis transaction atomicity
Do not place “charge card” or “ship package” into the mental transaction. The Redis transaction covers only its Redis commands. A production reservation workflow needs idempotency keys, durable event/state transitions, reconciliation, and explicit behavior when Redis succeeds but an external call fails—or vice versa.
11. Persistence, replication, failover, and Cluster boundaries
The stress test proves the invariant on one running standalone node. AOF everysec and RDB govern crash recovery; asynchronous replication governs replica freshness; Sentinel/Cluster failover can expose separate loss/retry windows; Cluster requires same-slot keys. Those properties are not proven by a zero-negative-stock assertion. Later chapters test them explicitly.
12. Security and tenant isolation
All keys are synthetic and the lab uses the broad disposable admin identity because Chapter 01’s app ACL does not include the transaction category. A production app should use a least-privilege user constrained to the relevant tenant/SKU key patterns and commands. Key prefixes/hash tags are naming/routing tools, not tenant authorization by themselves.
13. Verification checklist
- The script version prints redis-py 8.1.0 and the server remains Redis 8.10.1.
- No run leaves stock below zero.
- For every run, final_stock + exact_sold equals the reset initial stock.
- Retry/conflict and latency distributions come from actual runs, not example numbers.
- Each concurrency run resets only the two Chapter 13 capstone keys.
- Any retry-exhausted result is visible rather than silently retried forever.
14. Cleanup
UNLINK atlasmart:ch13:l5:{sku-42}:stock atlasmart:ch13:l5:{sku-42}:sold
15. Production judgment
WATCH-based optimistic updates are strongest when conflicts are genuinely occasional. Treat conflict rate and tail latency as capacity/data-model signals. Keep retry budgets finite, use hash-slot-compatible keys where Cluster is a target, and separate Redis correctness from durability/failover and external side effects. When contention is persistently high, compare a server-side function/script, partitioned ownership, or a different business coordination model rather than simply increasing retries.
Check your understanding
- What is the capstone’s exact invariant?
- Why collect watch_conflict separately from retry_exhausted?
- Why use p95/p99 instead of only an average?
- Does a passing standalone stress test prove Cluster/failover correctness?
- What is the likely next step if conflicts remain high at realistic load?
Review the answers
Remaining stock is never negative, and remaining stock plus exact sold reservations equals the initial stock.
Conflicts are normal optimistic contention; exhaustion is a caller-visible inability to complete within the retry budget.
Contention often appears in the latency tail; averages can hide slow retry-heavy operations.
No. Slot locality, topology changes, replication, and failover introduce separate behavior that must be tested.
Reduce coordination scope or compare a small bounded server-side atomic script/function or another data-model design—not unlimited retries.
16. Chapter summary and bridge to Chapter 14
Chapter 13 established four distinct ideas: MULTI/EXEC gives uninterrupted queued execution but no SQL rollback; WATCH supplies optimistic conditional execution; error classes leave different final states; and the simplest atomic command should be preferred when it already expresses the invariant. The capstone turned these semantics into measurable contention evidence. Chapter 14 now asks whether small race-prone multi-round-trip logic should move into Lua scripts or Redis Functions—and what blocking, key declaration, ACL, deployment, and versioning responsibilities come with that move.
Authoritative references
- Redis Open Source 8.10 release notes
- Redis 8.10 command reference
- Redis transactions
- MULTI
- EXEC
- DISCARD
- WATCH
- UNWATCH
- Redis multi-key operations
- Redis pipelining
- redis-py pipelines and transactions
- redis-py documentation
- SET
- DELEX
- MSETNX
- INCR
- HINCRBY
- EVAL
- FCALL
- Redis ACLs
- Redis persistence
- Redis replication
- Redis licenses