Chapter 23 · Application Patterns: Caching, Rate Limiting, Locks, Idempotency, and Hot-Key Control

Distributed Locks: SET NX PX, Ownership Tokens, Safe Unlock, Leases, and Fencing Limitations

Use expiring ownership-token locks safely, prove the stale-owner failure after lease expiry, and separate mutual exclusion from fencing of external side effects.

Advanced190–280 minutesSET NX PX, owner tokens, DELEX/Lua release, fencingRedis Open Source 8.10.1redis-py 8.1.0 where Python is usedDocker + redis-cli + Python stdlibStandalone loopback lab · DB 0AOF everysec + RDB · maxmemory 0/noeviction baselineNamed ACL users · TLS off only on loopbackFree/local-firstLast reviewed: September 6, 2026

Learning outcomes

This lesson turns Distributed Locks: SET NX PX, Ownership Tokens, Safe Unlock, Leases, and Fencing Limitations into an observable AtlasMart workflow with explicit correctness, failure, and production boundaries.

01

Explain the mechanisms and terminology behind Distributed Locks: SET NX PX, Ownership Tokens, Safe Unlock, Leases, and Fencing Limitations.

02

Collect Redis, client, configuration, and workload evidence before drawing operational conclusions.

03

Reproduce the lesson's deliberately incorrect or failure-prone case, diagnose the mechanism, and verify the repair.

04

Relate the design to memory, persistence, replication/Sentinel/Cluster, security, latency, and client behavior where applicable.

05

Apply the pattern to AtlasMart and state clearly what the implementation guarantees and what it does not guarantee.

Lab prerequisite: start the isolated Chapter 23 node

Run the Chapter 23 setup from Lesson 1 once. Before continuing, verify redis_version:8.10.1, authenticated identity atlasmart-app, database 0, AOF state, and that the endpoint is 127.0.0.1:6431. Do not point these failure/concurrency exercises at production.

Shell · verify existing Chapter 23 lab
docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-App-Lab-Only-2026 atlasmart-redis-ch23 redis-cli --user atlasmart-app PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Ch23-Admin-Lab-Only-2026 atlasmart-redis-ch23 redis-cli --user academy-admin INFO serverdocker exec -e REDISCLI_AUTH=AtlasMart-Ch23-App-Lab-Only-2026 atlasmart-redis-ch23 redis-cli --user atlasmart-app ACL WHOAMI
Reproducible Chapter 23 baseline

Redis Open Source 8.10.1 using the pinned redis:8.10.1 image, exposed only at 127.0.0.1:6431. Standalone topology, logical database 0, AOF everysec plus an RDB save rule, maxmemory 0 unless a lesson explicitly changes a setting on this disposable node, default ACL user disabled, named academy-admin and atlasmart-app users, and fixture prefix atlasmart:ch23:*. TLS is intentionally off only because this mandatory lab is loopback-local; production traffic must follow Chapter 22 network/TLS guidance. Python examples target redis==8.1.0. Search/JSON/vector/time-series/probabilistic features are not required.

1. Problem: a lock holder can pause longer than its lease

AtlasMart uses one worker at a time to rebuild a product feed. A Redis lock can reduce concurrent execution, but a process may pause for garbage collection, network delay, CPU starvation, or a long downstream call. If its lease expires, another worker can acquire the same lock while the first worker is still alive. That is the central distributed-lock boundary.

2. Acquire atomically with SET NX PX and a unique owner token

SET lock token NX PX lease-ms combines “only if absent” with an expiry. The token must be unique per acquisition, not just per process. Expiry prevents an indefinitely abandoned lock, but it also means ownership is temporary. Treat the lease duration as a failure assumption to measure, not a proof that work always finishes before expiry. When you also issue a fencing sequence number, acquire the lock and increment that sequence inside one Lua execution; a separate client-side SET followed later by INCR can pause between those operations and obtain a newer fence after it has already lost ownership.

Shell · one bounded lock acquisition
docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-App-Lab-Only-2026 atlasmart-redis-ch23   redis-cli --user atlasmart-app SET atlasmart:ch23:lock:feed token-A NX PX 5000docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-App-Lab-Only-2026 atlasmart-redis-ch23   redis-cli --user atlasmart-app PTTL atlasmart:ch23:lock:feed

3. Never release with an unconditional DEL

Suppose worker A pauses until its lease expires and worker B acquires the key. If A later executes DEL lock, it deletes B’s lock. Release must compare the current value with A’s token and delete only if they still match.

4. Redis 8.4+ provides conditional DELEX; Lua is the portable fallback

On the course baseline, DELEX key IFEQ token is one atomic conditional deletion. For older Redis versions, use the classic tiny compare-delete Lua script. Feature-gate current syntax if your deployment spans versions.

Shell · Redis 8.4+ safe release
docker exec -e REDISCLI_AUTH=AtlasMart-Ch23-App-Lab-Only-2026 atlasmart-redis-ch23   redis-cli --user atlasmart-app DELEX atlasmart:ch23:lock:feed IFEQ token-A
Lua · portable compare-delete release
if redis.call('GET', KEYS[1]) == ARGV[1] then  return redis.call('DEL', KEYS[1])endreturn 0

5. Lease renewal has the same ownership requirement

A watchdog that extends TTL must first prove the caller still owns the current token and then extend atomically. A plain PEXPIRE after reading the token in a separate round trip has a race. Renewal can improve liveness for known workloads but also lengthens failure recovery; set maximum work/retry time and instrument renewals.

6. A lease is not a fencing token

Even with perfect owner-aware release, worker A can continue writing to an external database after its Redis lease expired. Redis cannot retroactively stop that side effect. A fencing token is a monotonically increasing number attached to each ownership epoch; the protected downstream resource must reject operations carrying an older token than one it has already accepted. The demonstration acquires the owner token and increments the fencing counter atomically, and the {sku42} hash tag keeps those two keys in one slot if the pattern is later moved to Redis Cluster. If the downstream resource cannot enforce fencing, the Redis lease alone cannot provide that guarantee. Asynchronous failover can also lose recently issued Redis fencing increments, so stronger fencing requirements may need a sequence source with stronger failover guarantees.

7. Reproduce the stale-writer failure and fencing repair

The local Python harness intentionally pauses worker A beyond a 300-ms lease, lets B acquire a later fencing token, then attempts B’s and A’s external writes. The “external resource” is an in-memory mock so the safety mechanism is visible without modifying a real database.

Python · pause beyond lease and reject stale writer
import threading, time, uuidimport redisr = redis.Redis(host="127.0.0.1", port=6431, username="atlasmart-app",                password="AtlasMart-Ch23-App-Lab-Only-2026", decode_responses=True)LOCK="atlasmart:ch23:lock:{sku42}:owner"FENCE="atlasmart:ch23:lock:{sku42}:fence"r.delete(LOCK, FENCE)external={"last_fence":0, "value":None}mu=threading.Lock()ACQUIRE=r.register_script("""local lock_key = KEYS[1]local fence_key = KEYS[2]local token = ARGV[1]local lease_ms = tonumber(ARGV[2])if not redis.call('SET', lock_key, token, 'NX', 'PX', lease_ms) then return 0 endreturn redis.call('INCR', fence_key)""")def acquire(ms):    token=str(uuid.uuid4())    fence=int(ACQUIRE(keys=[LOCK, FENCE], args=[token, ms]))    return None if fence == 0 else (token, fence)def release(token):    return r.execute_command("DELEX", LOCK, "IFEQ", token)  # Redis 8.4+def fenced_write(fence, value):    with mu:        if fence <= external["last_fence"]:            return False        external["last_fence"] = fence        external["value"] = value        return Truea=acquire(300)assert aprint("A", a, "pttl", r.pttl(LOCK))time.sleep(0.45)                   # A pauses beyond its leaseb=acquire(1000)assert bprint("B", b, "pttl", r.pttl(LOCK))print("B write", fenced_write(b[1], "worker-B"))print("A stale write rejected", fenced_write(a[1], "worker-A"))print("A release after expiry", release(a[0]))print("B still owns", r.get(LOCK) == b[0])print("B release", release(b[0]), "final", external)

8. Acquisition retries must be bounded and randomized

Infinite retry loops convert contention or Redis failure into hung requests and synchronized traffic. Use a deadline, exponential/backoff with jitter where appropriate, and a clear “could not acquire” business outcome. Measure acquisition wait, lease renewals, expired-before-release count, contention rate, and downstream operation duration relative to lease.

9. Replication/failover changes lock guarantees

A lock written only to one primary can be lost during failover before asynchronous replication reaches the promoted replica. Redis’s distributed-lock documentation discusses stronger multi-instance algorithms, but no Redis lock removes the need to define what failure rate the business can tolerate. WAIT can improve replication acknowledgment but does not make the system universally linearizable or substitute for fencing.

10. Deliberately wrong: “if I hold the Redis key, I may write forever”

Holding the key at one instant proves only current Redis ownership. Repair by bounding work with a lease, checking owner token on release/renew, making side effects idempotent where possible, and using a fencing-capable downstream resource when stale owners must be rejected. Do not claim a lease itself fences external systems.

Check your understanding

  1. Why is DEL lock unsafe after a lease can expire?
  2. What does DELEX IFEQ add on Redis 8.4+?
  3. Does a unique token prevent an old owner from writing to PostgreSQL after expiry?
  4. Why bound acquisition retries?
Review the answers

The key may now belong to a different client, so unconditional deletion can remove another owner’s lock.

An atomic delete conditioned on the current string value matching the owner token.

No. The downstream system must enforce a fencing token or another stale-write defense.

To prevent indefinite request hangs and retry storms during contention or Redis failure.

11. Production judgment and bridge

Redis locks are useful when occasional coordination failure is acceptable or when downstream fencing/idempotency supplies the required safety. Keep critical sections short, measure pause tails, do not use one global lock for unrelated work, and test Redis failover, lease expiry, duplicate workers, and client timeouts. Lesson 4 shifts from mutual exclusion to retry-safe API semantics: duplicates should return the same logical result rather than relying on “only one request ever arrived.”

Summary and next step

Distributed Locks: SET NX PX, Ownership Tokens, Safe Unlock, Leases, and Fencing Limitations is now connected to observable Redis behavior, bounded failure cases, and production tradeoffs. Keep the evidence and cleanup state from this lesson; next, continue with Idempotency Keys, Deduplication, Request State Machines, and Retry-Safe APIs.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.