Chapter 14 · Lua Scripting and Redis Functions
Scripting Constraints: Blocking Time, Determinism, Replication, and Cluster Key Access
Measure why atomic Lua execution can block every client, keep scripts deterministic and bounded, and design declared key access that remains routable in Redis Cluster.
Learning outcomes
AtlasMart engineers now know how to write a correct small script, but correctness inside the script is not enough. Because Lua executes atomically on the Redis server, a slow or badly bounded script can turn one client's convenience into every client's tail-latency problem. Cluster and replication also impose explicit key and determinism disciplines.
Explain why atomic scripts block other client command execution and measure bounded blocking with SLOWLOG evidence.
Describe the configured Lua execution-time threshold and the safe limits of SCRIPT KILL.
Keep script inputs deterministic and understand current write-command replication behavior without relying on old verbatim-script folklore.
Validate Cluster hash-slot locality with declared keys and hash tags before a multi-key script is deployed.
Reject hidden/dynamic key discovery, unbounded loops, and large scans as server-side scripting patterns.
All Chapter 14 mandatory labs reuse the disposable Chapter 01
environment: Redis Open Source 8.10.1 from pinned
Docker image redis:8.10.1, container
atlasmart-redis-ch01, standalone topology, host
endpoint 127.0.0.1:6379, TLS disabled only
because traffic remains on loopback, default ACL user
disabled, named academy-admin and
atlasmart-app users, logical database 0, AOF with
appendfsync everysec plus RDB snapshots,
persistent /data, and no explicit
maxmemory/eviction policy. Client examples that
need Python use redis-py 8.1.0. Chapter-specific
fixtures stay under atlasmart:ch14:.
1. Atomicity has a blocking cost
A Redis script is not a background worker. Its Redis-side execution is atomic, and while it runs other client commands wait. That property is what prevents interleaving, but it also means script runtime directly contributes to server tail latency. The right question is not “can Lua do this?” but “is this bounded enough to execute in the Redis request path?”
| Script shape | Correctness | Latency risk | Better home |
|---|---|---|---|
| 2–5 constant-time key operations | Often a good fit | Usually small; still measure | Redis script/function |
| Loop over caller-bounded 10 items | Potentially acceptable | Scales with bound | Measure and cap |
| SCAN entire keyspace | Mechanically possible | Unbounded/global blocking | Client-side incremental SCAN |
| Remote HTTP call | Sandbox disallows it | Not applicable | Application/service layer |
| Large analytics loop | May be logically correct | Severe blocking | Batch/analytics system |
2. Measure a bounded slow script without changing host settings
The lab temporarily lowers the Redis Slow Log threshold inside
the disposable container, runs a bounded CPU loop, captures
SLOWLOG GET, and restores the exact previous
threshold in finally. It does not change host
kernel/network settings and does not use an infinite loop.
import redis, timer = redis.Redis(host="127.0.0.1", port=6379, username="academy-admin", password="AtlasMart-Admin-Lab-Only-2026", decode_responses=True)old = int(r.config_get("slowlog-log-slower-than")["slowlog-log-slower-than"])try: r.config_set("slowlog-log-slower-than", 100) r.slowlog_reset() script = "local x=0; for i=1,200000 do x=x+i end; return x" t0 = time.perf_counter() result = r.eval(script, 0) elapsed_ms = (time.perf_counter() - t0) * 1000 print("result", result, "client_ms", round(elapsed_ms, 3)) print("slowlog", r.slowlog_get(5))finally: r.config_set("slowlog-log-slower-than", old)
The exact latency is environment-dependent. The lesson requires recording the measured client latency and Slow Log duration; it does not provide a fabricated “normal” number.
3. lua-time-limit is a safety threshold, not a performance target
Redis exposes lua-time-limit (historically
defaulting to seconds-scale protection) to detect very long
scripts. A script crossing the threshold is not automatically
rolled back or magically preempted; Redis changes how it serves
requests so operators can diagnose/kill where allowed.
Production scripts should be designed to finish orders of
magnitude faster than any emergency threshold and validated
under realistic load.
CONFIG GET lua-time-limitSLOWLOG GET 10# Record the configured value. Do not raise it merely to make a slow script "work".
4. SCRIPT KILL has an atomicity boundary
SCRIPT KILL can terminate a running Eval script
only if it has not already performed writes. Once a script has
modified the dataset, killing it would expose a partially
applied atomic operation, so Redis refuses that route. This is
why “we can kill it if it gets slow” is not a design strategy.
Do not test an infinite write script and plan to recover with
SCRIPT KILL. A bounded read-only loop is enough
to study latency and kill semantics without creating a
dangerous server state.
5. Determinism and replication: teach the current mechanism
Current Redis propagates the write commands performed by a script to replicas/AOF rather than requiring applications to reason from obsolete “replicate the script source everywhere” assumptions. Scripts still need deterministic, database-local behavior because application correctness must not depend on hidden host time, network, files, or process state—the Lua environment is sandboxed specifically to keep execution tied to Redis data and explicit arguments.
SET atlasmart:ch14:l2:price 100EVAL "local p=tonumber(redis.call('GET',KEYS[1])); local pct=tonumber(ARGV[1]); return p*(100-pct)/100" 1 atlasmart:ch14:l2:price 15# Same Redis state + same ARGV gives the same business calculation. External network/filesystem access is not part of the Lua sandbox.
6. Cluster routing begins before Lua runs
Redis Cluster has 16,384 hash slots. For a single script
invocation, declared keys must be routable together; the normal
design is to make all involved keys share a hash tag. The lab
remains standalone, so CLUSTER KEYSLOT may be
unavailable as a command against non-cluster mode; instead, this
section provides commands to run later in the Chapter 21 Cluster
lab and explains the data-model requirement now.
CLUSTER KEYSLOT atlasmart:ch14:{order:9001}:stockCLUSTER KEYSLOT atlasmart:ch14:{order:9001}:reservationCLUSTER KEYSLOT atlasmart:ch14:{order:9002}:reservation# The first two should match because the hash tag is {order:9001}; the third intentionally differs.
Redis Lua supports advanced script flags, including
allow-cross-slot-keys, but declared input keys
are still required to hash to one slot and cross-slot
scripting is discouraged because it conflicts with
movable-slot architecture. This course treats same-slot key
design as the default.
7. Hidden key access is a deployment bug
A script that reads a “next key name” from a Redis value and then calls that key dynamically may work on standalone Redis, but Redis cannot reliably route or reason about undeclared keys in Cluster. The same problem appears when a script concatenates an ARGV prefix with an ID to invent a key.
local hidden = ARGV[1] .. ':' .. ARGV[2]return redis.call('GET', hidden)-- Wrong: hidden is a key name but was not supplied in KEYS.
return redis.call('GET', KEYS[1])-- Call with numkeys=1 and pass the complete key name explicitly.
8. Large scans and loops belong outside atomic server-side logic
Even if each individual Redis command is legal from Lua,
combining SCAN-like discovery or iterating large
collections inside one atomic script defeats Redis's
incremental/nonblocking design. Fetch bounded identifiers in the
application, use incremental iteration, and reserve scripts for
the small final state transition that actually needs atomicity.
| Need | Do in client/application | Do atomically in Redis |
|---|---|---|
| Find candidate orders | SCAN/Search/query with pagination | No |
| Compute large recommendation set | Application/analytics | No |
| Check stock + reserve one order | Pass explicit keys/qty | Yes, small script/function |
| Call payment API | Application | Never from Redis Lua |
9. Replication/failover evidence has topology prerequisites
The mandatory Chapter 14 lab is standalone, so it cannot honestly prove replica propagation or failover behavior. It records the persistence/replication assumptions and defers actual primary-replica/Sentinel failure tests to Chapters 19–20. What this chapter can prove locally is that Eval cache state is volatile and Functions are persisted with the dataset.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin PINGdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin INFO serverdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin CONFIG GET appendonly appendfsync maxmemory maxmemory-policy lua-time-limitdocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin ACL WHOAMI# Record redis_version, persistence settings, maxmemory policy, and ACL identity before interpreting results.
10. Verification checklist and cleanup
- The bounded loop produced a measured duration and a Slow Log entry after the temporary threshold change.
- The original Slow Log threshold was restored even if the experiment failed.
- You can explain why SCRIPT KILL is not a universal escape hatch after writes.
- Your multi-key design names all keys explicitly and gives related keys a common Cluster hash tag.
- No global key scan, external network call, unbounded loop, or fabricated failover result is part of the mandatory lab.
UNLINK atlasmart:ch14:l2:price
In production, bound script work by input size and command complexity, watch p95/p99 latency and Slow Log/latency events, keep external work out of Lua, and test Cluster routing before rollout. Server-side atomicity is valuable precisely because it blocks interleaving; that same mechanism makes large work dangerous.
Check your understanding
- Why can a correct script still be operationally unsafe?
- Can SCRIPT KILL always stop a slow script?
- What is the default Cluster design for multi-key scripts?
- Should lua-time-limit be treated as an acceptable runtime target?
- Did this standalone lab prove failover behavior?
Review the answers
Atomic execution blocks other Redis command execution for the duration, so long work creates tail-latency and availability risk.
No. It is allowed only when the script has not already performed writes.
Declare every key and place all keys for one invocation in the same hash slot, commonly with a shared hash tag.
No. It is an emergency safety threshold; application scripts should be measured and much shorter.
No. It states that limitation and defers topology failure evidence to the replication/Sentinel chapters.
11. Summary and next step
Atomic Lua is safe only when its work is explicit, bounded, routable, and measured. Lesson 3 now moves that small logic into Redis Functions so deployment, persistence, naming, introspection, and versioning become first-class operational concerns.
Authoritative references
- Redis Open Source 8.10 release notes
- Redis programmability
- Scripting with Lua
- Redis Lua API reference
- Redis Functions introduction
- EVAL
- EVAL_RO
- EVALSHA
- EVALSHA_RO
- SCRIPT EXISTS
- SCRIPT LOAD
- SCRIPT FLUSH
- SCRIPT KILL
- FCALL
- FCALL_RO
- FUNCTION LOAD
- FUNCTION LIST
- FUNCTION STATS
- FUNCTION DUMP
- FUNCTION RESTORE
- FUNCTION DELETE
- Redis ACLs
- ACL DRYRUN
- Redis Cluster specification
- Scale with Redis Cluster
- Redis latency diagnosis
- SLOWLOG
- Redis persistence
- Redis replication
- redis-py documentation
- redis-py 8.1.0 on PyPI
- Redis licenses