Chapter 16 · Pub/Sub, Keyspace Notifications, and Messaging Tradeoffs
Design Real-Time Notifications Without Mistaking Delivery Signals for Business State
Design AtlasMart real-time notifications so durable business state and replayable events remain authoritative while Pub/Sub is used only as a best-effort wake-up/fan-out signal.
Learning outcomes
AtlasMart's final Chapter 16 design must satisfy two requirements simultaneously: online dashboards should react quickly, and no order state transition may depend on a notification being delivered. The architecture therefore separates authoritative state, a replayable event, and a best-effort signal.
Design state-first notification flow where Pub/Sub never becomes the source of truth.
Use versioned notification payloads and re-read current state to survive duplicate, stale, or missed signals.
Keep a replayable Stream event for workflows that need catch-up while using Pub/Sub only for low-latency wake-up.
Apply least-privilege key and channel ACLs to a temporary Chapter 16 application user.
Specify reconnect, retry, failure-injection, observability, Cluster/Sentinel, and rollback behavior for production rollout.
All Chapter 16 mandatory labs reuse the disposable Chapter 01
environment: Redis Open Source 8.10.1 from pinned
Docker image redis:8.10.1, container
atlasmart-redis-ch01, standalone topology,
endpoint 127.0.0.1:6379, TLS disabled only
because traffic is loopback-local, default ACL user disabled,
named academy-admin and
atlasmart-app users, logical database 0, AOF with
appendfsync everysec plus RDB snapshots,
persistent /data, and no explicit
maxmemory/eviction policy. Python examples target
redis-py 8.1.0. The Chapter 01 application ACL
intentionally does not assume Pub/Sub command permissions, so
Pub/Sub/configuration demonstrations use
academy-admin until Lesson 5 creates a temporary
least-privilege Chapter 16 user. Fixtures stay under
atlasmart:ch16:.
1. Three layers, three different jobs
The order Hash answers “what is true now?” The Stream answers “what durable application events were recorded and can be replayed?” The Pub/Sub channel answers “which connected clients should wake up now?” Keeping those contracts separate makes disconnect loss harmless to correctness: a missed wake-up can be repaired by state read or Stream replay.
| Layer | AtlasMart example | Correctness role |
|---|---|---|
| Authoritative state | order Hash / primary database record | current truth |
| Replayable event | Redis Stream entry | catch-up/audit-like application processing within retention envelope |
| Ephemeral signal | Pub/Sub notification | reduce time-to-observe for online clients |
2. Commit durable state before publishing the hint
A robust sequence changes authoritative state and records the
durable event first. Then it publishes a lightweight hint. If
PUBLISH fails or reports zero live subscribers, the
order is still correct and the event remains replayable. If the
durable write fails, do not publish a success hint.
HSET atlasmart:ch16:l5:order:{1001} status paid version 1 updated_at 2026-09-06T12:00:00ZXADD atlasmart:ch16:l5:events:{1001} '*' event_id evt-1001-1 order_id 1001 version 1 status paidPUBLISH atlasmart:ch16:l5:notify:1001 '{"order_id":"1001","version":1}'HGETALL atlasmart:ch16:l5:order:{1001}XRANGE atlasmart:ch16:l5:events:{1001} - +# PUBLISH recipient count may be zero; state and event still prove the committed transition.
The three commands above are shown as a semantic sequence, not claimed to be one atomic unit. If your invariant requires the state mutation and Stream append to be atomic, use the Chapter 13/14 transaction/function techniques with Cluster same-slot keys; publish the best-effort hint only after that durable unit succeeds.
3. Version the signal so stale deliveries are harmless
Pub/Sub can be duplicated by overlapping subscriptions or application retries, and a client can process an older notification after newer state exists. Put an entity ID and monotonically meaningful application version in the hint. The receiver compares that hint with its local version, then re-reads authoritative state when necessary. The payload should not contain secrets merely because the channel is restricted.
HSET atlasmart:ch16:l5:order:{1001} status shipped version 2PUBLISH atlasmart:ch16:l5:notify:1001 '{"order_id":"1001","version":1}'HGET atlasmart:ch16:l5:order:{1001} versionHGET atlasmart:ch16:l5:order:{1001} status# A version=1 notification is stale; the authoritative record already says version=2 shipped.
4. Reconnect policy is part of the cache/notification protocol
On Pub/Sub reconnect, assume the gap may contain missed signals. The receiver should invalidate any state whose freshness depends on uninterrupted signals, re-fetch a checkpoint/snapshot, and—where exact missed events matter—replay the Stream from a stored ID/group state. Do not silently keep a local cache “because Redis reconnected successfully.”
| Reconnect step | Purpose |
|---|---|
| Mark signal-derived cache uncertain | Do not trust freshness across invalidation gap |
| Re-read current entity/version | Recover latest truth |
| Replay Stream if workflow needs every event | Recover durable missed history |
| Resume Pub/Sub only after baseline restored | Return to low-latency hints safely |
5. Least privilege spans commands, keys, and channels
Redis ACLs authorize key patterns and Pub/Sub channel patterns
separately. Since Redis 7.0, new Redis Open Source ACL users are
restrictive for Pub/Sub channels by default. The temporary user
below can read/write only Chapter 16 order/event keys and
publish/subscribe only Chapter 16 notification channels. It does
not receive CONFIG, ACL, or
all-channel access.
ACL SETUSER ch16-app reset on '>AtlasMart-Ch16-App-Lab-Only-2026' '~atlasmart:ch16:l5:*' 'resetchannels' '&atlasmart:ch16:l5:notify:*' '+@read' '+@write' '+@pubsub' '+ping' '+client|setname'ACL GETUSER ch16-appACL DRYRUN ch16-app PUBLISH atlasmart:ch16:l5:notify:1001 probeACL DRYRUN ch16-app PUBLISH other:tenant:notify probe# First DRYRUN should be allowed; second should be denied by channel pattern.# PSUBSCRIBE ACL matching has stricter literal-pattern considerations; test the exact pattern your client will use.
Channel ACLs prevent unauthorized Redis Pub/Sub access; they do not replace application authorization. A user allowed to receive a tenant channel can see every payload published there, so minimize payload sensitivity and tenant blast radius.
6. Cluster and Sentinel change reconnect/routing, not the source-of-truth rule
In Redis Cluster, place state and Stream keys that must
participate in one transaction/function into the same slot using
a deliberate hash tag such as {1001}. Ordinary
Pub/Sub is globally propagated across the Cluster; sharded
Pub/Sub can reduce Cluster-bus fan-out when channel slot design
matches traffic. Under Sentinel or Cluster failover, clients can
disconnect; therefore the post-reconnect reconciliation step
remains mandatory.
The Chapter 01 lab proves application semantics only. It does not measure Sentinel failover gaps, Cluster message propagation, shard hot spots, replica behavior, or managed-service routing. Build those later on the topology chapters and re-run the notification failure drill.
7. Durable processing stays idempotent
If a background service consumes
atlasmart:ch16:l5:events:{1001} through a Stream
consumer group, it can be redelivered after a crash before
acknowledgment. Use event_id or a business
idempotency key so replay is safe. The Pub/Sub hint should never
be used as the idempotency ledger.
XRANGE atlasmart:ch16:l5:events:{1001} - +XINFO STREAM atlasmart:ch16:l5:events:{1001}# If a consumer group is added, XPENDING/XACK become processing-state evidence.# The Pub/Sub channel itself exposes no historical event IDs.
8. Failure-injection matrix before production
| Injected failure | Expected correct behavior | Evidence |
|---|---|---|
| Subscriber offline during publish | state/event correct; UI catches up on reconnect | HGETALL + XRANGE + reconnect log |
| Duplicate notification | no duplicate business side effect | version/idempotency counters |
| Stale notification after newer write | receiver displays/re-reads newer version | version comparison + state read |
| Redis failover/reconnect | signal gap reconciled | client reconnect logs + snapshot/replay |
| Slow subscriber | bounded buffers; no memory runaway | CLIENT LIST + app queue depth |
| Unauthorized channel | ACL denial | ACL DRYRUN/real NOPERM test |
9. Observability that matches the contract
Measure publication rate and recipient counts only as live transport signals. Pair them with subscriber reconnect count, callback latency, output-buffer indicators, local queue depth, state-reconciliation failures, Stream lag/pending counts where durable processing exists, and entity-version mismatch counts. Alert on correctness-recovery failures rather than on “PUBLISH returned zero” alone; zero can simply mean no dashboard is online.
10. Cleanup and rollback
Delete only the Chapter 16 state/event fixtures and the temporary ACL user. No global notification configuration is changed in this lesson. If the user is still connected, deleting/disabling it may disconnect or deny subsequent commands; close test clients first.
UNLINK atlasmart:ch16:l5:order:{1001} atlasmart:ch16:l5:events:{1001}ACL DELUSER ch16-appACL USERS# Verify ch16-app is gone and no unrelated user/config was changed.
11. Production judgment and chapter synthesis
The safe architecture is intentionally asymmetric: durable state and replayable events carry correctness; Pub/Sub carries immediacy. Use explicit versions, idempotency, bounded retries/backoff, reconnection reconciliation, least-privilege channel ACLs, and topology-aware routing. If event retention/scale/independent failure domains exceed Redis operational goals, move the durable-event role to a purpose-built broker while keeping the same state-versus-signal separation.
Check your understanding
- Which layer is authoritative after a Pub/Sub gap?
- Why include a version in the notification?
- Should PUBLISH recipient count determine whether the order update succeeded?
- What should a client do after reconnect?
- What three ACL dimensions matter here?
Review the answers
The durable business state, and the replayable event log where event history is required.
To detect duplicate/stale hints and decide whether to re-read newer authoritative state.
No. The business update must succeed independently of whether any subscriber is online.
Treat signal-derived freshness as uncertain, re-read current state, and replay durable events if required.
Allowed commands, key patterns, and Pub/Sub channel patterns; application authorization is still separate.
12. Chapter summary and bridge
Chapter 16 separated live signals from durable processing. Pub/Sub is excellent for ephemeral fan-out when missed messages are recoverable; keyspace notifications are similarly lossy hints; Streams add replay and consumer state; external brokers may add broader event-platform capabilities. Chapter 17 moves underneath those mechanisms to persistence—RDB snapshots, AOF, fsync, rewrite, and crash recovery—where “data exists in Redis” must be translated into a tested recovery envelope.
Authoritative references
- Redis Open Source 8.10 release notes
- Redis Pub/Sub
- Redis Pub/Sub use case
- Redis Pub/Sub with redis-py
- PUBLISH
- SUBSCRIBE
- PSUBSCRIBE
- UNSUBSCRIBE
- PUNSUBSCRIBE
- SSUBSCRIBE
- SPUBLISH
- SUNSUBSCRIBE
- PUBSUB CHANNELS
- PUBSUB NUMSUB
- PUBSUB NUMPAT
- PUBSUB SHARDCHANNELS
- PUBSUB SHARDNUMSUB
- Redis keyspace notifications
- Redis subkey notifications
- CONFIG GET
- CONFIG SET
- CLIENT LIST
- CLIENT SETNAME
- Redis client output-buffer limits
- Redis Streams
- Redis streaming use case
- XADD
- XREADGROUP
- XPENDING
- XACK
- Redis ACLs
- ACL SETUSER
- ACL DRYRUN
- Redis Cluster specification
- Redis replication
- Redis Sentinel
- Redis persistence
- redis-py documentation
- redis-py 8.1.0 on PyPI
- Redis licenses