Chapter 16 · Pub/Sub, Keyspace Notifications, and Messaging Tradeoffs

Design Real-Time Notifications Without Mistaking Delivery Signals for Business State

Design AtlasMart real-time notifications so durable business state and replayable events remain authoritative while Pub/Sub is used only as a best-effort wake-up/fan-out signal.

Advanced170–230 minutesDurable state plus best-effort real-time signalsRedis Open Source 8.10.1redis-py 8.1.0Free/local-firstLast reviewed: September 6, 2026

Learning outcomes

AtlasMart's final Chapter 16 design must satisfy two requirements simultaneously: online dashboards should react quickly, and no order state transition may depend on a notification being delivered. The architecture therefore separates authoritative state, a replayable event, and a best-effort signal.

01

Design state-first notification flow where Pub/Sub never becomes the source of truth.

02

Use versioned notification payloads and re-read current state to survive duplicate, stale, or missed signals.

03

Keep a replayable Stream event for workflows that need catch-up while using Pub/Sub only for low-latency wake-up.

04

Apply least-privilege key and channel ACLs to a temporary Chapter 16 application user.

05

Specify reconnect, retry, failure-injection, observability, Cluster/Sentinel, and rollback behavior for production rollout.

Exact lab baseline

All Chapter 16 mandatory labs reuse the disposable Chapter 01 environment: Redis Open Source 8.10.1 from pinned Docker image redis:8.10.1, container atlasmart-redis-ch01, standalone topology, endpoint 127.0.0.1:6379, TLS disabled only because traffic is loopback-local, default ACL user disabled, named academy-admin and atlasmart-app users, logical database 0, AOF with appendfsync everysec plus RDB snapshots, persistent /data, and no explicit maxmemory/eviction policy. Python examples target redis-py 8.1.0. The Chapter 01 application ACL intentionally does not assume Pub/Sub command permissions, so Pub/Sub/configuration demonstrations use academy-admin until Lesson 5 creates a temporary least-privilege Chapter 16 user. Fixtures stay under atlasmart:ch16:.

1. Three layers, three different jobs

The order Hash answers “what is true now?” The Stream answers “what durable application events were recorded and can be replayed?” The Pub/Sub channel answers “which connected clients should wake up now?” Keeping those contracts separate makes disconnect loss harmless to correctness: a missed wake-up can be repaired by state read or Stream replay.

Layer AtlasMart example Correctness role
Authoritative state order Hash / primary database record current truth
Replayable event Redis Stream entry catch-up/audit-like application processing within retention envelope
Ephemeral signal Pub/Sub notification reduce time-to-observe for online clients

2. Commit durable state before publishing the hint

A robust sequence changes authoritative state and records the durable event first. Then it publishes a lightweight hint. If PUBLISH fails or reports zero live subscribers, the order is still correct and the event remains replayable. If the durable write fails, do not publish a success hint.

redis-cli · deterministic state/event/signal sequence
HSET atlasmart:ch16:l5:order:{1001} status paid version 1 updated_at 2026-09-06T12:00:00ZXADD atlasmart:ch16:l5:events:{1001} '*' event_id evt-1001-1 order_id 1001 version 1 status paidPUBLISH atlasmart:ch16:l5:notify:1001 '{"order_id":"1001","version":1}'HGETALL atlasmart:ch16:l5:order:{1001}XRANGE atlasmart:ch16:l5:events:{1001} - +# PUBLISH recipient count may be zero; state and event still prove the committed transition.
Atomicity boundary

The three commands above are shown as a semantic sequence, not claimed to be one atomic unit. If your invariant requires the state mutation and Stream append to be atomic, use the Chapter 13/14 transaction/function techniques with Cluster same-slot keys; publish the best-effort hint only after that durable unit succeeds.

3. Version the signal so stale deliveries are harmless

Pub/Sub can be duplicated by overlapping subscriptions or application retries, and a client can process an older notification after newer state exists. Put an entity ID and monotonically meaningful application version in the hint. The receiver compares that hint with its local version, then re-reads authoritative state when necessary. The payload should not contain secrets merely because the channel is restricted.

redis-cli · stale-signal thought experiment with observable state
HSET atlasmart:ch16:l5:order:{1001} status shipped version 2PUBLISH atlasmart:ch16:l5:notify:1001 '{"order_id":"1001","version":1}'HGET atlasmart:ch16:l5:order:{1001} versionHGET atlasmart:ch16:l5:order:{1001} status# A version=1 notification is stale; the authoritative record already says version=2 shipped.

4. Reconnect policy is part of the cache/notification protocol

On Pub/Sub reconnect, assume the gap may contain missed signals. The receiver should invalidate any state whose freshness depends on uninterrupted signals, re-fetch a checkpoint/snapshot, and—where exact missed events matter—replay the Stream from a stored ID/group state. Do not silently keep a local cache “because Redis reconnected successfully.”

Reconnect step Purpose
Mark signal-derived cache uncertain Do not trust freshness across invalidation gap
Re-read current entity/version Recover latest truth
Replay Stream if workflow needs every event Recover durable missed history
Resume Pub/Sub only after baseline restored Return to low-latency hints safely

5. Least privilege spans commands, keys, and channels

Redis ACLs authorize key patterns and Pub/Sub channel patterns separately. Since Redis 7.0, new Redis Open Source ACL users are restrictive for Pub/Sub channels by default. The temporary user below can read/write only Chapter 16 order/event keys and publish/subscribe only Chapter 16 notification channels. It does not receive CONFIG, ACL, or all-channel access.

redis-cli · create and inspect a temporary Chapter 16 user
ACL SETUSER ch16-app reset on '>AtlasMart-Ch16-App-Lab-Only-2026' '~atlasmart:ch16:l5:*' 'resetchannels' '&atlasmart:ch16:l5:notify:*' '+@read' '+@write' '+@pubsub' '+ping' '+client|setname'ACL GETUSER ch16-appACL DRYRUN ch16-app PUBLISH atlasmart:ch16:l5:notify:1001 probeACL DRYRUN ch16-app PUBLISH other:tenant:notify probe# First DRYRUN should be allowed; second should be denied by channel pattern.# PSUBSCRIBE ACL matching has stricter literal-pattern considerations; test the exact pattern your client will use.
Security boundary

Channel ACLs prevent unauthorized Redis Pub/Sub access; they do not replace application authorization. A user allowed to receive a tenant channel can see every payload published there, so minimize payload sensitivity and tenant blast radius.

6. Cluster and Sentinel change reconnect/routing, not the source-of-truth rule

In Redis Cluster, place state and Stream keys that must participate in one transaction/function into the same slot using a deliberate hash tag such as {1001}. Ordinary Pub/Sub is globally propagated across the Cluster; sharded Pub/Sub can reduce Cluster-bus fan-out when channel slot design matches traffic. Under Sentinel or Cluster failover, clients can disconnect; therefore the post-reconnect reconciliation step remains mandatory.

Do not infer topology results from standalone

The Chapter 01 lab proves application semantics only. It does not measure Sentinel failover gaps, Cluster message propagation, shard hot spots, replica behavior, or managed-service routing. Build those later on the topology chapters and re-run the notification failure drill.

7. Durable processing stays idempotent

If a background service consumes atlasmart:ch16:l5:events:{1001} through a Stream consumer group, it can be redelivered after a crash before acknowledgment. Use event_id or a business idempotency key so replay is safe. The Pub/Sub hint should never be used as the idempotency ledger.

redis-cli · inspect replayable event state
XRANGE atlasmart:ch16:l5:events:{1001} - +XINFO STREAM atlasmart:ch16:l5:events:{1001}# If a consumer group is added, XPENDING/XACK become processing-state evidence.# The Pub/Sub channel itself exposes no historical event IDs.

8. Failure-injection matrix before production

Injected failure Expected correct behavior Evidence
Subscriber offline during publish state/event correct; UI catches up on reconnect HGETALL + XRANGE + reconnect log
Duplicate notification no duplicate business side effect version/idempotency counters
Stale notification after newer write receiver displays/re-reads newer version version comparison + state read
Redis failover/reconnect signal gap reconciled client reconnect logs + snapshot/replay
Slow subscriber bounded buffers; no memory runaway CLIENT LIST + app queue depth
Unauthorized channel ACL denial ACL DRYRUN/real NOPERM test

9. Observability that matches the contract

Measure publication rate and recipient counts only as live transport signals. Pair them with subscriber reconnect count, callback latency, output-buffer indicators, local queue depth, state-reconciliation failures, Stream lag/pending counts where durable processing exists, and entity-version mismatch counts. Alert on correctness-recovery failures rather than on “PUBLISH returned zero” alone; zero can simply mean no dashboard is online.

10. Cleanup and rollback

Delete only the Chapter 16 state/event fixtures and the temporary ACL user. No global notification configuration is changed in this lesson. If the user is still connected, deleting/disabling it may disconnect or deny subsequent commands; close test clients first.

redis-cli · bounded cleanup and ACL rollback
UNLINK atlasmart:ch16:l5:order:{1001} atlasmart:ch16:l5:events:{1001}ACL DELUSER ch16-appACL USERS# Verify ch16-app is gone and no unrelated user/config was changed.

11. Production judgment and chapter synthesis

The safe architecture is intentionally asymmetric: durable state and replayable events carry correctness; Pub/Sub carries immediacy. Use explicit versions, idempotency, bounded retries/backoff, reconnection reconciliation, least-privilege channel ACLs, and topology-aware routing. If event retention/scale/independent failure domains exceed Redis operational goals, move the durable-event role to a purpose-built broker while keeping the same state-versus-signal separation.

Check your understanding

  1. Which layer is authoritative after a Pub/Sub gap?
  2. Why include a version in the notification?
  3. Should PUBLISH recipient count determine whether the order update succeeded?
  4. What should a client do after reconnect?
  5. What three ACL dimensions matter here?
Review the answers

The durable business state, and the replayable event log where event history is required.

To detect duplicate/stale hints and decide whether to re-read newer authoritative state.

No. The business update must succeed independently of whether any subscriber is online.

Treat signal-derived freshness as uncertain, re-read current state, and replay durable events if required.

Allowed commands, key patterns, and Pub/Sub channel patterns; application authorization is still separate.

12. Chapter summary and bridge

Chapter 16 separated live signals from durable processing. Pub/Sub is excellent for ephemeral fan-out when missed messages are recoverable; keyspace notifications are similarly lossy hints; Streams add replay and consumer state; external brokers may add broader event-platform capabilities. Chapter 17 moves underneath those mechanisms to persistence—RDB snapshots, AOF, fsync, rewrite, and crash recovery—where “data exists in Redis” must be translated into a tested recovery envelope.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.