Chapter 07 · Redis Streams and Consumer Groups

XPENDING, XCLAIM/XAUTOCLAIM, Idle Consumers, Retries, Poison Events, and Recovery

Recover abandoned Stream work with XPENDING, XCLAIM/XAUTOCLAIM, retry evidence, poison-event quarantine, and explicit slow-versus-dead reasoning.

Advanced165–200 minutesPending-entry recovery labRedis Open Source 8.10.1Free/local-firstLast reviewed: September 6, 2026

Learning outcomes

AtlasMart now has a realistic failure: a worker received an event and disappeared without acknowledging it. Redis did not lose the pending reference, but another worker must decide when it is safe to take ownership. Recovery uses PEL inspection, idle time, claim operations, delivery counts, and an application retry/poison-event policy.

01

Inspect PEL summary and per-entry owner/idle/delivery-count evidence with XPENDING.

02

Use XCLAIM for explicit IDs and XAUTOCLAIM for cursor-like stale-entry recovery.

03

Explain how claiming changes ownership, idle time, and retry counters.

04

Design poison-event handling that stops infinite retries without pretending dead-lettering is magically atomic.

05

Distinguish idle consumers from failed consumers and monitor recovery without claim thrashing.

Exact lab baseline

All Chapter 07 mandatory labs reuse the disposable Chapter 01 environment: Redis Open Source 8.10.1 from Docker Official Image redis:8.10.1, container atlasmart-redis-ch01, standalone topology, host publication 127.0.0.1:6379, TLS disabled only because traffic stays on loopback, default ACL user disabled, named ACL users atlasmart-app and academy-admin, logical database 0, AOF with appendfsync everysec plus RDB snapshots, persistent /data volume, and no explicit Redis maxmemory limit or eviction policy. The primary interface is the redis-cli shipped in the same pinned image. Mandatory examples use only bounded synthetic keys under atlasmart:ch07:*. No managed service, paid broker, or external API is required.

1. Build deterministic abandoned pending state

Explicit Stream IDs make claim examples reproducible. Worker A receives two entries; one is acknowledged, one is intentionally left pending.

redis-cli · create an abandoned PEL entry
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch07:recovery atlasmart:ch07:dead-letterdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XADD atlasmart:ch07:recovery 1000-0 kind valid order 6001docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XADD atlasmart:ch07:recovery 2000-0 kind poison order 6002docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XGROUP CREATE atlasmart:ch07:recovery workers 0-0docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XREADGROUP GROUP workers worker-a COUNT 2 STREAMS atlasmart:ch07:recovery '>'docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XACK atlasmart:ch07:recovery workers 1000-0docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XPENDING atlasmart:ch07:recovery workers

Expected summary: one pending entry owned by worker-a. No process is actually killed in this lab; “worker-a crashed” is simulated simply by leaving 2000-0 unacknowledged.

2. XPENDING summary versus detail

The summary shows total pending count, min/max pending IDs, and counts by consumer. The extended form adds each entry's owner, idle milliseconds, and delivery count.

redis-cli · inspect recovery evidence
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XPENDING atlasmart:ch07:recovery workersdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XPENDING atlasmart:ch07:recovery workers - + 10docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XINFO CONSUMERS atlasmart:ch07:recovery workers

Idle time means “time since the last delivery/claim,” not proof that a process is dead. A slow but healthy consumer can be idle long enough to satisfy an aggressive claim threshold.

3. XCLAIM transfers ownership only after min-idle-time

XCLAIM targets known IDs. Redis transfers a pending entry only if its idle time is at least the supplied minimum. A successful claim resets the idle timer and normally increments the delivery count.

redis-cli · explicit claim boundary
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XCLAIM atlasmart:ch07:recovery workers recovery-worker 999999999 2000-0docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XCLAIM atlasmart:ch07:recovery workers recovery-worker 0 2000-0docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XPENDING atlasmart:ch07:recovery workers - + 10

The huge idle threshold should claim nothing in a fresh lab. The zero threshold then transfers 2000-0 immediately. Production thresholds must reflect observed processing-time distributions and failure detection—not a copied constant.

4. Claim races reduce trivial double ownership, not all duplicate processing

Claiming resets idle time, so two consumers attempting the same stale entry at nearly the same moment will not both successfully claim it under the same positive min-idle rule. But duplicate business processing is still possible: the previous owner may merely be slow and can finish after another consumer claims the event. Fencing/idempotency at the business layer is still required.

5. XAUTOCLAIM combines PEL scanning with stale claiming

XAUTOCLAIM starts at a PEL ID, scans for entries old enough to claim, and returns a next-start ID for another call. It is conceptually like XPENDING + XCLAIM with cursor-like progress. COUNT limits attempted claims, but the number actually returned can be lower because Redis scans candidates and filters by idle time.

redis-cli · cursor-like autoclaim
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XAUTOCLAIM atlasmart:ch07:recovery workers rescue-b 0 0-0 COUNT 10docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XPENDING atlasmart:ch07:recovery workers - + 10

With min idle zero, 2000-0 is eligible immediately. The response includes the next start ID, claimed entries, and—on modern Redis—IDs of deleted entries whose PEL references were cleaned during scanning where applicable.

6. XAUTOCLAIM completion does not mean “never call again”

A returned next-start ID of 0-0 means the current scan reached the end of the PEL. Older pending entries that were not idle enough can become eligible later, so a recovery loop may start again at 0-0 after time passes.

7. Delivery count is a poison-event signal, not a universal policy

XPENDING's delivery count increases when an entry is redelivered through pending history or claimed (except JUSTID claim modes that suppress the increment). A high count is useful evidence that processing repeatedly fails, but the threshold for “poison” depends on operation cost, transient failure patterns, and business risk.

redis-cli · increase attempt evidence safely
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XCLAIM atlasmart:ch07:recovery workers rescue-c 0 2000-0docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XCLAIM atlasmart:ch07:recovery workers rescue-d 0 2000-0docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XPENDING atlasmart:ch07:recovery workers - + 10

Do not run tight claim loops in production: they manufacture retries, reset idle time, and can amplify load.

8. Dead-lettering is an application transition

For a permanently invalid event, AtlasMart can copy diagnostic context to a separate Stream and then acknowledge the original. These are two commands, so a crash between them can duplicate the dead-letter entry or leave the original pending. A later Redis Function/transaction or external durable workflow may be needed if the transition must be atomic.

redis-cli · bounded poison-event quarantine
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XADD atlasmart:ch07:dead-letter '*' original-id 2000-0 reason invalid-schema source atlasmart:ch07:recoverydocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XACK atlasmart:ch07:recovery workers 2000-0docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XPENDING atlasmart:ch07:recovery workersdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XRANGE atlasmart:ch07:dead-letter - +

This ends the retry loop in the lab. In production, include enough immutable context for investigation without leaking sensitive payloads into an over-retained dead-letter stream.

9. Deleted or trimmed bodies can outlive pending references

With default KEEPREF trimming in Redis 8.2+, a consumer group can retain a PEL reference after the Stream body entry has been removed. Recovery may then encounter a pending ID whose payload is no longer available. Retention and PEL-reference policy must therefore be tested together; DELREF and ACKED change this tradeoff.

10. Idle consumer is not the same as dead consumer

Use XINFO CONSUMERS together with application heartbeats, connection/process telemetry, and operation-specific deadlines. A consumer working on a 90-second video transformation may legitimately exceed a 30-second idle threshold. Claiming it at 30 seconds can create concurrent duplicate work.

11. Recovery checklist

Signal Question Action
XPENDING count is pending growing faster than acknowledgments? inspect consumers, downstream health, and lag
Idle time has the attempt exceeded a justified processing deadline? consider claim; do not equate idle with dead
Delivery count is the same ID failing repeatedly? apply retry/backoff/quarantine policy
XINFO consumer idle/pending is one consumer accumulating ownership? recover only after failure evidence
Stream retention is the payload still retained? align trim policy with recovery horizon

12. Reproducible recovery verification

At the end of the lab, pending should be zero and the dead-letter Stream should contain the quarantined event.

redis-cli · final recovery state
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XPENDING atlasmart:ch07:recovery workersdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XINFO GROUPS atlasmart:ch07:recoverydocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XINFO CONSUMERS atlasmart:ch07:recovery workersdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XLEN atlasmart:ch07:dead-letterdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app XRANGE atlasmart:ch07:dead-letter - + COUNT 10

Cleanup later can delete only atlasmart:ch07:* fixtures. Do not use FLUSHDB/FLUSHALL.

13. Production judgment

Recovery is a control loop. Choose min-idle from real processing-time and failure-detection data; bound XAUTOCLAIM COUNT; add retry backoff; detect poison entries; keep side effects idempotent; and monitor claim rate, delivery counts, pending age, group lag, consumer idle, retention, and downstream saturation. Test a slow-but-alive worker as well as a dead worker so your recovery policy does not create duplicate storms.

14. Summary and next step

PEL inspection plus claim/autoclaim lets workers recover abandoned attempts, but safe retries still depend on idempotent side effects and bounded backpressure. The final lesson combines these pieces into an end-to-end pipeline.

Check your understanding

  1. What four per-entry facts does extended XPENDING show?
  2. What does a successful XCLAIM normally do to idle time and delivery count?
  3. What does XAUTOCLAIM return for continuation?
  4. Does a consumer idle longer than the claim threshold prove it is dead?
  5. Why can dead-lettering with XADD then XACK still duplicate?
Review the answers

ID, owner, idle milliseconds, and delivery count.

It resets idle time and normally increments delivery count.

A next-start Stream ID for the following scan; 0-0 indicates the current scan reached the end.

No. It may simply be slow or working on a long task.

A crash can occur between the two commands, so the transition is not automatically atomic.

Authoritative references

  • Redis Streams — stream entries, consumer groups, pending entries, acknowledgments, and recovery
  • XADD — entry IDs, trimming, Redis 8.2 reference policies, and Redis 8.6 idempotent production options
  • XRANGE — ordered range inspection and exclusive continuation IDs
  • XREAD — blocking/nonblocking reads, explicit IDs, and the special dollar ID
  • XGROUP CREATE — consumer-group starting position and MKSTREAM
  • XREADGROUP — group delivery, pending history, and new-message marker
  • XPENDING — PEL summary/details, idle time, owner, and delivery count
  • XACK — acknowledgment semantics
  • XCLAIM — explicit ownership transfer and retry-count behavior
  • XAUTOCLAIM — cursor-like stale-pending recovery
  • XINFO GROUPS — pending count, last-delivered ID, entries-read, and lag
  • XTRIM — exact/approximate retention and PEL reference policies
  • Redis 8.10 release notes — pinned server release family

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.