Chapter 02 · Keys, Namespaces, Expiration, TTL, Scanning, and Data-Type Introspection

KEYS vs SCAN: Incremental Iteration, Cursor Semantics, and Production Safety

Choose KEYS or SCAN with a precise model of blocking work, cursor iteration, duplicates, concurrent mutation, COUNT/MATCH/TYPE hints, and production-safe processing.

Intermediate105–135 minutesCursor + safety labRedis Open Source 8.10.1Free/local-firstLast reviewed: September 6, 2026

Learning outcomes

AtlasMart support needs to find every temporary checkout key after a deployment. One engineer proposes KEYS atlasmart:checkout:* on production because it is one command; another replaces it with one SCAN 0 call and assumes the returned page is complete; a third sends a refund side effect once for every key returned and assumes SCAN is exactly-once. These are three different mistakes. Safe key discovery requires a model of server work, cursor state, duplicates, concurrent mutation, and topology.

01

Explain why KEYS performs a full keyspace pattern scan and why its O(N) server work can create latency risk.

02

Run a complete SCAN iteration by feeding each returned cursor into the next call until Redis returns cursor 0.

03

Use MATCH, COUNT, and TYPE as filtering/work hints without turning them into guarantees that Redis does not provide.

04

Design duplicate-safe/idempotent processing and state clearly what SCAN guarantees when the keyspace changes during iteration.

05

Adapt discovery reasoning for Redis Cluster instead of assuming one standalone cursor is a cluster-wide snapshot.

Exact lab baseline

All Chapter 02 labs reuse the disposable Chapter 01 environment: Redis Open Source 8.10.1 from Docker Official Image redis:8.10.1, container atlasmart-redis-ch01, standalone topology, host publication 127.0.0.1:6379, TLS disabled because traffic stays on loopback, default user disabled, named ACL users atlasmart-app and academy-admin, logical database 0, AOF with appendfsync everysec plus RDB snapshots, and a persistent /data Docker volume. No explicit maxmemory limit or eviction policy is introduced in this chapter. Search, JSON, vector, time-series, and probabilistic features are not required for these exercises. The only KEYS demonstration in this lesson is against a deliberately small, bounded disposable prefix and is explicitly not a benchmark or production recommendation. The production pattern is SCAN plus duplicate-safe processing or, better, application-maintained indexes/registries when exact inventory is a business requirement.

1. KEYS answers a global question in one server command

KEYS pattern searches the selected keyspace for every key matching a glob-style pattern and returns all matches in one reply. Its time complexity is O(N) in the number of keys in that database. Because Redis must do that work before the command finishes, a large keyspace can extend the time before other work is serviced and can inflate tail latency. The problem is not that KEYS is “forbidden syntax”; the problem is unbounded synchronous work on a shared latency-sensitive server.

On a tiny disposable database, KEYS is useful for debugging and tests. On production, the command documentation itself warns to use it with extreme care. If the business requires an exact list, build an explicit data structure or external inventory as part of the write path rather than repeatedly deriving global truth with a keyspace-wide command.

redis-cli · bounded KEYS example — lab only
# Create only 20 synthetic keys.docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 sh -lc 'for i in $(seq 1 20); do redis-cli --user atlasmart-app SET "atlasmart:ch02:keys-demo:$i" "$i" >/dev/null; done'# Safe only because this is a bounded disposable lab fixture.docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 \  redis-cli --user academy-admin KEYS 'atlasmart:ch02:keys-demo:*'# Do NOT extrapolate the observed runtime to a production keyspace.

2. SCAN decomposes discovery into cursor-based incremental work

SCAN returns two things: a new opaque cursor and a batch of keys. Start with cursor 0; pass the returned cursor to the next call; finish only when a call returns cursor 0. Each call performs a bounded portion of the iteration, so the client can interleave discovery with other work. A single call is not “the scan.” The complete cursor cycle is the scan.

The cursor is continuation state, not an offset or page number. Do not increment it, persist arithmetic assumptions about it, or infer percentage complete from its numeric value. A call may return zero keys while the cursor remains nonzero, so “empty batch” also does not mean “done.”

redis-cli · manual cursor protocol
# First call.docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 \  redis-cli --user academy-admin SCAN 0 MATCH 'atlasmart:ch02:scan:*' COUNT 25# Example shape only:# 1) "96"          <- next cursor; your value will vary# 2) 1) "atlasmart:ch02:scan:item:17"#    2) "atlasmart:ch02:scan:item:88"# Continue with the exact cursor Redis returned, for example:docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 \  redis-cli --user academy-admin SCAN 96 MATCH 'atlasmart:ch02:scan:*' COUNT 25# Repeat until the returned cursor is "0".

The specific cursor and batch membership depend on the keyspace and internal hash-table state; hard-coding 96 in automation would be wrong. Code must parse the response and feed back the actual returned cursor.

3. MATCH, COUNT, and TYPE are controls, not snapshot guarantees

MATCH pattern filters returned names using Redis glob matching. COUNT n is a hint about how much work Redis should attempt per call; it is not a page-size guarantee. TYPE type asks SCAN to return only keys of a given Redis data type. Because filtering happens within incremental iteration, a call may return fewer matches than COUNT or even none while the cursor is nonzero.

Option What it controls What it does not promise
MATCH atlasmart:ch02:* Which discovered names are returned That every call returns a match or that the pattern is indexed.
COUNT 100 Approximate amount of iteration work per call Exactly 100 keys, fixed latency, or fixed network payload.
TYPE string Filters returned keys by Redis type A schema lock; a key can change type between discovery and later use.
cursor 0 Start/end marker for a full iteration Snapshot timestamp, offset, or percentage complete.
redis-cli · filter by name and type
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 \  redis-cli --user academy-admin SCAN 0 MATCH 'atlasmart:ch02:scan:*' COUNT 50 TYPE string# Continue until cursor 0 even if an intermediate batch contains no keys.

4. Full-iteration guarantees still allow duplicates and undefined concurrent changes

For a key that remains present for the entire duration of a full iteration, SCAN guarantees that the full iteration returns it. A key that is absent for the entire duration is not returned. But a key may be returned more than once, and the behavior for keys added, deleted, or otherwise not continuously present during the iteration is undefined. This is why SCAN is not a consistent point-in-time snapshot and not an exactly-once enumeration.

“Duplicate-safe” means repeating the processing does not cause an incorrect additional effect. Computing memory statistics twice and deduplicating by key name is easy. Charging a customer, sending email, or issuing a refund directly for each SCAN result is not. For non-idempotent business effects, scan into a staging set/queue with explicit idempotency keys or use authoritative application records rather than treating key discovery as a transaction log.

Correctness boundary

SCAN is excellent for incremental maintenance, diagnostics, migration helpers, and best-effort discovery when the consumer follows its guarantees. It is the wrong primitive for an exact historical snapshot unless you create a separate snapshot/registry with the required semantics.

5. Deliberately wrong processor: one page, exactly-once side effects

The following design fails twice: it makes only one SCAN call and treats that page as complete, then performs a non-idempotent side effect per returned key. A larger keyspace can require many cursors; a full iteration can repeat a key. Concurrent writes add another ambiguity.

pseudo-code · wrong versus repaired scan consumer
# WRONGcursor, keys = SCAN(0, MATCH="atlasmart:order:*")for key in keys:    issue_refund(key)       # not idempotent# ignores cursor and possible duplicates# REPAIRED SHAPEcursor = 0seen = set()                # fine for a bounded maintenance batch; choose scalable state in productionwhile True:    cursor, keys = SCAN(cursor, MATCH="atlasmart:order:*", COUNT=200)    for key in keys:        if key not in seen:            seen.add(key)            stage_for_review_with_idempotency_key(key)    if cursor == 0:        break# authoritative business state is re-read before any external effect

The repaired flow separates discovery from business action. For very large inventories, an in-memory seen set may itself be unbounded; use a bounded batch design, a durable registry, or idempotent target operation instead. Correctness moves to an explicit layer rather than being accidentally inferred from SCAN.

6. Cluster changes the scope of “the keyspace”

In standalone Redis, SCAN iterates the selected logical database. In Redis Cluster, keys are partitioned across nodes by hash slot, and client/tool behavior can involve per-node scanning. Redis 8 can optimize some pattern operations when a pattern implies a single hash slot through a hash tag, but that does not turn a multi-node Cluster into one atomic snapshot. A cluster-aware inventory must define whether it scans primaries only, how it handles resharding/failover, and how duplicates or moving slots are reconciled.

This is another reason to maintain application-level indexes for exact business inventory. Chapter 21 will make the topology observable; Chapter 02 only establishes the semantic boundary so a standalone scan script is not blindly carried into Cluster.

7. Hands-on lab: bounded scan with mutation

Create 400 strings under one prefix. That is large enough to require multiple cursor calls on most runs but intentionally tiny compared with production. We then mutate the keyspace mid-iteration to make the undefined concurrent-change boundary concrete. Do not interpret this as a performance benchmark.

shell · create a bounded scan fixture
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 sh -lc 'for i in $(seq 1 400); do   redis-cli --user atlasmart-app SET "atlasmart:ch02:scan:item:$i" "$i" >/dev/null; done'docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 \  redis-cli --user academy-admin DBSIZE
redis-cli · start iteration, then mutate between pages
# 1) Start. Record the returned cursor C1 and keys.docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 \  redis-cli --user academy-admin SCAN 0 MATCH 'atlasmart:ch02:scan:item:*' COUNT 25# 2) While the iteration is in progress, add and remove disposable keys.docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 \  redis-cli --user atlasmart-app SET atlasmart:ch02:scan:item:401 401docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 \  redis-cli --user atlasmart-app UNLINK atlasmart:ch02:scan:item:2# 3) Continue with C1, then every returned cursor, until 0.# The presence/absence of keys modified during the iteration is intentionally not asserted.

For a reproducible automated exercise, write results to a local file, sort/unique them, and compare only the keys guaranteed to have existed for the whole scan. Do not expect key 401 or removed key 2 to prove a universal rule; their result is allowed to vary.

redis-cli · observe command statistics without inventing a benchmark
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 sh -lc 'redis-cli --user academy-admin INFO commandstats | grep -E "cmdstat_(scan|keys)"'# Fields such as calls/usec/usec_per_call are cumulative for this server.# They can show that commands executed, but this tiny lab is not a production latency benchmark.

Verification checklist:

  • You completed SCAN only when the returned cursor reached 0, not when a page happened to be empty.
  • You can explain why COUNT is a work hint rather than an exact page size.
  • You can state the stable-element guarantees and the duplicate/modified-element caveats from the SCAN contract.
  • Your processing plan is duplicate-safe or explicitly deduplicates before any non-idempotent action.
  • You can explain why a standalone SCAN loop is not automatically a cluster-wide consistent inventory.
shell · cleanup 401 known fixture names without KEYS
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 sh -lc 'for i in $(seq 1 401); do redis-cli --user atlasmart-app UNLINK "atlasmart:ch02:scan:item:$i" >/dev/null; done; for i in $(seq 1 20); do redis-cli --user atlasmart-app UNLINK "atlasmart:ch02:keys-demo:$i" >/dev/null; done' 

8. Production judgment

Use KEYS only when the keyspace is known and bounded enough that one O(N) server operation and one result payload are acceptable—typically diagnostics against disposable/dev data. Use SCAN for incremental production maintenance when approximate per-call work, duplicate handling, and non-snapshot semantics fit. For exact business inventory, maintain an explicit index or authoritative system of record.

Measure p50/p95/p99 application latency while maintenance runs, not merely SCAN command averages. COUNT, payload size, network round-trip time, concurrent clients, persistence activity, replication, Cluster topology, and key distribution all affect the outcome. A timeout/retry can cause a client to repeat a cursor call, reinforcing the need for idempotency. Failure tests should include process restart, client disconnect/retry, concurrent inserts/deletes, and Cluster resharding once that topology exists. Any migration tool must have a checkpoint/restart strategy and a rollback path.

9. Summary and next step

KEYS is a one-command O(N) scan whose unbounded synchronous work can hurt shared latency. SCAN spreads iteration across cursor calls, but it is not a snapshot or exactly-once feed. MATCH/COUNT/TYPE control filtering/work without changing those guarantees. Safe consumers continue until cursor 0 and tolerate duplicates and concurrent-change uncertainty. Next, you will inspect what each discovered key actually contains, how much memory Redis attributes to it, how its serialized form moves, and how deletion strategy affects latency.

Check your understanding

  1. Why can an empty SCAN batch still require another call?
  2. What does COUNT 100 guarantee?
  3. Can a full SCAN iteration return a key twice?
  4. What is guaranteed for a key present for the entire full iteration?
  5. When should AtlasMart build an explicit inventory instead of scanning?
Review the answers

1. SCAN completion is indicated only by a returned cursor of 0. Filtering and internal iteration can produce zero returned elements while the cursor remains nonzero.

2. It is a hint about the amount of work per call, not a guarantee of 100 returned keys or fixed latency.

3. Yes. Consumers must tolerate duplicates or make processing idempotent.

4. A full iteration returns it at least once. Keys that change presence during the iteration have undefined inclusion behavior.

5. When exact membership, snapshot semantics, durable progress, non-idempotent actions, or predictable query latency are business requirements.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.