Chapter 04 · Hashes and Object-Like Records
HSCAN, Field Cardinality, Large Hashes, and Incremental Processing
Iterate large hashes incrementally with HSCAN, bound per-call work, tolerate documented cursor semantics, and measure cardinality/memory instead of treating HGETALL as harmless.
Learning outcomes
AtlasMart's merchandising pipeline now uses a hash with tens of
thousands of synthetic attribute fields during a migration
rehearsal. An operator wants to inventory matching fields
without creating one giant reply or blocking the server with a
whole-record read. HSCAN is the incremental
primitive, but it is not pagination over an immutable snapshot.
The cursor is opaque iteration state, COUNT is a work-size hint,
and collection changes during iteration have documented but
limited guarantees.
Use HSCAN as cursor-based incremental iteration with MATCH, COUNT, and current NOVALUES support where appropriate.
Explain the full-iteration guarantees and duplicate/change semantics without describing HSCAN as snapshot pagination or exactly-once enumeration.
Measure HLEN, HGETALL response cost, HSCAN batch behavior, MEMORY USAGE, and current command metadata on a bounded synthetic large hash.
Design duplicate-safe/idempotent processing for administrative scans and know when a secondary index or different data model is required.
Reason about one very large hash as a hot-key/cardinality/backup/replication/Cluster failure-domain decision, not merely a memory trick.
All Chapter 04 mandatory labs reuse the disposable Chapter 01
environment: Redis Open Source 8.10.1 from
Docker Official Image redis:8.10.1, container
atlasmart-redis-ch01, standalone topology, host
publication 127.0.0.1:6379, TLS disabled only
because traffic stays on loopback, default ACL user disabled,
named ACL users atlasmart-app and
academy-admin, logical database 0, AOF with
appendfsync everysec plus RDB snapshots,
persistent /data volume, and no explicit Redis
maxmemory limit or eviction policy. The
application ACL is restricted to ~atlasmart:* and
normal read/write/connection categories. The primary interface
is the redis-cli shipped in the same 8.10.1
image, so server and CLI versions stay aligned. The lab
creates 5,000 small synthetic fields under one disposable
key—large enough to show incremental behavior but
intentionally far below stress-test scale. Increase
cardinality only on an isolated environment with memory/time
budgets.
1. HSCAN is incremental iteration, not page 1/page 2
HSCAN key cursor returns a new cursor plus a batch
of field/value pairs. Start with cursor 0 and
continue until Redis returns cursor 0 again. The
cursor is opaque: it is not an offset, row number, stable page
token, or count of fields already processed.
COUNT is a hint about work per call, not an exact
batch-size promise.
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch04:scan:smalldocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HSET atlasmart:ch04:scan:small color blue size M material cotton price_cents 1999 active 1docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HSCAN atlasmart:ch04:scan:small 0 COUNT 2# Redis may return all fields in one call for compact encodings.# Always use the returned cursor; never invent cursor arithmetic.
For compact internal representations, Redis may return the whole collection in one call despite a small COUNT hint. That is another reason to avoid interpreting COUNT as SQL-style LIMIT.
2. The guarantees are useful but narrower than a snapshot
HSCAN is not a snapshot. For a full SCAN-family iteration, Redis guarantees that an element present for the entire duration is returned at least once, and an element absent for the entire duration is not returned. Elements added or removed while the iteration is in progress have undefined visibility, and duplicate returns are possible. Therefore a processor that performs irreversible side effects must tolerate duplicates or use a stronger source of truth.
| Assumption | Reality | Design consequence |
|---|---|---|
| Cursor = offset | Cursor is opaque iteration state | Persist the exact returned cursor only for operational continuation; do not calculate it. |
| COUNT = exact page size | COUNT is a hint | Do not allocate business workflows around exact batch cardinality. |
| One full scan = snapshot | Concurrent changes have limited/undefined visibility | Re-read authoritative state when correctness depends on it. |
| Every field returned once | Duplicates are allowed | Make processing idempotent or deduplicate by field/version. |
| HSCAN solves query filtering | MATCH filters names during iteration | Use a real index/query structure when selection is a core workload. |
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin COMMAND INFO HSCANdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HSCAN atlasmart:ch04:scan:small 0 MATCH p* COUNT 10docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HSCAN atlasmart:ch04:scan:small 0 NOVALUES# NOVALUES is current syntax on modern Redis; verify COMMAND INFO/client support before relying on wrappers.
3. HGETALL and HSCAN solve different operational jobs
HGETALL is O(N) and returns the entire hash in one response. HSCAN spreads a complete O(N) iteration across multiple calls, each designed for incremental work. That can reduce individual blocking/reply spikes, but it does not make the total work disappear. Network round trips, client processing, mutation churn, and total fields still matter.
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch04:scan:large# The loop itself runs inside the Linux container. Paste this as one host-shell line.docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 sh -lc 'for i in $(seq 1 5000); do redis-cli --user atlasmart-app HSET atlasmart:ch04:scan:large field:$i value:$i >/dev/null; done'docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HLEN atlasmart:ch04:scan:largedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin MEMORY USAGE atlasmart:ch04:scan:largedocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HSCAN atlasmart:ch04:scan:large 0 MATCH "field:*" COUNT 200# Record returned cursor and actual batch length. Do not assume it equals 200.
The fixture uses a shell inside the Linux Redis container, so the command is independent of whether the host shell is PowerShell, cmd, Bash, or zsh. For a real load generator, pipelining would avoid 5,000 host-to-container process/round-trip costs; Chapter 15 teaches that optimization explicitly.
4. Duplicate-safe processing: use idempotent work or a seen/version set
Administrative inventory often tolerates reprocessing: collecting field names into a set, recomputing a report row keyed by field name, or writing an idempotent destination record. If the action charges a card, sends an email, or deletes authoritative data, HSCAN is the wrong workflow source unless the side effect has its own idempotency/reconciliation design.
# Optional client example; mandatory lab remains redis-cli only.# pip install redis==8.0.1import redisr = redis.Redis(host="127.0.0.1", port=6379, db=0, username="atlasmart-app", password="AtlasMart-App-Lab-Only-2026", decode_responses=True)cursor = 0seen = set()while True: cursor, batch = r.hscan("atlasmart:ch04:scan:large", cursor=cursor, match="field:*", count=200) for field, value in batch.items(): if field in seen: continue # duplicates are allowed by SCAN semantics seen.add(field) # idempotent processing keyed by field goes here if cursor == 0: breakprint("unique fields observed:", len(seen))# Mutation during the scan can still make visibility undefined for changed elements.
The optional client is pinned so response behavior can be reproduced. Current redis-py 8.x uses RESP3 by default, but application response shapes are designed to remain compatible; still test the exact client/server pair before an upgrade.
5. Mutation during a scan: observe without inventing a deterministic outcome
A controlled mutation drill should demonstrate the absence of a snapshot guarantee, not promise that one specific field will or will not appear. Start a multi-call scan, then add/delete fields before it finishes. The expected lesson outcome is that Redis documents limited guarantees for continuously present/absent elements and undefined visibility for changing elements.
# 1) Begin and record the nonzero cursor returned:docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HSCAN atlasmart:ch04:scan:large 0 COUNT 50# 2) Before continuing, mutate the collection:docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HSET atlasmart:ch04:scan:large field:new during-scandocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HDEL atlasmart:ch04:scan:large field:10# 3) Continue with the exact cursor from step 1 until cursor returns 0.# Do NOT assert that field:new must appear or that field:10 must be absent.
Treating HSCAN as an exactly-once export cursor can silently duplicate or miss fields that change during the run. Repair by quiescing writes for a true administrative snapshot, exporting from an authoritative snapshot/backup, versioning records, or making destination processing idempotent and reconcilable.
6. Large hash design: cardinality is an operational boundary
A large hash concentrates all fields under one Redis key. In standalone Redis this can be convenient; in Cluster the key maps to one hash slot, so the whole hash lives on one primary and can become a hot slot. Persistence and replication serialize/propagate changes to that key, backup/restore must handle its size, and deletion can create reclamation work. Search indexing can query hash documents, but one giant “all products as fields” hash is not equivalent to one hash key per product document for indexing or lifecycle.
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app TYPE atlasmart:ch04:scan:largedocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app HLEN atlasmart:ch04:scan:largedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin OBJECT ENCODING atlasmart:ch04:scan:largedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin MEMORY USAGE atlasmart:ch04:scan:largedocker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin INFO memory# MEMORY USAGE for one key is not total process RSS or index/replication overhead.
7. Hands-on lab: incremental inventory with an acceptance contract
Use the existing 5,000-field fixture. Complete a full HSCAN iteration manually for a few calls or with the optional client, record cursor transitions and total unique fields observed when no mutation occurs, then repeat with one controlled mutation and document why exact visibility is not guaranteed. The acceptance criterion is correct reasoning, not a fabricated latency number.
Verification checklist:
- HLEN reports 5,000 before the mutation drill, and MEMORY USAGE is recorded with Redis 8.10.1/environment context.
- At least one HSCAN call records the returned cursor and actual item count; COUNT is described as a hint.
- A stable full iteration is processed duplicate-safely and can reconcile unique fields against HLEN when no writes occur during the run.
- The mutation drill does not claim deterministic inclusion/exclusion for fields changed during the iteration.
- HGETALL is not used as the production-safe replacement for incremental processing of the large fixture.
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app UNLINK atlasmart:ch04:scan:large atlasmart:ch04:scan:smalldocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app EXISTS atlasmart:ch04:scan:large atlasmart:ch04:scan:small# Expected EXISTS count: 0; UNLINK removes keys from keyspace before background reclamation.
8. Production judgment
HSCAN is appropriate for incremental administrative work, maintenance, migration assistance, diagnostics, and background jobs that can tolerate its semantics. It is not a stable business pagination API, not a consistent snapshot, and not a substitute for an index. Measure total iteration time, per-call latency, reply bytes, mutation rate, client retry behavior, and memory/cardinality. Under Cluster, scan each relevant key/node according to topology rather than assuming a standalone keyspace model transfers unchanged.
For destructive maintenance, combine SCAN-family enumeration with re-validation before deletion. For side effects, require idempotency. For very large hashes, set design budgets around field count/value size and test HGETALL/HSCAN/backup/restore/failover behavior before production. Avoid universal COUNT values: workload, network, server load, and latency SLOs determine the appropriate tradeoff.
9. Summary and next step
HSCAN breaks a complete O(N) hash iteration into cursor-driven calls. The cursor is opaque, COUNT is a hint, duplicates can occur, and elements changed during the iteration do not have snapshot semantics. HLEN and MEMORY USAGE make cardinality/cost observable, while idempotent processing makes duplicate tolerance explicit. Next, Redis 7.4+ field expiration and Redis 8.0+ get/set-with-expiry commands will add a second lifecycle dimension inside the hash.
Check your understanding
- Why is an HSCAN cursor not a page number?
- Does COUNT 200 guarantee 200 fields?
- Can a full HSCAN return duplicates?
- What is guaranteed for a field present for the entire full iteration?
- When is HSCAN the wrong source for a side effect?
Review the answers
1. It is opaque server iteration state; clients must feed back the exact returned cursor until it becomes zero.
2. No. COUNT is a hint, and compact encodings or server iteration details can return different batch sizes.
3. Yes. Consumers must tolerate duplicates when processing requires correctness.
4. It will be returned at least once during that full iteration.
5. When the side effect cannot tolerate duplicates/undefined mutation visibility and has no idempotency or reconciliation mechanism.
Authoritative references
- SCAN family — full-iteration guarantees, duplicates, COUNT and HSCAN behavior
- HSCAN — hash cursor syntax, MATCH/COUNT/NOVALUES, and complexity
- HGETALL — whole-hash O(N) comparison
- HLEN — field cardinality evidence
- MEMORY USAGE — key-attributed memory diagnostics
- redis-py — optional pinned client example and current RESP/client compatibility context