Chapter 09 · Redis Search: Indexes, Text/Numeric/Tag/Geo Fields, Querying, and Aggregation
FT.AGGREGATE Pipelines, Grouping, Reducers, Filters, Apply, and Result Shaping
Transform Search result sets with FT.AGGREGATE grouping, reducers, APPLY/FILTER, bounded loading, sorting, cursors, and profiling.
Learning outcomes
Product discovery often needs more than a list of documents. AtlasMart wants counts per brand, average price, in-stock filtering, computed display values, and top groups. FT.AGGREGATE turns a Search result stream into a pipeline of grouping, reduction, transformation, filtering, sorting, limiting, and optional source loading. The order of those stages is part of correctness and performance.
Explain FT.AGGREGATE as an ordered pipeline over Search matches rather than SQL with different spelling.
Use GROUPBY and REDUCE for counts and numeric summaries over indexed/sortable attributes.
Use APPLY and FILTER in the correct pipeline order and understand when values become available.
Shape results with SORTBY/LIMIT and recognize when LOAD forces source-document access.
Profile aggregation and choose cursors/bounded results for large pipelines instead of returning unbounded arrays.
All Chapter 09 mandatory labs reuse the disposable Chapter 01
environment: Redis Open Source 8.10.1 from Docker
Official Image redis:8.10.1, container
atlasmart-redis-ch01, standalone topology, host
publication 127.0.0.1:6379, TLS disabled only
because traffic stays on loopback, default ACL user disabled,
named users atlasmart-app and
academy-admin, logical database 0, AOF with
appendfsync everysec plus RDB snapshots,
persistent /data, and no explicit
maxmemory limit or eviction policy. Redis 8
integrates Search and JSON into Redis Open Source; the
mandatory path needs no separate historical Redis Stack image
or paid service. Source writes use the restricted application
user where possible; Search/index administration and evidence
commands use the disposable academy-admin user
because the Chapter 01 application ACL was intentionally not
broadened for Search administration. Fixtures stay under
atlasmart:ch09:*. The mandatory lab stays on the
three-document JSON fixture; production cost conclusions must
be re-measured at representative scale.
1. Re-create the source and index
docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch09:product:1001 atlasmart:ch09:product:1002 atlasmart:ch09:product:1003 atlasmart:ch09:outside:product:9001docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app JSON.SET atlasmart:ch09:product:1001 '$' '{"schemaVersion":2,"sku":"SKU-1001","name":"Trail Running Backpack","description":"Lightweight running backpack for trail and city travel","category":["outdoor","travel"],"brand":"AtlasPeak","priceCents":12990,"stock":12,"location":"49.8671,40.4093"}'docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app JSON.SET atlasmart:ch09:product:1002 '$' '{"schemaVersion":2,"sku":"SKU-1002","name":"City Runner Pack","description":"Compact runner backpack for commuting and urban travel","category":["urban","travel"],"brand":"AtlasPeak","priceCents":9990,"stock":0,"location":"49.8920,40.3777"}'docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app JSON.SET atlasmart:ch09:product:1003 '$' '{"schemaVersion":2,"sku":"SKU-1003","name":"Trail Bottle","description":"Insulated bottle for hiking and trail running","category":["outdoor","hydration"],"brand":"NorthSpring","priceCents":3490,"stock":40,"location":"49.8500,40.4000"}'docker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app JSON.SET atlasmart:ch09:outside:product:9001 '$' '{"schemaVersion":2,"sku":"SKU-9001","name":"Outside Prefix Backpack","description":"This key proves prefix scope","category":["outdoor"],"brand":"ScopeTest","priceCents":1,"stock":1,"location":"49.8671,40.4093"}'
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.CREATE atlasmart-ch09-products-json-idx ON JSON PREFIX 1 atlasmart:ch09:product: SCHEMA '$.sku' AS sku TAG SORTABLE '$.name' AS name TEXT WEIGHT 2.0 '$.description' AS description TEXT '$.category[*]' AS category TAG SORTABLE '$.brand' AS brand TAG SORTABLE '$.priceCents' AS priceCents NUMERIC SORTABLE '$.stock' AS stock NUMERIC SORTABLE '$.location' AS location GEO '$.schemaVersion' AS schemaVersion NUMERIC
2. FT.SEARCH selects/projections; FT.AGGREGATE transforms sets
Use FT.SEARCH when the API needs matching documents
or selected attributes. Use FT.AGGREGATE when
results must be grouped, reduced, mapped, or filtered through a
pipeline. The base query uses the same Search query language,
then later stages operate on properties available in the
pipeline.
| Need | Prefer | Reason |
|---|---|---|
| Return matching product docs | FT.SEARCH | selection/projection only |
| Count products by brand | FT.AGGREGATE | GROUPBY + COUNT reducer |
| Compute priceCents / 100 for output | FT.AGGREGATE | APPLY creates derived property |
| Stream a very large aggregation | FT.AGGREGATE WITHCURSOR | bounded batches instead of one huge reply |
3. GROUPBY partitions the pipeline; REDUCE summarizes each group
Grouping by @brand partitions matching records by
the single-valued brand property, then reducers collapse each
group. A group should have at least one reducer. Using a
single-valued fixture avoids silently turning multi-value
category membership into a KPI definition. If a business metric
needs a canonical category, model that canonical value
explicitly.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.AGGREGATE atlasmart-ch09-products-json-idx '*' GROUPBY 1 @brand REDUCE COUNT 0 AS productCount REDUCE AVG 1 @priceCents AS avgPriceCents SORTBY 2 @productCount DESC DIALECT 2
The two AtlasPeak products should group together while NorthSpring forms its own group. This verifies grouping on a single-valued indexed attribute without relying on multi-value TAG aggregation behavior.
4. GROUPBY 0 reduces the whole result set
GROUPBY 0 applies reducers to the entire upstream
set. It is useful for one-row summaries such as overall count,
average price, min/max, or totals—provided the chosen reducers
match the numeric/business semantics.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.AGGREGATE atlasmart-ch09-products-json-idx '*' GROUPBY 0 REDUCE COUNT 0 AS products REDUCE AVG 1 @priceCents AS avgPriceCents REDUCE MIN 1 @priceCents AS minPriceCents REDUCE MAX 1 @priceCents AS maxPriceCents DIALECT 2
5. APPLY adds a derived property to later pipeline stages
APPLY evaluates an expression for each record and
stores the result under a pipeline property name. Because
priceCents is indexed as SORTABLE, it can be used
efficiently in the pipeline. The calculated major-unit price is
display/reporting logic, not a replacement for the integer-cent
source value.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.AGGREGATE atlasmart-ch09-products-json-idx '*' APPLY '@priceCents/100.0' AS priceMajor SORTBY 2 @priceMajor ASC LIMIT 0 3 DIALECT 2
Expressions, types, and null behavior are part of the current aggregation expression language. Test edge values before relying on a derived field in business decisions.
6. FILTER in an aggregation is a pipeline stage, not FT.CREATE FILTER
There are two different FILTER concepts.
FT.CREATE ... FILTER decides which source documents
become members of an index.
FT.AGGREGATE ... FILTER removes records at that
point in an aggregation pipeline. Mixing them produces
difficult-to-debug missing data.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.AGGREGATE atlasmart-ch09-products-json-idx '*' APPLY '@priceCents/100.0' AS priceMajor FILTER '@stock > 0 && @priceMajor < 100' SORTBY 2 @priceMajor ASC DIALECT 2
The filter can reference priceMajor only after
APPLY has created it. Pipeline stage order is therefore part of
query correctness.
7. LOAD can be the hidden source-read cost
Aggregation properties declared SORTABLE are available to the
pipeline efficiently. LOAD fetches attributes from
the source document; Redis documentation warns that loading
values over large result sets can be expensive because each
processed record may require source access. Do not write
LOAD * as a convenience default.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.AGGREGATE atlasmart-ch09-products-json-idx '@category:{outdoor}' LOAD 3 '$.description' AS description LIMIT 0 2 DIALECT 2
For a three-document lab the cost is negligible. In production, compare a schema that keeps frequently aggregated attributes SORTABLE against a LOAD-heavy design and include the added sortable-memory cost in the tradeoff.
8. Result shaping happens after the work you ask earlier stages to do
SORTBY orders the current pipeline;
LIMIT trims the returned range. For top-k sorts,
SORTBY ... MAX can avoid sorting more records than
necessary. A late LIMIT does not automatically erase the cost of
earlier GROUPBY/REDUCE/LOAD stages that needed to process many
matches.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.AGGREGATE atlasmart-ch09-products-json-idx '*' GROUPBY 1 @brand REDUCE COUNT 0 AS n SORTBY 2 @n DESC MAX 3 LIMIT 0 3 DIALECT 2
9. Wrong approach: LOAD every field before grouping millions of rows
A common migration from SQL is to load all source fields first and then group, even when grouping only needs indexed/sortable attributes. Redis documentation explicitly warns that LOAD can hurt aggregate performance substantially. Diagnose by profiling the query, identifying source loads, and redesigning the schema/pipeline so only needed fields are available at the stage that needs them.
Start from the output contract. Index only fields needed for selection/filter/sort/group. Mark truly hot aggregate properties SORTABLE after measuring its memory cost. LOAD only the small set of source values that must be returned or transformed and cannot reasonably live in the index.
10. WITHCURSOR is for incremental aggregate result consumption
For very large result sets, WITHCURSOR lets the
client fetch bounded batches rather than one unbounded reply.
Cursor lifecycle, idle timeout, retries, and client cleanup
become operational concerns; this is not the same as a
consistent database snapshot. The lab keeps results tiny, so
cursors are explained rather than required.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.AGGREGATE atlasmart-ch09-products-json-idx '*' LOAD 2 @sku @name WITHCURSOR COUNT 2 DIALECT 2
The first reply includes rows plus a cursor ID. Continue with
FT.CURSOR READ atlasmart-ch09-products-json-idx
<cursor-id> COUNT 2
using the actual returned ID, then delete/read to completion as
documented. Do not copy a placeholder cursor ID.
11. Profile the aggregate pipeline, not just the base match count
FT.PROFILE can profile SEARCH or AGGREGATE
execution. It exposes iterator/pipeline timing useful for
finding expensive filters, sorts, loads, and reducers. Profile
output is diagnostic and version-sensitive, so capture it as
evidence rather than parsing it as a stable application
response.
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.PROFILE atlasmart-ch09-products-json-idx AGGREGATE LIMITED QUERY '*' GROUPBY 1 @brand REDUCE COUNT 0 AS n SORTBY 2 @n DESC DIALECT 2
12. Aggregation does not create durable materialized business state
An FT.AGGREGATE reply is a query result over the current searchable data, not a persisted ledger or reconciliation artifact. If AtlasMart needs an auditable daily revenue total, retain authoritative order/payment records and compute/store governed aggregates using an appropriate pipeline. Search aggregates are excellent for interactive discovery/reporting, but they do not replace transactional/accounting invariants.
13. Reproducible lesson lab and verification
Run one GROUPBY/COUNT, one whole-set reducer, one APPLY/FILTER, one LOAD, and one profile. Record the exact output and identify where each property enters the pipeline.
| Check | Pass condition |
|---|---|
| GROUPBY | Brand groups and counts match the three known fixtures. |
| REDUCE | Whole-set count/min/max/avg are explainable from source cents. |
| APPLY/FILTER | Derived price exists before the FILTER references it. |
| LOAD | You can identify the source-read step and why it is not free. |
| PROFILE | Diagnostic output is captured with Redis version and query text. |
docker exec -e REDISCLI_AUTH=AtlasMart-Admin-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user academy-admin FT.DROPINDEX atlasmart-ch09-products-json-idxdocker exec -e REDISCLI_AUTH=AtlasMart-App-Lab-Only-2026 atlasmart-redis-ch01 redis-cli --user atlasmart-app DEL atlasmart:ch09:product:1001 atlasmart:ch09:product:1002 atlasmart:ch09:product:1003 atlasmart:ch09:outside:product:9001
14. Production judgment
Use FT.AGGREGATE for interactive grouping/transformation over Search result sets when the index already matches the access pattern. Large groups, sorts, LOAD operations, and multi-stage pipelines can raise CPU and tail latency; SORTABLE attributes trade memory for pipeline/sort efficiency; cursors require client lifecycle management; timeouts/retries must not duplicate application side effects because aggregation should remain read-only; and numeric summaries must not be mistaken for financial truth. Profile representative cardinality and concurrency, then cap result sizes and operational blast radius.
15. Summary and next step
Aggregation is an ordered dataflow: base Search query → available properties → APPLY/FILTER → GROUPBY/REDUCE → SORTBY/LIMIT or cursor. Lesson 5 closes the chapter by measuring index memory/write cost, profiling queries, planning schema/index lifecycle, and deciding when Search is worth using instead of direct key access.
Check your understanding
- What is the difference between FT.CREATE FILTER and FT.AGGREGATE FILTER?
- Why can LOAD be expensive?
- What does GROUPBY 0 mean?
- Does LIMIT make earlier expensive stages free?
- When is WITHCURSOR useful?
Review the answers
The first controls index membership at indexing time; the second filters records at a particular aggregation-pipeline stage.
It may fetch source-document attributes for every processed record instead of using values already stored in the index pipeline.
Apply reducers to the entire upstream result set as one group.
No. Earlier grouping/loading/sorting may already have processed many records.
When a large aggregation result should be consumed in bounded batches rather than one huge reply.
Authoritative references
These references were re-checked for the Redis 8.10.1 course snapshot. Command output and planner details can change between Redis releases, clients, RESP modes, and topologies; prefer the target-version reference when reproducing the lab.