Chapter 05 · Cypher Expressions, Aggregation, UNWIND, Collections, Maps, and Subqueries
List Expressions, List Comprehensions, reduce, any/all/none/single, and Collection Transformations
Shape bounded collections deliberately: understand list values, comprehensions, predicate truth tables and reduce before turning lists into rows or API payloads.
Learning outcomes
AtlasMart now wants compact API fields such as “products over a price threshold,” “all orders paid,” and “exactly one camera in this candidate set” without turning every transformation into another graph expansion. Lists let one row carry an ordered or unordered collection of values, but list expressions still have null and cardinality semantics that must be understood.
Use list literals, indexing, slicing and list comprehensions as value transformations rather than row multipliers.
Apply any, all,
none and single with correct
three-valued/null reasoning.
Use reduce() to fold a bounded list into one
scalar without confusing it with graph aggregation.
Build map/list API projections while controlling duplicates and payload size.
Explain when a list should be transformed in Cypher versus streamed and processed by the application.
The mandatory lab continues the course baseline: Neo4j
Community 2026.07.1, database neo4j,
explicit CYPHER 25 on version-sensitive queries,
authentication enabled, no mandatory APOC/GDS plugin, and the
reusable AtlasMart Customer/Order/Product/Category fixture
from Chapter 03. Neo4j 5.26.30 remains the LTS
comparison line. Because GROUP BY was introduced
for Cypher 25 in Neo4j 2026.07, examples that use it also show
the implicit-grouping form needed for Cypher 5 compatibility.
This generation environment does not run Neo4j or Docker. Commands and semantics were checked against current Neo4j documentation, but plan operators, estimated rows, memory figures and wall-clock timings must be captured on the learner's own pinned instance. Deterministic fixture counts are stated only where they follow directly from the fixture.
Re-establish the deterministic AtlasMart read fixture
Every lesson can stand alone. The fixture is idempotent, so
rerunning it is the safest reset: it restores four customers,
five orders, four products, two categories, five
PLACED, seven CONTAINS, and four
IN_CATEGORY relationships without deleting the
Chapter 04 operations graph. If your course sandbox already
contains the fixture, this simply normalizes the properties used
here.
CYPHER 25CREATE CONSTRAINT customer_id IF NOT EXISTS FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;CREATE CONSTRAINT product_id IF NOT EXISTS FOR (p:Product) REQUIRE p.productId IS UNIQUE;CREATE CONSTRAINT order_id IF NOT EXISTS FOR (o:Order) REQUIRE o.orderId IS UNIQUE;CREATE CONSTRAINT category_id IF NOT EXISTS FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;MERGE (c1:Customer {customerId:'C-1001'}) SET c1.name='Ava Chen', c1.tier='gold', c1.region='eu'MERGE (c2:Customer {customerId:'C-1002'}) SET c2.name='Noah Smith', c2.tier='silver', c2.region='us'MERGE (c3:Customer {customerId:'C-1003'}) SET c3.name='Mina Rahimi', c3.tier='gold', c3.region='me'MERGE (c4:Customer {customerId:'C-1004'}) SET c4.name='Leo Martin', c4.region='eu'MERGE (cat1:Category {categoryId:'CAT-CAMERAS'}) SET cat1.name='Cameras'MERGE (cat2:Category {categoryId:'CAT-AUDIO'}) SET cat2.name='Audio'MERGE (p1:Product {productId:'P-1001'}) SET p1.name='Trail Camera', p1.price=129.90, p1.rating=4.7MERGE (p2:Product {productId:'P-1002'}) SET p2.name='Studio Headphones', p2.price=89.00, p2.rating=4.7MERGE (p3:Product {productId:'P-1003'}) SET p3.name='Action Camera', p3.price=219.00, p3.rating=4.5MERGE (p4:Product {productId:'P-1004'}) SET p4.name='USB Microphone', p4.price=75.00MERGE (p1)-[:IN_CATEGORY]->(cat1)MERGE (p3)-[:IN_CATEGORY]->(cat1)MERGE (p2)-[:IN_CATEGORY]->(cat2)MERGE (p4)-[:IN_CATEGORY]->(cat2)MERGE (o1:Order {orderId:'O-2001'}) SET o1.placedAt=datetime('2026-08-01T09:00:00Z'), o1.status='paid', o1.total=218.90MERGE (o2:Order {orderId:'O-2002'}) SET o2.placedAt=datetime('2026-08-02T10:30:00Z'), o2.status='paid', o2.total=219.00MERGE (o3:Order {orderId:'O-2003'}) SET o3.placedAt=datetime('2026-08-03T12:00:00Z'), o3.status='shipped', o3.total=129.90MERGE (o4:Order {orderId:'O-2004'}) SET o4.placedAt=datetime('2026-08-04T14:15:00Z'), o4.status='paid', o4.total=164.00MERGE (o5:Order {orderId:'O-2005'}) SET o5.placedAt=datetime('2026-08-04T14:15:00Z'), o5.status='paid', o5.total=75.00MERGE (c1)-[:PLACED]->(o1)MERGE (c2)-[:PLACED]->(o2)MERGE (c1)-[:PLACED]->(o3)MERGE (c3)-[:PLACED]->(o4)MERGE (c4)-[:PLACED]->(o5)MERGE (o1)-[:CONTAINS {quantity:1}]->(p1)MERGE (o1)-[:CONTAINS {quantity:1}]->(p2)MERGE (o2)-[:CONTAINS {quantity:1}]->(p3)MERGE (o3)-[:CONTAINS {quantity:1}]->(p1)MERGE (o4)-[:CONTAINS {quantity:1}]->(p2)MERGE (o4)-[:CONTAINS {quantity:1}]->(p4)MERGE (o5)-[:CONTAINS {quantity:1}]->(p4);
Before analytical work, verify the relevant slice rather than assuming the whole database contains only these entities:
CYPHER 25MATCH (c:Customer) WITH count(c) AS customersMATCH (o:Order) WITH customers, count(o) AS ordersMATCH (p:Product) WITH customers, orders, count(p) AS productsMATCH (:Customer)-[placed:PLACED]->(:Order)WITH customers, orders, products, count(placed) AS placedMATCH (:Order)-[line:CONTAINS]->(:Product)RETURN customers, orders, products, placed, count(line) AS contains;
For the fixture above the invariant is
4 / 5 / 4 / 5 / 7. If Chapter 04 remains loaded,
the total database node/relationship counts will be larger;
label- and type-specific counts are the contract for this
chapter.
1. A list is one value on one row
A Cypher list is a value containing zero or
more elements. Returning a list does not itself multiply rows.
This is different from UNWIND, which turns elements
into rows. Lists are zero-indexed and support slicing. They may
contain heterogeneous query-time values, although
property-storage rules are stricter than general expression
rules.
CYPHER 25WITH ['P-1001','P-1002','P-1003','P-1004'] AS idsRETURN ids[0] AS first, ids[-1] AS last, ids[1..3] AS middle, size(ids) AS listSize;
2. List comprehensions combine filter and transform
A list comprehension has the shape
[x IN list WHERE predicate | expression]. It maps
each selected element to a new list element. Because it produces
one list value, it is useful when the API contract truly wants a
bounded collection.
CYPHER 25MATCH (o:Order {orderId:$orderId})-[line:CONTAINS]->(p:Product)WITH o, collect({productId:p.productId, price:p.price, quantity:line.quantity}) AS itemsRETURN o.orderId AS orderId, [item IN items WHERE item.price >= $minPrice | item.productId] AS expensiveProductIds, size(items) AS itemRows;
For O-2001 the fixture has two line relationships.
The comprehension never adds graph matches; it transforms the
list already in memory. If the relationship expansion produced
duplicates, the list will preserve them unless the query
deliberately deduplicates.
3. any/all/none/single are predicates, not counts
List predicate functions answer logical questions. They can
return null when null elements make the answer
undecidable. Empty-list results follow logic identities: “all
elements satisfy” and “no elements satisfy” can both be true on
an empty list, while “any” and “exactly one” are false.
CYPHER 25RETURN any(x IN [1,2,3] WHERE x > 2) AS anyGt2, all(x IN [1,2,3] WHERE x > 0) AS allPositive, none(x IN [1,2,3] WHERE x < 0) AS noneNegative, single(x IN [1,2,3] WHERE x = 2) AS exactlyOneTwo, any(x IN [1,null] WHERE x = 2) AS undecidableAny;
The last expression is the edge case: one element is false and one is unknown, so Cypher cannot prove either true or false. Do not mechanically map these predicates to two-valued application booleans without deciding how the API treats unknown.
4. reduce folds a list; aggregation folds rows
reduce() iterates over the elements of one list
value and updates an accumulator. That is conceptually different
from sum() across incoming rows, although both can
produce a scalar.
CYPHER 25WITH [218.90,129.90] AS orderTotalsRETURN reduce(gross = 0.0, value IN orderTotals | gross + value) AS gross;
Use reduce when the list itself is the data
structure you intend to fold. If you first
collect() millions of rows only to call
reduce(), you have added a large intermediate list
for no reason; aggregate the rows directly instead.
5. Deliberately wrong: unbounded collect as an API convenience
A developer might collect every product ever purchased by a high-volume customer into one response because a list is easy to serialize. The final query returns one row, but the server still has to build the list and the driver still has to receive it.
CYPHER 25MATCH (c:Customer {customerId:$customerId})-[:PLACED]->(:Order)-[:CONTAINS]->(p:Product)RETURN c.customerId AS customerId, collect(p{.productId,.name,.price}) AS allPurchasedProducts;
The safer contract is bounded: latest N orders, a deduplicated small set, a summary count, or a paginated endpoint. Which one is correct depends on the consumer; there is no universal list-size threshold.
Lab: transform values while recording row count separately
CYPHER 25MATCH (c:Customer)-[:PLACED]->(o:Order)WITH c, o ORDER BY o.placedAt DESC, o.orderId DESCWITH c, collect(o{.orderId,.status,.total})[0..3] AS latestOrdersRETURN c.customerId AS customerId, latestOrders, [x IN latestOrders WHERE x.total >= $minTotal | x.orderId] AS highValueOrderIds, all(x IN latestOrders WHERE x.status IN ['paid','shipped']) AS allAcceptedStatusesORDER BY customerId;
Verify that the final cardinality is one row per customer and
that each nested list is bounded to at most three order maps.
Then change $minTotal and prove that row count
remains stable while only list contents change.
Reset
No graph mutation is required. Rerun the fixture if prior
experiments changed relevant data; do not delete Chapter 04
OpsPoint nodes merely to reset an analytical
lesson.
Check your understanding
- What is the cardinality difference between a list comprehension and UNWIND?
- Why can any() return null?
- When should reduce() be preferred over sum()?
- Why is collect() potentially memory-sensitive even when the final result has one row?
- How can a bounded map list make an API contract clearer?
Review the answers
1. A comprehension transforms a list value on a row; UNWIND emits one row per element and can multiply cardinality.
2. If no element proves the predicate true but a null element leaves the outcome unknown, the result is null rather than false.
3. When you intentionally have one bounded list value to fold; sum() is usually better when the input is already a row stream.
4. The list itself must be materialized and then serialized, so memory and payload size grow with collected cardinality.
5. It states both the maximum collection size and exactly which fields cross the database/application boundary.
Summary and next step
Lists let Cypher reshape values without changing row count, but
predicates and nulls still matter and unbounded collections
remain a resource risk. Next, invert the operation with
UNWIND: turn input lists back into rows and reason
about multiplication, duplicates, empties and order.
Authoritative references
- Current Neo4j versions — Release/LTS snapshot used to pin the course baseline.
- Cypher Manual — Introduction — Cypher 25 status and language-version policy.
- Select Cypher version — Database/default and per-query Cypher version behavior.
- List expressions — List comprehensions, access, slicing and transformations.
- List functions — reduce() and related list functions.
- Working with null — Three-valued behavior for list predicates and null expressions.
- Functions — predicate functions — any/all/none/single signatures and current Cypher 25 functions.