Chapter 03 · Cypher Fundamentals: MATCH, RETURN, WHERE, Parameters, Ordering, and Pagination
Filter with WHERE, Predicates, Labels, Types, Property Access, Null Semantics, and Parameters
Make predicate truth and placement explicit: only true rows pass, missing properties become null, and user values stay separate from executable Cypher text.
Learning outcomes
AtlasMart now understands binding rows, but correctness depends
on which rows/patterns are allowed to survive. Cypher predicates
can evaluate true, false, or
null; missing properties evaluate to null; and
WHERE attached to a MATCH is part of
that pattern constraint, not merely a universal post-processing
filter.
Place WHERE with the MATCH/WITH it is intended to constrain.
Use labels, relationship types, property access, comparison/string/list/type predicates and parameters safely.
Predict three-valued logic for missing properties and use IS NULL / IS NOT NULL correctly.
Distinguish parameterized values from unsafe query-text concatenation.
Diagnose a predicate that is logically correct but applied after unnecessary row expansion.
Continue Chapters 01–02 with Neo4j Community
2026.07.1, database neo4j, explicit
CYPHER 25 in version-sensitive examples, local
container atlasmart-neo4j, Bolt
127.0.0.1:7687, HTTP 127.0.0.1:7474,
and the existing constraint-backed AtlasMart domain IDs. This
chapter does not redesign the graph; it adds a deterministic
read fixture and treats Cypher as a pipeline of row bindings
produced and transformed by patterns, predicates and
projections.
The current Cypher Manual covers Cypher 25. Cypher 5 is frozen
while new language features since Neo4j 2025.06 are added to
Cypher 25; current 2026.02+ newly created databases explicitly
default to Cypher 25, while existing deployments may differ.
The mandatory examples therefore use
CYPHER 25 and avoid assuming a server-wide
default. This environment has no running Neo4j/Docker runtime,
so expected results are deterministic fixture invariants
rather than fabricated captured output.
1. WHERE belongs to a query context
With MATCH, WHERE constrains the
pattern described by the directly preceding match. After
WITH, it filters rows produced by the previous
query part. That distinction becomes crucial with multiple
matches and optional patterns. Chapter 04 will cover
OPTIONAL MATCH deeply; here the rule is to read
WHERE together with the clause it qualifies.
CYPHER 25MATCH (c:Customer)WHERE c.region = $region AND c.tier = $tierMATCH (c)-[:PLACED]->(o:Order)RETURN c.customerId, o.orderIdORDER BY c.customerId, o.orderId;
With $region='eu' and $tier='gold',
only Ava qualifies in the deterministic fixture, then her two
orders expand to two rows.
2. Missing properties participate in three-valued logic
CYPHER 25CREATE CONSTRAINT customer_id IF NOT EXISTS FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;CREATE CONSTRAINT product_id IF NOT EXISTS FOR (p:Product) REQUIRE p.productId IS UNIQUE;CREATE CONSTRAINT order_id IF NOT EXISTS FOR (o:Order) REQUIRE o.orderId IS UNIQUE;CREATE CONSTRAINT category_id IF NOT EXISTS FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;MERGE (c1:Customer {customerId:'C-1001'}) SET c1.name='Ava Chen', c1.tier='gold', c1.region='eu'MERGE (c2:Customer {customerId:'C-1002'}) SET c2.name='Noah Smith', c2.tier='silver', c2.region='us'MERGE (c3:Customer {customerId:'C-1003'}) SET c3.name='Mina Rahimi', c3.tier='gold', c3.region='me'MERGE (c4:Customer {customerId:'C-1004'}) SET c4.name='Leo Martin', c4.region='eu'MERGE (cat1:Category {categoryId:'CAT-CAMERAS'}) SET cat1.name='Cameras'MERGE (cat2:Category {categoryId:'CAT-AUDIO'}) SET cat2.name='Audio'MERGE (p1:Product {productId:'P-1001'}) SET p1.name='Trail Camera', p1.price=129.90, p1.rating=4.7MERGE (p2:Product {productId:'P-1002'}) SET p2.name='Studio Headphones', p2.price=89.00, p2.rating=4.7MERGE (p3:Product {productId:'P-1003'}) SET p3.name='Action Camera', p3.price=219.00, p3.rating=4.5MERGE (p4:Product {productId:'P-1004'}) SET p4.name='USB Microphone', p4.price=75.00MERGE (p1)-[:IN_CATEGORY]->(cat1)MERGE (p3)-[:IN_CATEGORY]->(cat1)MERGE (p2)-[:IN_CATEGORY]->(cat2)MERGE (p4)-[:IN_CATEGORY]->(cat2)MERGE (o1:Order {orderId:'O-2001'}) SET o1.placedAt=datetime('2026-08-01T09:00:00Z'), o1.status='paid', o1.total=218.90MERGE (o2:Order {orderId:'O-2002'}) SET o2.placedAt=datetime('2026-08-02T10:30:00Z'), o2.status='paid', o2.total=219.00MERGE (o3:Order {orderId:'O-2003'}) SET o3.placedAt=datetime('2026-08-03T12:00:00Z'), o3.status='shipped', o3.total=129.90MERGE (o4:Order {orderId:'O-2004'}) SET o4.placedAt=datetime('2026-08-04T14:15:00Z'), o4.status='paid', o4.total=164.00MERGE (o5:Order {orderId:'O-2005'}) SET o5.placedAt=datetime('2026-08-04T14:15:00Z'), o5.status='paid', o5.total=75.00MERGE (c1)-[:PLACED]->(o1)MERGE (c2)-[:PLACED]->(o2)MERGE (c1)-[:PLACED]->(o3)MERGE (c3)-[:PLACED]->(o4)MERGE (c4)-[:PLACED]->(o5)MERGE (o1)-[:CONTAINS {quantity:1}]->(p1)MERGE (o1)-[:CONTAINS {quantity:1}]->(p2)MERGE (o2)-[:CONTAINS {quantity:1}]->(p3)MERGE (o3)-[:CONTAINS {quantity:1}]->(p1)MERGE (o4)-[:CONTAINS {quantity:1}]->(p2)MERGE (o4)-[:CONTAINS {quantity:1}]->(p4)MERGE (o5)-[:CONTAINS {quantity:1}]->(p4);
CYPHER 25MATCH (c:Customer {customerId:'C-1004'})RETURN c.tier AS tier, c.tier = 'gold' AS equalsGold, c.tier <> 'gold' AS notGold, c.tier IS NULL AS tierMissing;
The expected invariant is tier=null, both
equality/inequality predicates evaluate to null, and
tierMissing=true. In a WHERE filter,
only true passes; false and
null do not. This is why “not gold” should not be
written as c.tier <> 'gold' when missing
tiers must also be included.
CYPHER 25MATCH (c:Customer)WHERE c.tier IS NULL OR c.tier <> $excludedTierRETURN c.customerId, c.tierORDER BY c.customerId;
3. Parameters separate values from query structure
Parameters are named values supplied separately from the Cypher text. They improve plan reuse and, more importantly, prevent user data from becoming executable syntax when used correctly. They do not magically parameterize every piece of Cypher structure; keep user-controlled values in parameters and expose only validated/allowlisted structural choices.
:param {region: 'eu', minimumTotal: 100.0};CYPHER 25MATCH (c:Customer)-[:PLACED]->(o:Order)WHERE c.region = $region AND o.total >= $minimumTotalRETURN c.customerId, o.orderId, o.totalORDER BY o.total DESC, o.orderId;
# Wrong: user text changes Cypher syntaxquery = "MATCH (c:Customer) WHERE c.name = '" + user_name + "' RETURN c"# Correct: query text is stable; data is a parameterquery = "MATCH (c:Customer) WHERE c.name = $name RETURN c.customerId, c.name"records, summary, keys = driver.execute_query(query, name=user_name, database_="neo4j")
The mandatory lab can stay entirely in cypher-shell; the driver snippet demonstrates the same contract for application code. Driver version remains the Chapter 01 continuity reference rather than a new requirement.
4. Wrong approach: filter after exploding rows
Suppose the API wants gold customers’ paid orders. Expanding every customer through order lines/products and only then filtering by customer tier can multiply rows before the selective predicate is applied. The planner can sometimes push predicates, but query writers should express intent early and inspect plans rather than rely on folklore.
CYPHER 25 EXPLAINMATCH (c:Customer)-[:PLACED]->(o:Order)-[:CONTAINS]->(p:Product)WHERE c.tier=$tier AND o.status=$statusRETURN c.customerId, o.orderId, p.productId;CYPHER 25 EXPLAINMATCH (c:Customer)WHERE c.tier=$tierMATCH (c)-[:PLACED]->(o:Order)WHERE o.status=$statusMATCH (o)-[:CONTAINS]->(p:Product)RETURN c.customerId, o.orderId, p.productId;
Do not promise one universal plan: statistics, indexes and runtime can change operator choices. The evidence question is whether selective predicates and indexes reduce starting cardinality before expensive expansion.
Lab: predicate truth table
CYPHER 25MATCH (c:Customer)RETURN c.customerId, c.region = $region AS regionMatches, c.tier = $tier AS tierMatches, c.tier IS NULL AS tierMissing, (c.region = $region AND c.tier = $tier) AS bothMatchORDER BY c.customerId;
Run with $region='eu', $tier='gold'.
Verify Ava is true/true/false/true; Leo is true/null/true/null.
Then use a WHERE clause and confirm only rows whose predicate is
true survive.
Check your understanding
- Is WHERE always a standalone post-filter?
- What does a missing node property evaluate to?
- Why does c.tier <> "gold" not include missing tiers?
- What is the main security benefit of parameters?
- Can EXPLAIN alone prove runtime latency improved?
Review the answers
1. No. With MATCH/OPTIONAL MATCH it constrains the directly preceding pattern; after WITH it filters rows.
2. null.
3. The comparison itself becomes null, and WHERE only passes true.
4. User-supplied data remains data rather than being concatenated into executable query text.
5. No. EXPLAIN does not execute; runtime evidence needs PROFILE/measurement under controlled conditions.
Production judgment
Predicate correctness is an API contract. Document whether missing values mean unknown, absent, or excluded; parameterize user data; and test edge cases where a predicate becomes null. Query shape and graph density determine how much data is touched before a selective condition can reduce rows. Next, we project those rows into stable response shapes without accidentally confusing DISTINCT with aggregation or storage deduplication.
Summary and next step
Filter with WHERE, Predicates, Labels, Types, Property Access, Null Semantics, and Parameters is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Project Results with RETURN, Aliases, DISTINCT, Expressions, Maps, and Pattern Comprehensions. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- Current Neo4j versions — Official current-release and 5.26 LTS patch snapshot.
- Cypher Manual introduction — Current Cypher 25 baseline and Cypher 5 compatibility framing.
- MATCH — Pattern matching and variable-binding semantics.
- Variables — Variable naming and scope across query parts.
- WHERE — Pattern constraints and post-WITH filtering semantics.
- Predicates — Boolean/comparison/string/list/type predicates and three-valued results.
- Working with null — Null propagation and missing-property semantics.
- RETURN — Projection, aliases, expressions and DISTINCT semantics.
- List expressions and pattern comprehension — Current fixed/variable pattern-comprehension syntax and behavior.
- ORDER BY — Only ORDER BY guarantees result ordering; sort semantics and index-backed ordering.
- SKIP / OFFSET — Offset semantics and OFFSET synonym.
- LIMIT — Row limiting semantics and the absence of ordering guarantees without ORDER BY.
- Cypher Shell — Bolt CLI, parameter support, script/file execution and output modes.
- Query tuning and plans — EXPLAIN/PROFILE and execution-plan evidence.