Chapter 10 · Query Planning and Profiling: Cardinality, Operators, Index Selection, and Plan Stability
Cardinality Estimation, Selectivity, Starting Points, Expand Operators, and Why Graph Shape Matters
Explain AtlasMart query cost from row growth, selectivity and degree distribution instead of blaming Cypher syntax in isolation.
Learning outcomes
Two AtlasMart customers can run the same parameterized query and experience very different work because graph shape is data. A selective customer seek may be perfect, yet a high-degree customer can still generate thousands of downstream rows. This lesson follows those rows.
Explain cardinality estimates, selectivity and why starting-point quality matters.
Distinguish Expand(All) from Expand(Into) conceptually and identify traversal fan-out.
Measure degree distributions and sparse/dense parameter shapes before making tuning claims.
Compare customer-first and product-first access paths without assuming one anchor is universally best.
Recognize when a model-level supernode/fan-out problem cannot be fixed by adding another index.
The mandatory lab continues Neo4j Community
2026.07.1, database neo4j, explicit
CYPHER 25 for version-sensitive examples,
authentication enabled, no mandatory APOC/GDS plugin, and the
AtlasMart identifiers/model established in Chapters 01–09.
Neo4j 5.26.30 remains the LTS comparison line.
Community uses the slotted runtime by default. Enterprise uses
the pipelined runtime by default; Aura uses pipelined, and the
parallel runtime is an Enterprise capability for supported
read workloads. Therefore plan columns such as pipeline timing
and page-cache hits/misses are not universal Community output.
This generation environment does not run Neo4j or Docker.
Commands were checked against current official documentation
but were not executed here. The lessons never invent DB-hit
counts, memory figures, operator timings or latency
improvements. Instead, they define deterministic fixture
invariants and tell you exactly which values to record from
your own EXPLAIN/PROFILE runs. All
disposable Chapter 10 entities use labTag='ch10';
disposable schema objects use ch10_*.
1. Cardinality is the currency of the plan
Every operator receives rows and emits rows. The planner
estimates those counts from graph statistics, index selectivity
and a selectivity model; PROFILE shows what
actually happened. Estimation errors matter because they
influence join order, starting points and operator choice, but
an error is diagnostic evidence—not proof of a specific remedy.
CYPHER 25MATCH (c:Customer {labTag:'ch10'})-[:PLACED]->(o:Order {labTag:'ch10'})WITH c,count(o) AS degreeRETURN min(degree) AS minOrders, avg(degree) AS avgOrders, max(degree) AS maxOrders, percentileDisc(degree,0.50) AS p50Orders, percentileDisc(degree,0.95) AS p95Orders;
2. A good seek can still lead to expensive expansion
A uniqueness-backed seek for customerId should bind
one customer efficiently. From that point,
Expand(All)-style traversal discovers its outgoing
PLACED relationships and emits one row per matching
relationship/end node. For the hot customer, that expansion is
deliberately much larger than for a normal customer. An index on
customerId cannot change the degree of the
already-bound customer.
CYPHER 25PROFILEMATCH (c:Customer {customerId:$customerId})-[:PLACED]->(o:Order)RETURN c.customerId,count(o) AS orders;
Run with $customerId='CH10-C-0001' and again with a
typical customer such as CH10-C-0050. Record actual
rows at the seek and expansion operators plus total DB hits. Do
not fabricate a universal ratio; the fixture defines only the
business-degree difference.
3. Starting point depends on the question and data distribution
Suppose the API asks for high-priced camera products purchased by one customer. Customer-first traversal is attractive when the customer is sparse. Product-first access can be attractive when the product predicate is very selective. On the hot customer, the best plan can differ from what your intuition based on a typical customer suggests.
CYPHER 25EXPLAINMATCH (c:Customer {customerId:$customerId})-[:PLACED]->(o:Order)-[:CONTAINS]->(p:Product)WHERE p.categoryCode=$category AND p.price >= $minPriceRETURN DISTINCT p.productId,p.price;
CYPHER 25EXPLAINMATCH (p:Product)WHERE p.categoryCode=$category AND p.price >= $minPriceMATCH (p)<-[:CONTAINS]-(o:Order)<-[:PLACED]-(c:Customer {customerId:$customerId})RETURN DISTINCT p.productId,p.price;
4. Selectivity must be measured per predicate population
CYPHER 25MATCH (p:Product {labTag:'ch10'})WITH count(p) AS totalMATCH (p:Product {labTag:'ch10'})WHERE p.categoryCode=$category AND p.price >= $minPriceRETURN total,count(p) AS matches,toFloat(count(p))/total AS fraction;
| Shape | Likely implication | Do not overgeneralize |
|---|---|---|
| Unique ID seek → one node | Excellent anchor lookup | Downstream degree may still dominate |
| Low-selectivity category | Many candidate products | Index is useless; combined predicates may still be selective |
| Hot customer/supplier/store | Large expansion fan-out | Every high-degree node is a modeling error |
| Sparse parameter | Small actual rows | Same query is cheap for dense parameters |
5. Deliberately wrong: force an index to fix a supernode
An index finds starting entities; it does not shrink the number of relationships attached to an already-found node. If the workload repeatedly traverses a high-degree hub, consider query predicates, relationship semantics, time/window partitioning, precomputed summary relationships, intermediate fact nodes, or a different read model—but only after proving the model change preserves business semantics.
Check your understanding
- What is cardinality in a plan?
- Can a unique index remove traversal fan-out after a node is found?
- Why test both sparse and dense parameters?
- What does Expand(All) conceptually do?
- When should the model change?
Review the answers
1. The number of rows flowing from an operator; EXPLAIN estimates it and PROFILE measures it.
2. No. It improves the starting lookup, not the node’s degree.
3. The same plan can perform very different amounts of work across degree/selectivity distributions.
4. From a bound start node, it traverses matching relationships and emits rows for matching end nodes/relationships.
5. When evidence shows graph shape itself causes recurring work and the proposed refactor preserves required semantics and write behavior.
Summary and next step
Cardinality connects graph shape to plan cost. Next, we inspect the concrete access, join/apply, eager, sort and aggregation operators that transform those rows.
Authoritative references
- Current Neo4j versions — Release/LTS snapshot used for this chapter.
- Understanding query plans — EXPLAIN/PROFILE, plan columns, rows, DB hits, memory and plan reading.
- Operators — Current operator families and operator semantics.
- Operators in detail — Current seek/scan, expand, apply/join, eager, sort and aggregation operators.
- Statistics and execution plans — Statistics collection, selectivity, replanning thresholds and manual preparation.
- Query caches — Per-database query caches, cache sizing and current query-size behavior.
- Cypher runtimes — Slotted, pipelined and parallel runtime boundaries.
- Index impact on performance — Planner use of search-performance indexes and selectivity.
- Traversal operators — Expand and related traversal operator semantics.