Chapter 10 · Query Planning and Profiling: Cardinality, Operators, Index Selection, and Plan Stability

Cardinality Estimation, Selectivity, Starting Points, Expand Operators, and Why Graph Shape Matters

Explain AtlasMart query cost from row growth, selectivity and degree distribution instead of blaming Cypher syntax in isolation.

Advanced145–180 minutesCardinality + graph-shape labNeo4j 2026.07.1 Community · Cypher 25Last reviewed: September 2026

Learning outcomes

Two AtlasMart customers can run the same parameterized query and experience very different work because graph shape is data. A selective customer seek may be perfect, yet a high-degree customer can still generate thousands of downstream rows. This lesson follows those rows.

01

Explain cardinality estimates, selectivity and why starting-point quality matters.

02

Distinguish Expand(All) from Expand(Into) conceptually and identify traversal fan-out.

03

Measure degree distributions and sparse/dense parameter shapes before making tuning claims.

04

Compare customer-first and product-first access paths without assuming one anchor is universally best.

05

Recognize when a model-level supernode/fan-out problem cannot be fixed by adding another index.

Chapter 10 baseline · reviewed 9 September 2026

The mandatory lab continues Neo4j Community 2026.07.1, database neo4j, explicit CYPHER 25 for version-sensitive examples, authentication enabled, no mandatory APOC/GDS plugin, and the AtlasMart identifiers/model established in Chapters 01–09. Neo4j 5.26.30 remains the LTS comparison line. Community uses the slotted runtime by default. Enterprise uses the pipelined runtime by default; Aura uses pipelined, and the parallel runtime is an Enterprise capability for supported read workloads. Therefore plan columns such as pipeline timing and page-cache hits/misses are not universal Community output.

Evidence and measurement note

This generation environment does not run Neo4j or Docker. Commands were checked against current official documentation but were not executed here. The lessons never invent DB-hit counts, memory figures, operator timings or latency improvements. Instead, they define deterministic fixture invariants and tell you exactly which values to record from your own EXPLAIN/PROFILE runs. All disposable Chapter 10 entities use labTag='ch10'; disposable schema objects use ch10_*.

1. Cardinality is the currency of the plan

Every operator receives rows and emits rows. The planner estimates those counts from graph statistics, index selectivity and a selectivity model; PROFILE shows what actually happened. Estimation errors matter because they influence join order, starting points and operator choice, but an error is diagnostic evidence—not proof of a specific remedy.

Cypher · inspect the intended skew
CYPHER 25MATCH (c:Customer {labTag:'ch10'})-[:PLACED]->(o:Order {labTag:'ch10'})WITH c,count(o) AS degreeRETURN min(degree) AS minOrders,       avg(degree) AS avgOrders,       max(degree) AS maxOrders,       percentileDisc(degree,0.50) AS p50Orders,       percentileDisc(degree,0.95) AS p95Orders;

2. A good seek can still lead to expensive expansion

A uniqueness-backed seek for customerId should bind one customer efficiently. From that point, Expand(All)-style traversal discovers its outgoing PLACED relationships and emits one row per matching relationship/end node. For the hot customer, that expansion is deliberately much larger than for a normal customer. An index on customerId cannot change the degree of the already-bound customer.

Cypher · compare dense and sparse parameters
CYPHER 25PROFILEMATCH (c:Customer {customerId:$customerId})-[:PLACED]->(o:Order)RETURN c.customerId,count(o) AS orders;

Run with $customerId='CH10-C-0001' and again with a typical customer such as CH10-C-0050. Record actual rows at the seek and expansion operators plus total DB hits. Do not fabricate a universal ratio; the fixture defines only the business-degree difference.

3. Starting point depends on the question and data distribution

Suppose the API asks for high-priced camera products purchased by one customer. Customer-first traversal is attractive when the customer is sparse. Product-first access can be attractive when the product predicate is very selective. On the hot customer, the best plan can differ from what your intuition based on a typical customer suggests.

Cypher · customer-first candidate
CYPHER 25EXPLAINMATCH (c:Customer {customerId:$customerId})-[:PLACED]->(o:Order)-[:CONTAINS]->(p:Product)WHERE p.categoryCode=$category AND p.price >= $minPriceRETURN DISTINCT p.productId,p.price;
Cypher · product-first candidate
CYPHER 25EXPLAINMATCH (p:Product)WHERE p.categoryCode=$category AND p.price >= $minPriceMATCH (p)<-[:CONTAINS]-(o:Order)<-[:PLACED]-(c:Customer {customerId:$customerId})RETURN DISTINCT p.productId,p.price;

4. Selectivity must be measured per predicate population

Cypher · measure category and price selectivity in the fixture
CYPHER 25MATCH (p:Product {labTag:'ch10'})WITH count(p) AS totalMATCH (p:Product {labTag:'ch10'})WHERE p.categoryCode=$category AND p.price >= $minPriceRETURN total,count(p) AS matches,toFloat(count(p))/total AS fraction;
Shape Likely implication Do not overgeneralize
Unique ID seek → one node Excellent anchor lookup Downstream degree may still dominate
Low-selectivity category Many candidate products Index is useless; combined predicates may still be selective
Hot customer/supplier/store Large expansion fan-out Every high-degree node is a modeling error
Sparse parameter Small actual rows Same query is cheap for dense parameters

5. Deliberately wrong: force an index to fix a supernode

An index finds starting entities; it does not shrink the number of relationships attached to an already-found node. If the workload repeatedly traverses a high-degree hub, consider query predicates, relationship semantics, time/window partitioning, precomputed summary relationships, intermediate fact nodes, or a different read model—but only after proving the model change preserves business semantics.

Check your understanding

  1. What is cardinality in a plan?
  2. Can a unique index remove traversal fan-out after a node is found?
  3. Why test both sparse and dense parameters?
  4. What does Expand(All) conceptually do?
  5. When should the model change?
Review the answers

1. The number of rows flowing from an operator; EXPLAIN estimates it and PROFILE measures it.

2. No. It improves the starting lookup, not the node’s degree.

3. The same plan can perform very different amounts of work across degree/selectivity distributions.

4. From a bound start node, it traverses matching relationships and emits rows for matching end nodes/relationships.

5. When evidence shows graph shape itself causes recurring work and the proposed refactor preserves required semantics and write behavior.

Summary and next step

Cardinality connects graph shape to plan cost. Next, we inspect the concrete access, join/apply, eager, sort and aggregation operators that transform those rows.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.