Chapter 10 · Query Planning and Profiling: Cardinality, Operators, Index Selection, and Plan Stability

Index Seeks/Scans, Label/Type Lookups, Join/Apply Operators, Eagerization, and Sort/Aggregation Costs

Translate Neo4j plan operators into concrete row, storage and memory mechanisms that can be measured and challenged.

Advanced150–185 minutesOperator-mechanism profiling labNeo4j 2026.07.1 Community · Cypher 25Last reviewed: September 2026

Learning outcomes

Plans become actionable only when operator names are tied to mechanisms. This lesson groups the operators by what they do to rows: locate starts, traverse, combine streams, materialize, sort and aggregate.

01

Distinguish index seek/scan and token lookup starting operators by workload intent.

02

Explain Expand, Apply and join behavior from row flow rather than memorizing names.

03

Recognize Eager and blocking operators as correctness/memory boundaries.

04

Relate Sort and aggregation memory to input cardinality and ordering opportunities.

05

Use plan changes as causal evidence while allowing operator names/runtime details to evolve by release.

Chapter 10 baseline · reviewed 9 September 2026

The mandatory lab continues Neo4j Community 2026.07.1, database neo4j, explicit CYPHER 25 for version-sensitive examples, authentication enabled, no mandatory APOC/GDS plugin, and the AtlasMart identifiers/model established in Chapters 01–09. Neo4j 5.26.30 remains the LTS comparison line. Community uses the slotted runtime by default. Enterprise uses the pipelined runtime by default; Aura uses pipelined, and the parallel runtime is an Enterprise capability for supported read workloads. Therefore plan columns such as pipeline timing and page-cache hits/misses are not universal Community output.

Evidence and measurement note

This generation environment does not run Neo4j or Docker. Commands were checked against current official documentation but were not executed here. The lessons never invent DB-hit counts, memory figures, operator timings or latency improvements. Instead, they define deterministic fixture invariants and tell you exactly which values to record from your own EXPLAIN/PROFILE runs. All disposable Chapter 10 entities use labTag='ch10'; disposable schema objects use ch10_*.

1. Leaf operators choose where work begins

Operator family Mechanism Typical AtlasMart evidence
NodeUniqueIndexSeek / NodeIndexSeek Locate nodes through property indexes Customer by constrained customerId; Product category/price predicate
NodeIndexScan / range scan variants Read index entries over a broader predicate/value set Property existence/range-like access when a seek is not exact
NodeByLabelScan / relationship-type lookup Use label/type token lookup to enumerate a class All Products or all CONTAINS relationships
AllNodesScan Enumerate every node Usually evidence that the query has no usable selective anchor

Names can change or be specialized by release/runtime. The stable question is: how did the plan locate the first candidate entities, and how many rows emerged?

Cypher · compare a selective index predicate with a broad label query
CYPHER 25EXPLAIN MATCH (p:Product)WHERE p.categoryCode='CAM' AND p.price >= 150RETURN p.productId,p.price;EXPLAIN MATCH (p:Product)RETURN p.productId,p.price;

2. Expand transforms bound nodes into relationship rows

Expand(All) follows relationships from a bound node and introduces the relationship/end node. Expand(Into) is used when both endpoints are already bound and Neo4j needs to test/find relationships between them. Look at rows before and after expansion: that delta is graph-shape work.

Cypher · expose expansion cardinality directly
CYPHER 25PROFILEMATCH (c:Customer {customerId:$customerId})-[placed:PLACED]->(o:Order)-[line:CONTAINS]->(p:Product)RETURN count(placed) AS placedRows,count(line) AS lineRows;

3. Apply and joins combine row streams

Apply is a nested-loop family: for each left-side row, the right-side subtree is evaluated with imported variables. Semi/anti variants answer existence/non-existence questions without returning the right-side rows. Hash joins combine streams around common bound nodes when the planner judges that shape useful. Do not rewrite simply to eliminate the word “Apply”; measure the rows and repeated right-side work.

Cypher · correlated subquery that may plan with Apply-family operators
CYPHER 25EXPLAINMATCH (c:Customer {customerId:$customerId})CALL (c) {  MATCH (c)-[:PLACED]->(o:Order)  RETURN count(o) AS orderCount}RETURN c.customerId,orderCount;

4. Eagerization protects semantics but buffers work

Cypher is generally lazy: rows can flow upward as soon as they are produced. Some operations must consume/materialize upstream rows first. Eager can be inserted to preserve isolation when reads and writes could otherwise interfere; EagerAggregation groups rows; Sort must retain rows to order them. Blocking behavior matters when the upstream cardinality is large.

Cypher · profile a sort/aggregation boundary
CYPHER 25PROFILEMATCH (c:Customer {labTag:'ch10'})-[:PLACED]->(o:Order {labTag:'ch10'})WITH c.segment AS segment,count(o) AS ordersRETURN segment,ordersORDER BY orders DESC;

5. Deliberately wrong: optimize the final row count

A query returning three aggregate rows can still expand through thousands of intermediate rows and buffer them before aggregation/sort. The repair is to inspect the entire tree and especially the transitions where estimated/actual rows jump, DB hits accumulate, or eager memory appears.

Operator question Useful diagnostic
Seek/scan How many candidate entities enter the plan?
Expand How much does degree multiply rows?
Apply/join How often is the right side evaluated / how large are both inputs?
Eager/Sort/Aggregation How many rows must be materialized, and what memory does PROFILE expose?
ProduceResults How much data leaves the server compared with intermediate work?

Check your understanding

  1. Is NodeIndexSeek always faster than every scan?
  2. What is Expand(Into) for?
  3. Does Apply automatically mean a bad plan?
  4. Why can Eager be necessary?
  5. Can a three-row result hide a large plan?
Review the answers

1. No. It is a selective access mechanism; total plan cost depends on predicate selectivity, downstream work and data shape.

2. Finding matching relationships when both endpoint nodes are already bound.

3. No. It is a nested-loop mechanism used for correlated work; inspect cardinality and repeated right-side cost.

4. To preserve correct semantics/isolation by fully materializing prior work before later operations proceed.

5. Yes. Aggregation can collapse a very large intermediate row set into a tiny final result.

Summary and next step

Operators are mechanisms, not grades. Next, we add the time dimension: statistics evolve, plans are cached, parameters can be skewed, and one captured plan is not automatically representative.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.