Chapter 10 · Query Planning and Profiling: Cardinality, Operators, Index Selection, and Plan Stability
Index Seeks/Scans, Label/Type Lookups, Join/Apply Operators, Eagerization, and Sort/Aggregation Costs
Translate Neo4j plan operators into concrete row, storage and memory mechanisms that can be measured and challenged.
Learning outcomes
Plans become actionable only when operator names are tied to mechanisms. This lesson groups the operators by what they do to rows: locate starts, traverse, combine streams, materialize, sort and aggregate.
Distinguish index seek/scan and token lookup starting operators by workload intent.
Explain Expand, Apply and join behavior from row flow rather than memorizing names.
Recognize Eager and blocking operators as correctness/memory boundaries.
Relate Sort and aggregation memory to input cardinality and ordering opportunities.
Use plan changes as causal evidence while allowing operator names/runtime details to evolve by release.
The mandatory lab continues Neo4j Community
2026.07.1, database neo4j, explicit
CYPHER 25 for version-sensitive examples,
authentication enabled, no mandatory APOC/GDS plugin, and the
AtlasMart identifiers/model established in Chapters 01–09.
Neo4j 5.26.30 remains the LTS comparison line.
Community uses the slotted runtime by default. Enterprise uses
the pipelined runtime by default; Aura uses pipelined, and the
parallel runtime is an Enterprise capability for supported
read workloads. Therefore plan columns such as pipeline timing
and page-cache hits/misses are not universal Community output.
This generation environment does not run Neo4j or Docker.
Commands were checked against current official documentation
but were not executed here. The lessons never invent DB-hit
counts, memory figures, operator timings or latency
improvements. Instead, they define deterministic fixture
invariants and tell you exactly which values to record from
your own EXPLAIN/PROFILE runs. All
disposable Chapter 10 entities use labTag='ch10';
disposable schema objects use ch10_*.
1. Leaf operators choose where work begins
| Operator family | Mechanism | Typical AtlasMart evidence |
|---|---|---|
| NodeUniqueIndexSeek / NodeIndexSeek | Locate nodes through property indexes | Customer by constrained customerId; Product category/price predicate |
| NodeIndexScan / range scan variants | Read index entries over a broader predicate/value set | Property existence/range-like access when a seek is not exact |
| NodeByLabelScan / relationship-type lookup | Use label/type token lookup to enumerate a class | All Products or all CONTAINS relationships |
| AllNodesScan | Enumerate every node | Usually evidence that the query has no usable selective anchor |
Names can change or be specialized by release/runtime. The stable question is: how did the plan locate the first candidate entities, and how many rows emerged?
CYPHER 25EXPLAIN MATCH (p:Product)WHERE p.categoryCode='CAM' AND p.price >= 150RETURN p.productId,p.price;EXPLAIN MATCH (p:Product)RETURN p.productId,p.price;
2. Expand transforms bound nodes into relationship rows
Expand(All) follows relationships from a bound node
and introduces the relationship/end node.
Expand(Into) is used when both endpoints are
already bound and Neo4j needs to test/find relationships between
them. Look at rows before and after expansion: that delta is
graph-shape work.
CYPHER 25PROFILEMATCH (c:Customer {customerId:$customerId})-[placed:PLACED]->(o:Order)-[line:CONTAINS]->(p:Product)RETURN count(placed) AS placedRows,count(line) AS lineRows;
3. Apply and joins combine row streams
Apply is a nested-loop family: for each left-side
row, the right-side subtree is evaluated with imported
variables. Semi/anti variants answer existence/non-existence
questions without returning the right-side rows. Hash joins
combine streams around common bound nodes when the planner
judges that shape useful. Do not rewrite simply to eliminate the
word “Apply”; measure the rows and repeated right-side work.
CYPHER 25EXPLAINMATCH (c:Customer {customerId:$customerId})CALL (c) { MATCH (c)-[:PLACED]->(o:Order) RETURN count(o) AS orderCount}RETURN c.customerId,orderCount;
4. Eagerization protects semantics but buffers work
Cypher is generally lazy: rows can flow upward as soon as they
are produced. Some operations must consume/materialize upstream
rows first. Eager can be inserted to preserve
isolation when reads and writes could otherwise interfere;
EagerAggregation groups rows;
Sort must retain rows to order them. Blocking
behavior matters when the upstream cardinality is large.
CYPHER 25PROFILEMATCH (c:Customer {labTag:'ch10'})-[:PLACED]->(o:Order {labTag:'ch10'})WITH c.segment AS segment,count(o) AS ordersRETURN segment,ordersORDER BY orders DESC;
5. Deliberately wrong: optimize the final row count
A query returning three aggregate rows can still expand through thousands of intermediate rows and buffer them before aggregation/sort. The repair is to inspect the entire tree and especially the transitions where estimated/actual rows jump, DB hits accumulate, or eager memory appears.
| Operator question | Useful diagnostic |
|---|---|
| Seek/scan | How many candidate entities enter the plan? |
| Expand | How much does degree multiply rows? |
| Apply/join | How often is the right side evaluated / how large are both inputs? |
| Eager/Sort/Aggregation | How many rows must be materialized, and what memory does PROFILE expose? |
| ProduceResults | How much data leaves the server compared with intermediate work? |
Check your understanding
- Is NodeIndexSeek always faster than every scan?
- What is Expand(Into) for?
- Does Apply automatically mean a bad plan?
- Why can Eager be necessary?
- Can a three-row result hide a large plan?
Review the answers
1. No. It is a selective access mechanism; total plan cost depends on predicate selectivity, downstream work and data shape.
2. Finding matching relationships when both endpoint nodes are already bound.
3. No. It is a nested-loop mechanism used for correlated work; inspect cardinality and repeated right-side cost.
4. To preserve correct semantics/isolation by fully materializing prior work before later operations proceed.
5. Yes. Aggregation can collapse a very large intermediate row set into a tiny final result.
Summary and next step
Operators are mechanisms, not grades. Next, we add the time dimension: statistics evolve, plans are cached, parameters can be skewed, and one captured plan is not automatically representative.
Authoritative references
- Current Neo4j versions — Release/LTS snapshot used for this chapter.
- Understanding query plans — EXPLAIN/PROFILE, plan columns, rows, DB hits, memory and plan reading.
- Operators — Current operator families and operator semantics.
- Operators in detail — Current seek/scan, expand, apply/join, eager, sort and aggregation operators.
- Statistics and execution plans — Statistics collection, selectivity, replanning thresholds and manual preparation.
- Query caches — Per-database query caches, cache sizing and current query-size behavior.
- Cypher runtimes — Slotted, pipelined and parallel runtime boundaries.
- Operator summary — Leaf, eager, updating and other current operator families.
- Advanced query tuning — Current examples of plan/operator changes and index-backed optimizations.