Chapter 10 · Query Planning and Profiling: Cardinality, Operators, Index Selection, and Plan Stability
Parameter Values, Statistics, Plan Cache Behavior, and Testing Representative Workload Shapes
Test AtlasMart plans across parameter skew, statistics changes and cache state instead of assuming one captured plan represents production.
Learning outcomes
AtlasMart deploys one parameterized query, not one query per customer. The hot customer and a normal customer therefore exercise the same query contract with very different graph shapes. Meanwhile statistics and cached plans evolve as the database changes. This lesson turns “plan stability” into a testable operating property.
Explain what statistics the planner uses and why estimates can become stale or imprecise.
Describe current automatic replanning behavior without treating default thresholds as tuning commandments.
Understand per-database query caches and why parameters improve both safety and cache reuse.
Test sparse, typical and dense parameter populations under declared warmup/cache conditions.
Use forced replanning/manual statistics procedures only as controlled diagnostics, not routine production folklore.
The mandatory lab continues Neo4j Community
2026.07.1, database neo4j, explicit
CYPHER 25 for version-sensitive examples,
authentication enabled, no mandatory APOC/GDS plugin, and the
AtlasMart identifiers/model established in Chapters 01–09.
Neo4j 5.26.30 remains the LTS comparison line.
Community uses the slotted runtime by default. Enterprise uses
the pipelined runtime by default; Aura uses pipelined, and the
parallel runtime is an Enterprise capability for supported
read workloads. Therefore plan columns such as pipeline timing
and page-cache hits/misses are not universal Community output.
This generation environment does not run Neo4j or Docker.
Commands were checked against current official documentation
but were not executed here. The lessons never invent DB-hit
counts, memory figures, operator timings or latency
improvements. Instead, they define deterministic fixture
invariants and tell you exactly which values to record from
your own EXPLAIN/PROFILE runs. All
disposable Chapter 10 entities use labTag='ch10';
disposable schema objects use ch10_*.
1. The planner sees statistics, not your business story
Current Neo4j statistics include label counts, relationship-type
counts, relationship counts between labels/types, and index
selectivity. Index sampling is maintained in the background. The
cost planner combines those statistics with a selectivity model;
it does not receive runtime feedback from the specific execution
represented by Rows.
CYPHER 25UNWIND ['CH10-C-0001','CH10-C-0050','CH10-C-0099'] AS customerIdMATCH (c:Customer {customerId:customerId})-[:PLACED]->(o:Order)RETURN customerId,count(o) AS orderDegreeORDER BY orderDegree DESC;
2. Cached plan does not mean immutable plan
Query caches are initialized per database by default. Current
documentation exposes a default per-database cache capacity of
1000 entries when cache sharing is disabled. Since Neo4j
2026.01, query text larger than 128 KiB is skipped by the query
cache by default unless CYPHER cache=force is used;
using parameters rather than embedding data in query text
remains the recommended pattern.
| Mechanism | Current documented behavior | Operational consequence |
|---|---|---|
| Query cache | Per database by default | Identical/normalized query shapes can avoid repeat planning work |
| Parameters | Values supplied separately from query text | Safer input handling and better plan/cache reuse than string-built literals |
| Statistics divergence | Plans become stale when relevant statistics change enough | Replanning can occur later rather than on every write |
| Schema change | Can invalidate/replan affected queries | Index/constraint migrations can change plan shape |
| Large query text | 128 KiB cache eligibility limit by default since 2026.01 | Generated literal-heavy statements may bypass cache |
3. Replanning has thresholds and latency cost
In the current Operations Manual,
dbms.cypher.statistics_divergence_threshold
defaults to 0.75 and
dbms.cypher.min_replan_interval defaults to
10s. The threshold decays over time so moderately
changing graphs are eventually reconsidered. These are defaults
to understand, not values to copy into a tuning recipe.
CYPHER 25SHOW SETTINGSYIELD name,value,dynamic,descriptionWHERE name IN ['dbms.cypher.statistics_divergence_threshold', 'dbms.cypher.min_replan_interval', 'server.memory.query_cache.per_db_cache_num_entries']RETURN name,value,dynamic,description;
4. Force replanning only when the experiment requires it
CYPHER replan=force EXPLAIN ... can rebuild a plan
without executing the query.
db.clearQueryCaches() clears caches but does not
recalculate statistics.
db.prepareForReplanning() resamples indexes, waits,
then clears query caches so subsequent planning uses refreshed
statistics. These are administrative diagnostic tools with
workload impact; Aura and permission boundaries differ.
CYPHER 25 replan=forceEXPLAINMATCH (c:Customer {customerId:$customerId})-[:PLACED]->(o:Order)RETURN count(o) AS orders;
5. Representative workload matrix beats one golden parameter
| Case | Parameter example | Record |
|---|---|---|
| Dense degree | CH10-C-0001 | Plan, actual rows, DB hits, memory, result rows, latency distribution |
| Typical degree | CH10-C-0050 | Same metrics with identical query text |
| Selective product filter | CAM + high minPrice | Index operator and candidate rows |
| Broad product filter | CAM + low minPrice | Candidate rows, expansion work and final result |
| Warm vs declared cold-ish start | Same parameters | Do not compare without documenting cache/state protocol |
case,run,customerId,category,minPrice,planFingerprint,rows,dbHits,peakMemoryBytes,latencyMs,resultRowshot,1,CH10-C-0001,CAM,80,,,,,,hot,2,CH10-C-0001,CAM,80,,,,,,typical,1,CH10-C-0050,CAM,80,,,,,,typical,2,CH10-C-0050,CAM,80,,,,,,
6. Deliberately wrong: treat one cached plan as universal truth
One plan can be a reasonable compromise for a parameterized workload while individual parameter populations still differ greatly in execution cost. Conversely, repeatedly forcing replans can create planning overhead without fixing graph-shape skew. The repair is a regression matrix that includes representative degree/selectivity classes and records plan plus runtime evidence.
Check your understanding
- Are query caches global across every database by default?
- Does db.clearQueryCaches recalculate statistics?
- Why are parameters important beyond injection safety?
- Should you routinely force replan on every request?
- What makes a workload representative?
Review the answers
1. No. Current default behavior initializes query caches per database.
2. No. It clears query caches; db.prepareForReplanning performs index resampling and then clears caches.
3. They keep data values separate from reusable query text and support query-cache reuse.
4. No. Replanning has cost; use it for controlled diagnostics or specific operational needs.
5. It covers the actual distributions that change work: sparse/dense degrees, selective/broad predicates, result sizes, concurrency and declared cache/warmup states.
Summary and next step
Plan stability is not “the operator tree never changes.” It is the ability to detect when data/statistics/schema/parameters change work enough to threaten the latency or resource contract. The final lesson applies that discipline to one slow traversal.
Authoritative references
- Current Neo4j versions — Release/LTS snapshot used for this chapter.
- Understanding query plans — EXPLAIN/PROFILE, plan columns, rows, DB hits, memory and plan reading.
- Operators — Current operator families and operator semantics.
- Operators in detail — Current seek/scan, expand, apply/join, eager, sort and aggregation operators.
- Statistics and execution plans — Statistics collection, selectivity, replanning thresholds and manual preparation.
- Query caches — Per-database query caches, cache sizing and current query-size behavior.
- Cypher runtimes — Slotted, pipelined and parallel runtime boundaries.
- Built-in procedures — db.clearQueryCaches and db.prepareForReplanning behavior.
- Configuration settings — Current query-cache and replanning settings.