Chapter 10 · Query Planning and Profiling: Cardinality, Operators, Index Selection, and Plan Stability

Parameter Values, Statistics, Plan Cache Behavior, and Testing Representative Workload Shapes

Test AtlasMart plans across parameter skew, statistics changes and cache state instead of assuming one captured plan represents production.

Advanced145–180 minutesPlan-cache/statistics regression labNeo4j 2026.07.1 Community · Cypher 25Last reviewed: September 2026

Learning outcomes

AtlasMart deploys one parameterized query, not one query per customer. The hot customer and a normal customer therefore exercise the same query contract with very different graph shapes. Meanwhile statistics and cached plans evolve as the database changes. This lesson turns “plan stability” into a testable operating property.

01

Explain what statistics the planner uses and why estimates can become stale or imprecise.

02

Describe current automatic replanning behavior without treating default thresholds as tuning commandments.

03

Understand per-database query caches and why parameters improve both safety and cache reuse.

04

Test sparse, typical and dense parameter populations under declared warmup/cache conditions.

05

Use forced replanning/manual statistics procedures only as controlled diagnostics, not routine production folklore.

Chapter 10 baseline · reviewed 9 September 2026

The mandatory lab continues Neo4j Community 2026.07.1, database neo4j, explicit CYPHER 25 for version-sensitive examples, authentication enabled, no mandatory APOC/GDS plugin, and the AtlasMart identifiers/model established in Chapters 01–09. Neo4j 5.26.30 remains the LTS comparison line. Community uses the slotted runtime by default. Enterprise uses the pipelined runtime by default; Aura uses pipelined, and the parallel runtime is an Enterprise capability for supported read workloads. Therefore plan columns such as pipeline timing and page-cache hits/misses are not universal Community output.

Evidence and measurement note

This generation environment does not run Neo4j or Docker. Commands were checked against current official documentation but were not executed here. The lessons never invent DB-hit counts, memory figures, operator timings or latency improvements. Instead, they define deterministic fixture invariants and tell you exactly which values to record from your own EXPLAIN/PROFILE runs. All disposable Chapter 10 entities use labTag='ch10'; disposable schema objects use ch10_*.

1. The planner sees statistics, not your business story

Current Neo4j statistics include label counts, relationship-type counts, relationship counts between labels/types, and index selectivity. Index sampling is maintained in the background. The cost planner combines those statistics with a selectivity model; it does not receive runtime feedback from the specific execution represented by Rows.

Cypher · capture sparse, typical and dense degrees
CYPHER 25UNWIND ['CH10-C-0001','CH10-C-0050','CH10-C-0099'] AS customerIdMATCH (c:Customer {customerId:customerId})-[:PLACED]->(o:Order)RETURN customerId,count(o) AS orderDegreeORDER BY orderDegree DESC;

2. Cached plan does not mean immutable plan

Query caches are initialized per database by default. Current documentation exposes a default per-database cache capacity of 1000 entries when cache sharing is disabled. Since Neo4j 2026.01, query text larger than 128 KiB is skipped by the query cache by default unless CYPHER cache=force is used; using parameters rather than embedding data in query text remains the recommended pattern.

Mechanism Current documented behavior Operational consequence
Query cache Per database by default Identical/normalized query shapes can avoid repeat planning work
Parameters Values supplied separately from query text Safer input handling and better plan/cache reuse than string-built literals
Statistics divergence Plans become stale when relevant statistics change enough Replanning can occur later rather than on every write
Schema change Can invalidate/replan affected queries Index/constraint migrations can change plan shape
Large query text 128 KiB cache eligibility limit by default since 2026.01 Generated literal-heavy statements may bypass cache

3. Replanning has thresholds and latency cost

In the current Operations Manual, dbms.cypher.statistics_divergence_threshold defaults to 0.75 and dbms.cypher.min_replan_interval defaults to 10s. The threshold decays over time so moderately changing graphs are eventually reconsidered. These are defaults to understand, not values to copy into a tuning recipe.

Self-managed diagnostic only · inspect relevant settings
CYPHER 25SHOW SETTINGSYIELD name,value,dynamic,descriptionWHERE name IN ['dbms.cypher.statistics_divergence_threshold',               'dbms.cypher.min_replan_interval',               'server.memory.query_cache.per_db_cache_num_entries']RETURN name,value,dynamic,description;

4. Force replanning only when the experiment requires it

CYPHER replan=force EXPLAIN ... can rebuild a plan without executing the query. db.clearQueryCaches() clears caches but does not recalculate statistics. db.prepareForReplanning() resamples indexes, waits, then clears query caches so subsequent planning uses refreshed statistics. These are administrative diagnostic tools with workload impact; Aura and permission boundaries differ.

Cypher · controlled plan rebuild without query execution
CYPHER 25 replan=forceEXPLAINMATCH (c:Customer {customerId:$customerId})-[:PLACED]->(o:Order)RETURN count(o) AS orders;

5. Representative workload matrix beats one golden parameter

Case Parameter example Record
Dense degree CH10-C-0001 Plan, actual rows, DB hits, memory, result rows, latency distribution
Typical degree CH10-C-0050 Same metrics with identical query text
Selective product filter CAM + high minPrice Index operator and candidate rows
Broad product filter CAM + low minPrice Candidate rows, expansion work and final result
Warm vs declared cold-ish start Same parameters Do not compare without documenting cache/state protocol
Suggested benchmark record · fill with your observed values
case,run,customerId,category,minPrice,planFingerprint,rows,dbHits,peakMemoryBytes,latencyMs,resultRowshot,1,CH10-C-0001,CAM,80,,,,,,hot,2,CH10-C-0001,CAM,80,,,,,,typical,1,CH10-C-0050,CAM,80,,,,,,typical,2,CH10-C-0050,CAM,80,,,,,,

6. Deliberately wrong: treat one cached plan as universal truth

One plan can be a reasonable compromise for a parameterized workload while individual parameter populations still differ greatly in execution cost. Conversely, repeatedly forcing replans can create planning overhead without fixing graph-shape skew. The repair is a regression matrix that includes representative degree/selectivity classes and records plan plus runtime evidence.

Check your understanding

  1. Are query caches global across every database by default?
  2. Does db.clearQueryCaches recalculate statistics?
  3. Why are parameters important beyond injection safety?
  4. Should you routinely force replan on every request?
  5. What makes a workload representative?
Review the answers

1. No. Current default behavior initializes query caches per database.

2. No. It clears query caches; db.prepareForReplanning performs index resampling and then clears caches.

3. They keep data values separate from reusable query text and support query-cache reuse.

4. No. Replanning has cost; use it for controlled diagnostics or specific operational needs.

5. It covers the actual distributions that change work: sparse/dense degrees, selective/broad predicates, result sizes, concurrency and declared cache/warmup states.

Summary and next step

Plan stability is not “the operator tree never changes.” It is the ability to detect when data/statistics/schema/parameters change work enough to threaten the latency or resource contract. The final lesson applies that discipline to one slow traversal.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.