Chapter 09 · Constraints and Indexes: Range, Text, Point, Token, Full-Text, and Schema Enforcement
Composite Index Order, Selectivity, Labels/Types, Relationship Indexes, and Index Population States
Reason about composite and relationship indexes using cardinality, order and lifecycle evidence.
Learning outcomes
Single-property indexes are not enough for every AtlasMart read. Catalog APIs commonly filter a category and a price range together, while order analytics may filter relationship properties. This lesson treats composite schema, property order, selectivity and build state as observable planner inputs rather than folklore.
Explain when a composite range index is eligible and why property order matters.
Measure selectivity/cardinality instead of assuming a field is selective.
Use relationship indexes as first-class access paths.
Interpret POPULATING, populationPercent, ONLINE and owningConstraint evidence.
Detect overlapping structures without deleting a needed access path blindly.
The mandatory lab continues Neo4j Community
2026.07.1, database neo4j, explicit
CYPHER 25 for version-sensitive examples,
authentication enabled, no mandatory APOC/GDS plugin, and the
AtlasMart identifiers/model established in Chapters 01–08.
Neo4j 5.26.30 remains the LTS comparison line.
Community currently supports node and relationship
property-uniqueness constraints. Property-existence,
property-type and key constraints—and Cypher 25 graph
types—are Enterprise-only; Enterprise examples in this chapter
are optional and are never presented as Community output.
This generation environment does not run Neo4j or Docker.
Commands were checked against current official documentation
but were not executed here. Expected plans, counts, index
states and scores are therefore described by invariant rather
than fabricated as captured output. All disposable data uses
labTag='ch09'; all disposable schema objects are
named ch09_*. Never drop an index/constraint
merely because its name looks similar to a lab object—verify
SHOW INDEXES/SHOW CONSTRAINTS first.
1. Composite indexes are one schema, not several independent indexes
A composite range index on
(categoryCode, price) stores that property
combination. Current planner guidance requires the query to
constrain all properties of a composite index for it to be
usable, and the definition order affects how different predicate
shapes can be solved. A single-property
categoryCode index may still be necessary if
category-only queries are part of the workload.
CYPHER 25CREATE RANGE INDEX ch09_product_category_price IF NOT EXISTSFOR (p:Product) ON (p.categoryCode,p.price);SHOW INDEXES YIELD name,state,populationPercent,type,propertiesWHERE name='ch09_product_category_price'RETURN name,state,populationPercent,type,properties;CALL db.awaitIndex('ch09_product_category_price',300);
CYPHER 25EXPLAIN MATCH (p:Product)WHERE p.categoryCode=$category AND p.price >= $minPriceRETURN p.productId,p.price;EXPLAIN MATCH (p:Product)WHERE p.price >= $minPriceRETURN p.productId,p.price;
2. Selectivity is data, not a naming convention
A property is selective when a predicate narrows the candidate set substantially. Category codes with only two values are less selective than unique product IDs. The correct unit of reasoning is the actual distribution and query parameter population.
CYPHER 25MATCH (p:Product) WHERE p.labTag='ch09'WITH count(p) AS totalMATCH (p:Product) WHERE p.labTag='ch09'RETURN total,p.categoryCode AS category,count(*) AS rows, toFloat(count(*))/total AS fractionORDER BY rows DESC;
| Signal | Interpretation | Do not conclude |
|---|---|---|
| Low fraction for a predicate | Potentially selective starting point | That the same field is selective for every value |
| High degree from the anchor | Traversal can fan out after a good seek | That index choice alone fixes graph-density cost |
| Small test fixture | Useful for correctness and plan shape | Production latency or cache behavior |
3. Relationship indexes and labels/types are distinct planner dimensions
CYPHER 25CREATE RANGE INDEX ch09_contains_unitprice_range IF NOT EXISTSFOR ()-[r:CONTAINS]-() ON (r.unitPrice);CALL db.awaitIndex('ch09_contains_unitprice_range',300);EXPLAIN MATCH ()-[r:CONTAINS]->(p:Product)WHERE r.unitPrice >= 100RETURN r.lineId,p.productId,r.unitPrice;
The relationship type CONTAINS can be found through
the relationship token lookup index; the
unitPrice predicate can be solved by the
relationship range index. These are complementary access
structures.
4. POPULATING is a real operational state
Immediately after creation an index can be
POPULATING; populationPercent reports
progress and the index cannot serve queries until
ONLINE. On a tiny lab graph the transition may be
too fast to observe. That does not make the state irrelevant on
a large production store.
CYPHER 25SHOW INDEXESYIELD name,state,populationPercent,type,entityType,labelsOrTypes,properties,owningConstraintWHERE name STARTS WITH 'ch09_' OR owningConstraint IS NOT NULLRETURN name,state,populationPercent,type,entityType,labelsOrTypes,properties,owningConstraintORDER BY state,name;
5. Deliberately wrong: benchmark immediately after CREATE INDEX
If the index is still populating, the planner cannot use it. If
it just became online, caches and statistics may also differ
from steady state. The repair is to wait for
ONLINE, record the graph/index state, warm
according to a declared protocol, and compare plan/runtime
distributions—not one lucky request.
CYPHER 25CALL db.awaitIndexes(300);SHOW INDEXES YIELD name,state,populationPercentWHERE name STARTS WITH 'ch09_'RETURN name,state,populationPercent ORDER BY name;
Check your understanding
- Must a composite range query constrain every indexed property?
- Why does property order matter?
- Can a tiny lab reliably expose POPULATING?
- What does owningConstraint mean?
- Does a selective anchor guarantee a cheap query?
Review the answers
1. Under current documented behavior, yes, for the composite index to be used.
2. The ordered composite schema changes which predicate combinations/orderings the planner can solve efficiently.
3. Not necessarily; the build may finish before SHOW INDEXES is run.
4. The index is backing an index-backed constraint rather than being an independently managed access-path index.
5. No. Subsequent expansions, result cardinality, sorting and aggregation can still dominate cost.
Summary and next step
Composite design is workload- and distribution-specific, and index state is part of correctness when measuring. Next, full-text search introduces a different semantic-index lifecycle and explicit query procedure.
Authoritative references
- Current Neo4j versions — Release/LTS snapshot used for this chapter.
- Constraints — Current constraint types and edition boundaries.
- Create constraints — Current uniqueness, existence, type and key syntax and backing-index behavior.
- Search-performance indexes — Range, text, point and token lookup index semantics.
- Show indexes — Index lifecycle, state, population and usage evidence.
- Full-text indexes — Full-text schema, analyzers, query procedures and eventual-consistency behavior.
- Built-in index procedures — db.awaitIndex(es) and full-text refresh/analyzer procedures.
- Index syntax — Current composite and relationship index syntax plus SHOW INDEXES filters.
- Index impact on performance — Composite eligibility, property order and usage evidence.