Chapter 23 · Graph Data Science Foundations: Projections, Graph Catalog, Memory, and Execution Modes
Node/Relationship Properties, Orientation, Filtering, Graph Schema, and Projection Verification
Project only the labels, relationship types, properties, orientation, and filtered population an algorithm actually needs, then verify the graph schema and counts before trusting analytical output.
Learning outcomes
Explain how node/relationship properties become the analytical feature/weight schema.
Choose NATURAL, REVERSE, or UNDIRECTED orientation from algorithm semantics rather than convenience.
Use Cypher filtering to create a question-specific projection without deleting source data.
Verify graph schema, counts, degree distribution, and property availability before an algorithm.
Diagnose projection mismatches such as missing numeric properties, accidental direction changes, and parallel-edge assumptions.
1. The projection is the model the algorithm actually sees
AtlasMart has accumulated enough connected customer/product behavior that analysts want graph algorithms, but running “an algorithm” is not one operation. The analytical graph must first be selected, copied into a GDS representation, sized, verified, executed, and removed or refreshed deliberately. AtlasMart’s source graph contains customer segments, product margins, VIEWED and PURCHASED relationships. If the projection drops weights or reverses an edge, a perfectly implemented algorithm can still answer the wrong business question.
The safety rule for this chapter is therefore: treat projection definition, memory allocation, algorithm mode, and lifecycle as part of the analytical result. A score without those inputs is not reproducible evidence.
2. Properties: persisted source values become analytical inputs only if projected
Current Neo4j Database is 2026.07.1; current GDS
is 2026.07.0 for the Neo4j 2026.07 line.
Mandatory work uses self-managed
Neo4j Community 2026.07.1 + GDS Community 2026.07.0, database neo4j, user neo4j,
disposable password atlasmart-course-2026,
loopback HTTP 7474/Bolt 7687, and
CYPHER 25 for database-side fixture work. Neo4j
2026.x supports Java 21/25. GDS is the only plugin required in
this chapter; APOC is not required.
GDS Community includes the algorithm library needed for this course. Its execution concurrency is capped at 4 CPU cores and its model catalog is capped at 3 models. GDS Enterprise removes the CPU-core cap and adds capabilities such as graph backup/restore, Arrow import/export, cluster write support, capacity/load monitoring, and extended model-catalog persistence/sharing. Those Enterprise capabilities are discussed only as edition boundaries; no mandatory lab depends on them.
Two current projection styles matter.
Native projection uses the
gds.graph.project(...) procedure and remains
supported in 2026.07, but current GDS documentation says
native projection will be deprecated in a future release and
increasingly uses Cypher projection as the norm.
Current Cypher projection calls the
gds.graph.project(...) aggregation function from
a Cypher query. The older
gds.graph.project.cypher(...) procedure is
already deprecated. This chapter teaches native projection
because it makes schema/orientation configuration explicit,
then shows the current Cypher form learners should prefer for
new flexible projections.
| Assumption | Value / boundary |
|---|---|
| Server | Neo4j Community 2026.07.1; Java 21/25 supported. Disposable container atlasmart-gds. |
| GDS | GDS Community 2026.07.0. Verify with gds.version(); do not continue if compatibility differs. |
| Database/security | neo4j database; local disposable neo4j user; loopback transport only. Production credentials/TLS/RBAC differ. |
| Fixture | 6 Customer + 6 Product nodes; 14 VIEWED + 4 PURCHASED relationships; numeric weight/customerValue/margin. |
| Memory/concurrency | Learner records estimates/observations. Examples use concurrency=2; Community maximum is 4, but 2 is a lab choice, not a recommendation. |
| Plugins | GDS only. APOC is not required. |
| Runtime claims | Artifact generation does not execute Neo4j/GDS. Expected deterministic graph counts are fixture-derived; memory/timing/algorithm scores must be measured by the learner. |
GDS does not automatically expose every Neo4j property. Project only values the algorithm needs. Relationship properties used as weights must be supported numeric values; node properties used by downstream algorithms/features likewise need compatible types. A missing projected property is not “null from the database”—it is absent from the analytical schema.
| Source element | Projected role in this lab | Why it matters |
|---|---|---|
| Customer.customerValue | Node property | Demonstrates selective node-feature projection; not used by PageRank itself. |
| Product.margin | Node property | Demonstrates heterogeneous labels with different properties. |
| VIEWED.weight | Relationship property | Optional PageRank weight signal. |
| PURCHASED.weight | Relationship property | Higher-weight behavior path. |
| segment/name | Not projected in native graph | Still available in Neo4j through gds.util.asNode for inspection, but not an in-memory analytical property. |
3. Orientation changes graph semantics and memory
| Orientation | Analytical edge | Use when | Risk |
|---|---|---|---|
| NATURAL | source → target as stored | Direction is meaningful for the algorithm | May produce zero outgoing/incoming degree for one label in bipartite graphs |
| REVERSE | target → source | Question explicitly asks influence/flow opposite source relationship | Easy to reverse accidentally and reinterpret scores |
| UNDIRECTED | Connectivity available both ways | Algorithm/question is symmetric or needs reciprocal traversal | Can represent relationships twice and increase memory/work |
Do not choose UNDIRECTED merely because it “finds more.” Direction is part of the hypothesis. In a Customer→Product interaction graph, PageRank with NATURAL orientation largely transfers importance toward products; reversing it asks a different question.
4. Create and verify the full bounded schema
CYPHER 25
// Chapter 23 fixture: bounded and safe to delete by labTag.
CREATE CONSTRAINT ch23_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch23_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
UNWIND [
{id:'C-2301',segment:'LOYAL',value:92.0},
{id:'C-2302',segment:'LOYAL',value:78.0},
{id:'C-2303',segment:'NEW',value:35.0},
{id:'C-2304',segment:'NEW',value:28.0},
{id:'C-2305',segment:'B2B',value:96.0},
{id:'C-2306',segment:'B2B',value:84.0}
] AS row
MERGE (c:Customer {customerId:row.id})
SET c.segment=row.segment,c.customerValue=row.value,c.labTag='ch23';
UNWIND [
{id:'P-2301',name:'Trail Camera Pro',margin:0.31},
{id:'P-2302',name:'Action Camera 4K',margin:0.27},
{id:'P-2303',name:'Solar Trail Charger',margin:0.24},
{id:'P-2304',name:'Indoor Security Camera',margin:0.29},
{id:'P-2305',name:'Hydration Vest',margin:0.22},
{id:'P-2306',name:'Wildlife Field Guide',margin:0.35}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name,p.margin=row.margin,p.labTag='ch23';
MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WITH c,p WHERE
(c.customerId='C-2301' AND p.productId IN ['P-2301','P-2303','P-2306']) OR
(c.customerId='C-2302' AND p.productId IN ['P-2301','P-2302']) OR
(c.customerId='C-2303' AND p.productId IN ['P-2302','P-2305']) OR
(c.customerId='C-2304' AND p.productId IN ['P-2304','P-2305']) OR
(c.customerId='C-2305' AND p.productId IN ['P-2301','P-2303','P-2304']) OR
(c.customerId='C-2306' AND p.productId IN ['P-2303','P-2306'])
MERGE (c)-[r:VIEWED]->(p)
SET r.weight = CASE c.segment WHEN 'LOYAL' THEN 2.0 WHEN 'B2B' THEN 1.5 ELSE 1.0 END,
r.labTag='ch23';
MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WHERE (c.customerId='C-2301' AND p.productId='P-2301') OR
(c.customerId='C-2302' AND p.productId='P-2302') OR
(c.customerId='C-2305' AND p.productId='P-2303') OR
(c.customerId='C-2306' AND p.productId='P-2306')
MERGE (c)-[r:PURCHASED]->(p)
SET r.weight=3.0,r.labTag='ch23';
CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;
CALL gds.graph.project(
'atlasmart-ch23',
{
Customer:{properties:['customerValue']},
Product:{properties:['margin']}
},
{
VIEWED:{orientation:'NATURAL',properties:['weight']},
PURCHASED:{orientation:'NATURAL',properties:['weight']}
},
{readConcurrency:2}
)
YIELD graphName,nodeCount,relationshipCount,projectMillis,configuration
RETURN graphName,nodeCount,relationshipCount,projectMillis,configuration;
// Expected deterministic counts: nodeCount=12, relationshipCount=18.
CALL gds.graph.list('atlasmart-ch23')
YIELD graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration
RETURN graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration;
Acceptance criteria are structural: 12 nodes, 18 relationships, Customer/Product labels, VIEWED/PURCHASED relationship types, NATURAL direction, and projected numeric properties. The catalog listing is stronger evidence than an algorithm score because it proves the input contract.
5. Filtering is analytical scope, not source deletion
CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;
CYPHER 25
MATCH (source:Customer {labTag:'ch23'})-[r:VIEWED]->(target:Product {labTag:'ch23'})
WHERE source.segment IN ['LOYAL','B2B'] AND target.margin >= 0.24
RETURN gds.graph.project(
'atlasmart-ch23', source,target,
{
sourceNodeLabels:['Customer'],
targetNodeLabels:['Product'],
sourceNodeProperties:{customerValue:source.customerValue},
targetNodeProperties:{margin:target.margin},
relationshipType:'VIEWED',
relationshipProperties:{weight:r.weight}
},
{readConcurrency:2}
)
YIELD graphName,nodeCount,relationshipCount,projectMillis
RETURN *;
// Counts now prove projection scope, not source-store deletion.
CALL gds.graph.list('atlasmart-ch23')
YIELD graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration
RETURN graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration;
After the filtered projection, the transactional fixture still contains all Chapter 23 nodes and relationships. Only the in-memory analytical population changed. This distinction is crucial for cohort analysis: filters are a reproducible selection predicate, not a mutation of source truth.
6. Parallel relationships and aggregation are modeling choices
Neo4j can store multiple relationships between the same nodes.
GDS native projection preserves parallel relationships by
default. Some algorithms can use them directly; others require
or benefit from aggregation. Native projection supports
aggregation such as SINGLE, COUNT,
SUM, MIN, or
MAX depending on property configuration.
Aggregation changes analytical topology/weights and must be
treated as preprocessing, not a harmless optimization.
GDS projected relationships are identified by analytical source/target connectivity; the Neo4j relationship ID is not implicitly part of the analytical graph. If provenance back to a specific relationship matters, project an explicit relationship-id property.
7. Deliberately wrong: use one undirected projection for every algorithm
Make VIEWED and PURCHASED UNDIRECTED “so every node can reach every other node,” then reuse the projection for centrality, similarity and path algorithms.
Concrete problem: the algorithm now receives a
different graph from the stored customer→product behavior.
Relationship count/memory and degree distributions change, and a
centrality score may measure reciprocal connectivity rather than
directional influence. Repair: version
projection names/configs by analytical intent and assert
schemaWithOrientation before execution.
8. Verification checklist before any algorithm
| Check | Evidence | Failure interpretation |
|---|---|---|
| Identity/scope | nodeCount and relationshipCount | Unexpected counts mean filter/label/type mismatch or duplicate/parallel-edge behavior. |
| Schema | schemaWithOrientation | Missing properties/types/orientation means algorithm input differs from design. |
| Degree distribution | gds.graph.list degreeDistribution | Extreme hubs/isolates may be real or may expose a projection error. |
| Weights | graph property stream / algorithm configuration | Do not assume a weight property is projected just because it exists in Neo4j. |
| Freshness | projection creationTime + source-change window | Projection may be stale after transactional writes. |
Production judgment and bridge
| Decision surface | Evidence before increasing analytical scope |
|---|---|
| Projection scope | Exact labels/types/properties/filter predicates and projected node/relationship counts match the analytical question. |
| Memory | Projection estimate + algorithm estimate + observed Neo4j heap/process headroom leave safe room for transactional work and GC. |
| Concurrency | Measured throughput/tail latency/CPU under representative contention; never assume max concurrency is optimal. |
| Freshness | Document when the projection was created and what source updates occurred afterward; projections are not live materialized views. |
| Persistence | Explicit decision whether results belong only in stream output, only in the projected graph, or persisted into Neo4j. |
| Failure recovery | Graph can be reconstructed from versioned fixture/query/config; catalog loss on DBMS restart is expected unless an Enterprise persistence feature is deliberately used. |
| Edition/cost | Community core limits are accepted or Enterprise/Aura analytics capabilities are justified by workload, operations and licensing. |
The next lesson moves from logical correctness to resource correctness. Once the projection is semantically right, its memory, concurrency, lifetime and coexistence with the transactional workload must be proven.
Check your understanding
- Why can two projections of the same source graph produce different algorithm answers?
- What does UNDIRECTED change?
- Does filtering a Cypher projection delete Neo4j nodes?
- Why inspect schemaWithOrientation?
- When should relationship IDs be projected explicitly?
Review the answers
1. They can select different labels/types/properties, filters, orientation, or aggregation; the algorithm operates on the projected graph, not abstract source intent.
2. It changes analytical direction/connectivity and can increase stored analytical relationships/memory; it is a modeling decision.
3. No. It changes the in-memory projected population only.
4. It proves the actual analytical labels/types/properties and direction rather than trusting the projection command visually.
5. When algorithm outputs must be traced to individual persisted relationships, because Neo4j relationship IDs are not automatically part of GDS relationship identity.
Summary and next step
Node/Relationship Properties, Orientation, Filtering, Graph Schema, and Projection Verification is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Estimate Memory, Control Concurrency, Manage Graph Lifecycles, and Avoid Production Memory Starvation. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- GDS Manual v2026.07 — Current Graph Data Science manual and versioned feature baseline.
- GDS 2026.07 release notes — Current GDS release line; 2026.07.0 is compatible with Neo4j 2026.07.
- Supported Neo4j versions — Compatibility matrix between Neo4j Database and GDS.
- GDS editions and graph catalog — Community/Enterprise boundaries, graph catalog model, all-algorithms availability, concurrency and model-catalog limits.
- Neo4j Server GDS installation — Bundled products-to-plugins installation path and required GDS procedure security configuration.
- GDS on Docker — Container-based installation examples and plugin activation.
- System requirements — Heap/native-memory/CPU guidance and Community maximum concurrency of four.
- Native projection — Current gds.graph.project procedure, label/type/property/orientation configuration and lifecycle.
- Cypher projection — Current gds.graph.project aggregation-function projection from Cypher query context.
- Legacy Cypher projection — deprecated — Deprecated gds.graph.project.cypher procedure and migration boundary.
- Graph creation and schema — Supported projection types, relationship direction, properties, parallel relationships, and algorithm traits.
- Graph catalog operations — Current graph project/list/exists/drop and graph-property operations.
- Listing graphs — Graph metadata, schemaWithOrientation, degree distribution and configuration inspection.
- Memory estimation — Projection and algorithm estimate syntax, requiredMemory and byte-range evidence.
- Algorithm syntax and execution modes — stream/stats/mutate/write/estimate semantics.
- Running algorithms — Operational meaning and tradeoffs of algorithm execution modes.
- PageRank — Production-tier algorithm used only to demonstrate execution modes and memory estimates in this foundations chapter.
- Neo4j current versions — Current Neo4j Database release and 5.26 LTS line.