Chapter 23 · Graph Data Science Foundations: Projections, Graph Catalog, Memory, and Execution Modes
GDS Architecture, In-Memory Graph Catalog, Native/Cypher Projections, and Projection Scope
Build the correct mental model of GDS as a separate in-memory analytical graph catalog, then create bounded native and current Cypher projections without confusing them with the transactional store.
Learning outcomes
Explain why GDS uses a separate in-memory graph catalog instead of treating the transactional store as the algorithm graph.
Distinguish native projection, current Cypher projection, and the deprecated legacy Cypher-projection procedure.
Install/verify a compatible GDS Community plugin and build a bounded AtlasMart source fixture.
Estimate, project, list, verify, and drop a named analytical graph.
Make a production decision about scope, freshness, memory isolation, edition limits, and reconstruction.
1. Why AtlasMart needs an analytical graph boundary
AtlasMart has accumulated enough connected customer/product behavior that analysts want graph algorithms, but running “an algorithm” is not one operation. The analytical graph must first be selected, copied into a GDS representation, sized, verified, executed, and removed or refreshed deliberately. The first request—“find influential products”—looks harmless, yet projecting every customer, order, product, and relationship by default could duplicate far more data into heap than the question needs.
The safety rule for this chapter is therefore: treat projection definition, memory allocation, algorithm mode, and lifecycle as part of the analytical result. A score without those inputs is not reproducible evidence.
2. Transactional Neo4j vs the GDS graph catalog
| Term | Mechanism-first meaning |
|---|---|
| transactional store | The persisted Neo4j graph used by ordinary Cypher transactions. GDS does not automatically make algorithm computations operate directly on this store. |
| projection | A selected copy/transformation of nodes, relationships and numeric properties loaded into the GDS in-memory graph catalog. |
| graph catalog | Process-local catalog of named in-memory analytical graphs. A projection survives until explicitly dropped or until its source database/DBMS stops. |
| orientation | How a relationship is represented analytically: NATURAL keeps stored direction; REVERSE flips it; UNDIRECTED creates analytical connectivity in both directions and can increase memory. |
| graph schema | Projected labels, relationship types, properties and orientations—not the entire transactional database schema. |
| concurrency | Number of worker threads GDS may use for an operation. It is a resource budget, not a universal performance knob. |
| estimate | A dry-run-style memory calculation that reports requiredMemory/byte ranges without running the actual algorithm or graph projection. |
| stream | Algorithm mode returning per-entity results to the caller without modifying the projection or persistent store. |
| stats | Algorithm mode returning summary statistics only, without modifying the projection or store. |
| mutate | Algorithm mode writing computed results into the in-memory projected graph only. |
| write | Algorithm mode persisting computed results back into Neo4j database properties/relationships as defined by that algorithm. |
Neo4j’s transactional graph is designed for durable ACID reads/writes. GDS algorithms use a compact analytical representation loaded into heap. Projection is therefore a data contract: it decides which source entities exist analytically, which relationships count, which direction the algorithm sees, and which numeric values become weights/features.
| Layer | Durability | Freshness | Primary purpose |
|---|---|---|---|
| Neo4j store | Durable until modified/deleted | Current committed database state | Transactional graph queries and writes |
| GDS named graph | In-memory catalog; disappears when dropped or source DB/DBMS stops | Snapshot at projection/mutation time | High-throughput graph analytics |
| Algorithm stream result | Client result only | Computation-time view | Exploration/inspection |
| Algorithm write result | Persisted back to Neo4j | Durable after committed write | Operationalize selected analytical outputs |
3. Current projection APIs: choose deliberately
Two current projection styles matter.
Native projection uses the
gds.graph.project(...) procedure and remains
supported in 2026.07, but current GDS documentation says
native projection will be deprecated in a future release and
increasingly uses Cypher projection as the norm.
Current Cypher projection calls the
gds.graph.project(...) aggregation function from
a Cypher query. The older
gds.graph.project.cypher(...) procedure is
already deprecated. This chapter teaches native projection
because it makes schema/orientation configuration explicit,
then shows the current Cypher form learners should prefer for
new flexible projections.
| Projection path | Current status | Best fit | Important boundary |
|---|---|---|---|
| Native gds.graph.project procedure | Supported in 2026.07; docs say future deprecation is planned | Simple label/type/property projections | Configuration-driven; less expressive filtering than Cypher |
| Cypher + gds.graph.project aggregation function | Current flexible Cypher projection | Filtering, transformed data, multi-source/query-derived projection | Projection query itself becomes part of reproducibility and performance |
| gds.graph.project.cypher procedure | Deprecated legacy API | Migration-reading only | Do not introduce into new course code |
4. Install and prove the compatibility boundary
Current Neo4j Database is 2026.07.1; current GDS
is 2026.07.0 for the Neo4j 2026.07 line.
Mandatory work uses self-managed
Neo4j Community 2026.07.1 + GDS Community 2026.07.0, database neo4j, user neo4j,
disposable password atlasmart-course-2026,
loopback HTTP 7474/Bolt 7687, and
CYPHER 25 for database-side fixture work. Neo4j
2026.x supports Java 21/25. GDS is the only plugin required in
this chapter; APOC is not required.
GDS Community includes the algorithm library needed for this course. Its execution concurrency is capped at 4 CPU cores and its model catalog is capped at 3 models. GDS Enterprise removes the CPU-core cap and adds capabilities such as graph backup/restore, Arrow import/export, cluster write support, capacity/load monitoring, and extended model-catalog persistence/sharing. Those Enterprise capabilities are discussed only as edition boundaries; no mandatory lab depends on them.
Two current projection styles matter.
Native projection uses the
gds.graph.project(...) procedure and remains
supported in 2026.07, but current GDS documentation says
native projection will be deprecated in a future release and
increasingly uses Cypher projection as the norm.
Current Cypher projection calls the
gds.graph.project(...) aggregation function from
a Cypher query. The older
gds.graph.project.cypher(...) procedure is
already deprecated. This chapter teaches native projection
because it makes schema/orientation configuration explicit,
then shows the current Cypher form learners should prefer for
new flexible projections.
| Assumption | Value / boundary |
|---|---|
| Server | Neo4j Community 2026.07.1; Java 21/25 supported. Disposable container atlasmart-gds. |
| GDS | GDS Community 2026.07.0. Verify with gds.version(); do not continue if compatibility differs. |
| Database/security | neo4j database; local disposable neo4j user; loopback transport only. Production credentials/TLS/RBAC differ. |
| Fixture | 6 Customer + 6 Product nodes; 14 VIEWED + 4 PURCHASED relationships; numeric weight/customerValue/margin. |
| Memory/concurrency | Learner records estimates/observations. Examples use concurrency=2; Community maximum is 4, but 2 is a lab choice, not a recommendation. |
| Plugins | GDS only. APOC is not required. |
| Runtime claims | Artifact generation does not execute Neo4j/GDS. Expected deterministic graph counts are fixture-derived; memory/timing/algorithm scores must be measured by the learner. |
# PowerShell-oriented disposable local lab.
# Recreate only this course container; named data/log volumes stay separate from earlier labs.
docker rm -f atlasmart-gds 2>$null
docker run -d `
--name atlasmart-gds `
-p 127.0.0.1:7474:7474 `
-p 127.0.0.1:7687:7687 `
-v atlasmart-gds-data:/data `
-v atlasmart-gds-logs:/logs `
-e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
-e 'NEO4J_PLUGINS=["graph-data-science"]' `
-e NEO4J_dbms_security_procedures_unrestricted=gds.* `
-e NEO4J_dbms_security_procedures_allowlist=gds.* `
neo4j:2026.07.1
# Wait for the database to become ready, then verify the plugin from cypher-shell/Browser.
CYPHER 25
RETURN gds.version() AS gdsVersion;
// Expected for this chapter baseline: 2026.07.0.
SHOW PROCEDURES YIELD name
WHERE name STARTS WITH 'gds.'
RETURN count(*) AS gdsProcedureCount;
If gds.version() is not the compatible 2026.07
line, do not “try the commands anyway.” Resolve server/GDS
compatibility first. GDS uses Neo4j internals; compatibility
is an operational dependency.
5. Build the smallest useful AtlasMart source graph
CYPHER 25
// Chapter 23 fixture: bounded and safe to delete by labTag.
CREATE CONSTRAINT ch23_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch23_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
UNWIND [
{id:'C-2301',segment:'LOYAL',value:92.0},
{id:'C-2302',segment:'LOYAL',value:78.0},
{id:'C-2303',segment:'NEW',value:35.0},
{id:'C-2304',segment:'NEW',value:28.0},
{id:'C-2305',segment:'B2B',value:96.0},
{id:'C-2306',segment:'B2B',value:84.0}
] AS row
MERGE (c:Customer {customerId:row.id})
SET c.segment=row.segment,c.customerValue=row.value,c.labTag='ch23';
UNWIND [
{id:'P-2301',name:'Trail Camera Pro',margin:0.31},
{id:'P-2302',name:'Action Camera 4K',margin:0.27},
{id:'P-2303',name:'Solar Trail Charger',margin:0.24},
{id:'P-2304',name:'Indoor Security Camera',margin:0.29},
{id:'P-2305',name:'Hydration Vest',margin:0.22},
{id:'P-2306',name:'Wildlife Field Guide',margin:0.35}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name,p.margin=row.margin,p.labTag='ch23';
MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WITH c,p WHERE
(c.customerId='C-2301' AND p.productId IN ['P-2301','P-2303','P-2306']) OR
(c.customerId='C-2302' AND p.productId IN ['P-2301','P-2302']) OR
(c.customerId='C-2303' AND p.productId IN ['P-2302','P-2305']) OR
(c.customerId='C-2304' AND p.productId IN ['P-2304','P-2305']) OR
(c.customerId='C-2305' AND p.productId IN ['P-2301','P-2303','P-2304']) OR
(c.customerId='C-2306' AND p.productId IN ['P-2303','P-2306'])
MERGE (c)-[r:VIEWED]->(p)
SET r.weight = CASE c.segment WHEN 'LOYAL' THEN 2.0 WHEN 'B2B' THEN 1.5 ELSE 1.0 END,
r.labTag='ch23';
MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WHERE (c.customerId='C-2301' AND p.productId='P-2301') OR
(c.customerId='C-2302' AND p.productId='P-2302') OR
(c.customerId='C-2305' AND p.productId='P-2303') OR
(c.customerId='C-2306' AND p.productId='P-2306')
MERGE (c)-[r:PURCHASED]->(p)
SET r.weight=3.0,r.labTag='ch23';
CYPHER 25
MATCH (n {labTag:'ch23'})
RETURN labels(n) AS labels,count(*) AS nodes ORDER BY labels;
MATCH ()-[r]->() WHERE r.labTag='ch23'
RETURN type(r) AS type,count(*) AS relationships,sum(r.weight) AS totalWeight ORDER BY type;
// Deterministic invariants: 6 Customer + 6 Product nodes; 14 VIEWED; 4 PURCHASED.
The fixture is intentionally tiny because this chapter teaches lifecycle mechanics, not benchmark bragging. Counts are deterministic; PageRank scores and memory/timing are measured outputs.
6. Estimate before allocating; then project
CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;
CALL gds.graph.project.estimate(
{
Customer:{properties:['customerValue']},
Product:{properties:['margin']}
},
{
VIEWED:{orientation:'NATURAL',properties:['weight']},
PURCHASED:{orientation:'NATURAL',properties:['weight']}
},
{readConcurrency:2}
)
YIELD nodeCount,relationshipCount,requiredMemory,bytesMin,bytesMax
RETURN *;
// The byte estimate is environment/version dependent; record it instead of copying a canned number.
CALL gds.graph.project(
'atlasmart-ch23',
{
Customer:{properties:['customerValue']},
Product:{properties:['margin']}
},
{
VIEWED:{orientation:'NATURAL',properties:['weight']},
PURCHASED:{orientation:'NATURAL',properties:['weight']}
},
{readConcurrency:2}
)
YIELD graphName,nodeCount,relationshipCount,projectMillis,configuration
RETURN graphName,nodeCount,relationshipCount,projectMillis,configuration;
// Expected deterministic counts: nodeCount=12, relationshipCount=18.
CALL gds.graph.list('atlasmart-ch23')
YIELD graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration
RETURN graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration;
The estimate is not a promise that future workload fits: algorithm working memory, concurrent transactions, JVM overhead, page cache, result materialization, and other projections also compete for resources. It is evidence for a budget.
7. The current Cypher projection equivalent
CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;
CYPHER 25
MATCH (source {labTag:'ch23'})-[r:VIEWED|PURCHASED]->(target {labTag:'ch23'})
RETURN gds.graph.project(
'atlasmart-ch23',
source,
target,
{
sourceNodeLabels:labels(source),
targetNodeLabels:labels(target),
sourceNodeProperties:CASE WHEN source:Customer THEN {customerValue:source.customerValue} ELSE {} END,
targetNodeProperties:CASE WHEN target:Product THEN {margin:target.margin} ELSE {} END,
relationshipType:type(r),
relationshipProperties:{weight:r.weight}
},
{readConcurrency:2}
)
YIELD graphName,nodeCount,relationshipCount,projectMillis
RETURN *;
// Current Cypher projection is the gds.graph.project aggregation function.
// Do not replace this with deprecated gds.graph.project.cypher(...).
Both projection styles can produce the same
12-node/18-relationship analytical graph, but their
configuration/query metadata differ. Verify
gds.graph.list() rather than assuming “same
counts” means “same schema/properties/orientation.”
8. Deliberately wrong: project everything because GDS is “just a query”
CALL gds.graph.project('prod-all','*','*') on a
production-sized graph without estimate/scope review copies a
broad graph into GDS heap. That can increase GC pressure or
fail allocation even if transactional queries were healthy.
Diagnosis: the analytical memory model is separate from page cache and transaction semantics. Repair: state the question, select exact labels/types/properties, estimate, cap concurrency, project under a disposable/versioned name, verify counts/schema, and drop it after the workflow.
Production judgment and bridge
| Decision surface | Evidence before increasing analytical scope |
|---|---|
| Projection scope | Exact labels/types/properties/filter predicates and projected node/relationship counts match the analytical question. |
| Memory | Projection estimate + algorithm estimate + observed Neo4j heap/process headroom leave safe room for transactional work and GC. |
| Concurrency | Measured throughput/tail latency/CPU under representative contention; never assume max concurrency is optimal. |
| Freshness | Document when the projection was created and what source updates occurred afterward; projections are not live materialized views. |
| Persistence | Explicit decision whether results belong only in stream output, only in the projected graph, or persisted into Neo4j. |
| Failure recovery | Graph can be reconstructed from versioned fixture/query/config; catalog loss on DBMS restart is expected unless an Enterprise persistence feature is deliberately used. |
| Edition/cost | Community core limits are accepted or Enterprise/Aura analytics capabilities are justified by workload, operations and licensing. |
Chapter 23 treats graph creation as a controlled analytical deployment. The next lesson tightens the contract: properties, orientation, filtering, schema, and projection verification decide what an algorithm is actually allowed to “see.”
Check your understanding
- Why is a GDS projection not a live view of Neo4j?
- Which Cypher projection API is deprecated?
- What does Community cap at four?
- What must be verified after projection besides node/relationship counts?
- Why estimate before projection?
Review the answers
1. Because it is an in-memory analytical graph created from source state at projection time; later source writes are not automatically reflected.
2. The legacy gds.graph.project.cypher procedure. Current flexible Cypher projection uses the gds.graph.project aggregation function inside a Cypher query.
3. GDS execution concurrency/CPU cores, independent of Neo4j Database edition.
4. Projected labels/types/properties, orientation, configuration/query metadata, and memory/lifecycle evidence.
5. To test the analytical memory budget before allocating the in-memory graph, while remembering algorithm and concurrent workload memory still need headroom.
Summary and next step
GDS Architecture, In-Memory Graph Catalog, Native/Cypher Projections, and Projection Scope is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Node/Relationship Properties, Orientation, Filtering, Graph Schema, and Projection Verification. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- GDS Manual v2026.07 — Current Graph Data Science manual and versioned feature baseline.
- GDS 2026.07 release notes — Current GDS release line; 2026.07.0 is compatible with Neo4j 2026.07.
- Supported Neo4j versions — Compatibility matrix between Neo4j Database and GDS.
- GDS editions and graph catalog — Community/Enterprise boundaries, graph catalog model, all-algorithms availability, concurrency and model-catalog limits.
- Neo4j Server GDS installation — Bundled products-to-plugins installation path and required GDS procedure security configuration.
- GDS on Docker — Container-based installation examples and plugin activation.
- System requirements — Heap/native-memory/CPU guidance and Community maximum concurrency of four.
- Native projection — Current gds.graph.project procedure, label/type/property/orientation configuration and lifecycle.
- Cypher projection — Current gds.graph.project aggregation-function projection from Cypher query context.
- Legacy Cypher projection — deprecated — Deprecated gds.graph.project.cypher procedure and migration boundary.
- Graph creation and schema — Supported projection types, relationship direction, properties, parallel relationships, and algorithm traits.
- Graph catalog operations — Current graph project/list/exists/drop and graph-property operations.
- Listing graphs — Graph metadata, schemaWithOrientation, degree distribution and configuration inspection.
- Memory estimation — Projection and algorithm estimate syntax, requiredMemory and byte-range evidence.
- Algorithm syntax and execution modes — stream/stats/mutate/write/estimate semantics.
- Running algorithms — Operational meaning and tradeoffs of algorithm execution modes.
- PageRank — Production-tier algorithm used only to demonstrate execution modes and memory estimates in this foundations chapter.
- Neo4j current versions — Current Neo4j Database release and 5.26 LTS line.