Chapter 23 · Graph Data Science Foundations: Projections, Graph Catalog, Memory, and Execution Modes

GDS Architecture, In-Memory Graph Catalog, Native/Cypher Projections, and Projection Scope

Build the correct mental model of GDS as a separate in-memory analytical graph catalog, then create bounded native and current Cypher projections without confusing them with the transactional store.

Advanced220–310 minutesProjection architecture labNeo4j 2026.07.1 · Community mandatoryGDS Community 2026.07.0 · Cypher 25Projection catalog · memory estimates · concurrency ≤4 CEJava 21/25 · GDS plugin requiredLast reviewed: September 2026

Learning outcomes

01

Explain why GDS uses a separate in-memory graph catalog instead of treating the transactional store as the algorithm graph.

02

Distinguish native projection, current Cypher projection, and the deprecated legacy Cypher-projection procedure.

03

Install/verify a compatible GDS Community plugin and build a bounded AtlasMart source fixture.

04

Estimate, project, list, verify, and drop a named analytical graph.

05

Make a production decision about scope, freshness, memory isolation, edition limits, and reconstruction.

1. Why AtlasMart needs an analytical graph boundary

AtlasMart has accumulated enough connected customer/product behavior that analysts want graph algorithms, but running “an algorithm” is not one operation. The analytical graph must first be selected, copied into a GDS representation, sized, verified, executed, and removed or refreshed deliberately. The first request—“find influential products”—looks harmless, yet projecting every customer, order, product, and relationship by default could duplicate far more data into heap than the question needs.

The safety rule for this chapter is therefore: treat projection definition, memory allocation, algorithm mode, and lifecycle as part of the analytical result. A score without those inputs is not reproducible evidence.

2. Transactional Neo4j vs the GDS graph catalog

Term Mechanism-first meaning
transactional store The persisted Neo4j graph used by ordinary Cypher transactions. GDS does not automatically make algorithm computations operate directly on this store.
projection A selected copy/transformation of nodes, relationships and numeric properties loaded into the GDS in-memory graph catalog.
graph catalog Process-local catalog of named in-memory analytical graphs. A projection survives until explicitly dropped or until its source database/DBMS stops.
orientation How a relationship is represented analytically: NATURAL keeps stored direction; REVERSE flips it; UNDIRECTED creates analytical connectivity in both directions and can increase memory.
graph schema Projected labels, relationship types, properties and orientations—not the entire transactional database schema.
concurrency Number of worker threads GDS may use for an operation. It is a resource budget, not a universal performance knob.
estimate A dry-run-style memory calculation that reports requiredMemory/byte ranges without running the actual algorithm or graph projection.
stream Algorithm mode returning per-entity results to the caller without modifying the projection or persistent store.
stats Algorithm mode returning summary statistics only, without modifying the projection or store.
mutate Algorithm mode writing computed results into the in-memory projected graph only.
write Algorithm mode persisting computed results back into Neo4j database properties/relationships as defined by that algorithm.

Neo4j’s transactional graph is designed for durable ACID reads/writes. GDS algorithms use a compact analytical representation loaded into heap. Projection is therefore a data contract: it decides which source entities exist analytically, which relationships count, which direction the algorithm sees, and which numeric values become weights/features.

Layer Durability Freshness Primary purpose
Neo4j store Durable until modified/deleted Current committed database state Transactional graph queries and writes
GDS named graph In-memory catalog; disappears when dropped or source DB/DBMS stops Snapshot at projection/mutation time High-throughput graph analytics
Algorithm stream result Client result only Computation-time view Exploration/inspection
Algorithm write result Persisted back to Neo4j Durable after committed write Operationalize selected analytical outputs

3. Current projection APIs: choose deliberately

Projection API status

Two current projection styles matter. Native projection uses the gds.graph.project(...) procedure and remains supported in 2026.07, but current GDS documentation says native projection will be deprecated in a future release and increasingly uses Cypher projection as the norm. Current Cypher projection calls the gds.graph.project(...) aggregation function from a Cypher query. The older gds.graph.project.cypher(...) procedure is already deprecated. This chapter teaches native projection because it makes schema/orientation configuration explicit, then shows the current Cypher form learners should prefer for new flexible projections.

Projection path Current status Best fit Important boundary
Native gds.graph.project procedure Supported in 2026.07; docs say future deprecation is planned Simple label/type/property projections Configuration-driven; less expressive filtering than Cypher
Cypher + gds.graph.project aggregation function Current flexible Cypher projection Filtering, transformed data, multi-source/query-derived projection Projection query itself becomes part of reproducibility and performance
gds.graph.project.cypher procedure Deprecated legacy API Migration-reading only Do not introduce into new course code

4. Install and prove the compatibility boundary

Chapter 23 baseline · reviewed 9 September 2026

Current Neo4j Database is 2026.07.1; current GDS is 2026.07.0 for the Neo4j 2026.07 line. Mandatory work uses self-managed Neo4j Community 2026.07.1 + GDS Community 2026.07.0, database neo4j, user neo4j, disposable password atlasmart-course-2026, loopback HTTP 7474/Bolt 7687, and CYPHER 25 for database-side fixture work. Neo4j 2026.x supports Java 21/25. GDS is the only plugin required in this chapter; APOC is not required.

Current GDS Community boundary

GDS Community includes the algorithm library needed for this course. Its execution concurrency is capped at 4 CPU cores and its model catalog is capped at 3 models. GDS Enterprise removes the CPU-core cap and adds capabilities such as graph backup/restore, Arrow import/export, cluster write support, capacity/load monitoring, and extended model-catalog persistence/sharing. Those Enterprise capabilities are discussed only as edition boundaries; no mandatory lab depends on them.

Projection API status

Two current projection styles matter. Native projection uses the gds.graph.project(...) procedure and remains supported in 2026.07, but current GDS documentation says native projection will be deprecated in a future release and increasingly uses Cypher projection as the norm. Current Cypher projection calls the gds.graph.project(...) aggregation function from a Cypher query. The older gds.graph.project.cypher(...) procedure is already deprecated. This chapter teaches native projection because it makes schema/orientation configuration explicit, then shows the current Cypher form learners should prefer for new flexible projections.

Assumption Value / boundary
Server Neo4j Community 2026.07.1; Java 21/25 supported. Disposable container atlasmart-gds.
GDS GDS Community 2026.07.0. Verify with gds.version(); do not continue if compatibility differs.
Database/security neo4j database; local disposable neo4j user; loopback transport only. Production credentials/TLS/RBAC differ.
Fixture 6 Customer + 6 Product nodes; 14 VIEWED + 4 PURCHASED relationships; numeric weight/customerValue/margin.
Memory/concurrency Learner records estimates/observations. Examples use concurrency=2; Community maximum is 4, but 2 is a lab choice, not a recommendation.
Plugins GDS only. APOC is not required.
Runtime claims Artifact generation does not execute Neo4j/GDS. Expected deterministic graph counts are fixture-derived; memory/timing/algorithm scores must be measured by the learner.
Create disposable Neo4j + GDS Community container
# PowerShell-oriented disposable local lab.
# Recreate only this course container; named data/log volumes stay separate from earlier labs.
docker rm -f atlasmart-gds 2>$null

docker run -d `
  --name atlasmart-gds `
  -p 127.0.0.1:7474:7474 `
  -p 127.0.0.1:7687:7687 `
  -v atlasmart-gds-data:/data `
  -v atlasmart-gds-logs:/logs `
  -e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
  -e 'NEO4J_PLUGINS=["graph-data-science"]' `
  -e NEO4J_dbms_security_procedures_unrestricted=gds.* `
  -e NEO4J_dbms_security_procedures_allowlist=gds.* `
  neo4j:2026.07.1

# Wait for the database to become ready, then verify the plugin from cypher-shell/Browser.
Verify exact GDS version and procedure surface
CYPHER 25
RETURN gds.version() AS gdsVersion;
// Expected for this chapter baseline: 2026.07.0.

SHOW PROCEDURES YIELD name
WHERE name STARTS WITH 'gds.'
RETURN count(*) AS gdsProcedureCount;
Stop on mismatch

If gds.version() is not the compatible 2026.07 line, do not “try the commands anyway.” Resolve server/GDS compatibility first. GDS uses Neo4j internals; compatibility is an operational dependency.

5. Build the smallest useful AtlasMart source graph

Create deterministic transactional fixture
CYPHER 25
// Chapter 23 fixture: bounded and safe to delete by labTag.
CREATE CONSTRAINT ch23_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch23_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;

UNWIND [
 {id:'C-2301',segment:'LOYAL',value:92.0},
 {id:'C-2302',segment:'LOYAL',value:78.0},
 {id:'C-2303',segment:'NEW',value:35.0},
 {id:'C-2304',segment:'NEW',value:28.0},
 {id:'C-2305',segment:'B2B',value:96.0},
 {id:'C-2306',segment:'B2B',value:84.0}
] AS row
MERGE (c:Customer {customerId:row.id})
SET c.segment=row.segment,c.customerValue=row.value,c.labTag='ch23';

UNWIND [
 {id:'P-2301',name:'Trail Camera Pro',margin:0.31},
 {id:'P-2302',name:'Action Camera 4K',margin:0.27},
 {id:'P-2303',name:'Solar Trail Charger',margin:0.24},
 {id:'P-2304',name:'Indoor Security Camera',margin:0.29},
 {id:'P-2305',name:'Hydration Vest',margin:0.22},
 {id:'P-2306',name:'Wildlife Field Guide',margin:0.35}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name,p.margin=row.margin,p.labTag='ch23';

MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WITH c,p WHERE
 (c.customerId='C-2301' AND p.productId IN ['P-2301','P-2303','P-2306']) OR
 (c.customerId='C-2302' AND p.productId IN ['P-2301','P-2302']) OR
 (c.customerId='C-2303' AND p.productId IN ['P-2302','P-2305']) OR
 (c.customerId='C-2304' AND p.productId IN ['P-2304','P-2305']) OR
 (c.customerId='C-2305' AND p.productId IN ['P-2301','P-2303','P-2304']) OR
 (c.customerId='C-2306' AND p.productId IN ['P-2303','P-2306'])
MERGE (c)-[r:VIEWED]->(p)
SET r.weight = CASE c.segment WHEN 'LOYAL' THEN 2.0 WHEN 'B2B' THEN 1.5 ELSE 1.0 END,
    r.labTag='ch23';

MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WHERE (c.customerId='C-2301' AND p.productId='P-2301') OR
      (c.customerId='C-2302' AND p.productId='P-2302') OR
      (c.customerId='C-2305' AND p.productId='P-2303') OR
      (c.customerId='C-2306' AND p.productId='P-2306')
MERGE (c)-[r:PURCHASED]->(p)
SET r.weight=3.0,r.labTag='ch23';
Verify source-store invariants
CYPHER 25
MATCH (n {labTag:'ch23'})
RETURN labels(n) AS labels,count(*) AS nodes ORDER BY labels;
MATCH ()-[r]->() WHERE r.labTag='ch23'
RETURN type(r) AS type,count(*) AS relationships,sum(r.weight) AS totalWeight ORDER BY type;
// Deterministic invariants: 6 Customer + 6 Product nodes; 14 VIEWED; 4 PURCHASED.

The fixture is intentionally tiny because this chapter teaches lifecycle mechanics, not benchmark bragging. Counts are deterministic; PageRank scores and memory/timing are measured outputs.

6. Estimate before allocating; then project

Drop stale projection from a previous run
CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;
Estimate native projection memory
CALL gds.graph.project.estimate(
  {
    Customer:{properties:['customerValue']},
    Product:{properties:['margin']}
  },
  {
    VIEWED:{orientation:'NATURAL',properties:['weight']},
    PURCHASED:{orientation:'NATURAL',properties:['weight']}
  },
  {readConcurrency:2}
)
YIELD nodeCount,relationshipCount,requiredMemory,bytesMin,bytesMax
RETURN *;
// The byte estimate is environment/version dependent; record it instead of copying a canned number.
Create bounded native projection
CALL gds.graph.project(
  'atlasmart-ch23',
  {
    Customer:{properties:['customerValue']},
    Product:{properties:['margin']}
  },
  {
    VIEWED:{orientation:'NATURAL',properties:['weight']},
    PURCHASED:{orientation:'NATURAL',properties:['weight']}
  },
  {readConcurrency:2}
)
YIELD graphName,nodeCount,relationshipCount,projectMillis,configuration
RETURN graphName,nodeCount,relationshipCount,projectMillis,configuration;
// Expected deterministic counts: nodeCount=12, relationshipCount=18.
Inspect catalog metadata and schema
CALL gds.graph.list('atlasmart-ch23')
YIELD graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration
RETURN graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration;

The estimate is not a promise that future workload fits: algorithm working memory, concurrent transactions, JVM overhead, page cache, result materialization, and other projections also compete for resources. It is evidence for a budget.

7. The current Cypher projection equivalent

Rebuild with current Cypher projection
CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;


CYPHER 25
MATCH (source {labTag:'ch23'})-[r:VIEWED|PURCHASED]->(target {labTag:'ch23'})
RETURN gds.graph.project(
  'atlasmart-ch23',
  source,
  target,
  {
    sourceNodeLabels:labels(source),
    targetNodeLabels:labels(target),
    sourceNodeProperties:CASE WHEN source:Customer THEN {customerValue:source.customerValue} ELSE {} END,
    targetNodeProperties:CASE WHEN target:Product THEN {margin:target.margin} ELSE {} END,
    relationshipType:type(r),
    relationshipProperties:{weight:r.weight}
  },
  {readConcurrency:2}
)
YIELD graphName,nodeCount,relationshipCount,projectMillis
RETURN *;
// Current Cypher projection is the gds.graph.project aggregation function.
// Do not replace this with deprecated gds.graph.project.cypher(...).
Observable equivalence vs identical mechanism

Both projection styles can produce the same 12-node/18-relationship analytical graph, but their configuration/query metadata differ. Verify gds.graph.list() rather than assuming “same counts” means “same schema/properties/orientation.”

8. Deliberately wrong: project everything because GDS is “just a query”

Wrong approach

CALL gds.graph.project('prod-all','*','*') on a production-sized graph without estimate/scope review copies a broad graph into GDS heap. That can increase GC pressure or fail allocation even if transactional queries were healthy.

Diagnosis: the analytical memory model is separate from page cache and transaction semantics. Repair: state the question, select exact labels/types/properties, estimate, cap concurrency, project under a disposable/versioned name, verify counts/schema, and drop it after the workflow.

Production judgment and bridge

Decision surface Evidence before increasing analytical scope
Projection scope Exact labels/types/properties/filter predicates and projected node/relationship counts match the analytical question.
Memory Projection estimate + algorithm estimate + observed Neo4j heap/process headroom leave safe room for transactional work and GC.
Concurrency Measured throughput/tail latency/CPU under representative contention; never assume max concurrency is optimal.
Freshness Document when the projection was created and what source updates occurred afterward; projections are not live materialized views.
Persistence Explicit decision whether results belong only in stream output, only in the projected graph, or persisted into Neo4j.
Failure recovery Graph can be reconstructed from versioned fixture/query/config; catalog loss on DBMS restart is expected unless an Enterprise persistence feature is deliberately used.
Edition/cost Community core limits are accepted or Enterprise/Aura analytics capabilities are justified by workload, operations and licensing.

Chapter 23 treats graph creation as a controlled analytical deployment. The next lesson tightens the contract: properties, orientation, filtering, schema, and projection verification decide what an algorithm is actually allowed to “see.”

Check your understanding

  1. Why is a GDS projection not a live view of Neo4j?
  2. Which Cypher projection API is deprecated?
  3. What does Community cap at four?
  4. What must be verified after projection besides node/relationship counts?
  5. Why estimate before projection?
Review the answers

1. Because it is an in-memory analytical graph created from source state at projection time; later source writes are not automatically reflected.

2. The legacy gds.graph.project.cypher procedure. Current flexible Cypher projection uses the gds.graph.project aggregation function inside a Cypher query.

3. GDS execution concurrency/CPU cores, independent of Neo4j Database edition.

4. Projected labels/types/properties, orientation, configuration/query metadata, and memory/lifecycle evidence.

5. To test the analytical memory budget before allocating the in-memory graph, while remembering algorithm and concurrent workload memory still need headroom.

Summary and next step

GDS Architecture, In-Memory Graph Catalog, Native/Cypher Projections, and Projection Scope is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Node/Relationship Properties, Orientation, Filtering, Graph Schema, and Projection Verification. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.