Chapter 23 · Graph Data Science Foundations: Projections, Graph Catalog, Memory, and Execution Modes

Build, Validate, Use, and Drop a GDS Projection with Reproducible Resource Measurements

Run a reproducible AtlasMart GDS lifecycle from installation and version checks through estimate, projection, algorithm modes, verification, resource evidence, cleanup, and production decision gates.

Advanced260–370 minutesEnd-to-end GDS labNeo4j 2026.07.1 · Community mandatoryGDS Community 2026.07.0 · Cypher 25Projection catalog · memory estimates · concurrency ≤4 CEJava 21/25 · GDS plugin requiredLast reviewed: September 2026

Learning outcomes

01

Execute the complete GDS projection lifecycle as an evidence-producing runbook.

02

Separate deterministic graph invariants from machine-dependent memory/timing/algorithm measurements.

03

Verify projected vs persisted state after stream/stats/mutate/write operations.

04

Recover safely from stale projections, plugin/version mismatch, and resource-admission failure.

05

Document a production operating envelope and cleanup/reconstruction procedure before moving to graph algorithms.

1. The final lab is a lifecycle test, not an algorithm demo

AtlasMart has accumulated enough connected customer/product behavior that analysts want graph algorithms, but running “an algorithm” is not one operation. The analytical graph must first be selected, copied into a GDS representation, sized, verified, executed, and removed or refreshed deliberately. AtlasMart wants a repeatable analytics job that another engineer can rerun after a restart, upgrade, data refresh, or failure. The deliverable is therefore the projection specification plus evidence and cleanup—not a screenshot of one score.

The safety rule for this chapter is therefore: treat projection definition, memory allocation, algorithm mode, and lifecycle as part of the analytical result. A score without those inputs is not reproducible evidence.

2. Phase 0 · record the exact software and edition boundary

Chapter 23 baseline · reviewed 9 September 2026

Current Neo4j Database is 2026.07.1; current GDS is 2026.07.0 for the Neo4j 2026.07 line. Mandatory work uses self-managed Neo4j Community 2026.07.1 + GDS Community 2026.07.0, database neo4j, user neo4j, disposable password atlasmart-course-2026, loopback HTTP 7474/Bolt 7687, and CYPHER 25 for database-side fixture work. Neo4j 2026.x supports Java 21/25. GDS is the only plugin required in this chapter; APOC is not required.

Current GDS Community boundary

GDS Community includes the algorithm library needed for this course. Its execution concurrency is capped at 4 CPU cores and its model catalog is capped at 3 models. GDS Enterprise removes the CPU-core cap and adds capabilities such as graph backup/restore, Arrow import/export, cluster write support, capacity/load monitoring, and extended model-catalog persistence/sharing. Those Enterprise capabilities are discussed only as edition boundaries; no mandatory lab depends on them.

Projection API status

Two current projection styles matter. Native projection uses the gds.graph.project(...) procedure and remains supported in 2026.07, but current GDS documentation says native projection will be deprecated in a future release and increasingly uses Cypher projection as the norm. Current Cypher projection calls the gds.graph.project(...) aggregation function from a Cypher query. The older gds.graph.project.cypher(...) procedure is already deprecated. This chapter teaches native projection because it makes schema/orientation configuration explicit, then shows the current Cypher form learners should prefer for new flexible projections.

Assumption Value / boundary
Server Neo4j Community 2026.07.1; Java 21/25 supported. Disposable container atlasmart-gds.
GDS GDS Community 2026.07.0. Verify with gds.version(); do not continue if compatibility differs.
Database/security neo4j database; local disposable neo4j user; loopback transport only. Production credentials/TLS/RBAC differ.
Fixture 6 Customer + 6 Product nodes; 14 VIEWED + 4 PURCHASED relationships; numeric weight/customerValue/margin.
Memory/concurrency Learner records estimates/observations. Examples use concurrency=2; Community maximum is 4, but 2 is a lab choice, not a recommendation.
Plugins GDS only. APOC is not required.
Runtime claims Artifact generation does not execute Neo4j/GDS. Expected deterministic graph counts are fixture-derived; memory/timing/algorithm scores must be measured by the learner.
Container/install baseline
# PowerShell-oriented disposable local lab.
# Recreate only this course container; named data/log volumes stay separate from earlier labs.
docker rm -f atlasmart-gds 2>$null

docker run -d `
  --name atlasmart-gds `
  -p 127.0.0.1:7474:7474 `
  -p 127.0.0.1:7687:7687 `
  -v atlasmart-gds-data:/data `
  -v atlasmart-gds-logs:/logs `
  -e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
  -e 'NEO4J_PLUGINS=["graph-data-science"]' `
  -e NEO4J_dbms_security_procedures_unrestricted=gds.* `
  -e NEO4J_dbms_security_procedures_allowlist=gds.* `
  neo4j:2026.07.1

# Wait for the database to become ready, then verify the plugin from cypher-shell/Browser.
Version evidence
CYPHER 25
RETURN gds.version() AS gdsVersion;
// Expected for this chapter baseline: 2026.07.0.

SHOW PROCEDURES YIELD name
WHERE name STARTS WITH 'gds.'
RETURN count(*) AS gdsProcedureCount;
Record Why
Neo4j 2026.07.1 image/runtime Store/runtime semantics and Java support.
gds.version() = 2026.07.0 Plugin compatibility and API behavior.
GDS Community Concurrency/model-catalog/Enterprise-feature boundary.
Projection query/config hash in your runbook Makes analytical input reconstructible.

3. Phase 1 · seed and reconcile the transactional fixture

Seed AtlasMart Chapter 23 source data
CYPHER 25
// Chapter 23 fixture: bounded and safe to delete by labTag.
CREATE CONSTRAINT ch23_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch23_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;

UNWIND [
 {id:'C-2301',segment:'LOYAL',value:92.0},
 {id:'C-2302',segment:'LOYAL',value:78.0},
 {id:'C-2303',segment:'NEW',value:35.0},
 {id:'C-2304',segment:'NEW',value:28.0},
 {id:'C-2305',segment:'B2B',value:96.0},
 {id:'C-2306',segment:'B2B',value:84.0}
] AS row
MERGE (c:Customer {customerId:row.id})
SET c.segment=row.segment,c.customerValue=row.value,c.labTag='ch23';

UNWIND [
 {id:'P-2301',name:'Trail Camera Pro',margin:0.31},
 {id:'P-2302',name:'Action Camera 4K',margin:0.27},
 {id:'P-2303',name:'Solar Trail Charger',margin:0.24},
 {id:'P-2304',name:'Indoor Security Camera',margin:0.29},
 {id:'P-2305',name:'Hydration Vest',margin:0.22},
 {id:'P-2306',name:'Wildlife Field Guide',margin:0.35}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name,p.margin=row.margin,p.labTag='ch23';

MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WITH c,p WHERE
 (c.customerId='C-2301' AND p.productId IN ['P-2301','P-2303','P-2306']) OR
 (c.customerId='C-2302' AND p.productId IN ['P-2301','P-2302']) OR
 (c.customerId='C-2303' AND p.productId IN ['P-2302','P-2305']) OR
 (c.customerId='C-2304' AND p.productId IN ['P-2304','P-2305']) OR
 (c.customerId='C-2305' AND p.productId IN ['P-2301','P-2303','P-2304']) OR
 (c.customerId='C-2306' AND p.productId IN ['P-2303','P-2306'])
MERGE (c)-[r:VIEWED]->(p)
SET r.weight = CASE c.segment WHEN 'LOYAL' THEN 2.0 WHEN 'B2B' THEN 1.5 ELSE 1.0 END,
    r.labTag='ch23';

MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WHERE (c.customerId='C-2301' AND p.productId='P-2301') OR
      (c.customerId='C-2302' AND p.productId='P-2302') OR
      (c.customerId='C-2305' AND p.productId='P-2303') OR
      (c.customerId='C-2306' AND p.productId='P-2306')
MERGE (c)-[r:PURCHASED]->(p)
SET r.weight=3.0,r.labTag='ch23';
Reconcile deterministic source counts
CYPHER 25
MATCH (n {labTag:'ch23'})
RETURN labels(n) AS labels,count(*) AS nodes ORDER BY labels;
MATCH ()-[r]->() WHERE r.labTag='ch23'
RETURN type(r) AS type,count(*) AS relationships,sum(r.weight) AS totalWeight ORDER BY type;
// Deterministic invariants: 6 Customer + 6 Product nodes; 14 VIEWED; 4 PURCHASED.

If these invariants fail, stop. An algorithm cannot validate an incorrect source fixture. Repair source identity/relationships first.

4. Phase 2 · estimate and admit the graph

Drop stale projection then estimate
CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;


CALL gds.graph.project.estimate(
  {
    Customer:{properties:['customerValue']},
    Product:{properties:['margin']}
  },
  {
    VIEWED:{orientation:'NATURAL',properties:['weight']},
    PURCHASED:{orientation:'NATURAL',properties:['weight']}
  },
  {readConcurrency:2}
)
YIELD nodeCount,relationshipCount,requiredMemory,bytesMin,bytesMax
RETURN *;
// The byte estimate is environment/version dependent; record it instead of copying a canned number.

Record the estimate, current heap/process/host headroom, and chosen concurrency. This chapter does not prescribe a universal “safe percentage.” Your admission gate must come from representative service behavior and an explicit reserve.

5. Phase 3 · project and prove schema/state

Create bounded projection
CALL gds.graph.project(
  'atlasmart-ch23',
  {
    Customer:{properties:['customerValue']},
    Product:{properties:['margin']}
  },
  {
    VIEWED:{orientation:'NATURAL',properties:['weight']},
    PURCHASED:{orientation:'NATURAL',properties:['weight']}
  },
  {readConcurrency:2}
)
YIELD graphName,nodeCount,relationshipCount,projectMillis,configuration
RETURN graphName,nodeCount,relationshipCount,projectMillis,configuration;
// Expected deterministic counts: nodeCount=12, relationshipCount=18.
List schema, memory and degree evidence
CALL gds.graph.list('atlasmart-ch23')
YIELD graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration
RETURN graphName,nodeCount,relationshipCount,schemaWithOrientation,degreeDistribution,memoryUsage,configuration;

Acceptance: nodeCount=12 and relationshipCount=18, expected labels/types/properties/orientation, and a catalog memory footprint that remains inside the approved envelope.

6. Phase 4 · estimate, stream, stats, mutate; prove persistence boundary

Estimate algorithm
CALL gds.pageRank.stream.estimate('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  concurrency:2,
  maxIterations:20,
  dampingFactor:0.85
})
YIELD nodeCount,relationshipCount,requiredMemory,bytesMin,bytesMax,heapPercentageMin,heapPercentageMax
RETURN *;
Stream results
CALL gds.pageRank.stream('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  concurrency:2,
  maxIterations:20,
  dampingFactor:0.85
})
YIELD nodeId,score
RETURN gds.util.asNode(nodeId).productId AS productId,
       gds.util.asNode(nodeId).customerId AS customerId,
       score
ORDER BY score DESC LIMIT 8;
// stream returns rows; it does not alter graph catalog properties or Neo4j properties.
Stats results
CALL gds.pageRank.stats('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  concurrency:2,
  maxIterations:20
})
YIELD ranIterations,didConverge,centralityDistribution,computeMillis
RETURN *;
// stats returns summary evidence without exposing per-node scores or writing state.
Mutate projected graph
CALL gds.pageRank.mutate('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  mutateProperty:'ch23PageRank',
  concurrency:2,
  maxIterations:20
})
YIELD nodePropertiesWritten,mutateMillis,ranIterations,didConverge
RETURN *;

CALL gds.graph.nodeProperty.stream('atlasmart-ch23','ch23PageRank')
YIELD nodeId,propertyValue
RETURN gds.util.asNode(nodeId).productId AS productId,
       gds.util.asNode(nodeId).customerId AS customerId,
       propertyValue
ORDER BY propertyValue DESC LIMIT 8;
// ch23PageRank exists only in the in-memory projected graph after mutate.
Verify mutate did not touch Neo4j properties
CYPHER 25
MATCH (n {labTag:'ch23'})
RETURN count(n) AS sourceNodes,
       count(n.ch23PageRank) AS persistedPageRankProperties;
// Expected persistedPageRankProperties=0 immediately after mutate.

7. Phase 5 · optional persistent write with explicit rollback

Optional but reproducible

Run write mode only to learn the persistence boundary, then remove the Chapter 23 fixture during cleanup. In a real application graph, use a namespaced/versioned property and migration/rollback plan rather than overwriting a business-owned field.

Write and verify persistent property
CALL gds.pageRank.write('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  writeProperty:'ch23PageRankWritten',
  concurrency:2,
  maxIterations:20
})
YIELD nodePropertiesWritten,writeMillis,ranIterations,didConverge
RETURN *;

MATCH (n {labTag:'ch23'})
RETURN labels(n) AS labels,count(n.ch23PageRankWritten) AS persisted
ORDER BY labels;
// write changes Neo4j; cleanup must remove this lab property.

8. Phase 6 · cleanup and prove resource release

Drop analytical graph and delete only lab data
// First free GDS heap.
CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;

// Then remove only Chapter 23 persistent lab data/properties.
MATCH (n {labTag:'ch23'}) DETACH DELETE n;
DROP CONSTRAINT ch23_customer_id IF EXISTS;
DROP CONSTRAINT ch23_product_id IF EXISTS;
Verify catalog no longer contains the lab graph
CALL gds.graph.list()
YIELD graphName,nodeCount,relationshipCount,memoryUsage,creationTime
RETURN graphName,nodeCount,relationshipCount,memoryUsage,creationTime
ORDER BY graphName;
// atlasmart-ch23 must be absent after cleanup.

Cleanup order matters: drop the graph first to release GDS heap, then remove persistent lab data. If an earlier step fails, the runbook’s finally/cleanup path should still attempt gds.graph.drop.

9. Failure-injection matrix

Injected condition Expected evidence Safe recovery
GDS version mismatch gds.version differs / procedures unavailable Stop; install the compatible GDS line before running projections.
Stale graph already exists gds.graph.exists true / project name collision Inspect ownership/freshness; drop only the disposable lab graph, then reconstruct.
Projection count mismatch gds.graph.list counts/schema differ Do not run algorithm; repair source/filter/projection configuration.
Memory admission fails Estimate + headroom check exceeds policy Reduce scope/concurrency, queue/move workload; do not provoke OOM.
Mutate mistaken for persistence Neo4j count(n.ch23PageRank)=0 Use write/export only when durable output is intended.
DBMS restart Catalog graph absent after restart Rebuild from versioned source/projection spec; disappearance is expected Community lifecycle.

10. Production operating envelope

Decision surface Evidence before increasing analytical scope
Projection scope Exact labels/types/properties/filter predicates and projected node/relationship counts match the analytical question.
Memory Projection estimate + algorithm estimate + observed Neo4j heap/process headroom leave safe room for transactional work and GC.
Concurrency Measured throughput/tail latency/CPU under representative contention; never assume max concurrency is optimal.
Freshness Document when the projection was created and what source updates occurred afterward; projections are not live materialized views.
Persistence Explicit decision whether results belong only in stream output, only in the projected graph, or persisted into Neo4j.
Failure recovery Graph can be reconstructed from versioned fixture/query/config; catalog loss on DBMS restart is expected unless an Enterprise persistence feature is deliberately used.
Edition/cost Community core limits are accepted or Enterprise/Aura analytics capabilities are justified by workload, operations and licensing.
Runbook field Example evidence—not a universal threshold
Projection identity atlasmart-ch23 + source/filter/config hash + creation timestamp.
Correctness gate Expected labels/types/properties/orientation and reconciled counts.
Memory gate Measured graph/algorithm estimates plus explicit transactional/JVM reserve.
Concurrency gate Chosen from coexistence test; <=4 in GDS Community.
Freshness gate Maximum source-update age appropriate to the analytical decision.
Persistence gate stream/stats/mutate by default; write only with schema/change owner.
Cleanup gate Graph absent from catalog after workflow/failure unless deliberate reuse is documented.

With these foundations in place, Chapter 24 can teach centrality, community detection, similarity and path algorithms without hiding the projection and resource model underneath them.

Check your understanding

  1. What is the first acceptance gate before running an algorithm?
  2. Which measurements are deterministic in this fixture?
  3. What proves mutate is not persistent?
  4. What should happen if memory admission fails?
  5. What does a successful cleanup prove?
Review the answers

1. Source/projection correctness: expected identity, counts, schema, orientation and properties.

2. Fixture-derived node/relationship counts. Memory, timing and PageRank score values are runtime-dependent and must be measured.

3. After mutate, streaming the GDS property succeeds while ordinary Neo4j Cypher reports zero persisted properties with that name.

4. Reject or reshape the job; reduce scope/concurrency or move/queue analytics instead of forcing allocation.

5. The named graph is absent from the catalog and the disposable persistent fixture is removed, so the runbook does not leak analytical or database state.

Summary and next step

Build, Validate, Use, and Drop a GDS Projection with Reproducible Resource Measurements is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Degree, PageRank, Betweenness, Eigenvector-Like Centrality Concepts and Business Interpretation. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.