Chapter 23 · Graph Data Science Foundations: Projections, Graph Catalog, Memory, and Execution Modes
Estimate Memory, Control Concurrency, Manage Graph Lifecycles, and Avoid Production Memory Starvation
Estimate projection and algorithm memory before allocation, bound concurrency, observe heap/process evidence, and manage graph lifecycles so analytical work cannot silently starve transactional workloads.
Learning outcomes
Estimate graph and algorithm memory before allocation/execution.
Distinguish GDS heap use from Neo4j page cache, transactional/query memory, and OS/native memory.
Treat concurrency as a bounded resource budget and respect the Community maximum of four.
Define projection lifecycle/freshness/cleanup ownership so catalog graphs do not become invisible resource leaks.
Build a production admission check that rejects unsafe analytical work instead of starving the DBMS.
1. AtlasMart has a memory budget, not “free RAM”
AtlasMart has accumulated enough connected customer/product behavior that analysts want graph algorithms, but running “an algorithm” is not one operation. The analytical graph must first be selected, copied into a GDS representation, sized, verified, executed, and removed or refreshed deliberately. A second analyst wants a larger projection while customer-facing queries are already running. The safe question is not “will PageRank finish?” but “can projection + algorithm + transaction/query/JVM overhead coexist within the service SLO?”
The safety rule for this chapter is therefore: treat projection definition, memory allocation, algorithm mode, and lifecycle as part of the analytical result. A score without those inputs is not reproducible evidence.
2. Where GDS memory fits
Current Neo4j Database is 2026.07.1; current GDS
is 2026.07.0 for the Neo4j 2026.07 line.
Mandatory work uses self-managed
Neo4j Community 2026.07.1 + GDS Community 2026.07.0, database neo4j, user neo4j,
disposable password atlasmart-course-2026,
loopback HTTP 7474/Bolt 7687, and
CYPHER 25 for database-side fixture work. Neo4j
2026.x supports Java 21/25. GDS is the only plugin required in
this chapter; APOC is not required.
GDS Community includes the algorithm library needed for this course. Its execution concurrency is capped at 4 CPU cores and its model catalog is capped at 3 models. GDS Enterprise removes the CPU-core cap and adds capabilities such as graph backup/restore, Arrow import/export, cluster write support, capacity/load monitoring, and extended model-catalog persistence/sharing. Those Enterprise capabilities are discussed only as edition boundaries; no mandatory lab depends on them.
Two current projection styles matter.
Native projection uses the
gds.graph.project(...) procedure and remains
supported in 2026.07, but current GDS documentation says
native projection will be deprecated in a future release and
increasingly uses Cypher projection as the norm.
Current Cypher projection calls the
gds.graph.project(...) aggregation function from
a Cypher query. The older
gds.graph.project.cypher(...) procedure is
already deprecated. This chapter teaches native projection
because it makes schema/orientation configuration explicit,
then shows the current Cypher form learners should prefer for
new flexible projections.
| Assumption | Value / boundary |
|---|---|
| Server | Neo4j Community 2026.07.1; Java 21/25 supported. Disposable container atlasmart-gds. |
| GDS | GDS Community 2026.07.0. Verify with gds.version(); do not continue if compatibility differs. |
| Database/security | neo4j database; local disposable neo4j user; loopback transport only. Production credentials/TLS/RBAC differ. |
| Fixture | 6 Customer + 6 Product nodes; 14 VIEWED + 4 PURCHASED relationships; numeric weight/customerValue/margin. |
| Memory/concurrency | Learner records estimates/observations. Examples use concurrency=2; Community maximum is 4, but 2 is a lab choice, not a recommendation. |
| Plugins | GDS only. APOC is not required. |
| Runtime claims | Artifact generation does not execute Neo4j/GDS. Expected deterministic graph counts are fixture-derived; memory/timing/algorithm scores must be measured by the learner. |
| Memory/resource surface | What it does | Why GDS changes the picture |
|---|---|---|
| JVM heap | Objects, query/runtime state, GDS in-memory graph and algorithm working memory | GDS graph models/algorithms operate on heap and can become a dominant resident workload. |
| Page cache | Caches Neo4j store/index pages outside the ordinary object heap model | Shrinking it blindly to make room for GDS can increase storage I/O for transactions. |
| Transaction/query memory | Execution state for concurrent Cypher work | Analytical activity can overlap with request spikes; estimates are not exclusive reservations. |
| Native/direct memory | JVM/native subsystems; Arrow uses native memory when enabled | Enterprise Arrow paths require an additional budget; not used in this mandatory lab. |
| OS CPU/I/O | Shared host resources | More GDS workers can reduce tail latency headroom for database/driver activity. |
3. Estimate graph memory before projection
CYPHER 25
// Chapter 23 fixture: bounded and safe to delete by labTag.
CREATE CONSTRAINT ch23_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch23_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
UNWIND [
{id:'C-2301',segment:'LOYAL',value:92.0},
{id:'C-2302',segment:'LOYAL',value:78.0},
{id:'C-2303',segment:'NEW',value:35.0},
{id:'C-2304',segment:'NEW',value:28.0},
{id:'C-2305',segment:'B2B',value:96.0},
{id:'C-2306',segment:'B2B',value:84.0}
] AS row
MERGE (c:Customer {customerId:row.id})
SET c.segment=row.segment,c.customerValue=row.value,c.labTag='ch23';
UNWIND [
{id:'P-2301',name:'Trail Camera Pro',margin:0.31},
{id:'P-2302',name:'Action Camera 4K',margin:0.27},
{id:'P-2303',name:'Solar Trail Charger',margin:0.24},
{id:'P-2304',name:'Indoor Security Camera',margin:0.29},
{id:'P-2305',name:'Hydration Vest',margin:0.22},
{id:'P-2306',name:'Wildlife Field Guide',margin:0.35}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name,p.margin=row.margin,p.labTag='ch23';
MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WITH c,p WHERE
(c.customerId='C-2301' AND p.productId IN ['P-2301','P-2303','P-2306']) OR
(c.customerId='C-2302' AND p.productId IN ['P-2301','P-2302']) OR
(c.customerId='C-2303' AND p.productId IN ['P-2302','P-2305']) OR
(c.customerId='C-2304' AND p.productId IN ['P-2304','P-2305']) OR
(c.customerId='C-2305' AND p.productId IN ['P-2301','P-2303','P-2304']) OR
(c.customerId='C-2306' AND p.productId IN ['P-2303','P-2306'])
MERGE (c)-[r:VIEWED]->(p)
SET r.weight = CASE c.segment WHEN 'LOYAL' THEN 2.0 WHEN 'B2B' THEN 1.5 ELSE 1.0 END,
r.labTag='ch23';
MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WHERE (c.customerId='C-2301' AND p.productId='P-2301') OR
(c.customerId='C-2302' AND p.productId='P-2302') OR
(c.customerId='C-2305' AND p.productId='P-2303') OR
(c.customerId='C-2306' AND p.productId='P-2306')
MERGE (c)-[r:PURCHASED]->(p)
SET r.weight=3.0,r.labTag='ch23';
CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;
CALL gds.graph.project.estimate(
{
Customer:{properties:['customerValue']},
Product:{properties:['margin']}
},
{
VIEWED:{orientation:'NATURAL',properties:['weight']},
PURCHASED:{orientation:'NATURAL',properties:['weight']}
},
{readConcurrency:2}
)
YIELD nodeCount,relationshipCount,requiredMemory,bytesMin,bytesMax
RETURN *;
// The byte estimate is environment/version dependent; record it instead of copying a canned number.
Record requiredMemory, bytesMin, and
bytesMax. The estimate is version/config/data-shape
dependent. Do not turn a tutorial’s number into a capacity
constant.
4. Estimate algorithm memory before computation
CALL gds.graph.project(
'atlasmart-ch23',
{
Customer:{properties:['customerValue']},
Product:{properties:['margin']}
},
{
VIEWED:{orientation:'NATURAL',properties:['weight']},
PURCHASED:{orientation:'NATURAL',properties:['weight']}
},
{readConcurrency:2}
)
YIELD graphName,nodeCount,relationshipCount,projectMillis,configuration
RETURN graphName,nodeCount,relationshipCount,projectMillis,configuration;
// Expected deterministic counts: nodeCount=12, relationshipCount=18.
CALL gds.pageRank.stream.estimate('atlasmart-ch23', {
relationshipTypes:['VIEWED','PURCHASED'],
relationshipWeightProperty:'weight',
concurrency:2,
maxIterations:20,
dampingFactor:0.85
})
YIELD nodeCount,relationshipCount,requiredMemory,bytesMin,bytesMax,heapPercentageMin,heapPercentageMax
RETURN *;
Projection memory and algorithm working memory are distinct. For production admission, compare the combined analytical budget with observed free/headroom under representative transactional load—not an idle laptop.
5. Concurrency: limit, measure, then justify
| Value | Meaning in this chapter |
|---|---|
| concurrency=2 | Deliberate conservative lab setting that keeps resource coupling visible; not a universal recommendation. |
| Community max=4 | License-enforced GDS ceiling. A host with 32 cores does not remove it. |
| Enterprise | Can use more CPU cores, but “unlimited” license capacity is not unlimited safe capacity. |
Higher concurrency can reduce one job’s compute time while increasing CPU contention, allocation rate, GC, I/O pressure, and tail latency elsewhere. Measure end-to-end service behavior, not only algorithm milliseconds.
6. Lifecycle and freshness are operational contracts
CALL gds.graph.list()
YIELD graphName,nodeCount,relationshipCount,memoryUsage,creationTime
RETURN graphName,nodeCount,relationshipCount,memoryUsage,creationTime
ORDER BY graphName;
// atlasmart-ch23 must be absent after cleanup.
A named graph stays in the catalog until it is dropped, its source database is stopped/dropped, or the DBMS stops. That is convenient for repeated algorithms but creates ownership questions: Who refreshes it? Who drops it? Which source commit/time does it represent? Which service is allowed to allocate a new copy?
| Lifecycle state | Required control |
|---|---|
| Create | Versioned projection spec + estimate + owner + expected counts. |
| Reuse | Check graph existence, creation time, schema, freshness budget and source-model compatibility. |
| Refresh | Drop/reproject or use an explicitly supported update workflow; never silently call stale data current. |
| Drop | Free catalog heap after workflow or on failure path; verify absence. |
7. Controlled failure injection: reject an unsafe admission request
Do not deliberately exhaust heap. Instead, create an admission worksheet: measured free/headroom = H; graph estimate max = G; algorithm estimate max = A; required safety reserve for transactional/JVM work = R. Reject the job whenever G + A + R exceeds H or concurrency/tail-latency tests show unacceptable interference.
This failure injection teaches the real control without risking the machine. A production scheduler can encode the same gate using capacity telemetry and queueing rather than attempting an allocation and hoping for an OOM exception.
8. Deliberately wrong: fix pressure by increasing heap until the projection fits
A projection estimate does not fit, so increase Neo4j heap until it does, without checking page cache, container/host limits, GC behavior, concurrent transactions or other graphs.
Diagnosis: one component was optimized in isolation. Repair: reduce projection scope first, estimate graph + algorithm, measure workload coexistence, choose a bounded concurrency, reserve headroom, and separate analytics from the transactional path when the resource envelope cannot be made safe.
Production judgment and bridge
| Decision surface | Evidence before increasing analytical scope |
|---|---|
| Projection scope | Exact labels/types/properties/filter predicates and projected node/relationship counts match the analytical question. |
| Memory | Projection estimate + algorithm estimate + observed Neo4j heap/process headroom leave safe room for transactional work and GC. |
| Concurrency | Measured throughput/tail latency/CPU under representative contention; never assume max concurrency is optimal. |
| Freshness | Document when the projection was created and what source updates occurred afterward; projections are not live materialized views. |
| Persistence | Explicit decision whether results belong only in stream output, only in the projected graph, or persisted into Neo4j. |
| Failure recovery | Graph can be reconstructed from versioned fixture/query/config; catalog loss on DBMS restart is expected unless an Enterprise persistence feature is deliberately used. |
| Edition/cost | Community core limits are accepted or Enterprise/Aura analytics capabilities are justified by workload, operations and licensing. |
Once resource admission is safe, the remaining question is where results should live. Lesson 4 distinguishes five GDS execution modes so exploration cannot accidentally become persistent data mutation.
Check your understanding
- Does gds.graph.project.estimate include every future algorithm’s working memory?
- Why is maximum concurrency not automatically best?
- What is the Community GDS concurrency ceiling in this release?
- When does a projected graph disappear?
- What is the safest response if the analytical memory envelope does not coexist with production SLOs?
Review the answers
1. No. Projection estimate covers graph projection memory; algorithm estimates must be checked separately.
2. It can increase contention and tail latency for other work; choose it from measured resource/SLO evidence.
3. Four CPU cores.
4. When explicitly dropped, when its source database is stopped/dropped, or when the DBMS stops.
5. Reduce scope/concurrency, queue or move analytics off the transactional path; do not simply force allocation.
Summary and next step
Estimate Memory, Control Concurrency, Manage Graph Lifecycles, and Avoid Production Memory Starvation is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to stream, stats, mutate, write, and estimate Modes: Separating Exploration from Persistent Changes. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- GDS Manual v2026.07 — Current Graph Data Science manual and versioned feature baseline.
- GDS 2026.07 release notes — Current GDS release line; 2026.07.0 is compatible with Neo4j 2026.07.
- Supported Neo4j versions — Compatibility matrix between Neo4j Database and GDS.
- GDS editions and graph catalog — Community/Enterprise boundaries, graph catalog model, all-algorithms availability, concurrency and model-catalog limits.
- Neo4j Server GDS installation — Bundled products-to-plugins installation path and required GDS procedure security configuration.
- GDS on Docker — Container-based installation examples and plugin activation.
- System requirements — Heap/native-memory/CPU guidance and Community maximum concurrency of four.
- Native projection — Current gds.graph.project procedure, label/type/property/orientation configuration and lifecycle.
- Cypher projection — Current gds.graph.project aggregation-function projection from Cypher query context.
- Legacy Cypher projection — deprecated — Deprecated gds.graph.project.cypher procedure and migration boundary.
- Graph creation and schema — Supported projection types, relationship direction, properties, parallel relationships, and algorithm traits.
- Graph catalog operations — Current graph project/list/exists/drop and graph-property operations.
- Listing graphs — Graph metadata, schemaWithOrientation, degree distribution and configuration inspection.
- Memory estimation — Projection and algorithm estimate syntax, requiredMemory and byte-range evidence.
- Algorithm syntax and execution modes — stream/stats/mutate/write/estimate semantics.
- Running algorithms — Operational meaning and tradeoffs of algorithm execution modes.
- PageRank — Production-tier algorithm used only to demonstrate execution modes and memory estimates in this foundations chapter.
- Neo4j current versions — Current Neo4j Database release and 5.26 LTS line.