Chapter 23 · Graph Data Science Foundations: Projections, Graph Catalog, Memory, and Execution Modes

stream, stats, mutate, write, and estimate Modes: Separating Exploration from Persistent Changes

Use stream, stats, mutate, write, and estimate deliberately: understand exactly where each result exists, which modes persist nothing, which change only the catalog, and which write back to Neo4j.

Advanced240–340 minutesExecution-mode labNeo4j 2026.07.1 · Community mandatoryGDS Community 2026.07.0 · Cypher 25Projection catalog · memory estimates · concurrency ≤4 CEJava 21/25 · GDS plugin requiredLast reviewed: September 2026

Learning outcomes

01

State exactly where stream, stats, mutate, write, and estimate results exist.

02

Use PageRank only as a controlled vehicle for learning mode semantics, not as an unexplained business score.

03

Verify that mutate changes the projected graph but not Neo4j persistent properties.

04

Verify that write persists data and therefore requires transactional/change-management controls.

05

Choose the least-persistent mode that satisfies the analytical workflow.

1. Execution mode is a data-governance decision

AtlasMart has accumulated enough connected customer/product behavior that analysts want graph algorithms, but running “an algorithm” is not one operation. The analytical graph must first be selected, copied into a GDS representation, sized, verified, executed, and removed or refreshed deliberately. AtlasMart analysts need to explore centrality, chain algorithms, and eventually publish one feature. If every experiment writes properties to Product/Customer nodes, the operational graph becomes an accidental scratchpad.

The safety rule for this chapter is therefore: treat projection definition, memory allocation, algorithm mode, and lifecycle as part of the analytical result. A score without those inputs is not reproducible evidence.

2. Five modes, five state boundaries

Chapter 23 baseline · reviewed 9 September 2026

Current Neo4j Database is 2026.07.1; current GDS is 2026.07.0 for the Neo4j 2026.07 line. Mandatory work uses self-managed Neo4j Community 2026.07.1 + GDS Community 2026.07.0, database neo4j, user neo4j, disposable password atlasmart-course-2026, loopback HTTP 7474/Bolt 7687, and CYPHER 25 for database-side fixture work. Neo4j 2026.x supports Java 21/25. GDS is the only plugin required in this chapter; APOC is not required.

Current GDS Community boundary

GDS Community includes the algorithm library needed for this course. Its execution concurrency is capped at 4 CPU cores and its model catalog is capped at 3 models. GDS Enterprise removes the CPU-core cap and adds capabilities such as graph backup/restore, Arrow import/export, cluster write support, capacity/load monitoring, and extended model-catalog persistence/sharing. Those Enterprise capabilities are discussed only as edition boundaries; no mandatory lab depends on them.

Projection API status

Two current projection styles matter. Native projection uses the gds.graph.project(...) procedure and remains supported in 2026.07, but current GDS documentation says native projection will be deprecated in a future release and increasingly uses Cypher projection as the norm. Current Cypher projection calls the gds.graph.project(...) aggregation function from a Cypher query. The older gds.graph.project.cypher(...) procedure is already deprecated. This chapter teaches native projection because it makes schema/orientation configuration explicit, then shows the current Cypher form learners should prefer for new flexible projections.

Assumption Value / boundary
Server Neo4j Community 2026.07.1; Java 21/25 supported. Disposable container atlasmart-gds.
GDS GDS Community 2026.07.0. Verify with gds.version(); do not continue if compatibility differs.
Database/security neo4j database; local disposable neo4j user; loopback transport only. Production credentials/TLS/RBAC differ.
Fixture 6 Customer + 6 Product nodes; 14 VIEWED + 4 PURCHASED relationships; numeric weight/customerValue/margin.
Memory/concurrency Learner records estimates/observations. Examples use concurrency=2; Community maximum is 4, but 2 is a lab choice, not a recommendation.
Plugins GDS only. APOC is not required.
Runtime claims Artifact generation does not execute Neo4j/GDS. Expected deterministic graph counts are fixture-derived; memory/timing/algorithm scores must be measured by the learner.
Mode Runs computation? Returns Changes projected graph? Changes Neo4j store?
estimate No final computation Memory/size estimate No No
stream Yes Per-entity result rows No No
stats Yes Summary/statistical row No No
mutate Yes Summary + writes an analytical property/relationship into named graph Yes No
write Yes Summary + persists configured result Not the purpose Yes

3. Prepare and estimate before running

Fixture + projection
CYPHER 25
// Chapter 23 fixture: bounded and safe to delete by labTag.
CREATE CONSTRAINT ch23_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch23_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;

UNWIND [
 {id:'C-2301',segment:'LOYAL',value:92.0},
 {id:'C-2302',segment:'LOYAL',value:78.0},
 {id:'C-2303',segment:'NEW',value:35.0},
 {id:'C-2304',segment:'NEW',value:28.0},
 {id:'C-2305',segment:'B2B',value:96.0},
 {id:'C-2306',segment:'B2B',value:84.0}
] AS row
MERGE (c:Customer {customerId:row.id})
SET c.segment=row.segment,c.customerValue=row.value,c.labTag='ch23';

UNWIND [
 {id:'P-2301',name:'Trail Camera Pro',margin:0.31},
 {id:'P-2302',name:'Action Camera 4K',margin:0.27},
 {id:'P-2303',name:'Solar Trail Charger',margin:0.24},
 {id:'P-2304',name:'Indoor Security Camera',margin:0.29},
 {id:'P-2305',name:'Hydration Vest',margin:0.22},
 {id:'P-2306',name:'Wildlife Field Guide',margin:0.35}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name,p.margin=row.margin,p.labTag='ch23';

MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WITH c,p WHERE
 (c.customerId='C-2301' AND p.productId IN ['P-2301','P-2303','P-2306']) OR
 (c.customerId='C-2302' AND p.productId IN ['P-2301','P-2302']) OR
 (c.customerId='C-2303' AND p.productId IN ['P-2302','P-2305']) OR
 (c.customerId='C-2304' AND p.productId IN ['P-2304','P-2305']) OR
 (c.customerId='C-2305' AND p.productId IN ['P-2301','P-2303','P-2304']) OR
 (c.customerId='C-2306' AND p.productId IN ['P-2303','P-2306'])
MERGE (c)-[r:VIEWED]->(p)
SET r.weight = CASE c.segment WHEN 'LOYAL' THEN 2.0 WHEN 'B2B' THEN 1.5 ELSE 1.0 END,
    r.labTag='ch23';

MATCH (c:Customer {labTag:'ch23'}),(p:Product {labTag:'ch23'})
WHERE (c.customerId='C-2301' AND p.productId='P-2301') OR
      (c.customerId='C-2302' AND p.productId='P-2302') OR
      (c.customerId='C-2305' AND p.productId='P-2303') OR
      (c.customerId='C-2306' AND p.productId='P-2306')
MERGE (c)-[r:PURCHASED]->(p)
SET r.weight=3.0,r.labTag='ch23';


CALL gds.graph.exists('atlasmart-ch23') YIELD exists
WITH exists WHERE exists
CALL gds.graph.drop('atlasmart-ch23') YIELD graphName
RETURN graphName;


CALL gds.graph.project(
  'atlasmart-ch23',
  {
    Customer:{properties:['customerValue']},
    Product:{properties:['margin']}
  },
  {
    VIEWED:{orientation:'NATURAL',properties:['weight']},
    PURCHASED:{orientation:'NATURAL',properties:['weight']}
  },
  {readConcurrency:2}
)
YIELD graphName,nodeCount,relationshipCount,projectMillis,configuration
RETURN graphName,nodeCount,relationshipCount,projectMillis,configuration;
// Expected deterministic counts: nodeCount=12, relationshipCount=18.
Estimate PageRank stream mode
CALL gds.pageRank.stream.estimate('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  concurrency:2,
  maxIterations:20,
  dampingFactor:0.85
})
YIELD nodeCount,relationshipCount,requiredMemory,bytesMin,bytesMax,heapPercentageMin,heapPercentageMax
RETURN *;

PageRank is used here because it exposes all teaching modes. Chapter 24 will teach what centrality means and when it is appropriate. In this chapter, the score is just a result payload whose state boundary must be controlled.

4. stream: inspect individual results without changing graph state

PageRank stream
CALL gds.pageRank.stream('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  concurrency:2,
  maxIterations:20,
  dampingFactor:0.85
})
YIELD nodeId,score
RETURN gds.util.asNode(nodeId).productId AS productId,
       gds.util.asNode(nodeId).customerId AS customerId,
       score
ORDER BY score DESC LIMIT 8;
// stream returns rows; it does not alter graph catalog properties or Neo4j properties.

Stream is ideal for exploration, evaluation, external post-processing, or a service response when result cardinality is bounded. Large streams can themselves create network/client-memory pressure, so use top-N/filtering/consumer backpressure where appropriate.

5. stats: prove behavior without moving all scores

PageRank stats
CALL gds.pageRank.stats('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  concurrency:2,
  maxIterations:20
})
YIELD ranIterations,didConverge,centralityDistribution,computeMillis
RETURN *;
// stats returns summary evidence without exposing per-node scores or writing state.

Stats mode answers “what happened during the computation?” with summary evidence such as convergence/distribution/timing. It does not give per-node scores and does not write them anywhere.

6. mutate: chain analytics inside the catalog

Mutate PageRank into projected graph
CALL gds.pageRank.mutate('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  mutateProperty:'ch23PageRank',
  concurrency:2,
  maxIterations:20
})
YIELD nodePropertiesWritten,mutateMillis,ranIterations,didConverge
RETURN *;

CALL gds.graph.nodeProperty.stream('atlasmart-ch23','ch23PageRank')
YIELD nodeId,propertyValue
RETURN gds.util.asNode(nodeId).productId AS productId,
       gds.util.asNode(nodeId).customerId AS customerId,
       propertyValue
ORDER BY propertyValue DESC LIMIT 8;
// ch23PageRank exists only in the in-memory projected graph after mutate.
Prove mutate did not persist to Neo4j
CYPHER 25
MATCH (n {labTag:'ch23'})
RETURN count(n) AS sourceNodes,
       count(n.ch23PageRank) AS persistedPageRankProperties;
// Expected persistedPageRankProperties=0 immediately after mutate.

Mutate is powerful because a later GDS algorithm can consume the new in-memory property without round-tripping through the database. It is also ephemeral: dropping the graph or stopping the source database/DBMS removes that analytical state.

7. write: persist intentionally, then verify

Persist selected algorithm output
CALL gds.pageRank.write('atlasmart-ch23', {
  relationshipTypes:['VIEWED','PURCHASED'],
  relationshipWeightProperty:'weight',
  writeProperty:'ch23PageRankWritten',
  concurrency:2,
  maxIterations:20
})
YIELD nodePropertiesWritten,writeMillis,ranIterations,didConverge
RETURN *;

MATCH (n {labTag:'ch23'})
RETURN labels(n) AS labels,count(n.ch23PageRankWritten) AS persisted
ORDER BY labels;
// write changes Neo4j; cleanup must remove this lab property.
Write is a production change

A GDS write mode changes the transactional database. Apply normal schema/property ownership, security, backup/recovery, rollout/rollback, CDC/downstream-consumer, and query/index considerations. Never use write merely because it is “faster than exporting rows.”

8. Deliberately wrong: confuse mutate with durable feature engineering

Wrong approach

Run gds.pageRank.mutate(... mutateProperty:'rank'), restart Neo4j, and expect application Cypher to read n.rank from persisted nodes.

Concrete problem: mutate wrote only the named in-memory graph. After projection/DBMS loss the property is gone, and ordinary Cypher never had a persisted rank property. Repair: decide whether the feature is temporary (mutate/stream) or governed durable data (write/export + controlled transaction), then verify the correct storage layer.

9. Mode selection under production constraints

Need Prefer Reason
Explore top candidates stream No persistence; easy to compare/evaluate.
Check convergence/distribution stats Avoids transferring all node results.
Feed next GDS step mutate Keeps intermediate feature in analytical graph.
Publish governed feature to application graph write only after review Persists result and therefore creates operational/schema coupling.
Decide if job is admissible estimate Tests memory before computation.

Production judgment and bridge

Decision surface Evidence before increasing analytical scope
Projection scope Exact labels/types/properties/filter predicates and projected node/relationship counts match the analytical question.
Memory Projection estimate + algorithm estimate + observed Neo4j heap/process headroom leave safe room for transactional work and GC.
Concurrency Measured throughput/tail latency/CPU under representative contention; never assume max concurrency is optimal.
Freshness Document when the projection was created and what source updates occurred afterward; projections are not live materialized views.
Persistence Explicit decision whether results belong only in stream output, only in the projected graph, or persisted into Neo4j.
Failure recovery Graph can be reconstructed from versioned fixture/query/config; catalog loss on DBMS restart is expected unless an Enterprise persistence feature is deliberately used.
Edition/cost Community core limits are accepted or Enterprise/Aura analytics capabilities are justified by workload, operations and licensing.

Lesson 5 assembles these pieces into one auditable runbook: verify compatible plugin, seed source, estimate, project, inspect, run least-persistent modes, measure, prove state boundaries, and clean up.

Check your understanding

  1. Which mode changes only the GDS named graph?
  2. Which mode persists algorithm output into Neo4j?
  3. Does stats expose every node score?
  4. Why can stream still be expensive?
  5. What should happen before write mode in production?
Review the answers

1. mutate.

2. write.

3. No; it returns summary statistics.

4. Large result cardinality can consume network/client memory and serialization time even though it does not persist state.

5. Treat it as a governed database change: ownership, security, backup/recovery, rollback, downstream effects and verification.

Summary and next step

stream, stats, mutate, write, and estimate Modes: Separating Exploration from Persistent Changes is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Build, Validate, Use, and Drop a GDS Projection with Reproducible Resource Measurements. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.