Chapter 25 · Graph Embeddings and Machine Learning Pipelines with GDS

Node Embeddings from Topology/Properties, FastRP/Node2Vec-Like Concepts, Dimensions, and Reproducibility

Build and inspect AtlasMart FastRP embeddings, distinguish topology/property and inductive/transductive behavior, and make dimensions/randomness/reproducibility explicit before downstream ML.

Advanced260–380 minutesEmbeddings · FastRP · Node2VecNeo4j 2026.07.1 · Community mandatoryGDS Community 2026.07.0 · Cypher 25GDS ML · concurrency ≤4 CE · model catalog ≤3Java 21/25 · GDS plugin requiredLast reviewed: September 2026

Learning outcomes

01

Define graph embeddings as learned/computed numeric representations and separate them from semantic/content embeddings.

02

Explain how FastRP combines topology, optional numeric node properties, embedding dimension, iteration weights, and randomness.

03

Distinguish transductive Node2Vec from documented inductive embedding configurations and state the serving consequence.

04

Use a bounded AtlasMart projection and fixed random seeds without claiming that reproducibility makes an embedding objectively correct.

05

Choose dimensions and refresh strategy from measured downstream quality, memory, latency, and lifecycle constraints rather than folklore.

1. AtlasMart problem: a graph feature must earn its place

AtlasMart wants to predict whether a customer will purchase again in the next quarter. The existing application already has non-graph attributes such as account tenure and historical spend. A graph embedding is useful only if the pattern of customer–product connectivity adds predictive signal beyond those ordinary properties.

A node embedding is a numeric vector assigned to a node so that selected structural or property relationships are represented in a lower-dimensional feature space. It is not a business truth, a probability, or automatically a semantic language embedding. In this chapter, embeddings are candidate machine-learning features whose value is tested by controlled ablation.

Dimension Chapter 25 reproducible assumption
Neo4j 2026.07.1 Community in the disposable local/container lab; database neo4j.
Cypher Cypher 25 examples. GDS procedures are called from Cypher; no paid notebook or external ML service is required.
Java Java 21/25 supported by the Neo4j 2026 line; the chosen Neo4j distribution/container supplies the runtime.
Auth/TLS User neo4j, password atlasmart-course-2026; loopback Bolt without TLS only for this isolated lab. Production/remote deployments require authenticated encrypted transport.
Driver No application driver is required for mandatory procedure labs; Browser or cypher-shell is sufficient.
GDS GDS Community 2026.07.0. Community includes all algorithms, caps GDS concurrency at four CPU cores, and limits the model catalog to three models.
GDS ML quality tiers Node Classification and Link Prediction pipelines are Beta; Node Regression pipelines are Alpha. Treat tier as a production-risk input.
Persistence Community model/pipeline/graph catalogs are in-memory lifecycle objects. Model persistence to disk is Enterprise-only.
APOC Not required for mandatory Chapter 25 work.
Scope Only persisted entities with chapter25=true and in-memory names beginning atlas-ch25- are created.
Evidence This generation environment does not run Neo4j/GDS. Fixture counts and split rules are deterministic; training scores, memory estimates, timings and predictions must be measured locally and must not be copied as fabricated results.
Verify the Chapter 25 runtime
RETURN gds.version() AS gdsVersion;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;

// Inspect current in-memory catalogs before the lab.
CALL gds.graph.list() YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCount ORDER BY graphName;

CALL gds.pipeline.list() YIELD pipelineName, pipelineType
RETURN pipelineName, pipelineType ORDER BY pipelineName;

CALL gds.model.list() YIELD modelName, modelType, loaded, stored, published
RETURN modelName, modelType, loaded, stored, published ORDER BY modelName;

2. FastRP, Node2Vec, topology, properties, and dimensions

Construct Mechanism Boundary that changes the design
FastRP Starts from random/projection vectors and iteratively mixes neighbor information. iterationWeights determine how many hops contribute and their relative weight. More dimensions/iterations can encode more structure but cost memory/CPU and can blur local distinctions. Measure downstream quality rather than maximizing them.
FastRP feature properties featureProperties and propertyRatio reserve part of the vector for numeric node-property information. With propertyRatio=1.0 plus a fixed randomSeed, current GDS documents FastRP as inductive; topology-random-vector configurations depend on database node identity and graph continuity.
Node2Vec Uses biased random walks plus a skip-gram-style learning objective to place nodes with similar walk contexts near one another. Current GDS treats Node2Vec as transductive: do not assume a model trained on one graph can embed genuinely new nodes in a different graph with the same coordinate meaning.
embeddingDimension Length of the output vector. A larger vector can increase representational capacity but also feature width, training cost, memory, serving payload, and overfit risk.
randomSeed Controls documented randomness where supported. Same seed + same graph/config improves repeatability; it does not neutralize changes in graph contents, node identity, software version, concurrency-sensitive algorithms, or upstream feature engineering.
topology vs properties Topology comes from projected relationships; properties are projected numeric attributes explicitly selected as features. A feature derived from the target window is leakage even if it enters only through topology.
Create the bounded AtlasMart graph-ML fixture
// Remove only Chapter 25 persisted fixtures before recreating them.
MATCH (n) WHERE n.chapter25 = true DETACH DELETE n;

// 30 customers. The target is a synthetic future-horizon label for teaching only.
UNWIND range(1,30) AS i
CREATE (:Customer:CH25Customer {
  customerId: 'CH25-C-' + right('00' + toString(i), 2),
  seq: i,
  tenureMonths: toFloat(3 + (i % 18)),
  spendIndex: toFloat(20 + ((i * 17) % 80)),
  willRepeat: CASE WHEN i % 3 = 0 OR i % 7 = 0 THEN 0 ELSE 1 END,
  featureCutoff: date('2026-06-30'),
  labelWindow: '2026-Q3',
  chapter25: true
});

UNWIND range(1,8) AS j
CREATE (:Product:CH25Product {
  productId: 'CH25-P-' + right('00' + toString(j), 2),
  categoryCode: toFloat(j % 4),
  chapter25: true
});

// Feature-window interactions only. No Q3 target-window interaction is loaded.
MATCH (c:CH25Customer), (p:CH25Product)
WHERE ((c.seq + toInteger(right(p.productId,2))) % 4 = 0)
   OR (toInteger(right(p.productId,2)) = ((c.seq - 1) % 8) + 1)
CREATE (c)-[:VIEWED_CH25 {chapter25:true, asOf:'2026-Q2'}]->(p);

MATCH (c:CH25Customer)
RETURN count(c) AS customers,
       sum(CASE WHEN c.willRepeat = 1 THEN 1 ELSE 0 END) AS positiveLabels,
       sum(CASE WHEN c.willRepeat = 0 THEN 1 ELSE 0 END) AS negativeLabels;

MATCH (:CH25Customer)-[r:VIEWED_CH25]->(:CH25Product)
RETURN count(r) AS featureWindowRelationships;
Project only the pre-label-horizon graph
// Drop atlas-ch25-ml first if it already exists.
CALL gds.graph.project(
  'atlas-ch25-ml',
  {
    CH25Customer: {properties: ['tenureMonths','spendIndex','willRepeat']},
    CH25Product: {properties: ['categoryCode']}
  },
  {
    VIEWED_CH25: {orientation:'UNDIRECTED'}
  }
)
YIELD graphName, nodeCount, relationshipCount, projectMillis
RETURN graphName, nodeCount, relationshipCount, projectMillis;

CALL gds.graph.list('atlas-ch25-ml')
YIELD graphName, nodeCount, relationshipCount, schema
RETURN graphName, nodeCount, relationshipCount, schema;

3. Observe the embedding directly before putting it inside a model

FastRP stream: topology-only diagnostic
CALL gds.fastRP.stream('atlas-ch25-ml', {
  nodeLabels: ['CH25Customer'],
  relationshipTypes: ['VIEWED_CH25'],
  embeddingDimension: 16,
  iterationWeights: [0.0, 1.0, 1.0],
  randomSeed: 42,
  concurrency: 4
})
YIELD nodeId, embedding
RETURN gds.util.asNode(nodeId).customerId AS customerId,
       size(embedding) AS dimensions,
       embedding[0..4] AS firstFourCoordinates
ORDER BY customerId
LIMIT 10;
Expected invariant, not fabricated values

Every streamed CH25 customer embedding has length 16. The numeric coordinates must be measured on the learner’s installation; this lesson does not invent them. A changed graph/config/version may legitimately change coordinates.

4. Reproducibility is a contract, not just a seed

Record Why it belongs in experiment lineage
Neo4j/GDS versions Embedding algorithms and pipeline behavior can evolve.
projection query/schema Determines which nodes, relationships, directions, and properties exist in feature space.
feature cutoff Proves that target-window events were not available to the embedding.
embedding config Dimension, iteration weights, property ratio, feature properties, weights, filters and random seed define the representation.
entity identity/version Transductive embeddings can depend on graph/node identity; restore/reload/rebuild semantics matter.
downstream split An embedding can look stable while evaluation leaks labels or future edges.

5. Deliberately wrong approach: “Node2Vec scored well once, so serve it to new customers”

The error is not that Node2Vec is weak. The error is assuming an embedding coordinate system is automatically reusable for unseen nodes or a structurally different future graph. Current GDS documentation treats Node2Vec as transductive. A model that consumes those coordinates expects the training graph’s representation regime.

Repair: either keep prediction on the same graph snapshot/identity regime, retrain the embedding and downstream model together when the graph changes, or choose an inductive representation such as property-only FastRP with propertyRatio=1.0 and a fixed seed (when that representation matches the task). Verify the repaired design on an explicit future/holdout slice.

Production judgment

Embedding refresh consumes GDS memory and CPU while the transactional workload still needs page cache, heap, disk and network headroom. A 512-dimensional embedding is not “better” than 32 dimensions unless the measured validation/test benefit justifies the extra computation, model size, write cost, and serving latency. Protect customer/tenant boundaries in the projection; use business identifiers rather than internal node IDs for lineage; keep model and source backups separate from ephemeral in-memory catalogs; and shadow-test an embedding/config/version change before replacing the active feature pipeline.

Check your understanding

  1. What does an embedding coordinate prove by itself?
  2. Why is a fixed random seed insufficient for full reproducibility?
  3. Which current GDS embedding is the production-quality starting point?
  4. What serving limitation is especially important for Node2Vec?
  5. What is the bridge to Lesson 2?
Review the answers

1. Nothing about business outcome; it is one learned/computed representation coordinate under a specific graph/configuration.

2. Graph contents, projection, entity identity, versions, features and split protocol can still change.

3. FastRP.

4. It is transductive; prediction/generalization across genuinely different graphs/new nodes must not be assumed.

5. Embeddings become one feature source among many, so leakage and feature lineage become the primary correctness risk.

Summary and next step

Node Embeddings from Topology/Properties, FastRP/Node2Vec-Like Concepts, Dimensions, and Reproducibility is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Feature Engineering from Graph Algorithms, Properties, Labels, and Leakage Prevention. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.