Chapter 25 · Graph Embeddings and Machine Learning Pipelines with GDS
Node Embeddings from Topology/Properties, FastRP/Node2Vec-Like Concepts, Dimensions, and Reproducibility
Build and inspect AtlasMart FastRP embeddings, distinguish topology/property and inductive/transductive behavior, and make dimensions/randomness/reproducibility explicit before downstream ML.
Learning outcomes
Define graph embeddings as learned/computed numeric representations and separate them from semantic/content embeddings.
Explain how FastRP combines topology, optional numeric node properties, embedding dimension, iteration weights, and randomness.
Distinguish transductive Node2Vec from documented inductive embedding configurations and state the serving consequence.
Use a bounded AtlasMart projection and fixed random seeds without claiming that reproducibility makes an embedding objectively correct.
Choose dimensions and refresh strategy from measured downstream quality, memory, latency, and lifecycle constraints rather than folklore.
1. AtlasMart problem: a graph feature must earn its place
AtlasMart wants to predict whether a customer will purchase again in the next quarter. The existing application already has non-graph attributes such as account tenure and historical spend. A graph embedding is useful only if the pattern of customer–product connectivity adds predictive signal beyond those ordinary properties.
A node embedding is a numeric vector assigned to a node so that selected structural or property relationships are represented in a lower-dimensional feature space. It is not a business truth, a probability, or automatically a semantic language embedding. In this chapter, embeddings are candidate machine-learning features whose value is tested by controlled ablation.
| Dimension | Chapter 25 reproducible assumption |
|---|---|
| Neo4j |
2026.07.1 Community in the disposable local/container lab;
database neo4j.
|
| Cypher | Cypher 25 examples. GDS procedures are called from Cypher; no paid notebook or external ML service is required. |
| Java | Java 21/25 supported by the Neo4j 2026 line; the chosen Neo4j distribution/container supplies the runtime. |
| Auth/TLS |
User neo4j, password
atlasmart-course-2026; loopback Bolt without
TLS only for this isolated lab. Production/remote
deployments require authenticated encrypted transport.
|
| Driver | No application driver is required for mandatory procedure labs; Browser or cypher-shell is sufficient. |
| GDS | GDS Community 2026.07.0. Community includes all algorithms, caps GDS concurrency at four CPU cores, and limits the model catalog to three models. |
| GDS ML quality tiers | Node Classification and Link Prediction pipelines are Beta; Node Regression pipelines are Alpha. Treat tier as a production-risk input. |
| Persistence | Community model/pipeline/graph catalogs are in-memory lifecycle objects. Model persistence to disk is Enterprise-only. |
| APOC | Not required for mandatory Chapter 25 work. |
| Scope |
Only persisted entities with
chapter25=true and in-memory names beginning
atlas-ch25- are created.
|
| Evidence | This generation environment does not run Neo4j/GDS. Fixture counts and split rules are deterministic; training scores, memory estimates, timings and predictions must be measured locally and must not be copied as fabricated results. |
RETURN gds.version() AS gdsVersion;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;
// Inspect current in-memory catalogs before the lab.
CALL gds.graph.list() YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCount ORDER BY graphName;
CALL gds.pipeline.list() YIELD pipelineName, pipelineType
RETURN pipelineName, pipelineType ORDER BY pipelineName;
CALL gds.model.list() YIELD modelName, modelType, loaded, stored, published
RETURN modelName, modelType, loaded, stored, published ORDER BY modelName;
2. FastRP, Node2Vec, topology, properties, and dimensions
| Construct | Mechanism | Boundary that changes the design |
|---|---|---|
| FastRP |
Starts from random/projection vectors and iteratively
mixes neighbor information.
iterationWeights determine how many hops
contribute and their relative weight.
|
More dimensions/iterations can encode more structure but cost memory/CPU and can blur local distinctions. Measure downstream quality rather than maximizing them. |
| FastRP feature properties |
featureProperties and
propertyRatio reserve part of the vector
for numeric node-property information.
|
With propertyRatio=1.0 plus a fixed
randomSeed, current GDS documents FastRP as
inductive; topology-random-vector configurations depend
on database node identity and graph continuity.
|
| Node2Vec | Uses biased random walks plus a skip-gram-style learning objective to place nodes with similar walk contexts near one another. | Current GDS treats Node2Vec as transductive: do not assume a model trained on one graph can embed genuinely new nodes in a different graph with the same coordinate meaning. |
| embeddingDimension | Length of the output vector. | A larger vector can increase representational capacity but also feature width, training cost, memory, serving payload, and overfit risk. |
| randomSeed | Controls documented randomness where supported. | Same seed + same graph/config improves repeatability; it does not neutralize changes in graph contents, node identity, software version, concurrency-sensitive algorithms, or upstream feature engineering. |
| topology vs properties | Topology comes from projected relationships; properties are projected numeric attributes explicitly selected as features. | A feature derived from the target window is leakage even if it enters only through topology. |
// Remove only Chapter 25 persisted fixtures before recreating them.
MATCH (n) WHERE n.chapter25 = true DETACH DELETE n;
// 30 customers. The target is a synthetic future-horizon label for teaching only.
UNWIND range(1,30) AS i
CREATE (:Customer:CH25Customer {
customerId: 'CH25-C-' + right('00' + toString(i), 2),
seq: i,
tenureMonths: toFloat(3 + (i % 18)),
spendIndex: toFloat(20 + ((i * 17) % 80)),
willRepeat: CASE WHEN i % 3 = 0 OR i % 7 = 0 THEN 0 ELSE 1 END,
featureCutoff: date('2026-06-30'),
labelWindow: '2026-Q3',
chapter25: true
});
UNWIND range(1,8) AS j
CREATE (:Product:CH25Product {
productId: 'CH25-P-' + right('00' + toString(j), 2),
categoryCode: toFloat(j % 4),
chapter25: true
});
// Feature-window interactions only. No Q3 target-window interaction is loaded.
MATCH (c:CH25Customer), (p:CH25Product)
WHERE ((c.seq + toInteger(right(p.productId,2))) % 4 = 0)
OR (toInteger(right(p.productId,2)) = ((c.seq - 1) % 8) + 1)
CREATE (c)-[:VIEWED_CH25 {chapter25:true, asOf:'2026-Q2'}]->(p);
MATCH (c:CH25Customer)
RETURN count(c) AS customers,
sum(CASE WHEN c.willRepeat = 1 THEN 1 ELSE 0 END) AS positiveLabels,
sum(CASE WHEN c.willRepeat = 0 THEN 1 ELSE 0 END) AS negativeLabels;
MATCH (:CH25Customer)-[r:VIEWED_CH25]->(:CH25Product)
RETURN count(r) AS featureWindowRelationships;
// Drop atlas-ch25-ml first if it already exists.
CALL gds.graph.project(
'atlas-ch25-ml',
{
CH25Customer: {properties: ['tenureMonths','spendIndex','willRepeat']},
CH25Product: {properties: ['categoryCode']}
},
{
VIEWED_CH25: {orientation:'UNDIRECTED'}
}
)
YIELD graphName, nodeCount, relationshipCount, projectMillis
RETURN graphName, nodeCount, relationshipCount, projectMillis;
CALL gds.graph.list('atlas-ch25-ml')
YIELD graphName, nodeCount, relationshipCount, schema
RETURN graphName, nodeCount, relationshipCount, schema;
3. Observe the embedding directly before putting it inside a model
CALL gds.fastRP.stream('atlas-ch25-ml', {
nodeLabels: ['CH25Customer'],
relationshipTypes: ['VIEWED_CH25'],
embeddingDimension: 16,
iterationWeights: [0.0, 1.0, 1.0],
randomSeed: 42,
concurrency: 4
})
YIELD nodeId, embedding
RETURN gds.util.asNode(nodeId).customerId AS customerId,
size(embedding) AS dimensions,
embedding[0..4] AS firstFourCoordinates
ORDER BY customerId
LIMIT 10;
Every streamed CH25 customer embedding has length 16. The numeric coordinates must be measured on the learner’s installation; this lesson does not invent them. A changed graph/config/version may legitimately change coordinates.
4. Reproducibility is a contract, not just a seed
| Record | Why it belongs in experiment lineage |
|---|---|
| Neo4j/GDS versions | Embedding algorithms and pipeline behavior can evolve. |
| projection query/schema | Determines which nodes, relationships, directions, and properties exist in feature space. |
| feature cutoff | Proves that target-window events were not available to the embedding. |
| embedding config | Dimension, iteration weights, property ratio, feature properties, weights, filters and random seed define the representation. |
| entity identity/version | Transductive embeddings can depend on graph/node identity; restore/reload/rebuild semantics matter. |
| downstream split | An embedding can look stable while evaluation leaks labels or future edges. |
5. Deliberately wrong approach: “Node2Vec scored well once, so serve it to new customers”
The error is not that Node2Vec is weak. The error is assuming an embedding coordinate system is automatically reusable for unseen nodes or a structurally different future graph. Current GDS documentation treats Node2Vec as transductive. A model that consumes those coordinates expects the training graph’s representation regime.
Repair: either keep prediction on the same
graph snapshot/identity regime, retrain the embedding and
downstream model together when the graph changes, or choose an
inductive representation such as property-only FastRP with
propertyRatio=1.0 and a fixed seed (when that
representation matches the task). Verify the repaired design on
an explicit future/holdout slice.
Production judgment
Embedding refresh consumes GDS memory and CPU while the transactional workload still needs page cache, heap, disk and network headroom. A 512-dimensional embedding is not “better” than 32 dimensions unless the measured validation/test benefit justifies the extra computation, model size, write cost, and serving latency. Protect customer/tenant boundaries in the projection; use business identifiers rather than internal node IDs for lineage; keep model and source backups separate from ephemeral in-memory catalogs; and shadow-test an embedding/config/version change before replacing the active feature pipeline.
Check your understanding
- What does an embedding coordinate prove by itself?
- Why is a fixed random seed insufficient for full reproducibility?
- Which current GDS embedding is the production-quality starting point?
- What serving limitation is especially important for Node2Vec?
- What is the bridge to Lesson 2?
Review the answers
1. Nothing about business outcome; it is one learned/computed representation coordinate under a specific graph/configuration.
2. Graph contents, projection, entity identity, versions, features and split protocol can still change.
3. FastRP.
4. It is transductive; prediction/generalization across genuinely different graphs/new nodes must not be assumed.
5. Embeddings become one feature source among many, so leakage and feature lineage become the primary correctness risk.
Summary and next step
Node Embeddings from Topology/Properties, FastRP/Node2Vec-Like Concepts, Dimensions, and Reproducibility is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Feature Engineering from Graph Algorithms, Properties, Labels, and Leakage Prevention. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- GDS Manual v2026.07 — Current GDS manual and release baseline.
- GDS supported Neo4j versions — Compatibility matrix for Neo4j and GDS releases.
- GDS editions — Community/Enterprise limits including four-core concurrency and three-model catalog capacity.
- Node embeddings overview — Current embedding families, quality tiers, inductive/transductive guidance.
- Fast Random Projection — FastRP dimensions, propertyRatio, featureProperties, randomSeed, iterations and modes.
- Node2Vec — Current Node2Vec algorithm and transductive embedding behavior.
- Machine learning pipelines — Current node classification, link prediction and node regression pipeline tiers.
- Node classification pipelines — End-to-end node classification semantics and prediction model behavior.
- Link prediction pipelines — Feature/train/test split semantics and negative examples.
- Link prediction configuration — Split configuration, FastRP property steps, link features and candidate models.
- Link prediction training — Cross-validation, model selection and model-catalog registration semantics.
- Training methods — Supported classification/regression trainers and auto-tuning.
- Pipeline catalog — Pipeline catalog list/exists/drop lifecycle.
- Model catalog listing — Model metadata, training config, schema and loaded/stored/published state.
- Store models on disk — Enterprise-only persistent model storage semantics.
- Getting started ML pipeline — Current link-prediction pipeline example and split workflow.
- GDS server installation — Bundled plugin installation and procedure configuration.
- Neo4j GDS release notes — Current GDS 2026.07.0 release line.
- Neo4j current versions — Current Neo4j database release and LTS line.