Chapter 25 · Graph Embeddings and Machine Learning Pipelines with GDS
Model Catalog, Training, Hyperparameter Search, Prediction, Persistence, and Production Boundaries
Train bounded baseline and graph models in the current GDS model catalog, inspect tuning/prediction lifecycle, and separate Community in-memory state from Enterprise model persistence.
Learning outcomes
Distinguish the pipeline catalog from the trained model catalog and explain what survives restart in Community.
Configure bounded hyperparameter search and inspect winning parameters without confusing tuning with evaluation.
Train and list baseline/graph models within the Community three-model limit.
Use prediction stream/mutate/write modes deliberately and document persistent-write/idempotency consequences.
Explain Enterprise-only model persistence and define a production retraining/serving/rollback boundary.
1. Pipeline object, trained model, persisted model: three lifecycles
A training pipeline is configurable workflow state in the pipeline catalog. Training it creates a prediction model in the model catalog. The model catalog is in-memory lifecycle state; current GDS Community limits it to three models. Enterprise adds persistent model storage and sharing/publishing capabilities. A database backup does not magically turn an in-memory Community model into a durable artifact.
| Dimension | Chapter 25 reproducible assumption |
|---|---|
| Neo4j |
2026.07.1 Community in the disposable local/container lab;
database neo4j.
|
| Cypher | Cypher 25 examples. GDS procedures are called from Cypher; no paid notebook or external ML service is required. |
| Java | Java 21/25 supported by the Neo4j 2026 line; the chosen Neo4j distribution/container supplies the runtime. |
| Auth/TLS |
User neo4j, password
atlasmart-course-2026; loopback Bolt without
TLS only for this isolated lab. Production/remote
deployments require authenticated encrypted transport.
|
| Driver | No application driver is required for mandatory procedure labs; Browser or cypher-shell is sufficient. |
| GDS | GDS Community 2026.07.0. Community includes all algorithms, caps GDS concurrency at four CPU cores, and limits the model catalog to three models. |
| GDS ML quality tiers | Node Classification and Link Prediction pipelines are Beta; Node Regression pipelines are Alpha. Treat tier as a production-risk input. |
| Persistence | Community model/pipeline/graph catalogs are in-memory lifecycle objects. Model persistence to disk is Enterprise-only. |
| APOC | Not required for mandatory Chapter 25 work. |
| Scope |
Only persisted entities with
chapter25=true and in-memory names beginning
atlas-ch25- are created.
|
| Evidence | This generation environment does not run Neo4j/GDS. Fixture counts and split rules are deterministic; training scores, memory estimates, timings and predictions must be measured locally and must not be copied as fabricated results. |
RETURN gds.version() AS gdsVersion;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;
// Inspect current in-memory catalogs before the lab.
CALL gds.graph.list() YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCount ORDER BY graphName;
CALL gds.pipeline.list() YIELD pipelineName, pipelineType
RETURN pipelineName, pipelineType ORDER BY pipelineName;
CALL gds.model.list() YIELD modelName, modelType, loaded, stored, published
RETURN modelName, modelType, loaded, stored, published ORDER BY modelName;
// Remove only Chapter 25 persisted fixtures before recreating them.
MATCH (n) WHERE n.chapter25 = true DETACH DELETE n;
// 30 customers. The target is a synthetic future-horizon label for teaching only.
UNWIND range(1,30) AS i
CREATE (:Customer:CH25Customer {
customerId: 'CH25-C-' + right('00' + toString(i), 2),
seq: i,
tenureMonths: toFloat(3 + (i % 18)),
spendIndex: toFloat(20 + ((i * 17) % 80)),
willRepeat: CASE WHEN i % 3 = 0 OR i % 7 = 0 THEN 0 ELSE 1 END,
featureCutoff: date('2026-06-30'),
labelWindow: '2026-Q3',
chapter25: true
});
UNWIND range(1,8) AS j
CREATE (:Product:CH25Product {
productId: 'CH25-P-' + right('00' + toString(j), 2),
categoryCode: toFloat(j % 4),
chapter25: true
});
// Feature-window interactions only. No Q3 target-window interaction is loaded.
MATCH (c:CH25Customer), (p:CH25Product)
WHERE ((c.seq + toInteger(right(p.productId,2))) % 4 = 0)
OR (toInteger(right(p.productId,2)) = ((c.seq - 1) % 8) + 1)
CREATE (c)-[:VIEWED_CH25 {chapter25:true, asOf:'2026-Q2'}]->(p);
MATCH (c:CH25Customer)
RETURN count(c) AS customers,
sum(CASE WHEN c.willRepeat = 1 THEN 1 ELSE 0 END) AS positiveLabels,
sum(CASE WHEN c.willRepeat = 0 THEN 1 ELSE 0 END) AS negativeLabels;
MATCH (:CH25Customer)-[r:VIEWED_CH25]->(:CH25Product)
RETURN count(r) AS featureWindowRelationships;
// Drop atlas-ch25-ml first if it already exists.
CALL gds.graph.project(
'atlas-ch25-ml',
{
CH25Customer: {properties: ['tenureMonths','spendIndex','willRepeat']},
CH25Product: {properties: ['categoryCode']}
},
{
VIEWED_CH25: {orientation:'UNDIRECTED'}
}
)
YIELD graphName, nodeCount, relationshipCount, projectMillis
RETURN graphName, nodeCount, relationshipCount, projectMillis;
CALL gds.graph.list('atlas-ch25-ml')
YIELD graphName, nodeCount, relationshipCount, schema
RETURN graphName, nodeCount, relationshipCount, schema;
2. Train the two-model ablation without exceeding the Community catalog limit
CALL gds.beta.pipeline.nodeClassification.create('atlas-ch25-baseline-pipe');
CALL gds.beta.pipeline.nodeClassification.selectFeatures('atlas-ch25-baseline-pipe',['tenureMonths','spendIndex']);
CALL gds.beta.pipeline.nodeClassification.configureSplit('atlas-ch25-baseline-pipe',{testFraction:0.2,validationFolds:3});
CALL gds.beta.pipeline.nodeClassification.configureAutoTuning('atlas-ch25-baseline-pipe',{maxTrials:5});
CALL gds.beta.pipeline.nodeClassification.addLogisticRegression(
'atlas-ch25-baseline-pipe',
{penalty:{range:[0.0001,1.0]}}
);
CALL gds.beta.pipeline.nodeClassification.train.estimate(
'atlas-ch25-ml',
{
pipeline:'atlas-ch25-baseline-pipe',
modelName:'atlas-ch25-baseline-model',
targetNodeLabels:['CH25Customer'],
targetProperty:'willRepeat',
metrics:['F1_WEIGHTED'],
randomSeed:42,
concurrency:4
}
)
YIELD requiredMemory, treeView
RETURN requiredMemory, treeView;
CALL gds.beta.pipeline.nodeClassification.train(
'atlas-ch25-ml',
{
pipeline:'atlas-ch25-baseline-pipe',
modelName:'atlas-ch25-baseline-model',
targetNodeLabels:['CH25Customer'],
targetProperty:'willRepeat',
metrics:['F1_WEIGHTED'],
randomSeed:42,
concurrency:4
}
)
YIELD modelInfo, modelSelectionStats, trainMillis
RETURN modelInfo.bestParameters AS bestParameters,
modelInfo.metrics AS metrics,
size(modelSelectionStats.modelCandidates) AS evaluatedCandidates,
trainMillis;
3. Train the graph-augmented model
CALL gds.beta.pipeline.nodeClassification.create('atlas-ch25-graph-pipe');
CALL gds.beta.pipeline.nodeClassification.addNodeProperty(
'atlas-ch25-graph-pipe','fastRP',
{
mutateProperty:'ch25Embedding',
embeddingDimension:16,
iterationWeights:[0.0,1.0,1.0],
randomSeed:42,
contextNodeLabels:['CH25Customer','CH25Product'],
contextRelationshipTypes:['VIEWED_CH25']
}
);
CALL gds.beta.pipeline.nodeClassification.selectFeatures(
'atlas-ch25-graph-pipe',['tenureMonths','spendIndex','ch25Embedding']
);
CALL gds.beta.pipeline.nodeClassification.configureSplit('atlas-ch25-graph-pipe',{testFraction:0.2,validationFolds:3});
CALL gds.beta.pipeline.nodeClassification.configureAutoTuning('atlas-ch25-graph-pipe',{maxTrials:5});
CALL gds.beta.pipeline.nodeClassification.addLogisticRegression(
'atlas-ch25-graph-pipe',{penalty:{range:[0.0001,1.0]}}
);
CALL gds.beta.pipeline.nodeClassification.train.estimate(
'atlas-ch25-ml',
{
pipeline:'atlas-ch25-graph-pipe',
modelName:'atlas-ch25-graph-model',
targetNodeLabels:['CH25Customer'],
targetProperty:'willRepeat',
metrics:['F1_WEIGHTED'],
randomSeed:42,
concurrency:4
}
)
YIELD requiredMemory, treeView
RETURN requiredMemory, treeView;
CALL gds.beta.pipeline.nodeClassification.train(
'atlas-ch25-ml',
{
pipeline:'atlas-ch25-graph-pipe',
modelName:'atlas-ch25-graph-model',
targetNodeLabels:['CH25Customer'],
targetProperty:'willRepeat',
metrics:['F1_WEIGHTED'],
randomSeed:42,
concurrency:4
}
)
YIELD modelInfo, modelSelectionStats, trainMillis
RETURN modelInfo.bestParameters AS bestParameters,
modelInfo.metrics AS metrics,
size(modelSelectionStats.modelCandidates) AS evaluatedCandidates,
trainMillis;
Exactly two Chapter 25 prediction models should exist after the mandatory lab—within the Community three-model limit. Record actual test metrics and training times; do not substitute values from this HTML.
4. Inspect catalog state and prediction lifecycle
CALL gds.pipeline.list()
YIELD pipelineName, pipelineType, creationTime
WHERE pipelineName STARTS WITH 'atlas-ch25-'
RETURN pipelineName, pipelineType, creationTime ORDER BY pipelineName;
CALL gds.model.list()
YIELD modelName, modelType, modelInfo, trainConfig, graphSchema, loaded, stored, published
WHERE modelName STARTS WITH 'atlas-ch25-'
RETURN modelName, modelType, loaded, stored, published,
modelInfo.metrics AS metrics,
trainConfig,
graphSchema
ORDER BY modelName;
CALL gds.beta.pipeline.nodeClassification.predict.stream(
'atlas-ch25-ml',
{
modelName:'atlas-ch25-graph-model',
targetNodeLabels:['CH25Customer'],
includePredictedProbabilities:true,
concurrency:4
}
)
YIELD nodeId, predictedClass, predictedProbabilities
RETURN gds.util.asNode(nodeId).customerId AS customerId,
predictedClass,
predictedProbabilities
ORDER BY customerId
LIMIT 15;
5. stream, mutate, write, and persistence boundaries
| Operation | State changed | Use with care because… |
|---|---|---|
| predict.stream | No graph/store mutation; rows returned to client | Large result sets can move memory/network pressure to the client. |
| predict.mutate | Adds predicted properties to the in-memory GDS graph | Ephemeral state can be mistaken for persisted application truth. |
| predict.write | Writes predicted property/relationship output to Neo4j where the specific pipeline supports write mode | Retries, stale models, authorization, audit, indexes and rollback become database concerns. |
| gds.model.store | Stores supported model to disk | Enterprise-only; storage path, upgrade compatibility, backup and access control become operational concerns. |
| catalog model only | Model stays available while the instance/catalog lifecycle remains | Community restart loses it; production must have deterministic retraining or an edition-supported persistence plan. |
6. Deliberately wrong approach: auto-tune until the test set looks good
The test set is not a tuning feedback loop. Hyperparameter
ranges and maxTrials should be set before looking
at final test/holdout outcomes. Repeatedly changing the pipeline
after observing test results leaks test information into model
selection.
Repair: use validation/cross-validation for selection, bound search cost, log every candidate space, and reserve the test/future holdout for final evidence. If the experiment changes materially after test inspection, create a new untouched holdout period.
Production judgment
Serving design must specify who refreshes the projection, who retrains the model, what version/config is active, where predictions are stored, how stale predictions are invalidated, and how rollback occurs. Keep the graph feature cutoff and model training cutoff synchronized. Treat a model write as a production data mutation with idempotency/audit requirements. Community’s in-memory model lifecycle can be perfectly adequate for reproducible batch labs, but it is not durable serving by itself.
Check your understanding
- What does training a pipeline create?
- How many models can current GDS Community hold in its model catalog?
- Which edition can persist supported GDS models to disk?
- Why estimate training before running it?
- What is the bridge to Lesson 5?
Review the answers
1. A trained prediction model registered in the model catalog.
2. Three.
3. GDS Enterprise.
4. To catch memory/resource infeasibility before allocating the full training workload.
5. Now the baseline and graph models exist; the final lesson compares them under one evaluation protocol and writes an explicit graph-added-value conclusion.
Summary and next step
Model Catalog, Training, Hyperparameter Search, Prediction, Persistence, and Production Boundaries is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Evaluate a Graph ML Workflow Against Non-Graph Baselines and Document Where the Graph Actually Adds Signal. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- GDS Manual v2026.07 — Current GDS manual and release baseline.
- GDS supported Neo4j versions — Compatibility matrix for Neo4j and GDS releases.
- GDS editions — Community/Enterprise limits including four-core concurrency and three-model catalog capacity.
- Node embeddings overview — Current embedding families, quality tiers, inductive/transductive guidance.
- Fast Random Projection — FastRP dimensions, propertyRatio, featureProperties, randomSeed, iterations and modes.
- Node2Vec — Current Node2Vec algorithm and transductive embedding behavior.
- Machine learning pipelines — Current node classification, link prediction and node regression pipeline tiers.
- Node classification pipelines — End-to-end node classification semantics and prediction model behavior.
- Link prediction pipelines — Feature/train/test split semantics and negative examples.
- Link prediction configuration — Split configuration, FastRP property steps, link features and candidate models.
- Link prediction training — Cross-validation, model selection and model-catalog registration semantics.
- Training methods — Supported classification/regression trainers and auto-tuning.
- Pipeline catalog — Pipeline catalog list/exists/drop lifecycle.
- Model catalog listing — Model metadata, training config, schema and loaded/stored/published state.
- Store models on disk — Enterprise-only persistent model storage semantics.
- Getting started ML pipeline — Current link-prediction pipeline example and split workflow.
- GDS server installation — Bundled plugin installation and procedure configuration.
- Neo4j GDS release notes — Current GDS 2026.07.0 release line.
- Neo4j current versions — Current Neo4j database release and LTS line.