Chapter 25 · Graph Embeddings and Machine Learning Pipelines with GDS
Node Classification, Link Prediction, and Regression Pipeline Concepts, Splits, Metrics, and Baselines
Compare current GDS node-classification, link-prediction, and node-regression pipeline semantics, splits, negative sampling, metrics, baselines, and holdout limitations.
Learning outcomes
Distinguish node classification, link prediction, and node regression by supervised target and prediction unit.
Explain train/test and cross-validation roles, plus link-prediction feature/train/test relationship splits and negative sampling.
Match metrics to class imbalance, ranking, probability quality, or regression error rather than optimizing one convenient number.
Use a non-graph baseline and graph-feature ablation as mandatory evidence of graph-added value.
Identify temporal/entity split cases where built-in random splits need an external holdout design.
1. Three supervised tasks, three different units of prediction
GDS pipelines automate feature steps, candidate-model training, cross-validation/model selection, test evaluation, catalog registration, and prediction. They do not make the supervised question interchangeable. A node classifier predicts a class property on nodes. A node regressor predicts a numeric node property. A link predictor predicts whether a node pair should be adjacent under a target relationship definition.
| Dimension | Chapter 25 reproducible assumption |
|---|---|
| Neo4j |
2026.07.1 Community in the disposable local/container lab;
database neo4j.
|
| Cypher | Cypher 25 examples. GDS procedures are called from Cypher; no paid notebook or external ML service is required. |
| Java | Java 21/25 supported by the Neo4j 2026 line; the chosen Neo4j distribution/container supplies the runtime. |
| Auth/TLS |
User neo4j, password
atlasmart-course-2026; loopback Bolt without
TLS only for this isolated lab. Production/remote
deployments require authenticated encrypted transport.
|
| Driver | No application driver is required for mandatory procedure labs; Browser or cypher-shell is sufficient. |
| GDS | GDS Community 2026.07.0. Community includes all algorithms, caps GDS concurrency at four CPU cores, and limits the model catalog to three models. |
| GDS ML quality tiers | Node Classification and Link Prediction pipelines are Beta; Node Regression pipelines are Alpha. Treat tier as a production-risk input. |
| Persistence | Community model/pipeline/graph catalogs are in-memory lifecycle objects. Model persistence to disk is Enterprise-only. |
| APOC | Not required for mandatory Chapter 25 work. |
| Scope |
Only persisted entities with
chapter25=true and in-memory names beginning
atlas-ch25- are created.
|
| Evidence | This generation environment does not run Neo4j/GDS. Fixture counts and split rules are deterministic; training scores, memory estimates, timings and predictions must be measured locally and must not be copied as fabricated results. |
RETURN gds.version() AS gdsVersion;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;
// Inspect current in-memory catalogs before the lab.
CALL gds.graph.list() YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCount ORDER BY graphName;
CALL gds.pipeline.list() YIELD pipelineName, pipelineType
RETURN pipelineName, pipelineType ORDER BY pipelineName;
CALL gds.model.list() YIELD modelName, modelType, loaded, stored, published
RETURN modelName, modelType, loaded, stored, published ORDER BY modelName;
2. Current GDS pipeline quality tiers and evaluation surfaces
| Task | Current tier | Training unit | Typical evidence | Key leakage/sampling risk |
|---|---|---|---|---|
| Node Classification | Beta | Labeled target nodes | F1 family, accuracy/log-loss where appropriate, class distribution, calibration checks outside/alongside GDS | Future graph/property leakage; entity/group overlap; class imbalance. |
| Link Prediction | Beta | Positive target relationships + sampled/provided negative node pairs | AUCPR/ranking-quality evidence, negative sampling ratio/type, candidate coverage | Target relationships leaking into feature computation; unrealistic random negatives; temporal edge split. |
| Node Regression | Alpha | Nodes with numeric target property | MAE/MSE and residual/error distribution | Target-derived numeric features; non-stationary target scale; alpha-tier production risk. |
3. Link prediction has an extra graph split that ordinary tabular ML does not
Current GDS link-prediction training derives feature-input, train, and test relationship sets. Node-property steps are computed using the feature-input graph so held-out target edges are not simply exposed to the feature pipeline. Negative examples are sampled from non-adjacent node pairs unless an explicit negative relationship type is supplied.
CALL gds.beta.pipeline.linkPrediction.create('atlas-ch25-lp-demo');
CALL gds.beta.pipeline.linkPrediction.configureSplit(
'atlas-ch25-lp-demo',
{
testFraction:0.20,
trainFraction:0.60,
validationFolds:3,
negativeSamplingRatio:1.0
}
);
CALL gds.beta.pipeline.linkPrediction.addNodeProperty(
'atlas-ch25-lp-demo',
'fastRP',
{mutateProperty:'embedding', embeddingDimension:16, randomSeed:42}
);
CALL gds.beta.pipeline.linkPrediction.addFeature(
'atlas-ch25-lp-demo',
'hadamard',
{nodeProperties:['embedding']}
);
CALL gds.beta.pipeline.linkPrediction.addLogisticRegression('atlas-ch25-lp-demo');
CALL gds.pipeline.list('atlas-ch25-lp-demo')
YIELD pipelineInfo
RETURN pipelineInfo;
Creating a training pipeline does not itself consume one of the three Community model slots, but training creates a model. The mandatory experiment uses only two trained models so the model-catalog limit remains visible and controlled.
4. Metric choice follows the decision and class distribution
| Situation | Prefer/inspect | Why one metric is insufficient |
|---|---|---|
| imbalanced repeat-purchase classification | F1 by class/weighted F1 + confusion matrix + probability calibration | Accuracy can look high by predicting the majority class. |
| link recommendation with many non-links | AUCPR/ranking quality + candidate coverage + business acceptance | ROC-like separation can hide poor top-of-list usefulness when negatives dominate. |
| regression | MAE + MSE/RMSE-style error + residual slices | Averages hide outliers, heteroscedasticity and segment-specific error. |
| probability-triggered action | calibration/reliability + threshold utility | A ranking can be good while probabilities are badly calibrated. |
| graph-added value | same split/protocol baseline vs +graph feature, preferably repeated seeds | A higher training metric proves nothing about generalization or causal value. |
5. Class imbalance, negative sampling, and the real deployment population
Class weights or negative sampling change the effective training distribution. They can improve learning, but the resulting probability should not automatically be interpreted as deployment prevalence. For link prediction, sampled negatives are a computational approximation to a huge non-edge space; candidate-generation rules and production constraints define which negatives are actually plausible. Always record the sampling policy and evaluate on a population that resembles the decision surface.
6. Deliberately wrong approach: choose the model with the best training score
Training score rewards memorization and data leakage. GDS pipelines use validation/cross-validation for model selection and report test evidence for the winning configuration. Even then, a real temporal or entity-generalization problem may require an external holdout beyond the pipeline’s random split.
Repair: freeze the business split protocol first, use validation only for model/hyperparameter selection, inspect the untouched test/holdout once for selection confirmation, and compare against the same baseline protocol.
Production judgment
Prediction is one part of the serving path. A link model can generate far more candidate pairs than the application can rank or review; a classifier can emit probabilities that still need policy thresholds; a regression model can drift as prices or customer behavior shift. Budget GDS memory/CPU separately from transactional query capacity, bound prediction output, retain stable business IDs in results, version the projection/feature/model together, and monitor both statistical metrics and downstream business error/cost.
Check your understanding
- What is the supervised example in link prediction?
- Why can sampled negatives distort interpretation?
- What is the quality tier of node regression pipelines in current GDS?
- Why might a later-time holdout still be needed?
- What is the bridge to Lesson 4?
Review the answers
1. A node pair labeled positive when the target relationship exists and negative according to the configured sampling/negative-relationship scheme.
2. They change the effective class prevalence and may include implausible pairs that would never be candidates in production.
3. Alpha.
4. Built-in random splits do not automatically reproduce a temporal serving scenario.
5. After defining trustworthy evaluation, we need to control catalog lifecycle, auto-tuning, prediction, persistence and deployment boundaries.
Summary and next step
Node Classification, Link Prediction, and Regression Pipeline Concepts, Splits, Metrics, and Baselines is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Model Catalog, Training, Hyperparameter Search, Prediction, Persistence, and Production Boundaries. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- GDS Manual v2026.07 — Current GDS manual and release baseline.
- GDS supported Neo4j versions — Compatibility matrix for Neo4j and GDS releases.
- GDS editions — Community/Enterprise limits including four-core concurrency and three-model catalog capacity.
- Node embeddings overview — Current embedding families, quality tiers, inductive/transductive guidance.
- Fast Random Projection — FastRP dimensions, propertyRatio, featureProperties, randomSeed, iterations and modes.
- Node2Vec — Current Node2Vec algorithm and transductive embedding behavior.
- Machine learning pipelines — Current node classification, link prediction and node regression pipeline tiers.
- Node classification pipelines — End-to-end node classification semantics and prediction model behavior.
- Link prediction pipelines — Feature/train/test split semantics and negative examples.
- Link prediction configuration — Split configuration, FastRP property steps, link features and candidate models.
- Link prediction training — Cross-validation, model selection and model-catalog registration semantics.
- Training methods — Supported classification/regression trainers and auto-tuning.
- Pipeline catalog — Pipeline catalog list/exists/drop lifecycle.
- Model catalog listing — Model metadata, training config, schema and loaded/stored/published state.
- Store models on disk — Enterprise-only persistent model storage semantics.
- Getting started ML pipeline — Current link-prediction pipeline example and split workflow.
- GDS server installation — Bundled plugin installation and procedure configuration.
- Neo4j GDS release notes — Current GDS 2026.07.0 release line.
- Neo4j current versions — Current Neo4j database release and LTS line.