Chapter 25 · Graph Embeddings and Machine Learning Pipelines with GDS

Node Classification, Link Prediction, and Regression Pipeline Concepts, Splits, Metrics, and Baselines

Compare current GDS node-classification, link-prediction, and node-regression pipeline semantics, splits, negative sampling, metrics, baselines, and holdout limitations.

Advanced260–380 minutesPipelines · splits · metricsNeo4j 2026.07.1 · Community mandatoryGDS Community 2026.07.0 · Cypher 25GDS ML · concurrency ≤4 CE · model catalog ≤3Java 21/25 · GDS plugin requiredLast reviewed: September 2026

Learning outcomes

01

Distinguish node classification, link prediction, and node regression by supervised target and prediction unit.

02

Explain train/test and cross-validation roles, plus link-prediction feature/train/test relationship splits and negative sampling.

03

Match metrics to class imbalance, ranking, probability quality, or regression error rather than optimizing one convenient number.

04

Use a non-graph baseline and graph-feature ablation as mandatory evidence of graph-added value.

05

Identify temporal/entity split cases where built-in random splits need an external holdout design.

1. Three supervised tasks, three different units of prediction

GDS pipelines automate feature steps, candidate-model training, cross-validation/model selection, test evaluation, catalog registration, and prediction. They do not make the supervised question interchangeable. A node classifier predicts a class property on nodes. A node regressor predicts a numeric node property. A link predictor predicts whether a node pair should be adjacent under a target relationship definition.

Dimension Chapter 25 reproducible assumption
Neo4j 2026.07.1 Community in the disposable local/container lab; database neo4j.
Cypher Cypher 25 examples. GDS procedures are called from Cypher; no paid notebook or external ML service is required.
Java Java 21/25 supported by the Neo4j 2026 line; the chosen Neo4j distribution/container supplies the runtime.
Auth/TLS User neo4j, password atlasmart-course-2026; loopback Bolt without TLS only for this isolated lab. Production/remote deployments require authenticated encrypted transport.
Driver No application driver is required for mandatory procedure labs; Browser or cypher-shell is sufficient.
GDS GDS Community 2026.07.0. Community includes all algorithms, caps GDS concurrency at four CPU cores, and limits the model catalog to three models.
GDS ML quality tiers Node Classification and Link Prediction pipelines are Beta; Node Regression pipelines are Alpha. Treat tier as a production-risk input.
Persistence Community model/pipeline/graph catalogs are in-memory lifecycle objects. Model persistence to disk is Enterprise-only.
APOC Not required for mandatory Chapter 25 work.
Scope Only persisted entities with chapter25=true and in-memory names beginning atlas-ch25- are created.
Evidence This generation environment does not run Neo4j/GDS. Fixture counts and split rules are deterministic; training scores, memory estimates, timings and predictions must be measured locally and must not be copied as fabricated results.
Verify the Chapter 25 runtime
RETURN gds.version() AS gdsVersion;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;

// Inspect current in-memory catalogs before the lab.
CALL gds.graph.list() YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCount ORDER BY graphName;

CALL gds.pipeline.list() YIELD pipelineName, pipelineType
RETURN pipelineName, pipelineType ORDER BY pipelineName;

CALL gds.model.list() YIELD modelName, modelType, loaded, stored, published
RETURN modelName, modelType, loaded, stored, published ORDER BY modelName;

2. Current GDS pipeline quality tiers and evaluation surfaces

Task Current tier Training unit Typical evidence Key leakage/sampling risk
Node Classification Beta Labeled target nodes F1 family, accuracy/log-loss where appropriate, class distribution, calibration checks outside/alongside GDS Future graph/property leakage; entity/group overlap; class imbalance.
Link Prediction Beta Positive target relationships + sampled/provided negative node pairs AUCPR/ranking-quality evidence, negative sampling ratio/type, candidate coverage Target relationships leaking into feature computation; unrealistic random negatives; temporal edge split.
Node Regression Alpha Nodes with numeric target property MAE/MSE and residual/error distribution Target-derived numeric features; non-stationary target scale; alpha-tier production risk.

3. Link prediction has an extra graph split that ordinary tabular ML does not

Current GDS link-prediction training derives feature-input, train, and test relationship sets. Node-property steps are computed using the feature-input graph so held-out target edges are not simply exposed to the feature pipeline. Negative examples are sampled from non-adjacent node pairs unless an explicit negative relationship type is supplied.

Illustrative link-prediction split configuration
CALL gds.beta.pipeline.linkPrediction.create('atlas-ch25-lp-demo');
CALL gds.beta.pipeline.linkPrediction.configureSplit(
  'atlas-ch25-lp-demo',
  {
    testFraction:0.20,
    trainFraction:0.60,
    validationFolds:3,
    negativeSamplingRatio:1.0
  }
);
CALL gds.beta.pipeline.linkPrediction.addNodeProperty(
  'atlas-ch25-lp-demo',
  'fastRP',
  {mutateProperty:'embedding', embeddingDimension:16, randomSeed:42}
);
CALL gds.beta.pipeline.linkPrediction.addFeature(
  'atlas-ch25-lp-demo',
  'hadamard',
  {nodeProperties:['embedding']}
);
CALL gds.beta.pipeline.linkPrediction.addLogisticRegression('atlas-ch25-lp-demo');
CALL gds.pipeline.list('atlas-ch25-lp-demo')
YIELD pipelineInfo
RETURN pipelineInfo;
Do not run every demo pipeline into the Community model catalog

Creating a training pipeline does not itself consume one of the three Community model slots, but training creates a model. The mandatory experiment uses only two trained models so the model-catalog limit remains visible and controlled.

4. Metric choice follows the decision and class distribution

Situation Prefer/inspect Why one metric is insufficient
imbalanced repeat-purchase classification F1 by class/weighted F1 + confusion matrix + probability calibration Accuracy can look high by predicting the majority class.
link recommendation with many non-links AUCPR/ranking quality + candidate coverage + business acceptance ROC-like separation can hide poor top-of-list usefulness when negatives dominate.
regression MAE + MSE/RMSE-style error + residual slices Averages hide outliers, heteroscedasticity and segment-specific error.
probability-triggered action calibration/reliability + threshold utility A ranking can be good while probabilities are badly calibrated.
graph-added value same split/protocol baseline vs +graph feature, preferably repeated seeds A higher training metric proves nothing about generalization or causal value.

5. Class imbalance, negative sampling, and the real deployment population

Class weights or negative sampling change the effective training distribution. They can improve learning, but the resulting probability should not automatically be interpreted as deployment prevalence. For link prediction, sampled negatives are a computational approximation to a huge non-edge space; candidate-generation rules and production constraints define which negatives are actually plausible. Always record the sampling policy and evaluate on a population that resembles the decision surface.

6. Deliberately wrong approach: choose the model with the best training score

Training score rewards memorization and data leakage. GDS pipelines use validation/cross-validation for model selection and report test evidence for the winning configuration. Even then, a real temporal or entity-generalization problem may require an external holdout beyond the pipeline’s random split.

Repair: freeze the business split protocol first, use validation only for model/hyperparameter selection, inspect the untouched test/holdout once for selection confirmation, and compare against the same baseline protocol.

Production judgment

Prediction is one part of the serving path. A link model can generate far more candidate pairs than the application can rank or review; a classifier can emit probabilities that still need policy thresholds; a regression model can drift as prices or customer behavior shift. Budget GDS memory/CPU separately from transactional query capacity, bound prediction output, retain stable business IDs in results, version the projection/feature/model together, and monitor both statistical metrics and downstream business error/cost.

Check your understanding

  1. What is the supervised example in link prediction?
  2. Why can sampled negatives distort interpretation?
  3. What is the quality tier of node regression pipelines in current GDS?
  4. Why might a later-time holdout still be needed?
  5. What is the bridge to Lesson 4?
Review the answers

1. A node pair labeled positive when the target relationship exists and negative according to the configured sampling/negative-relationship scheme.

2. They change the effective class prevalence and may include implausible pairs that would never be candidates in production.

3. Alpha.

4. Built-in random splits do not automatically reproduce a temporal serving scenario.

5. After defining trustworthy evaluation, we need to control catalog lifecycle, auto-tuning, prediction, persistence and deployment boundaries.

Summary and next step

Node Classification, Link Prediction, and Regression Pipeline Concepts, Splits, Metrics, and Baselines is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Model Catalog, Training, Hyperparameter Search, Prediction, Persistence, and Production Boundaries. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.