Chapter 25 · Graph Embeddings and Machine Learning Pipelines with GDS

Model Catalog, Training, Hyperparameter Search, Prediction, Persistence, and Production Boundaries

Train bounded baseline and graph models in the current GDS model catalog, inspect tuning/prediction lifecycle, and separate Community in-memory state from Enterprise model persistence.

Advanced260–380 minutesModel catalog · tuning · servingNeo4j 2026.07.1 · Community mandatoryGDS Community 2026.07.0 · Cypher 25GDS ML · concurrency ≤4 CE · model catalog ≤3Java 21/25 · GDS plugin requiredLast reviewed: September 2026

Learning outcomes

01

Distinguish the pipeline catalog from the trained model catalog and explain what survives restart in Community.

02

Configure bounded hyperparameter search and inspect winning parameters without confusing tuning with evaluation.

03

Train and list baseline/graph models within the Community three-model limit.

04

Use prediction stream/mutate/write modes deliberately and document persistent-write/idempotency consequences.

05

Explain Enterprise-only model persistence and define a production retraining/serving/rollback boundary.

1. Pipeline object, trained model, persisted model: three lifecycles

A training pipeline is configurable workflow state in the pipeline catalog. Training it creates a prediction model in the model catalog. The model catalog is in-memory lifecycle state; current GDS Community limits it to three models. Enterprise adds persistent model storage and sharing/publishing capabilities. A database backup does not magically turn an in-memory Community model into a durable artifact.

Dimension Chapter 25 reproducible assumption
Neo4j 2026.07.1 Community in the disposable local/container lab; database neo4j.
Cypher Cypher 25 examples. GDS procedures are called from Cypher; no paid notebook or external ML service is required.
Java Java 21/25 supported by the Neo4j 2026 line; the chosen Neo4j distribution/container supplies the runtime.
Auth/TLS User neo4j, password atlasmart-course-2026; loopback Bolt without TLS only for this isolated lab. Production/remote deployments require authenticated encrypted transport.
Driver No application driver is required for mandatory procedure labs; Browser or cypher-shell is sufficient.
GDS GDS Community 2026.07.0. Community includes all algorithms, caps GDS concurrency at four CPU cores, and limits the model catalog to three models.
GDS ML quality tiers Node Classification and Link Prediction pipelines are Beta; Node Regression pipelines are Alpha. Treat tier as a production-risk input.
Persistence Community model/pipeline/graph catalogs are in-memory lifecycle objects. Model persistence to disk is Enterprise-only.
APOC Not required for mandatory Chapter 25 work.
Scope Only persisted entities with chapter25=true and in-memory names beginning atlas-ch25- are created.
Evidence This generation environment does not run Neo4j/GDS. Fixture counts and split rules are deterministic; training scores, memory estimates, timings and predictions must be measured locally and must not be copied as fabricated results.
Verify the Chapter 25 runtime
RETURN gds.version() AS gdsVersion;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;

// Inspect current in-memory catalogs before the lab.
CALL gds.graph.list() YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCount ORDER BY graphName;

CALL gds.pipeline.list() YIELD pipelineName, pipelineType
RETURN pipelineName, pipelineType ORDER BY pipelineName;

CALL gds.model.list() YIELD modelName, modelType, loaded, stored, published
RETURN modelName, modelType, loaded, stored, published ORDER BY modelName;
Create the bounded AtlasMart graph-ML fixture
// Remove only Chapter 25 persisted fixtures before recreating them.
MATCH (n) WHERE n.chapter25 = true DETACH DELETE n;

// 30 customers. The target is a synthetic future-horizon label for teaching only.
UNWIND range(1,30) AS i
CREATE (:Customer:CH25Customer {
  customerId: 'CH25-C-' + right('00' + toString(i), 2),
  seq: i,
  tenureMonths: toFloat(3 + (i % 18)),
  spendIndex: toFloat(20 + ((i * 17) % 80)),
  willRepeat: CASE WHEN i % 3 = 0 OR i % 7 = 0 THEN 0 ELSE 1 END,
  featureCutoff: date('2026-06-30'),
  labelWindow: '2026-Q3',
  chapter25: true
});

UNWIND range(1,8) AS j
CREATE (:Product:CH25Product {
  productId: 'CH25-P-' + right('00' + toString(j), 2),
  categoryCode: toFloat(j % 4),
  chapter25: true
});

// Feature-window interactions only. No Q3 target-window interaction is loaded.
MATCH (c:CH25Customer), (p:CH25Product)
WHERE ((c.seq + toInteger(right(p.productId,2))) % 4 = 0)
   OR (toInteger(right(p.productId,2)) = ((c.seq - 1) % 8) + 1)
CREATE (c)-[:VIEWED_CH25 {chapter25:true, asOf:'2026-Q2'}]->(p);

MATCH (c:CH25Customer)
RETURN count(c) AS customers,
       sum(CASE WHEN c.willRepeat = 1 THEN 1 ELSE 0 END) AS positiveLabels,
       sum(CASE WHEN c.willRepeat = 0 THEN 1 ELSE 0 END) AS negativeLabels;

MATCH (:CH25Customer)-[r:VIEWED_CH25]->(:CH25Product)
RETURN count(r) AS featureWindowRelationships;
Project only the pre-label-horizon graph
// Drop atlas-ch25-ml first if it already exists.
CALL gds.graph.project(
  'atlas-ch25-ml',
  {
    CH25Customer: {properties: ['tenureMonths','spendIndex','willRepeat']},
    CH25Product: {properties: ['categoryCode']}
  },
  {
    VIEWED_CH25: {orientation:'UNDIRECTED'}
  }
)
YIELD graphName, nodeCount, relationshipCount, projectMillis
RETURN graphName, nodeCount, relationshipCount, projectMillis;

CALL gds.graph.list('atlas-ch25-ml')
YIELD graphName, nodeCount, relationshipCount, schema
RETURN graphName, nodeCount, relationshipCount, schema;

2. Train the two-model ablation without exceeding the Community catalog limit

Create/configure the baseline pipeline if not already present
CALL gds.beta.pipeline.nodeClassification.create('atlas-ch25-baseline-pipe');
CALL gds.beta.pipeline.nodeClassification.selectFeatures('atlas-ch25-baseline-pipe',['tenureMonths','spendIndex']);
CALL gds.beta.pipeline.nodeClassification.configureSplit('atlas-ch25-baseline-pipe',{testFraction:0.2,validationFolds:3});
CALL gds.beta.pipeline.nodeClassification.configureAutoTuning('atlas-ch25-baseline-pipe',{maxTrials:5});
CALL gds.beta.pipeline.nodeClassification.addLogisticRegression(
  'atlas-ch25-baseline-pipe',
  {penalty:{range:[0.0001,1.0]}}
);
Estimate and train baseline model
CALL gds.beta.pipeline.nodeClassification.train.estimate(
  'atlas-ch25-ml',
  {
    pipeline:'atlas-ch25-baseline-pipe',
    modelName:'atlas-ch25-baseline-model',
    targetNodeLabels:['CH25Customer'],
    targetProperty:'willRepeat',
    metrics:['F1_WEIGHTED'],
    randomSeed:42,
    concurrency:4
  }
)
YIELD requiredMemory, treeView
RETURN requiredMemory, treeView;

CALL gds.beta.pipeline.nodeClassification.train(
  'atlas-ch25-ml',
  {
    pipeline:'atlas-ch25-baseline-pipe',
    modelName:'atlas-ch25-baseline-model',
    targetNodeLabels:['CH25Customer'],
    targetProperty:'willRepeat',
    metrics:['F1_WEIGHTED'],
    randomSeed:42,
    concurrency:4
  }
)
YIELD modelInfo, modelSelectionStats, trainMillis
RETURN modelInfo.bestParameters AS bestParameters,
       modelInfo.metrics AS metrics,
       size(modelSelectionStats.modelCandidates) AS evaluatedCandidates,
       trainMillis;

3. Train the graph-augmented model

Configure graph pipeline
CALL gds.beta.pipeline.nodeClassification.create('atlas-ch25-graph-pipe');
CALL gds.beta.pipeline.nodeClassification.addNodeProperty(
  'atlas-ch25-graph-pipe','fastRP',
  {
    mutateProperty:'ch25Embedding',
    embeddingDimension:16,
    iterationWeights:[0.0,1.0,1.0],
    randomSeed:42,
    contextNodeLabels:['CH25Customer','CH25Product'],
    contextRelationshipTypes:['VIEWED_CH25']
  }
);
CALL gds.beta.pipeline.nodeClassification.selectFeatures(
  'atlas-ch25-graph-pipe',['tenureMonths','spendIndex','ch25Embedding']
);
CALL gds.beta.pipeline.nodeClassification.configureSplit('atlas-ch25-graph-pipe',{testFraction:0.2,validationFolds:3});
CALL gds.beta.pipeline.nodeClassification.configureAutoTuning('atlas-ch25-graph-pipe',{maxTrials:5});
CALL gds.beta.pipeline.nodeClassification.addLogisticRegression(
  'atlas-ch25-graph-pipe',{penalty:{range:[0.0001,1.0]}}
);
Estimate and train graph model
CALL gds.beta.pipeline.nodeClassification.train.estimate(
  'atlas-ch25-ml',
  {
    pipeline:'atlas-ch25-graph-pipe',
    modelName:'atlas-ch25-graph-model',
    targetNodeLabels:['CH25Customer'],
    targetProperty:'willRepeat',
    metrics:['F1_WEIGHTED'],
    randomSeed:42,
    concurrency:4
  }
)
YIELD requiredMemory, treeView
RETURN requiredMemory, treeView;

CALL gds.beta.pipeline.nodeClassification.train(
  'atlas-ch25-ml',
  {
    pipeline:'atlas-ch25-graph-pipe',
    modelName:'atlas-ch25-graph-model',
    targetNodeLabels:['CH25Customer'],
    targetProperty:'willRepeat',
    metrics:['F1_WEIGHTED'],
    randomSeed:42,
    concurrency:4
  }
)
YIELD modelInfo, modelSelectionStats, trainMillis
RETURN modelInfo.bestParameters AS bestParameters,
       modelInfo.metrics AS metrics,
       size(modelSelectionStats.modelCandidates) AS evaluatedCandidates,
       trainMillis;
Expected state

Exactly two Chapter 25 prediction models should exist after the mandatory lab—within the Community three-model limit. Record actual test metrics and training times; do not substitute values from this HTML.

4. Inspect catalog state and prediction lifecycle

List models and pipelines
CALL gds.pipeline.list()
YIELD pipelineName, pipelineType, creationTime
WHERE pipelineName STARTS WITH 'atlas-ch25-'
RETURN pipelineName, pipelineType, creationTime ORDER BY pipelineName;

CALL gds.model.list()
YIELD modelName, modelType, modelInfo, trainConfig, graphSchema, loaded, stored, published
WHERE modelName STARTS WITH 'atlas-ch25-'
RETURN modelName, modelType, loaded, stored, published,
       modelInfo.metrics AS metrics,
       trainConfig,
       graphSchema
ORDER BY modelName;
Stream predictions from the graph model
CALL gds.beta.pipeline.nodeClassification.predict.stream(
  'atlas-ch25-ml',
  {
    modelName:'atlas-ch25-graph-model',
    targetNodeLabels:['CH25Customer'],
    includePredictedProbabilities:true,
    concurrency:4
  }
)
YIELD nodeId, predictedClass, predictedProbabilities
RETURN gds.util.asNode(nodeId).customerId AS customerId,
       predictedClass,
       predictedProbabilities
ORDER BY customerId
LIMIT 15;

5. stream, mutate, write, and persistence boundaries

Operation State changed Use with care because…
predict.stream No graph/store mutation; rows returned to client Large result sets can move memory/network pressure to the client.
predict.mutate Adds predicted properties to the in-memory GDS graph Ephemeral state can be mistaken for persisted application truth.
predict.write Writes predicted property/relationship output to Neo4j where the specific pipeline supports write mode Retries, stale models, authorization, audit, indexes and rollback become database concerns.
gds.model.store Stores supported model to disk Enterprise-only; storage path, upgrade compatibility, backup and access control become operational concerns.
catalog model only Model stays available while the instance/catalog lifecycle remains Community restart loses it; production must have deterministic retraining or an edition-supported persistence plan.

6. Deliberately wrong approach: auto-tune until the test set looks good

The test set is not a tuning feedback loop. Hyperparameter ranges and maxTrials should be set before looking at final test/holdout outcomes. Repeatedly changing the pipeline after observing test results leaks test information into model selection.

Repair: use validation/cross-validation for selection, bound search cost, log every candidate space, and reserve the test/future holdout for final evidence. If the experiment changes materially after test inspection, create a new untouched holdout period.

Production judgment

Serving design must specify who refreshes the projection, who retrains the model, what version/config is active, where predictions are stored, how stale predictions are invalidated, and how rollback occurs. Keep the graph feature cutoff and model training cutoff synchronized. Treat a model write as a production data mutation with idempotency/audit requirements. Community’s in-memory model lifecycle can be perfectly adequate for reproducible batch labs, but it is not durable serving by itself.

Check your understanding

  1. What does training a pipeline create?
  2. How many models can current GDS Community hold in its model catalog?
  3. Which edition can persist supported GDS models to disk?
  4. Why estimate training before running it?
  5. What is the bridge to Lesson 5?
Review the answers

1. A trained prediction model registered in the model catalog.

2. Three.

3. GDS Enterprise.

4. To catch memory/resource infeasibility before allocating the full training workload.

5. Now the baseline and graph models exist; the final lesson compares them under one evaluation protocol and writes an explicit graph-added-value conclusion.

Summary and next step

Model Catalog, Training, Hyperparameter Search, Prediction, Persistence, and Production Boundaries is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Evaluate a Graph ML Workflow Against Non-Graph Baselines and Document Where the Graph Actually Adds Signal. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.