Chapter 25 · Graph Embeddings and Machine Learning Pipelines with GDS
Evaluate a Graph ML Workflow Against Non-Graph Baselines and Document Where the Graph Actually Adds Signal
Evaluate AtlasMart graph ML against a non-graph baseline under one protocol, run ablation/leakage checks, and document whether graph features add enough signal to justify production cost.
Learning outcomes
Compare baseline and graph-augmented models on the same untouched test protocol and report uncertainty/limits rather than a marketing conclusion.
Run a graph-feature ablation and inspect actual catalog metrics without fabricating improvement.
Repeat sensitivity checks over seeds/configurations without silently tuning on the test set.
Write an explicit “graph adds signal / graph does not add enough signal” decision with cost, drift, fairness, serving and rollback consequences.
Clean up projections, pipelines, models and synthetic source data while preserving unrelated academy/database state.
1. AtlasMart final question: does graph information improve the decision enough to pay for itself?
The chapter succeeds even if the graph model does not beat the baseline. The correct conclusion is empirical: compare the same supervised target, feature cutoff, split protocol, trainer family and metric; change only the graph-derived feature set. Then judge the observed lift against added complexity, compute, refresh cadence, serving dependencies, privacy/security scope, and operational risk.
| Dimension | Chapter 25 reproducible assumption |
|---|---|
| Neo4j |
2026.07.1 Community in the disposable local/container lab;
database neo4j.
|
| Cypher | Cypher 25 examples. GDS procedures are called from Cypher; no paid notebook or external ML service is required. |
| Java | Java 21/25 supported by the Neo4j 2026 line; the chosen Neo4j distribution/container supplies the runtime. |
| Auth/TLS |
User neo4j, password
atlasmart-course-2026; loopback Bolt without
TLS only for this isolated lab. Production/remote
deployments require authenticated encrypted transport.
|
| Driver | No application driver is required for mandatory procedure labs; Browser or cypher-shell is sufficient. |
| GDS | GDS Community 2026.07.0. Community includes all algorithms, caps GDS concurrency at four CPU cores, and limits the model catalog to three models. |
| GDS ML quality tiers | Node Classification and Link Prediction pipelines are Beta; Node Regression pipelines are Alpha. Treat tier as a production-risk input. |
| Persistence | Community model/pipeline/graph catalogs are in-memory lifecycle objects. Model persistence to disk is Enterprise-only. |
| APOC | Not required for mandatory Chapter 25 work. |
| Scope |
Only persisted entities with
chapter25=true and in-memory names beginning
atlas-ch25- are created.
|
| Evidence | This generation environment does not run Neo4j/GDS. Fixture counts and split rules are deterministic; training scores, memory estimates, timings and predictions must be measured locally and must not be copied as fabricated results. |
RETURN gds.version() AS gdsVersion;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;
// Inspect current in-memory catalogs before the lab.
CALL gds.graph.list() YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCount ORDER BY graphName;
CALL gds.pipeline.list() YIELD pipelineName, pipelineType
RETURN pipelineName, pipelineType ORDER BY pipelineName;
CALL gds.model.list() YIELD modelName, modelType, loaded, stored, published
RETURN modelName, modelType, loaded, stored, published ORDER BY modelName;
2. Acceptance criteria before reading any scores
| Gate | Pass condition |
|---|---|
| feature time boundary | All VIEWED_CH25 relationships and raw features predate the Q3 label horizon. |
| same population |
Baseline and graph models target the same CH25Customer
nodes and willRepeat property.
|
| same split/search protocol | Same testFraction, validationFolds, trainer family, tuning budget and randomSeed for the direct ablation. |
| catalog capacity | No more than three Community models; mandatory comparison uses two. |
| resource evidence | Training estimates + actual trainMillis recorded for both models. |
| no fabricated output |
Test metrics are copied only from the learner’s actual
gds.model.list/training result.
|
3. Read the two models from the catalog side by side
CALL gds.model.list()
YIELD modelName, modelType, modelInfo, trainConfig, graphSchema
WHERE modelName IN ['atlas-ch25-baseline-model','atlas-ch25-graph-model']
RETURN modelName,
modelType,
modelInfo.metrics.F1_WEIGHTED.test AS testF1Weighted,
modelInfo.bestParameters AS bestParameters,
trainConfig,
graphSchema
ORDER BY modelName;
Do not assume graph > baseline. If the graph
model is lower, equal within noise, or only trivially higher,
the correct production decision may be to keep the simpler
baseline.
4. Ablation matrix: what changed?
| Experiment | Features | What it isolates |
|---|---|---|
| Baseline | tenureMonths + spendIndex | Predictive power available without graph computation. |
| Graph augmented | tenureMonths + spendIndex + FastRP embedding | Incremental value of graph-derived representation under the same split/trainer protocol. |
| Topology only (optional) | FastRP embedding only | Whether graph structure can carry the signal without raw properties; not required for mandatory two-model limit. |
| Property-only inductive FastRP (optional) | FastRP with numeric featureProperties, propertyRatio=1.0, fixed seed | Generalization/representation alternative; should be a separately registered experiment, not silently substituted. |
5. Sensitivity without test-set fishing
After the primary comparison is frozen, run a planned sensitivity study. For example, retrain one experiment at a time with seeds 41, 42, and 43, dropping each temporary model before the next so Community model capacity is not exceeded. Compare validation variability and only use a newly reserved holdout period if the experiment design changes after inspecting the original test.
// Example process, not a loop to paste blindly:
// 1. Keep baseline + active graph model in catalog.
// 2. Train at most one temporary sensitivity model (third Community slot).
// 3. Record validation/test evidence under the predeclared protocol.
// 4. Drop the temporary model before the next seed/config.
CALL gds.model.list()
YIELD modelName, modelType
RETURN modelName, modelType ORDER BY modelName;
6. Decision template: graph-added-value statement
| Question | Evidence to write down |
|---|---|
| Did graph add signal? | Measured test/holdout delta under the same protocol, plus seed/config sensitivity. |
| Is the lift material? | Translate metric change into business threshold/decision utility; avoid “statistically nonzero = operationally worth it.” |
| What does it cost? | Projection + embedding + training memory/time, refresh frequency, write/storage/serving overhead. |
| What can drift? | Graph degree/distribution, product mix, raw properties, class prevalence, embedding coordinate regime, model calibration. |
| What can leak? | Future edges, target-derived properties, tenant connections, identity duplicates, post-outcome data. |
| Can it be served safely? | Model/projection lifecycle, prediction latency/output size, auth/TLS, retries/idempotency, stale-result policy. |
| What is rollback? | Retain baseline/previous model/config and the source cutoff needed to reproduce it. |
7. Failure injection: prove the evaluation catches a leaked graph
In a disposable copy only, deliberately add one target-window
relationship type such as FUTURE_PURCHASE_CH25 that
directly follows willRepeat, then project it into a
throwaway graph and observe how easily model metrics can
inflate. Do not use the contaminated graph for the real
comparison. The educational point is that a “better” score can
be a data-integrity alarm rather than progress.
// DISPOSABLE DEMONSTRATION ONLY. Do not add this relationship to atlas-ch25-ml.
MATCH (c:CH25Customer), (p:CH25Product {productId:'CH25-P-01'})
WHERE c.willRepeat = 1
CREATE (c)-[:FUTURE_PURCHASE_CH25 {chapter25:true, syntheticLeak:true}]->(p);
// Verify the contamination exists, explain why it is forbidden, then delete it.
MATCH (:CH25Customer)-[r:FUTURE_PURCHASE_CH25]->(:CH25Product)
RETURN count(r) AS leakedFutureEdges;
MATCH ()-[r:FUTURE_PURCHASE_CH25]->() DELETE r;
The controlled fault proves that graph feature correctness depends on temporal/semantic source integrity. It does not prove a specific amount of metric inflation, because this artifact does not execute training.
8. Production-readiness matrix
| Area | Chapter 25 gate |
|---|---|
| Graph/workload fit | Graph feature has measured incremental value for the actual supervised decision. |
| Correctness | Cutoff, target, projection, identity, split, negative sampling and leakage checks are documented. |
| Cardinality/degree | Projection size/degree distributions measured; hub/outlier behavior reviewed. |
| Transactions/concurrency | GDS batch work isolated from latency-critical transactions; Community concurrency ≤4. |
| Memory/CPU/disk/network | Projection/training estimates + actual runtime/resource evidence; prediction output bounded. |
| Indexes/constraints | Source business identities constrained/indexed; ML does not replace source invariants. |
| Driver/retries/idempotency | If orchestrated by an app, retries cannot duplicate persistent prediction writes or create multiple model versions accidentally. |
| Security/tenant risk | Projection is authorization-safe; sensitive labels/features and model outputs have explicit access controls. |
| Backup/recovery | Source graph and model/config lineage are recoverable; Community in-memory model is reproducibly retrainable. |
| Observability | Model version, feature age, drift, class balance, metrics, training failures, prediction latency/output monitored. |
| Version/edition/tier | Neo4j/GDS compatibility and Beta/Alpha pipeline tier recorded; Enterprise-only persistence not assumed in Community. |
| Migration/rollback | Shadow evaluate new feature/model/version; retain last validated baseline/model until acceptance passes. |
9. Chapter 25 final verification checklist
-
Record
gds.version(), Neo4j version/edition, Cypher version, Java/runtime and deployment. -
Confirm only feature-window relationships enter
atlas-ch25-ml. - Record class counts and split configuration before training.
- Train the non-graph baseline and graph-augmented model under the same protocol.
- Copy actual test metrics, winning parameters, train time and estimate evidence into the experiment log.
- Run at least one predeclared sensitivity/ablation check without test-set fishing.
- Write an explicit graph-added-value conclusion including cost and non-guarantees.
- Drop Chapter 25 models, pipelines and projection and delete only Chapter 25 fixtures.
// Check names before dropping. These calls fail if a named object does not exist,
// so run the corresponding list() command first when rerunning a partial lab.
CALL gds.model.list() YIELD modelName
RETURN modelName ORDER BY modelName;
CALL gds.pipeline.list() YIELD pipelineName
RETURN pipelineName ORDER BY pipelineName;
CALL gds.graph.list() YIELD graphName
RETURN graphName ORDER BY graphName;
// Drop the two Chapter 25 models when present.
CALL gds.model.drop('atlas-ch25-baseline-model') YIELD modelName;
CALL gds.model.drop('atlas-ch25-graph-model') YIELD modelName;
// Drop the two training pipelines when present.
CALL gds.pipeline.drop('atlas-ch25-baseline-pipe') YIELD pipelineName;
CALL gds.pipeline.drop('atlas-ch25-graph-pipe') YIELD pipelineName;
// Drop the in-memory graph when present.
CALL gds.graph.drop('atlas-ch25-ml') YIELD graphName;
// Delete only persisted synthetic Chapter 25 data.
MATCH (n) WHERE n.chapter25 = true DETACH DELETE n;
10. Bridge to Chapter 26: once the model is correct, make the workload sustainable
Chapter 26 moves from analytical correctness to end-to-end performance engineering: workload mix, fan-out, traversal depth, hot nodes, result sizes, query selectivity, heap/page cache/transaction memory, storage, containers and saturation. A graph ML workflow that adds predictive signal but starves production queries is not production-ready.
Check your understanding
- What is the valid outcome if the graph model does not beat the baseline materially?
- Why must both models share one evaluation protocol?
- Why inject a future-edge leak in a disposable graph?
- What is the Community model-capacity implication for sensitivity tests?
- What does Chapter 26 add?
Review the answers
1. Keep the simpler baseline or redesign the graph hypothesis; “graph” is not a required winner.
2. Otherwise a score difference can come from split/tuning/population changes instead of the graph feature.
3. To demonstrate that excellent metrics can signal broken temporal feature integrity.
4. With two primary models retained, train at most one temporary third model at a time and drop it before the next run.
5. Capacity, latency, saturation and workload-isolation engineering around the now-validated graph/application design.
Summary and next step
Evaluate a Graph ML Workflow Against Non-Graph Baselines and Document Where the Graph Actually Adds Signal is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Profile Workload Mix: Read/Write Ratios, Traversal Depth, Fan-Out, Hot Nodes, Result Sizes, and Concurrency. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- GDS Manual v2026.07 — Current GDS manual and release baseline.
- GDS supported Neo4j versions — Compatibility matrix for Neo4j and GDS releases.
- GDS editions — Community/Enterprise limits including four-core concurrency and three-model catalog capacity.
- Node embeddings overview — Current embedding families, quality tiers, inductive/transductive guidance.
- Fast Random Projection — FastRP dimensions, propertyRatio, featureProperties, randomSeed, iterations and modes.
- Node2Vec — Current Node2Vec algorithm and transductive embedding behavior.
- Machine learning pipelines — Current node classification, link prediction and node regression pipeline tiers.
- Node classification pipelines — End-to-end node classification semantics and prediction model behavior.
- Link prediction pipelines — Feature/train/test split semantics and negative examples.
- Link prediction configuration — Split configuration, FastRP property steps, link features and candidate models.
- Link prediction training — Cross-validation, model selection and model-catalog registration semantics.
- Training methods — Supported classification/regression trainers and auto-tuning.
- Pipeline catalog — Pipeline catalog list/exists/drop lifecycle.
- Model catalog listing — Model metadata, training config, schema and loaded/stored/published state.
- Store models on disk — Enterprise-only persistent model storage semantics.
- Getting started ML pipeline — Current link-prediction pipeline example and split workflow.
- GDS server installation — Bundled plugin installation and procedure configuration.
- Neo4j GDS release notes — Current GDS 2026.07.0 release line.
- Neo4j current versions — Current Neo4j database release and LTS line.