Chapter 25 · Graph Embeddings and Machine Learning Pipelines with GDS

Evaluate a Graph ML Workflow Against Non-Graph Baselines and Document Where the Graph Actually Adds Signal

Evaluate AtlasMart graph ML against a non-graph baseline under one protocol, run ablation/leakage checks, and document whether graph features add enough signal to justify production cost.

Advanced260–380 minutesAblation · graph-added valueNeo4j 2026.07.1 · Community mandatoryGDS Community 2026.07.0 · Cypher 25GDS ML · concurrency ≤4 CE · model catalog ≤3Java 21/25 · GDS plugin requiredLast reviewed: September 2026

Learning outcomes

01

Compare baseline and graph-augmented models on the same untouched test protocol and report uncertainty/limits rather than a marketing conclusion.

02

Run a graph-feature ablation and inspect actual catalog metrics without fabricating improvement.

03

Repeat sensitivity checks over seeds/configurations without silently tuning on the test set.

04

Write an explicit “graph adds signal / graph does not add enough signal” decision with cost, drift, fairness, serving and rollback consequences.

05

Clean up projections, pipelines, models and synthetic source data while preserving unrelated academy/database state.

1. AtlasMart final question: does graph information improve the decision enough to pay for itself?

The chapter succeeds even if the graph model does not beat the baseline. The correct conclusion is empirical: compare the same supervised target, feature cutoff, split protocol, trainer family and metric; change only the graph-derived feature set. Then judge the observed lift against added complexity, compute, refresh cadence, serving dependencies, privacy/security scope, and operational risk.

Dimension Chapter 25 reproducible assumption
Neo4j 2026.07.1 Community in the disposable local/container lab; database neo4j.
Cypher Cypher 25 examples. GDS procedures are called from Cypher; no paid notebook or external ML service is required.
Java Java 21/25 supported by the Neo4j 2026 line; the chosen Neo4j distribution/container supplies the runtime.
Auth/TLS User neo4j, password atlasmart-course-2026; loopback Bolt without TLS only for this isolated lab. Production/remote deployments require authenticated encrypted transport.
Driver No application driver is required for mandatory procedure labs; Browser or cypher-shell is sufficient.
GDS GDS Community 2026.07.0. Community includes all algorithms, caps GDS concurrency at four CPU cores, and limits the model catalog to three models.
GDS ML quality tiers Node Classification and Link Prediction pipelines are Beta; Node Regression pipelines are Alpha. Treat tier as a production-risk input.
Persistence Community model/pipeline/graph catalogs are in-memory lifecycle objects. Model persistence to disk is Enterprise-only.
APOC Not required for mandatory Chapter 25 work.
Scope Only persisted entities with chapter25=true and in-memory names beginning atlas-ch25- are created.
Evidence This generation environment does not run Neo4j/GDS. Fixture counts and split rules are deterministic; training scores, memory estimates, timings and predictions must be measured locally and must not be copied as fabricated results.
Verify the Chapter 25 runtime
RETURN gds.version() AS gdsVersion;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;

// Inspect current in-memory catalogs before the lab.
CALL gds.graph.list() YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCount ORDER BY graphName;

CALL gds.pipeline.list() YIELD pipelineName, pipelineType
RETURN pipelineName, pipelineType ORDER BY pipelineName;

CALL gds.model.list() YIELD modelName, modelType, loaded, stored, published
RETURN modelName, modelType, loaded, stored, published ORDER BY modelName;

2. Acceptance criteria before reading any scores

Gate Pass condition
feature time boundary All VIEWED_CH25 relationships and raw features predate the Q3 label horizon.
same population Baseline and graph models target the same CH25Customer nodes and willRepeat property.
same split/search protocol Same testFraction, validationFolds, trainer family, tuning budget and randomSeed for the direct ablation.
catalog capacity No more than three Community models; mandatory comparison uses two.
resource evidence Training estimates + actual trainMillis recorded for both models.
no fabricated output Test metrics are copied only from the learner’s actual gds.model.list/training result.

3. Read the two models from the catalog side by side

Extract measured test evidence
CALL gds.model.list()
YIELD modelName, modelType, modelInfo, trainConfig, graphSchema
WHERE modelName IN ['atlas-ch25-baseline-model','atlas-ch25-graph-model']
RETURN modelName,
       modelType,
       modelInfo.metrics.F1_WEIGHTED.test AS testF1Weighted,
       modelInfo.bestParameters AS bestParameters,
       trainConfig,
       graphSchema
ORDER BY modelName;
Interpretation rule

Do not assume graph > baseline. If the graph model is lower, equal within noise, or only trivially higher, the correct production decision may be to keep the simpler baseline.

4. Ablation matrix: what changed?

Experiment Features What it isolates
Baseline tenureMonths + spendIndex Predictive power available without graph computation.
Graph augmented tenureMonths + spendIndex + FastRP embedding Incremental value of graph-derived representation under the same split/trainer protocol.
Topology only (optional) FastRP embedding only Whether graph structure can carry the signal without raw properties; not required for mandatory two-model limit.
Property-only inductive FastRP (optional) FastRP with numeric featureProperties, propertyRatio=1.0, fixed seed Generalization/representation alternative; should be a separately registered experiment, not silently substituted.

5. Sensitivity without test-set fishing

After the primary comparison is frozen, run a planned sensitivity study. For example, retrain one experiment at a time with seeds 41, 42, and 43, dropping each temporary model before the next so Community model capacity is not exceeded. Compare validation variability and only use a newly reserved holdout period if the experiment design changes after inspecting the original test.

Model-capacity-aware sensitivity pattern
// Example process, not a loop to paste blindly:
// 1. Keep baseline + active graph model in catalog.
// 2. Train at most one temporary sensitivity model (third Community slot).
// 3. Record validation/test evidence under the predeclared protocol.
// 4. Drop the temporary model before the next seed/config.
CALL gds.model.list()
YIELD modelName, modelType
RETURN modelName, modelType ORDER BY modelName;

6. Decision template: graph-added-value statement

Question Evidence to write down
Did graph add signal? Measured test/holdout delta under the same protocol, plus seed/config sensitivity.
Is the lift material? Translate metric change into business threshold/decision utility; avoid “statistically nonzero = operationally worth it.”
What does it cost? Projection + embedding + training memory/time, refresh frequency, write/storage/serving overhead.
What can drift? Graph degree/distribution, product mix, raw properties, class prevalence, embedding coordinate regime, model calibration.
What can leak? Future edges, target-derived properties, tenant connections, identity duplicates, post-outcome data.
Can it be served safely? Model/projection lifecycle, prediction latency/output size, auth/TLS, retries/idempotency, stale-result policy.
What is rollback? Retain baseline/previous model/config and the source cutoff needed to reproduce it.

7. Failure injection: prove the evaluation catches a leaked graph

In a disposable copy only, deliberately add one target-window relationship type such as FUTURE_PURCHASE_CH25 that directly follows willRepeat, then project it into a throwaway graph and observe how easily model metrics can inflate. Do not use the contaminated graph for the real comparison. The educational point is that a “better” score can be a data-integrity alarm rather than progress.

Safe contamination demonstration outline
// DISPOSABLE DEMONSTRATION ONLY. Do not add this relationship to atlas-ch25-ml.
MATCH (c:CH25Customer), (p:CH25Product {productId:'CH25-P-01'})
WHERE c.willRepeat = 1
CREATE (c)-[:FUTURE_PURCHASE_CH25 {chapter25:true, syntheticLeak:true}]->(p);

// Verify the contamination exists, explain why it is forbidden, then delete it.
MATCH (:CH25Customer)-[r:FUTURE_PURCHASE_CH25]->(:CH25Product)
RETURN count(r) AS leakedFutureEdges;

MATCH ()-[r:FUTURE_PURCHASE_CH25]->() DELETE r;
What this proves

The controlled fault proves that graph feature correctness depends on temporal/semantic source integrity. It does not prove a specific amount of metric inflation, because this artifact does not execute training.

8. Production-readiness matrix

Area Chapter 25 gate
Graph/workload fit Graph feature has measured incremental value for the actual supervised decision.
Correctness Cutoff, target, projection, identity, split, negative sampling and leakage checks are documented.
Cardinality/degree Projection size/degree distributions measured; hub/outlier behavior reviewed.
Transactions/concurrency GDS batch work isolated from latency-critical transactions; Community concurrency ≤4.
Memory/CPU/disk/network Projection/training estimates + actual runtime/resource evidence; prediction output bounded.
Indexes/constraints Source business identities constrained/indexed; ML does not replace source invariants.
Driver/retries/idempotency If orchestrated by an app, retries cannot duplicate persistent prediction writes or create multiple model versions accidentally.
Security/tenant risk Projection is authorization-safe; sensitive labels/features and model outputs have explicit access controls.
Backup/recovery Source graph and model/config lineage are recoverable; Community in-memory model is reproducibly retrainable.
Observability Model version, feature age, drift, class balance, metrics, training failures, prediction latency/output monitored.
Version/edition/tier Neo4j/GDS compatibility and Beta/Alpha pipeline tier recorded; Enterprise-only persistence not assumed in Community.
Migration/rollback Shadow evaluate new feature/model/version; retain last validated baseline/model until acceptance passes.

9. Chapter 25 final verification checklist

  1. Record gds.version(), Neo4j version/edition, Cypher version, Java/runtime and deployment.
  2. Confirm only feature-window relationships enter atlas-ch25-ml.
  3. Record class counts and split configuration before training.
  4. Train the non-graph baseline and graph-augmented model under the same protocol.
  5. Copy actual test metrics, winning parameters, train time and estimate evidence into the experiment log.
  6. Run at least one predeclared sensitivity/ablation check without test-set fishing.
  7. Write an explicit graph-added-value conclusion including cost and non-guarantees.
  8. Drop Chapter 25 models, pipelines and projection and delete only Chapter 25 fixtures.
Chapter 25 cleanup/reset
// Check names before dropping. These calls fail if a named object does not exist,
// so run the corresponding list() command first when rerunning a partial lab.
CALL gds.model.list() YIELD modelName
RETURN modelName ORDER BY modelName;
CALL gds.pipeline.list() YIELD pipelineName
RETURN pipelineName ORDER BY pipelineName;
CALL gds.graph.list() YIELD graphName
RETURN graphName ORDER BY graphName;

// Drop the two Chapter 25 models when present.
CALL gds.model.drop('atlas-ch25-baseline-model') YIELD modelName;
CALL gds.model.drop('atlas-ch25-graph-model') YIELD modelName;

// Drop the two training pipelines when present.
CALL gds.pipeline.drop('atlas-ch25-baseline-pipe') YIELD pipelineName;
CALL gds.pipeline.drop('atlas-ch25-graph-pipe') YIELD pipelineName;

// Drop the in-memory graph when present.
CALL gds.graph.drop('atlas-ch25-ml') YIELD graphName;

// Delete only persisted synthetic Chapter 25 data.
MATCH (n) WHERE n.chapter25 = true DETACH DELETE n;

10. Bridge to Chapter 26: once the model is correct, make the workload sustainable

Chapter 26 moves from analytical correctness to end-to-end performance engineering: workload mix, fan-out, traversal depth, hot nodes, result sizes, query selectivity, heap/page cache/transaction memory, storage, containers and saturation. A graph ML workflow that adds predictive signal but starves production queries is not production-ready.

Check your understanding

  1. What is the valid outcome if the graph model does not beat the baseline materially?
  2. Why must both models share one evaluation protocol?
  3. Why inject a future-edge leak in a disposable graph?
  4. What is the Community model-capacity implication for sensitivity tests?
  5. What does Chapter 26 add?
Review the answers

1. Keep the simpler baseline or redesign the graph hypothesis; “graph” is not a required winner.

2. Otherwise a score difference can come from split/tuning/population changes instead of the graph feature.

3. To demonstrate that excellent metrics can signal broken temporal feature integrity.

4. With two primary models retained, train at most one temporary third model at a time and drop it before the next run.

5. Capacity, latency, saturation and workload-isolation engineering around the now-validated graph/application design.

Summary and next step

Evaluate a Graph ML Workflow Against Non-Graph Baselines and Document Where the Graph Actually Adds Signal is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Profile Workload Mix: Read/Write Ratios, Traversal Depth, Fan-Out, Hot Nodes, Result Sizes, and Concurrency. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.