Chapter 25 · Graph Embeddings and Machine Learning Pipelines with GDS
Feature Engineering from Graph Algorithms, Properties, Labels, and Leakage Prevention
Engineer AtlasMart raw and graph-derived features with explicit cutoff/lineage controls, diagnose graph-specific leakage, and configure matched baseline versus graph-augmented feature pipelines.
Learning outcomes
Build a feature lineage table that distinguishes raw properties, graph-algorithm features, embeddings, labels, and forbidden future information.
Explain feature leakage, target leakage, temporal leakage, identity leakage, and graph-topology leakage with concrete AtlasMart examples.
Use graph projections whose feature cutoff precedes the label horizon and document what information each node-property step can see.
Compare raw/tabular and graph-derived feature sets through ablation rather than assuming graph features help.
Define production controls for feature refresh, lineage, tenant boundaries, fairness/risk, and rollback.
1. AtlasMart problem: the easiest way to “improve” a model is to accidentally show it the answer
Suppose willRepeat means “customer purchased in
2026-Q3.” A graph feature computed from relationships that
include 2026-Q3 purchases is not predictive evidence; it is a
disguised copy of the target. Graph ML makes leakage easier
because information can travel through neighbors, paths,
aggregates, embeddings, and communities even when the target
property itself is excluded.
The safe design begins with an information boundary: every feature must be computable at the decision time. The Chapter 25 fixture therefore loads only a Q2 interaction graph and keeps the Q3 synthetic label as the supervised target.
| Dimension | Chapter 25 reproducible assumption |
|---|---|
| Neo4j |
2026.07.1 Community in the disposable local/container lab;
database neo4j.
|
| Cypher | Cypher 25 examples. GDS procedures are called from Cypher; no paid notebook or external ML service is required. |
| Java | Java 21/25 supported by the Neo4j 2026 line; the chosen Neo4j distribution/container supplies the runtime. |
| Auth/TLS |
User neo4j, password
atlasmart-course-2026; loopback Bolt without
TLS only for this isolated lab. Production/remote
deployments require authenticated encrypted transport.
|
| Driver | No application driver is required for mandatory procedure labs; Browser or cypher-shell is sufficient. |
| GDS | GDS Community 2026.07.0. Community includes all algorithms, caps GDS concurrency at four CPU cores, and limits the model catalog to three models. |
| GDS ML quality tiers | Node Classification and Link Prediction pipelines are Beta; Node Regression pipelines are Alpha. Treat tier as a production-risk input. |
| Persistence | Community model/pipeline/graph catalogs are in-memory lifecycle objects. Model persistence to disk is Enterprise-only. |
| APOC | Not required for mandatory Chapter 25 work. |
| Scope |
Only persisted entities with
chapter25=true and in-memory names beginning
atlas-ch25- are created.
|
| Evidence | This generation environment does not run Neo4j/GDS. Fixture counts and split rules are deterministic; training scores, memory estimates, timings and predictions must be measured locally and must not be copied as fabricated results. |
RETURN gds.version() AS gdsVersion;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;
// Inspect current in-memory catalogs before the lab.
CALL gds.graph.list() YIELD graphName, nodeCount, relationshipCount
RETURN graphName, nodeCount, relationshipCount ORDER BY graphName;
CALL gds.pipeline.list() YIELD pipelineName, pipelineType
RETURN pipelineName, pipelineType ORDER BY pipelineName;
CALL gds.model.list() YIELD modelName, modelType, loaded, stored, published
RETURN modelName, modelType, loaded, stored, published ORDER BY modelName;
// Remove only Chapter 25 persisted fixtures before recreating them.
MATCH (n) WHERE n.chapter25 = true DETACH DELETE n;
// 30 customers. The target is a synthetic future-horizon label for teaching only.
UNWIND range(1,30) AS i
CREATE (:Customer:CH25Customer {
customerId: 'CH25-C-' + right('00' + toString(i), 2),
seq: i,
tenureMonths: toFloat(3 + (i % 18)),
spendIndex: toFloat(20 + ((i * 17) % 80)),
willRepeat: CASE WHEN i % 3 = 0 OR i % 7 = 0 THEN 0 ELSE 1 END,
featureCutoff: date('2026-06-30'),
labelWindow: '2026-Q3',
chapter25: true
});
UNWIND range(1,8) AS j
CREATE (:Product:CH25Product {
productId: 'CH25-P-' + right('00' + toString(j), 2),
categoryCode: toFloat(j % 4),
chapter25: true
});
// Feature-window interactions only. No Q3 target-window interaction is loaded.
MATCH (c:CH25Customer), (p:CH25Product)
WHERE ((c.seq + toInteger(right(p.productId,2))) % 4 = 0)
OR (toInteger(right(p.productId,2)) = ((c.seq - 1) % 8) + 1)
CREATE (c)-[:VIEWED_CH25 {chapter25:true, asOf:'2026-Q2'}]->(p);
MATCH (c:CH25Customer)
RETURN count(c) AS customers,
sum(CASE WHEN c.willRepeat = 1 THEN 1 ELSE 0 END) AS positiveLabels,
sum(CASE WHEN c.willRepeat = 0 THEN 1 ELSE 0 END) AS negativeLabels;
MATCH (:CH25Customer)-[r:VIEWED_CH25]->(:CH25Product)
RETURN count(r) AS featureWindowRelationships;
// Drop atlas-ch25-ml first if it already exists.
CALL gds.graph.project(
'atlas-ch25-ml',
{
CH25Customer: {properties: ['tenureMonths','spendIndex','willRepeat']},
CH25Product: {properties: ['categoryCode']}
},
{
VIEWED_CH25: {orientation:'UNDIRECTED'}
}
)
YIELD graphName, nodeCount, relationshipCount, projectMillis
RETURN graphName, nodeCount, relationshipCount, projectMillis;
CALL gds.graph.list('atlas-ch25-ml')
YIELD graphName, nodeCount, relationshipCount, schema
RETURN graphName, nodeCount, relationshipCount, schema;
2. Feature engineering inventory
| Feature family | AtlasMart example | Leakage question |
|---|---|---|
| raw node property |
tenureMonths, spendIndex
|
Was the value known by the feature cutoff, and is its definition stable between training and serving? |
| topology count | degree / number of viewed-product neighbors | Do relationships include target-window/future events or restricted tenant data? |
| centrality/community | PageRank/community ID from a Q2 graph | Was the algorithm run only on the allowed snapshot, and are opaque community IDs being treated as durable semantics? |
| embedding | FastRP vector | Which labels/types/properties and cutoff went into the projection? Is the representation inductive or transductive? |
| target property | willRepeat |
Must be available to training evaluation only, never selected as a feature. |
| post-outcome attribute | refund reason recorded after Q3 purchase | Forbidden: directly downstream of the target outcome. |
3. Leakage mechanisms unique to graphs
| Leakage mode | Concrete mechanism | Repair |
|---|---|---|
| future-edge leakage | Projection contains purchase/view edges after the decision timestamp. | Materialize or filter an as-of snapshot before projection. |
| label propagation leakage | A target or proxy label is used as a node property step or neighbor aggregate. | Remove label-derived properties from all feature-producing steps; audit transitive dependencies. |
| edge-target leakage | For link prediction, the relationship being predicted remains available to an embedding/community step. | Use the pipeline feature-input split or a pre-cut graph that removes held-out target relationships. |
| entity leakage | Same real-world entity appears in train/test through duplicate nodes/accounts/devices. | Split/group by stable business entity when the use case requires entity-generalization. |
| tenant leakage | One tenant’s graph connections become features for another tenant. | Project only authorized scope; security filtering after feature computation is too late. |
| restore/ID leakage | Model relies on transductive node identity after store reload/import changes identity semantics. | Retrain together or use a documented inductive representation. |
4. Build a baseline feature set and a graph-augmented feature set
CALL gds.beta.pipeline.nodeClassification.create('atlas-ch25-baseline-pipe');
CALL gds.beta.pipeline.nodeClassification.selectFeatures(
'atlas-ch25-baseline-pipe',
['tenureMonths','spendIndex']
);
CALL gds.beta.pipeline.nodeClassification.configureSplit(
'atlas-ch25-baseline-pipe',
{testFraction:0.2, validationFolds:3}
);
CALL gds.beta.pipeline.nodeClassification.configureAutoTuning(
'atlas-ch25-baseline-pipe',
{maxTrials:5}
);
CALL gds.beta.pipeline.nodeClassification.addLogisticRegression(
'atlas-ch25-baseline-pipe',
{penalty:{range:[0.0001,1.0]}}
);
CALL gds.pipeline.list('atlas-ch25-baseline-pipe')
YIELD pipelineName, pipelineType, pipelineInfo
RETURN pipelineName, pipelineType, pipelineInfo;
CALL gds.beta.pipeline.nodeClassification.create('atlas-ch25-graph-pipe');
CALL gds.beta.pipeline.nodeClassification.addNodeProperty(
'atlas-ch25-graph-pipe',
'fastRP',
{
mutateProperty:'ch25Embedding',
embeddingDimension:16,
iterationWeights:[0.0,1.0,1.0],
randomSeed:42,
contextNodeLabels:['CH25Customer','CH25Product'],
contextRelationshipTypes:['VIEWED_CH25']
}
);
CALL gds.beta.pipeline.nodeClassification.selectFeatures(
'atlas-ch25-graph-pipe',
['tenureMonths','spendIndex','ch25Embedding']
);
CALL gds.beta.pipeline.nodeClassification.configureSplit(
'atlas-ch25-graph-pipe',
{testFraction:0.2, validationFolds:3}
);
CALL gds.beta.pipeline.nodeClassification.configureAutoTuning(
'atlas-ch25-graph-pipe',
{maxTrials:5}
);
CALL gds.beta.pipeline.nodeClassification.addLogisticRegression(
'atlas-ch25-graph-pipe',
{penalty:{range:[0.0001,1.0]}}
);
A node-property step added to the pipeline is recomputed as part of training/prediction. That is safer than manually creating an embedding once and silently reusing it after the graph has changed—but only if the graph snapshot and feature lineage remain valid.
5. Deliberately wrong approach: random split on a graph containing future edges
A random node split does not repair temporal leakage. If the projection contains Q3 interactions and the target is a Q3 outcome, every fold is contaminated even though train and test node IDs are disjoint.
Repair: enforce the time boundary before the pipeline sees the graph. For a real temporal deployment, also reserve a later-time holdout outside ordinary model selection. The GDS pipeline’s internal test/cross-validation machinery is useful model-selection machinery; it is not a substitute for business-specific time/entity split design.
Production judgment
Feature lineage should identify owner, definition, source system, cutoff, refresh cadence, projection query, GDS algorithm/config/version, tenant/security scope, and the models consuming it. Treat fairness/risk as domain-specific: a structurally predictive feature can still encode undesirable proxies. Keep graph-feature generation off latency-sensitive transactional paths unless measured capacity shows headroom; observe drift in both raw feature distributions and graph structure; and retain a rollback path to the last validated feature set/model.
Check your understanding
- Can a feature leak the target without containing the target property?
- What does a random train/test split guarantee about time leakage?
- Why compare baseline versus graph-augmented features?
- When should tenant/security filtering occur?
- What is the bridge to Lesson 3?
Review the answers
1. Yes. Future edges, neighbor properties, aggregates, communities or embeddings can transmit target/proxy information.
2. Nothing. The graph and properties must already respect the feature cutoff.
3. To prove whether graph information adds measurable out-of-sample signal beyond ordinary features.
4. Before or during projection/feature computation, not only after the model has consumed cross-tenant information.
5. Once features are trustworthy, the next question is how each supervised pipeline splits, samples, trains and evaluates its task.
Summary and next step
Feature Engineering from Graph Algorithms, Properties, Labels, and Leakage Prevention is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Node Classification, Link Prediction, and Regression Pipeline Concepts, Splits, Metrics, and Baselines. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- GDS Manual v2026.07 — Current GDS manual and release baseline.
- GDS supported Neo4j versions — Compatibility matrix for Neo4j and GDS releases.
- GDS editions — Community/Enterprise limits including four-core concurrency and three-model catalog capacity.
- Node embeddings overview — Current embedding families, quality tiers, inductive/transductive guidance.
- Fast Random Projection — FastRP dimensions, propertyRatio, featureProperties, randomSeed, iterations and modes.
- Node2Vec — Current Node2Vec algorithm and transductive embedding behavior.
- Machine learning pipelines — Current node classification, link prediction and node regression pipeline tiers.
- Node classification pipelines — End-to-end node classification semantics and prediction model behavior.
- Link prediction pipelines — Feature/train/test split semantics and negative examples.
- Link prediction configuration — Split configuration, FastRP property steps, link features and candidate models.
- Link prediction training — Cross-validation, model selection and model-catalog registration semantics.
- Training methods — Supported classification/regression trainers and auto-tuning.
- Pipeline catalog — Pipeline catalog list/exists/drop lifecycle.
- Model catalog listing — Model metadata, training config, schema and loaded/stored/published state.
- Store models on disk — Enterprise-only persistent model storage semantics.
- Getting started ML pipeline — Current link-prediction pipeline example and split workflow.
- GDS server installation — Bundled plugin installation and procedure configuration.
- Neo4j GDS release notes — Current GDS 2026.07.0 release line.
- Neo4j current versions — Current Neo4j database release and LTS line.