Chapter 26 · Performance Engineering: Data Model, Traversal Shape, Page Cache, Memory, and Workload Isolation
Model-Level Optimization: Relationship Direction/Type, Intermediate Nodes, Denormalization, and Precomputation
Change one AtlasMart graph-model factor at a time and measure how relationship direction/type, intermediate nodes, denormalization, and precomputation alter traversal work, write cost, freshness, and operational risk.
Learning outcomes
Explain how relationship direction and relationship type constrain traversal work without claiming that Cypher clause order controls the planner.
Decide when an intermediate node represents real domain identity/lifecycle and when it merely adds traversal work.
Measure denormalization as a read-versus-write/freshness tradeoff instead of treating duplication as automatically good or bad.
Use precomputation only when the refresh/invalidation contract is explicit and the measured read benefit justifies it.
Compare model variants with equivalent semantics, identical workload envelopes, and before/after plan/latency evidence.
1. AtlasMart problem: the graph model is part of the query plan
The AtlasMart team can buy more CPU, but the recommendation query still begins from a popular product and expands through thousands of historical orders. Hardware cannot erase unnecessary graph work. Model-level optimization changes what relationships/nodes/properties exist so that common questions can be answered with fewer or more selective expansions.
Relationship direction is stored semantics;
Cypher can match without specifying direction, but a direction
that matches the domain lets queries express intended traversal
precisely. Relationship types are also
selectivity dimensions: one generic
INTERACTED_WITH {kind:...} edge forces a different
filter shape than separate meaningful VIEWED,
PURCHASED, or RETURNED relationship
types. Neither design is universally faster; model semantics and
measured workload decide.
| Dimension | Chapter 26 reproducible assumption |
|---|---|
| Neo4j |
2026.07.1 Community. The continuity database remains
neo4j; performance experiments use a separate
disposable container named
atlasmart-neo4j-perf so tuning/failure tests
do not disturb earlier labs.
|
| Cypher | Cypher 25 examples. Planner/operator names and numeric PROFILE values are evidence to capture locally, not constants to memorize. |
| Java | Neo4j 2026 line with a supported Java 21/25 runtime as supplied/required by the chosen distribution. |
| Driver | Neo4j Python driver 6.3.x for the optional load harness; the driver object is shared, sessions/transactions are not shared between worker threads. |
| Auth/TLS |
User neo4j, password
atlasmart-course-2026. Loopback Bolt without
TLS only for the isolated disposable lab;
remote/production traffic should use validated TLS.
|
| Ports |
Performance container maps HTTP
17474→7474 and Bolt 17687→7687,
avoiding the continuity container on 7474/7687.
|
| Initial resource envelope |
Exercise baseline: Docker limit 2 GiB,
2 CPUs, explicit heap 512 MiB,
explicit page cache 512 MiB. These are lab
controls, not production recommendations.
|
| Dataset | 200 CH26 customers, 300 products, 2,000 orders, 6,000 CONTAINS relationships, 1,200 VIEWED_CH26 relationships, six categories, plus an isolated hot-counter node. Recount locally after setup. |
| Observability | Community labs rely on PROFILE, SHOW commands, Docker/OS counters, driver timing, container stats and logs. Neo4j metrics exporters and query.log are Enterprise surfaces and are optional, clearly labeled. |
| Evidence rule | This generated material does not execute your Docker host. Latencies, DB Hits, page-cache behavior, saturation points, GC, throughput, errors and recovery time must be measured locally; illustrative tables are labeled as templates. |
// Run against bolt://localhost:17687, database neo4j.
// Cleanup only the Chapter 26 fixture.
MATCH (n) WHERE n.chapter26 = true DETACH DELETE n;
CREATE CONSTRAINT ch26_customer_id IF NOT EXISTS
FOR (c:CH26Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch26_product_id IF NOT EXISTS
FOR (p:CH26Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch26_order_id IF NOT EXISTS
FOR (o:CH26Order) REQUIRE o.orderId IS UNIQUE;
CREATE INDEX ch26_order_day IF NOT EXISTS
FOR (o:CH26Order) ON (o.createdDay);
UNWIND range(1,6) AS i
CREATE (:Category:CH26Category {categoryId:'CH26-CAT-'+toString(i), chapter26:true});
UNWIND range(1,200) AS i
CREATE (:Customer:CH26Customer {
customerId:'CH26-C-'+right('0000'+toString(i),4),
segment: CASE WHEN i % 5 = 0 THEN 'business' ELSE 'consumer' END,
chapter26:true
});
UNWIND range(1,300) AS i
CREATE (:Product:CH26Product {
productId:'CH26-P-'+right('0000'+toString(i),4),
name:'Atlas Product '+toString(i),
price:toFloat(10 + (i % 90)),
chapter26:true
});
MATCH (p:CH26Product), (cat:CH26Category)
WHERE toInteger(right(p.productId,4)) % 6 + 1 = toInteger(replace(cat.categoryId,'CH26-CAT-',''))
CREATE (p)-[:IN_CATEGORY {chapter26:true}]->(cat);
UNWIND range(1,2000) AS i
MATCH (c:CH26Customer {customerId:'CH26-C-'+right('0000'+toString(((i-1)%200)+1),4)})
CREATE (o:Order:CH26Order {
orderId:'CH26-O-'+right('00000'+toString(i),5),
createdDay:i % 120,
status:CASE WHEN i % 9 = 0 THEN 'RETURNED' ELSE 'COMPLETE' END,
total:toFloat(40 + (i % 260)),
chapter26:true
})
CREATE (c)-[:PLACED {chapter26:true}]->(o);
MATCH (o:CH26Order)
WITH o, toInteger(right(o.orderId,5)) AS i
UNWIND range(0,2) AS j
WITH o, i, j,
CASE WHEN j=0 AND i % 5 = 0 THEN 1
ELSE ((i*17 + j*43) % 300)+1 END AS pno
MATCH (p:CH26Product {productId:'CH26-P-'+right('0000'+toString(pno),4)})
CREATE (o)-[:CONTAINS {quantity:1+(i+j)%3, chapter26:true}]->(p);
MATCH (c:CH26Customer)
WITH c, toInteger(right(c.customerId,4)) AS i
UNWIND range(0,5) AS j
WITH c, ((i*13 + j*29) % 300)+1 AS pno
MATCH (p:CH26Product {productId:'CH26-P-'+right('0000'+toString(pno),4)})
CREATE (c)-[:VIEWED_CH26 {chapter26:true}]->(p);
MERGE (:CH26HotCounter {id:'global', n:0, chapter26:true});
MATCH (c:CH26Customer) WITH count(c) AS customers
MATCH (p:CH26Product) WITH customers, count(p) AS products
MATCH (o:CH26Order) WITH customers, products, count(o) AS orders
MATCH (:CH26Order)-[r:CONTAINS]->(:CH26Product)
WITH customers, products, orders, count(r) AS contains
MATCH (:CH26Customer)-[v:VIEWED_CH26]->(:CH26Product)
RETURN customers, products, orders, contains, count(v) AS viewed;
2. Direction and type: model the question you ask
| Model choice | Potential benefit | Potential cost / boundary |
|---|---|---|
| Specific relationship type | Traversal can target one semantic edge family; clearer invariants. | Type explosion is a modeling smell if values are unbounded/dynamic rather than semantic categories. |
| Stored direction aligned to domain | Queries and integrity rules communicate source/target roles clearly. | Cypher can match undirected; reversing duplicate edges solely “for speed” doubles write/storage/invariant work unless evidence justifies it. |
| Intermediate node | Can give an event/line/agreement its own identity, properties, multiple relationships and lifecycle. | Adds another entity/hop when the relationship itself already models the fact sufficiently. |
| Denormalized property | Avoids repeated traversal/aggregation for a stable, hot read. | Creates duplicate state that must be updated atomically or reconciled; stale reads become a correctness risk. |
| Precomputed relationship/property | Can materialize expensive derived structure for repeated queries. | Refresh cost, invalidation, storage and temporal semantics may outweigh saved read work. |
3. Experiment A: a generic interaction edge versus a typed edge
Create a deliberately isolated generic representation for the same six Chapter 26 views per customer. Do not mix it into production labels. Compare equivalent queries: both must return the same customer/product pairs before timing.
MATCH (c:CH26Customer)-[v:VIEWED_CH26]->(p:CH26Product)
MERGE (c)-[g:CH26_INTERACTED_WITH {kind:'VIEWED'}]->(p)
SET g.chapter26 = true;
MATCH (:CH26Customer)-[a:VIEWED_CH26]->(:CH26Product)
WITH count(a) AS typedCount
MATCH (:CH26Customer)-[b:CH26_INTERACTED_WITH {kind:'VIEWED'}]->(:CH26Product)
RETURN typedCount, count(b) AS genericCount;
PROFILE
MATCH (c:CH26Customer {customerId:$cid})-[:VIEWED_CH26]->(p:CH26Product)
RETURN p.productId ORDER BY p.productId;
PROFILE
MATCH (c:CH26Customer {customerId:$cid})-[r:CH26_INTERACTED_WITH]->(p:CH26Product)
WHERE r.kind='VIEWED'
RETURN p.productId ORDER BY p.productId;
Do not assert that the typed version is faster merely because it “looks selective.” Compare equivalent result sets, Rows/DB Hits/plan shape and latency distributions on your graph. The primary reason for a type must still be semantic correctness.
4. Experiment B: precompute an aggregate with a freshness contract
AtlasMart repeatedly displays each customer’s completed-order
count. The normalized answer expands PLACED and
filters orders every time. A cached property can avoid this
work, but only if there is a defined writer/reconciliation path.
MATCH (c:CH26Customer)
OPTIONAL MATCH (c)-[:PLACED]->(o:CH26Order)
WHERE o.status='COMPLETE'
WITH c, count(o) AS n
SET c.ch26CompleteOrderCount=n;
// Compare normalized and precomputed answers; mismatch must be zero.
MATCH (c:CH26Customer)
OPTIONAL MATCH (c)-[:PLACED]->(o:CH26Order)
WHERE o.status='COMPLETE'
WITH c, count(o) AS truth
WITH c, truth, c.ch26CompleteOrderCount AS cached
WHERE truth <> cached
RETURN count(*) AS mismatches;
PROFILE
MATCH (c:CH26Customer {customerId:$cid})
RETURN c.ch26CompleteOrderCount AS completed;
The precomputed property changes the write contract: each order-status transition must update/rebuild it consistently, or the UI must tolerate bounded staleness. In a multi-writer system, recomputation jobs and transactional increments have different contention and recovery tradeoffs.
5. Intermediate nodes: identity before speed
An OrderLine node is justified when line items need
independent identity, lifecycle, tax/refund relationships,
fulfillment splits, or cross-system references. If the only
facts are quantity and unit price between one order
and one product, a relationship such as
CONTAINS with properties may be simpler.
Performance follows from actual traversal patterns, not a rule
that “nodes are faster than relationship properties” or the
opposite.
When comparing representations, validate invariants first: same order totals, same number of logical lines, same product references, same temporal semantics. Then measure read/write amplification, store size, import/update complexity, and the query classes that matter.
6. Deliberately wrong approach: duplicate both relationship directions “for speed”
Creating both
(customer)-[:VIEWED]->(product) and
(product)-[:VIEWED_BY]->(customer) for every
event may seem to make both traversal directions fast. Neo4j
already stores relationship direction and Cypher can traverse in
either direction. The duplicate representation creates two facts
that can diverge, doubles write/storage work, complicates
CDC/backups and can inflate counts if a query matches both.
Repair: keep one semantic fact unless an independently managed/materialized relationship has a measured use case. If you introduce a derived reverse/summary edge, name it as derived, define its refresh/rebuild procedure, exclude it from canonical counts, and add a reconciliation query. Verify performance only after correctness is established.
Production judgment
Model optimizations move cost; they do not delete it. Typed/directed edges can reduce irrelevant expansion, but high-degree hubs still dominate. Denormalization/precompute can reduce read latency but adds write amplification, lock contention, recovery/rebuild work and stale-state risk. Intermediate nodes can improve domain correctness even if they add a hop. Keep constraints/indexes on stable identifiers, preserve tenant/security boundaries, budget page cache/store growth, measure transaction-log/backup impact, and keep a migration/rollback path for model changes. Lesson 3 keeps the model fixed and optimizes the Cypher pipeline itself.
Check your understanding
- Why should relationship type be chosen before benchmarking?
- When is an intermediate node appropriate?
- What new invariant does a denormalized count create?
- Why is duplicating reverse relationships dangerous?
- How do you prove a model optimization helped?
Review the answers
1. Because it expresses domain semantics; performance is a secondary measured consequence, not the only justification.
2. When the concept needs its own identity, lifecycle, relationships or independently addressable properties—not merely because an extra node is presumed faster.
3. The cached value must remain consistent with or intentionally bounded relative to the normalized source of truth.
4. It duplicates a fact, increases write/store cost and can diverge or double-count unless managed as an explicit derived representation.
5. Run equivalent correctness checks and controlled before/after plan/latency/resource measurements under the same workload envelope.
Summary and next step
Model-Level Optimization: Relationship Direction/Type, Intermediate Nodes, Denormalization, and Precomputation is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Query-Level Optimization: Anchors, Selectivity, Predicate Placement, Subqueries, Aggregation, and Result Shaping. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- Neo4j current versions — Current database release and LTS baseline.
- Neo4j Operations Manual — Performance — Current performance topics and operational tuning surface.
- Memory configuration — Heap, page cache, transaction/native memory, OS headroom, and memory recommendation guidance.
- Disks, RAM and other tips — Page-cache warmup, storage and RAM behavior.
- Configuration settings — Authoritative current setting names and edition/dynamic boundaries.
- Docker configuration — Container configuration mapping and production configuration guidance.
- Cypher execution plans — EXPLAIN/PROFILE semantics and runtime evidence.
- Cypher operators in detail — Current operators, Rows, DB Hits, memory and plan behavior.
- Indexes for search performance — Current range/text/point/token index behavior and syntax.
- Cypher query tuning — Planner, statistics and query-tuning concepts.
- Neo4j Python driver performance — Driver-side result streaming, database selection and performance guidance.
- Neo4j Python driver API — Connection pool, timeout, retry and fetch-size configuration.
- Neo4j logging — Current debug/query/security/GC logging surfaces and edition boundaries.
- Neo4j metrics — Enterprise metrics surfaces and operational evidence.
- Transaction management — Transaction lifecycle and operational behavior.
- Java requirements — Current supported Java/runtime and platform requirements.
- Neo4j status codes — Classify transient/client/database failures rather than collapsing them into latency.