Chapter 26 · Performance Engineering: Data Model, Traversal Shape, Page Cache, Memory, and Workload Isolation

Profile Workload Mix: Read/Write Ratios, Traversal Depth, Fan-Out, Hot Nodes, Result Sizes, and Concurrency

Build an evidence-first AtlasMart workload profile that quantifies read/write mix, traversal depth, fan-out, hot nodes, result size, concurrency, latency distributions, and saturation before tuning anything.

Advanced260–380 minutesWorkload profile · fan-out · saturationNeo4j 2026.07.1 · Community mandatoryCypher 25 · Python driver 6.3.x optionalDisposable perf container · 2 GiB/2 CPU baselineJava 21/25 · no APOC/GDS requiredLast reviewed: September 2026

Learning outcomes

01

Turn an application description into a quantified workload mix: reads/writes, query classes, concurrency, result sizes, and SLO-relevant latency distributions.

02

Measure traversal depth and fan-out from the actual AtlasMart graph rather than assuming every relationship expansion has similar cost.

03

Recognize hot nodes and skew as concurrency/locking/cache phenomena that averages can hide.

04

Establish a reproducible baseline with dataset, indexes, warmup, client/pool settings, resource envelope, and failure counts recorded.

05

Use saturation curves and tail latency to decide what to optimize next instead of tuning heap, indexes, or hardware by intuition.

1. AtlasMart problem: “Neo4j is slow” is not a workload description

AtlasMart receives a complaint that the product recommendation endpoint is “slow after lunch.” The same Neo4j database also serves order history, catalog search joins, background writes, and an admin report. A performance engineer cannot responsibly tune the server until the workload is decomposed into observable classes. Workload mix means the proportion and shape of operations the system actually executes; concurrency means how many operations overlap in time; tail latency is the slow end of the latency distribution (for example p95/p99), not an average.

Graph cost is particularly sensitive to traversal depth (how many relationship hops a pattern can expand) and fan-out (how many relationships each step can branch into). A query with two hops over degree 3 is not the same workload as two hops from a product connected to hundreds of thousands of orders. A hot node is an entity touched disproportionately often; hot reads can dominate cache behavior and hot writes can serialize on locks.

Dimension Chapter 26 reproducible assumption
Neo4j 2026.07.1 Community. The continuity database remains neo4j; performance experiments use a separate disposable container named atlasmart-neo4j-perf so tuning/failure tests do not disturb earlier labs.
Cypher Cypher 25 examples. Planner/operator names and numeric PROFILE values are evidence to capture locally, not constants to memorize.
Java Neo4j 2026 line with a supported Java 21/25 runtime as supplied/required by the chosen distribution.
Driver Neo4j Python driver 6.3.x for the optional load harness; the driver object is shared, sessions/transactions are not shared between worker threads.
Auth/TLS User neo4j, password atlasmart-course-2026. Loopback Bolt without TLS only for the isolated disposable lab; remote/production traffic should use validated TLS.
Ports Performance container maps HTTP 17474→7474 and Bolt 17687→7687, avoiding the continuity container on 7474/7687.
Initial resource envelope Exercise baseline: Docker limit 2 GiB, 2 CPUs, explicit heap 512 MiB, explicit page cache 512 MiB. These are lab controls, not production recommendations.
Dataset 200 CH26 customers, 300 products, 2,000 orders, 6,000 CONTAINS relationships, 1,200 VIEWED_CH26 relationships, six categories, plus an isolated hot-counter node. Recount locally after setup.
Observability Community labs rely on PROFILE, SHOW commands, Docker/OS counters, driver timing, container stats and logs. Neo4j metrics exporters and query.log are Enterprise surfaces and are optional, clearly labeled.
Evidence rule This generated material does not execute your Docker host. Latencies, DB Hits, page-cache behavior, saturation points, GC, throughput, errors and recovery time must be measured locally; illustrative tables are labeled as templates.
Create the isolated Neo4j 2026.07.1 performance container
# PowerShell — isolated performance lab. Safe blast radius: this container/volumes only.
$Name = "atlasmart-neo4j-perf"
docker rm -f $Name 2>$null

docker run -d --name $Name `
  --memory=2g --cpus=2 `
  -p 127.0.0.1:17474:7474 -p 127.0.0.1:17687:7687 `
  -e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
  -e NEO4J_server_memory_heap_initial__size=512M `
  -e NEO4J_server_memory_heap_max__size=512M `
  -e NEO4J_server_memory_pagecache_size=512M `
  -v atlasmart-neo4j-perf-data:/data `
  -v atlasmart-neo4j-perf-logs:/logs `
  neo4j:2026.07.1

# Verify the container is healthy before loading data.
docker ps --filter "name=$Name"
docker logs --tail 80 $Name
Seed a deterministic, skewed AtlasMart fixture
// Run against bolt://localhost:17687, database neo4j.
// Cleanup only the Chapter 26 fixture.
MATCH (n) WHERE n.chapter26 = true DETACH DELETE n;

CREATE CONSTRAINT ch26_customer_id IF NOT EXISTS
FOR (c:CH26Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch26_product_id IF NOT EXISTS
FOR (p:CH26Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch26_order_id IF NOT EXISTS
FOR (o:CH26Order) REQUIRE o.orderId IS UNIQUE;
CREATE INDEX ch26_order_day IF NOT EXISTS
FOR (o:CH26Order) ON (o.createdDay);

UNWIND range(1,6) AS i
CREATE (:Category:CH26Category {categoryId:'CH26-CAT-'+toString(i), chapter26:true});

UNWIND range(1,200) AS i
CREATE (:Customer:CH26Customer {
  customerId:'CH26-C-'+right('0000'+toString(i),4),
  segment: CASE WHEN i % 5 = 0 THEN 'business' ELSE 'consumer' END,
  chapter26:true
});

UNWIND range(1,300) AS i
CREATE (:Product:CH26Product {
  productId:'CH26-P-'+right('0000'+toString(i),4),
  name:'Atlas Product '+toString(i),
  price:toFloat(10 + (i % 90)),
  chapter26:true
});

MATCH (p:CH26Product), (cat:CH26Category)
WHERE toInteger(right(p.productId,4)) % 6 + 1 = toInteger(replace(cat.categoryId,'CH26-CAT-',''))
CREATE (p)-[:IN_CATEGORY {chapter26:true}]->(cat);

UNWIND range(1,2000) AS i
MATCH (c:CH26Customer {customerId:'CH26-C-'+right('0000'+toString(((i-1)%200)+1),4)})
CREATE (o:Order:CH26Order {
  orderId:'CH26-O-'+right('00000'+toString(i),5),
  createdDay:i % 120,
  status:CASE WHEN i % 9 = 0 THEN 'RETURNED' ELSE 'COMPLETE' END,
  total:toFloat(40 + (i % 260)),
  chapter26:true
})
CREATE (c)-[:PLACED {chapter26:true}]->(o);

MATCH (o:CH26Order)
WITH o, toInteger(right(o.orderId,5)) AS i
UNWIND range(0,2) AS j
WITH o, i, j,
     CASE WHEN j=0 AND i % 5 = 0 THEN 1
          ELSE ((i*17 + j*43) % 300)+1 END AS pno
MATCH (p:CH26Product {productId:'CH26-P-'+right('0000'+toString(pno),4)})
CREATE (o)-[:CONTAINS {quantity:1+(i+j)%3, chapter26:true}]->(p);

MATCH (c:CH26Customer)
WITH c, toInteger(right(c.customerId,4)) AS i
UNWIND range(0,5) AS j
WITH c, ((i*13 + j*29) % 300)+1 AS pno
MATCH (p:CH26Product {productId:'CH26-P-'+right('0000'+toString(pno),4)})
CREATE (c)-[:VIEWED_CH26 {chapter26:true}]->(p);

MERGE (:CH26HotCounter {id:'global', n:0, chapter26:true});

MATCH (c:CH26Customer) WITH count(c) AS customers
MATCH (p:CH26Product) WITH customers, count(p) AS products
MATCH (o:CH26Order) WITH customers, products, count(o) AS orders
MATCH (:CH26Order)-[r:CONTAINS]->(:CH26Product)
WITH customers, products, orders, count(r) AS contains
MATCH (:CH26Customer)-[v:VIEWED_CH26]->(:CH26Product)
RETURN customers, products, orders, contains, count(v) AS viewed;

2. Profile the graph before profiling the query

Record graph size, relationship counts and degree distribution before any timing run. The average degree alone is insufficient: one flagship product is intentionally included in many orders, so the high-percentile degree tells a different story. Keep the same dataset and index state while comparing one change.

Measure counts and degree skew
MATCH (p:CH26Product)
OPTIONAL MATCH (p)<-[r:CONTAINS]-(:CH26Order)
WITH p, count(r) AS degree
RETURN count(*) AS products,
       min(degree) AS minDegree,
       percentileCont(degree,0.50) AS p50Degree,
       percentileCont(degree,0.95) AS p95Degree,
       max(degree) AS maxDegree;

MATCH (c:CH26Customer)-[:PLACED]->(o:CH26Order)-[:CONTAINS]->(p:CH26Product)
RETURN count(DISTINCT c) AS customers,
       count(DISTINCT o) AS orders,
       count(*) AS orderProductEdges;
Expected invariant, not a fabricated measurement

The fixture should reconcile to 200 customers, 300 products, 2,000 orders and 6,000 order→product relationships. The degree percentiles and exact flagship-product degree should be captured from your run. If counts differ, stop and repair the fixture before benchmarking.

3. Build a workload matrix

Question Evidence to record Why it changes performance
Read/write ratio Operations per class over a representative interval Writes add locking, transaction-log and store-update cost; read-only tests can miss contention.
Traversal depth Bounded hop count and pattern shape per query Branching can multiply rows before the final projection shrinks them.
Fan-out/skew p50/p95/max degree for key labels/types The hot tail can dominate work even when the average looks small.
Result size Rows and approximate serialized bytes returned Large responses consume server/driver memory and network time after graph matching.
Concurrency Active requests, driver pool utilization, queued acquisition Single-client latency does not reveal saturation or queueing.
SLO p50/p95/p99, throughput, errors/timeouts An optimization that improves mean latency but worsens p99 may violate the user-visible objective.

A ratio such as “90% reads / 10% writes” is not enough. Keep separate query classes because a one-row indexed lookup and a dense multi-hop read are both reads but have very different resource footprints.

Capture plan evidence for representative query classes
// Do not copy numeric values from another machine. Run both EXPLAIN and PROFILE locally.
EXPLAIN
MATCH (c:CH26Customer {customerId:$customerId})-[:VIEWED_CH26]->(p:CH26Product)-[:IN_CATEGORY]->(cat:CH26Category)
RETURN cat.categoryId, count(*) AS views
ORDER BY views DESC;

PROFILE
MATCH (p:CH26Product {productId:$productId})<-[:CONTAINS]-(o:CH26Order)<-[:PLACED]-(c:CH26Customer)
RETURN c.customerId, o.orderId
ORDER BY o.orderId LIMIT 20;

4. Warmup, cold tests, and saturation are different experiments

Neo4j page cache begins cold after startup and warms as store pages are touched. Therefore, a warm steady-state benchmark should have a separately documented warmup phase, while a restart/cold-cache test should be labeled as such. Never silently discard the cold behavior if restart recovery is part of the SLO.

A saturation curve increases offered concurrency while observing throughput, p95/p99 latency, CPU, I/O, pool wait and failure rate. Saturation is the region where more offered work produces little throughput gain while latency/queueing/error rates rise. The knee is workload- and hardware-specific; this chapter does not prescribe a universal concurrency number.

Run the optional free Python load harness
# pip install neo4j==6.3.0
from concurrent.futures import ThreadPoolExecutor, as_completed
from statistics import median
from time import perf_counter
from neo4j import GraphDatabase
import math, random

URI = "bolt://localhost:17687"
AUTH = ("neo4j", "atlasmart-course-2026")
DB = "neo4j"

QUERIES = {
    "product": ("""
        MATCH (p:CH26Product {productId:$id})<-[:CONTAINS]-(o:CH26Order)<-[:PLACED]-(c:CH26Customer)
        RETURN c.customerId AS customerId, o.orderId AS orderId
        ORDER BY o.orderId LIMIT 20
    """, {"id":"CH26-P-0001"}),
    "customer": ("""
        MATCH (c:CH26Customer {customerId:$id})-[:VIEWED_CH26]->(p:CH26Product)-[:IN_CATEGORY]->(cat:CH26Category)
        RETURN cat.categoryId AS categoryId, count(*) AS views
        ORDER BY views DESC, categoryId LIMIT 6
    """, {"id":"CH26-C-0042"}),
    "recent": ("""
        MATCH (o:CH26Order)
        WHERE o.createdDay >= $cutoff
        RETURN o.status AS status, count(*) AS orders, sum(o.total) AS revenue
        ORDER BY orders DESC
    """, {"cutoff":100}),
}

def pct(xs, p):
    ys=sorted(xs)
    if not ys: return float('nan')
    k=(len(ys)-1)*p/100
    lo=math.floor(k); hi=math.ceil(k)
    return ys[lo] if lo==hi else ys[lo]*(hi-k)+ys[hi]*(k-lo)

def one(driver, name):
    q, params = QUERIES[name]
    t0=perf_counter()
    try:
        records, summary, keys = driver.execute_query(q, parameters_=params, database_=DB)
        ms=(perf_counter()-t0)*1000
        return name, ms, True, len(records)
    except Exception as e:
        return name, (perf_counter()-t0)*1000, False, type(e).__name__

def run(concurrency=4, operations=400, seed=26):
    rnd=random.Random(seed)
    mix=["product"]*50+["customer"]*30+["recent"]*20
    lat=[]; failures=0; rows=0
    t0=perf_counter()
    with GraphDatabase.driver(
        URI, auth=AUTH,
        max_connection_pool_size=max(4, concurrency),
        connection_acquisition_timeout=5,
        max_transaction_retry_time=3,
    ) as driver:
        driver.verify_connectivity()
        with ThreadPoolExecutor(max_workers=concurrency) as ex:
            fs=[ex.submit(one, driver, rnd.choice(mix)) for _ in range(operations)]
            for f in as_completed(fs):
                _, ms, ok, detail=f.result(); lat.append(ms)
                if ok: rows += detail
                else: failures += 1
    elapsed=perf_counter()-t0
    print({
        "concurrency":concurrency,
        "operations":operations,
        "seconds":round(elapsed,3),
        "throughput_ops_s":round(operations/elapsed,2),
        "p50_ms":round(pct(lat,50),2),
        "p95_ms":round(pct(lat,95),2),
        "p99_ms":round(pct(lat,99),2),
        "failures":failures,
        "rows_materialized":rows,
    })

if __name__ == "__main__":
    for c in (1,2,4,8,16):
        run(concurrency=c)

5. Deliberately wrong approach: publish one QPS number

Suppose AtlasMart reports “Neo4j handles 850 QPS.” Without the graph size, query mix, result cardinality, warmup, concurrency, CPU/memory limits, driver pool, error rate and percentile latency, this number is not portable evidence. It can even hide a benchmark in which requests queue for seconds while throughput remains flat.

Repair: publish a benchmark envelope: dataset fingerprint/counts, indexes, query weights, warmup, resource controls, concurrency, throughput, p50/p95/p99, failures/timeouts, driver configuration and the exact server/driver versions. Repeat at increasing concurrency and keep the raw per-operation timings. Verify that a rerun under the same envelope stays within a noise band before using it as a regression baseline.

Production judgment

Workload profiling is the gate before every later optimization. Model/cardinality choices determine fan-out; transaction concurrency can turn a hot node into a lock convoy; page cache and heap share finite memory with native buffers and the OS; returning huge result sets can dominate network/client time; driver pool limits can create application-side queueing; retries must remain idempotent; and failure/recovery behavior belongs in the same capacity envelope. Preserve security, tenant isolation, backups and observability while testing. Aura exposes a managed resource surface, while self-managed deployments let operators control JVM/container/OS settings directly; do not compare their raw hardware counters as if they were identical environments.

Lesson 2 changes the graph model one factor at a time and asks whether fewer expansions or precomputed values justify extra write/freshness complexity.

Check your understanding

  1. Why is a read/write ratio insufficient by itself?
  2. What does p99 measure?
  3. Why measure p95/max degree as well as average degree?
  4. Why separate warmup from measurement?
  5. What is the next tuning step after a baseline?
Review the answers

1. Because reads/writes contain different query shapes, result sizes, fan-out, lock behavior and resource costs; preserve per-query-class distributions.

2. A tail-latency percentile: roughly 99% of measured operations completed at or below that latency under the documented experiment.

3. Graph workloads are skew-sensitive; a small number of high-degree nodes can dominate expansion, locking and cache behavior.

4. A cold page cache has materially different I/O behavior from steady state; mixing them makes comparisons ambiguous.

5. Choose one evidenced bottleneck or design variable and change only that factor before remeasuring.

Summary and next step

Profile Workload Mix: Read/Write Ratios, Traversal Depth, Fan-Out, Hot Nodes, Result Sizes, and Concurrency is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Model-Level Optimization: Relationship Direction/Type, Intermediate Nodes, Denormalization, and Precomputation. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.