Chapter 26 · Performance Engineering: Data Model, Traversal Shape, Page Cache, Memory, and Workload Isolation
Run Controlled Benchmarks and Capacity Tests with Tail Latency, Saturation, Failure, and Regression Thresholds
Run a controlled AtlasMart capacity experiment with warmup, p50/p95/p99, throughput, saturation, driver-pool evidence, bounded failure injection, recovery observations, and explicit regression gates.
Learning outcomes
Design a benchmark with immutable environment/dataset/query-mix metadata, explicit warmup, repeated runs, and raw latency distributions.
Find saturation by increasing offered concurrency while observing throughput, p95/p99, errors, pool wait, CPU, memory and I/O together.
Inject a bounded failure only into the disposable performance container and classify client outcomes without claiming zero ambiguity.
Define regression thresholds from SLOs and measured noise rather than universal percentage folklore.
Produce a capacity decision with headroom, cost, failure/recovery, rollback, and “do not extrapolate beyond tested envelope” caveats.
1. AtlasMart problem: a benchmark must survive contact with failure
A capacity test that runs one query on a warm database at concurrency 1 says little about an order service during a traffic spike or restart. This lesson treats benchmark design as an experiment: hold data/query semantics constant, vary one factor, capture distributions and resource saturation, then introduce one safe bounded failure to observe the application’s error/retry/recovery behavior.
| Dimension | Chapter 26 reproducible assumption |
|---|---|
| Neo4j |
2026.07.1 Community. The continuity database remains
neo4j; performance experiments use a separate
disposable container named
atlasmart-neo4j-perf so tuning/failure tests
do not disturb earlier labs.
|
| Cypher | Cypher 25 examples. Planner/operator names and numeric PROFILE values are evidence to capture locally, not constants to memorize. |
| Java | Neo4j 2026 line with a supported Java 21/25 runtime as supplied/required by the chosen distribution. |
| Driver | Neo4j Python driver 6.3.x for the optional load harness; the driver object is shared, sessions/transactions are not shared between worker threads. |
| Auth/TLS |
User neo4j, password
atlasmart-course-2026. Loopback Bolt without
TLS only for the isolated disposable lab;
remote/production traffic should use validated TLS.
|
| Ports |
Performance container maps HTTP
17474→7474 and Bolt 17687→7687,
avoiding the continuity container on 7474/7687.
|
| Initial resource envelope |
Exercise baseline: Docker limit 2 GiB,
2 CPUs, explicit heap 512 MiB,
explicit page cache 512 MiB. These are lab
controls, not production recommendations.
|
| Dataset | 200 CH26 customers, 300 products, 2,000 orders, 6,000 CONTAINS relationships, 1,200 VIEWED_CH26 relationships, six categories, plus an isolated hot-counter node. Recount locally after setup. |
| Observability | Community labs rely on PROFILE, SHOW commands, Docker/OS counters, driver timing, container stats and logs. Neo4j metrics exporters and query.log are Enterprise surfaces and are optional, clearly labeled. |
| Evidence rule | This generated material does not execute your Docker host. Latencies, DB Hits, page-cache behavior, saturation points, GC, throughput, errors and recovery time must be measured locally; illustrative tables are labeled as templates. |
2. Benchmark identity: make the result reproducible
| Artifact | Minimum content |
|---|---|
| Software | Neo4j 2026.07.1, Cypher version, Python driver 6.3.x, Docker/host platform. |
| Data | Entity/relationship counts, degree p50/p95/max, property/schema/index state, fixture hash/script revision. |
| Resources | CPU quota, container memory, heap/page cache/transaction limits, storage medium, network topology. |
| Workload | Query texts/hashes, parameter distribution, read/write weights, result caps, concurrency schedule. |
| Protocol | Warmup duration/operations, measurement duration/operations, repetitions, restart/cold-cache protocol. |
| Outputs | Throughput, p50/p95/p99/max, failures/timeouts/retries, result rows/bytes, CPU/memory/I/O/GC/pool evidence. |
| Decision | SLO, regression gate, saturation point/range, chosen operating headroom, rollback trigger. |
// Run against bolt://localhost:17687, database neo4j.
// Cleanup only the Chapter 26 fixture.
MATCH (n) WHERE n.chapter26 = true DETACH DELETE n;
CREATE CONSTRAINT ch26_customer_id IF NOT EXISTS
FOR (c:CH26Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch26_product_id IF NOT EXISTS
FOR (p:CH26Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch26_order_id IF NOT EXISTS
FOR (o:CH26Order) REQUIRE o.orderId IS UNIQUE;
CREATE INDEX ch26_order_day IF NOT EXISTS
FOR (o:CH26Order) ON (o.createdDay);
UNWIND range(1,6) AS i
CREATE (:Category:CH26Category {categoryId:'CH26-CAT-'+toString(i), chapter26:true});
UNWIND range(1,200) AS i
CREATE (:Customer:CH26Customer {
customerId:'CH26-C-'+right('0000'+toString(i),4),
segment: CASE WHEN i % 5 = 0 THEN 'business' ELSE 'consumer' END,
chapter26:true
});
UNWIND range(1,300) AS i
CREATE (:Product:CH26Product {
productId:'CH26-P-'+right('0000'+toString(i),4),
name:'Atlas Product '+toString(i),
price:toFloat(10 + (i % 90)),
chapter26:true
});
MATCH (p:CH26Product), (cat:CH26Category)
WHERE toInteger(right(p.productId,4)) % 6 + 1 = toInteger(replace(cat.categoryId,'CH26-CAT-',''))
CREATE (p)-[:IN_CATEGORY {chapter26:true}]->(cat);
UNWIND range(1,2000) AS i
MATCH (c:CH26Customer {customerId:'CH26-C-'+right('0000'+toString(((i-1)%200)+1),4)})
CREATE (o:Order:CH26Order {
orderId:'CH26-O-'+right('00000'+toString(i),5),
createdDay:i % 120,
status:CASE WHEN i % 9 = 0 THEN 'RETURNED' ELSE 'COMPLETE' END,
total:toFloat(40 + (i % 260)),
chapter26:true
})
CREATE (c)-[:PLACED {chapter26:true}]->(o);
MATCH (o:CH26Order)
WITH o, toInteger(right(o.orderId,5)) AS i
UNWIND range(0,2) AS j
WITH o, i, j,
CASE WHEN j=0 AND i % 5 = 0 THEN 1
ELSE ((i*17 + j*43) % 300)+1 END AS pno
MATCH (p:CH26Product {productId:'CH26-P-'+right('0000'+toString(pno),4)})
CREATE (o)-[:CONTAINS {quantity:1+(i+j)%3, chapter26:true}]->(p);
MATCH (c:CH26Customer)
WITH c, toInteger(right(c.customerId,4)) AS i
UNWIND range(0,5) AS j
WITH c, ((i*13 + j*29) % 300)+1 AS pno
MATCH (p:CH26Product {productId:'CH26-P-'+right('0000'+toString(pno),4)})
CREATE (c)-[:VIEWED_CH26 {chapter26:true}]->(p);
MERGE (:CH26HotCounter {id:'global', n:0, chapter26:true});
MATCH (c:CH26Customer) WITH count(c) AS customers
MATCH (p:CH26Product) WITH customers, count(p) AS products
MATCH (o:CH26Order) WITH customers, products, count(o) AS orders
MATCH (:CH26Order)-[r:CONTAINS]->(:CH26Product)
WITH customers, products, orders, count(r) AS contains
MATCH (:CH26Customer)-[v:VIEWED_CH26]->(:CH26Product)
RETURN customers, products, orders, contains, count(v) AS viewed;
// After seeding, run the Python harness from Lesson 1 with the same versioned file.
# pip install neo4j==6.3.0
from concurrent.futures import ThreadPoolExecutor, as_completed
from statistics import median
from time import perf_counter
from neo4j import GraphDatabase
import math, random
URI = "bolt://localhost:17687"
AUTH = ("neo4j", "atlasmart-course-2026")
DB = "neo4j"
QUERIES = {
"product": ("""
MATCH (p:CH26Product {productId:$id})<-[:CONTAINS]-(o:CH26Order)<-[:PLACED]-(c:CH26Customer)
RETURN c.customerId AS customerId, o.orderId AS orderId
ORDER BY o.orderId LIMIT 20
""", {"id":"CH26-P-0001"}),
"customer": ("""
MATCH (c:CH26Customer {customerId:$id})-[:VIEWED_CH26]->(p:CH26Product)-[:IN_CATEGORY]->(cat:CH26Category)
RETURN cat.categoryId AS categoryId, count(*) AS views
ORDER BY views DESC, categoryId LIMIT 6
""", {"id":"CH26-C-0042"}),
"recent": ("""
MATCH (o:CH26Order)
WHERE o.createdDay >= $cutoff
RETURN o.status AS status, count(*) AS orders, sum(o.total) AS revenue
ORDER BY orders DESC
""", {"cutoff":100}),
}
def pct(xs, p):
ys=sorted(xs)
if not ys: return float('nan')
k=(len(ys)-1)*p/100
lo=math.floor(k); hi=math.ceil(k)
return ys[lo] if lo==hi else ys[lo]*(hi-k)+ys[hi]*(k-lo)
def one(driver, name):
q, params = QUERIES[name]
t0=perf_counter()
try:
records, summary, keys = driver.execute_query(q, parameters_=params, database_=DB)
ms=(perf_counter()-t0)*1000
return name, ms, True, len(records)
except Exception as e:
return name, (perf_counter()-t0)*1000, False, type(e).__name__
def run(concurrency=4, operations=400, seed=26):
rnd=random.Random(seed)
mix=["product"]*50+["customer"]*30+["recent"]*20
lat=[]; failures=0; rows=0
t0=perf_counter()
with GraphDatabase.driver(
URI, auth=AUTH,
max_connection_pool_size=max(4, concurrency),
connection_acquisition_timeout=5,
max_transaction_retry_time=3,
) as driver:
driver.verify_connectivity()
with ThreadPoolExecutor(max_workers=concurrency) as ex:
fs=[ex.submit(one, driver, rnd.choice(mix)) for _ in range(operations)]
for f in as_completed(fs):
_, ms, ok, detail=f.result(); lat.append(ms)
if ok: rows += detail
else: failures += 1
elapsed=perf_counter()-t0
print({
"concurrency":concurrency,
"operations":operations,
"seconds":round(elapsed,3),
"throughput_ops_s":round(operations/elapsed,2),
"p50_ms":round(pct(lat,50),2),
"p95_ms":round(pct(lat,95),2),
"p99_ms":round(pct(lat,99),2),
"failures":failures,
"rows_materialized":rows,
})
if __name__ == "__main__":
for c in (1,2,4,8,16):
run(concurrency=c)
3. Interpret the saturation curve
At low concurrency, throughput may rise as more CPU/I/O/network capacity is used. Eventually another request adds queueing rather than useful throughput. The signature is not one magic number: p95/p99 rise sharply, connection acquisition grows, CPU or I/O approaches its effective limit, throughput gains flatten, and errors/timeouts may appear. The chosen operating point should leave capacity headroom for skew, background work and failure—not sit exactly at the highest measured throughput.
Create a table with one row per concurrency level: concurrency, operations, elapsed seconds, ops/s, p50/p95/p99, failures, container CPU %, memory, disk read/write and client pool-wait. Fill it from your machine. This generated lesson intentionally does not pre-populate fake latency/QPS values.
4. Bounded failure injection: pause only the disposable container
Blast radius: only
atlasmart-neo4j-perf on ports 17474/17687. Do not
perform this against the continuity container, Aura, a cluster,
or production. Start a longer harness run in one terminal. In a
second terminal, pause the container briefly and unpause it. The
exact client outcomes depend on in-flight transaction phase,
driver timeout/retry configuration and OS/network behavior.
# Terminal B — while Terminal A is running a measurement interval.
docker pause atlasmart-neo4j-perf
Start-Sleep -Seconds 5
docker unpause atlasmart-neo4j-perf
# Verify the server returns and inspect logs; do not assume every failed call was uncommitted.
docker ps --filter "name=atlasmart-neo4j-perf"
docker logs --tail 120 atlasmart-neo4j-perf
For read-only benchmark calls, retry risk is simpler. For writes, a connection loss around commit can leave the client uncertain whether the server committed. Production retry logic therefore requires idempotent transaction design/business keys; “the driver retried” is not proof that a side effect occurred exactly once.
5. Add a deliberate hot-write experiment
The fixture contains one CH26HotCounter.
Concurrently incrementing it creates an intentionally shared
write point and can expose lock serialization. Run it only in
the disposable lab, then compare with updates distributed across
per-customer counters if that better matches the business model.
// Execute concurrently from several workers only in the disposable lab.
MATCH (h:CH26HotCounter {id:'global'})
SET h.n = h.n + 1
RETURN h.n;
// Inspect active transactions while the test runs.
SHOW TRANSACTIONS YIELD *;
The lesson is not “never use counters.” It is to recognize that a globally shared write entity imposes serialization/lock behavior. If the business invariant requires a single counter, correctness may outweigh throughput; if not, partition or derive the aggregate and reconcile it.
6. Regression gates: derive them from SLO + noise
A gate should be tied to a stable baseline and service objective. For this course lab, you may choose an exercise-only rule such as “fail the experiment if p95 regresses more than the measured repeat-run noise band or if any unexpected failures appear.” In a real service, quantify the noise through repeated runs and use an SLO-driven threshold. A fixed “10% slower = fail” rule copied between machines/workloads is vendor folklore.
| Gate dimension | Good question |
|---|---|
| Correctness | Are result sets/invariants identical before comparing speed? |
| Latency | Did p95/p99 move beyond baseline variance and SLO budget? |
| Throughput | Did useful successful throughput change at the chosen offered load? |
| Errors | Did timeouts/connection failures/retries increase? |
| Resources | Did the change shift cost to CPU, heap, page cache, storage, network or client memory? |
| Failure recovery | Does the application recover within the tested objective, and are ambiguous writes reconciled? |
7. Deliberately wrong approach: change model, query, heap and CPU together
If a run gets faster after four simultaneous changes, AtlasMart cannot identify the causal factor, the cost tradeoff, or a safe rollback. It also cannot tell whether the result came from cache warmup.
Repair: freeze the baseline, change one factor, repeat runs, compare plans/resources/latency, revert, and rerun if necessary. When interactions matter, add a deliberately designed multi-factor experiment after single-factor effects are understood. Preserve all scripts/configs and raw measurements.
Production judgment and bridge to the capstone
A capacity statement is valid only inside its tested envelope. Document graph size/density/skew, query mix, concurrency, resource limits, pool/timeouts, versions, indexes, security/TLS, backup/background workload and failure conditions. Leave headroom for growth and single-component degradation; include cost per useful workload, not just maximum QPS. Separate Community lab findings from Enterprise/Aura topology/observability features. Keep rollback criteria for model/index/config changes and a restore/upgrade path from Chapters 16–18.
Chapter 27 is the production capstone: it combines graph fit, modeling, import, Cypher, search/vector/GraphRAG, GDS, security, backup/recovery, failure handling, observability and the evidence-driven performance envelope created here.
Check your understanding
- What makes a benchmark reproducible?
- How do you recognize saturation?
- Why can a write failure around commit be ambiguous?
- Why should regression gates include correctness?
- Why leave capacity headroom?
Review the answers
1. The environment, software, data, query mix/parameters, warmup, concurrency, resources and raw outputs are recorded so another run can recreate the envelope.
2. Throughput gains flatten while queueing/tail latency/resource utilization and possibly errors rise as offered concurrency increases.
3. The client can lose the connection after the server commits but before the acknowledgement is observed; retries therefore need idempotency/reconciliation.
4. A faster query/model is not an improvement if it changes the business result or violates invariants.
5. Production must absorb skew, background work, growth and degraded/failure conditions without operating continuously at the saturation knee.
# PowerShell — delete only the disposable Chapter 26 performance lab.
docker rm -f atlasmart-neo4j-perf 2>$null
docker volume rm atlasmart-neo4j-perf-data atlasmart-neo4j-perf-logs 2>$null
# The continuity container atlasmart-neo4j is not touched.
Summary and next step
Run Controlled Benchmarks and Capacity Tests with Tail Latency, Saturation, Failure, and Regression Thresholds is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Define Requirements, Workloads, Availability/RPO/RTO, Data Sources, Security Boundaries, and Graph-Fit Criteria. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- Neo4j current versions — Current database release and LTS baseline.
- Neo4j Operations Manual — Performance — Current performance topics and operational tuning surface.
- Memory configuration — Heap, page cache, transaction/native memory, OS headroom, and memory recommendation guidance.
- Disks, RAM and other tips — Page-cache warmup, storage and RAM behavior.
- Configuration settings — Authoritative current setting names and edition/dynamic boundaries.
- Docker configuration — Container configuration mapping and production configuration guidance.
- Cypher execution plans — EXPLAIN/PROFILE semantics and runtime evidence.
- Cypher operators in detail — Current operators, Rows, DB Hits, memory and plan behavior.
- Indexes for search performance — Current range/text/point/token index behavior and syntax.
- Cypher query tuning — Planner, statistics and query-tuning concepts.
- Neo4j Python driver performance — Driver-side result streaming, database selection and performance guidance.
- Neo4j Python driver API — Connection pool, timeout, retry and fetch-size configuration.
- Neo4j logging — Current debug/query/security/GC logging surfaces and edition boundaries.
- Neo4j metrics — Enterprise metrics surfaces and operational evidence.
- Transaction management — Transaction lifecycle and operational behavior.
- Java requirements — Current supported Java/runtime and platform requirements.
- Neo4j status codes — Classify transient/client/database failures rather than collapsing them into latency.