Chapter 26 · Performance Engineering: Data Model, Traversal Shape, Page Cache, Memory, and Workload Isolation

Run Controlled Benchmarks and Capacity Tests with Tail Latency, Saturation, Failure, and Regression Thresholds

Run a controlled AtlasMart capacity experiment with warmup, p50/p95/p99, throughput, saturation, driver-pool evidence, bounded failure injection, recovery observations, and explicit regression gates.

Advanced320–480 minutesBenchmarking · p95/p99 · failure · regression gatesNeo4j 2026.07.1 · Community mandatoryCypher 25 · Python driver 6.3.x optionalDisposable perf container · 2 GiB/2 CPU baselineJava 21/25 · no APOC/GDS requiredLast reviewed: September 2026

Learning outcomes

01

Design a benchmark with immutable environment/dataset/query-mix metadata, explicit warmup, repeated runs, and raw latency distributions.

02

Find saturation by increasing offered concurrency while observing throughput, p95/p99, errors, pool wait, CPU, memory and I/O together.

03

Inject a bounded failure only into the disposable performance container and classify client outcomes without claiming zero ambiguity.

04

Define regression thresholds from SLOs and measured noise rather than universal percentage folklore.

05

Produce a capacity decision with headroom, cost, failure/recovery, rollback, and “do not extrapolate beyond tested envelope” caveats.

1. AtlasMart problem: a benchmark must survive contact with failure

A capacity test that runs one query on a warm database at concurrency 1 says little about an order service during a traffic spike or restart. This lesson treats benchmark design as an experiment: hold data/query semantics constant, vary one factor, capture distributions and resource saturation, then introduce one safe bounded failure to observe the application’s error/retry/recovery behavior.

Dimension Chapter 26 reproducible assumption
Neo4j 2026.07.1 Community. The continuity database remains neo4j; performance experiments use a separate disposable container named atlasmart-neo4j-perf so tuning/failure tests do not disturb earlier labs.
Cypher Cypher 25 examples. Planner/operator names and numeric PROFILE values are evidence to capture locally, not constants to memorize.
Java Neo4j 2026 line with a supported Java 21/25 runtime as supplied/required by the chosen distribution.
Driver Neo4j Python driver 6.3.x for the optional load harness; the driver object is shared, sessions/transactions are not shared between worker threads.
Auth/TLS User neo4j, password atlasmart-course-2026. Loopback Bolt without TLS only for the isolated disposable lab; remote/production traffic should use validated TLS.
Ports Performance container maps HTTP 17474→7474 and Bolt 17687→7687, avoiding the continuity container on 7474/7687.
Initial resource envelope Exercise baseline: Docker limit 2 GiB, 2 CPUs, explicit heap 512 MiB, explicit page cache 512 MiB. These are lab controls, not production recommendations.
Dataset 200 CH26 customers, 300 products, 2,000 orders, 6,000 CONTAINS relationships, 1,200 VIEWED_CH26 relationships, six categories, plus an isolated hot-counter node. Recount locally after setup.
Observability Community labs rely on PROFILE, SHOW commands, Docker/OS counters, driver timing, container stats and logs. Neo4j metrics exporters and query.log are Enterprise surfaces and are optional, clearly labeled.
Evidence rule This generated material does not execute your Docker host. Latencies, DB Hits, page-cache behavior, saturation points, GC, throughput, errors and recovery time must be measured locally; illustrative tables are labeled as templates.

2. Benchmark identity: make the result reproducible

Artifact Minimum content
Software Neo4j 2026.07.1, Cypher version, Python driver 6.3.x, Docker/host platform.
Data Entity/relationship counts, degree p50/p95/max, property/schema/index state, fixture hash/script revision.
Resources CPU quota, container memory, heap/page cache/transaction limits, storage medium, network topology.
Workload Query texts/hashes, parameter distribution, read/write weights, result caps, concurrency schedule.
Protocol Warmup duration/operations, measurement duration/operations, repetitions, restart/cold-cache protocol.
Outputs Throughput, p50/p95/p99/max, failures/timeouts/retries, result rows/bytes, CPU/memory/I/O/GC/pool evidence.
Decision SLO, regression gate, saturation point/range, chosen operating headroom, rollback trigger.
Use the same deterministic fixture and load harness
// Run against bolt://localhost:17687, database neo4j.
// Cleanup only the Chapter 26 fixture.
MATCH (n) WHERE n.chapter26 = true DETACH DELETE n;

CREATE CONSTRAINT ch26_customer_id IF NOT EXISTS
FOR (c:CH26Customer) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch26_product_id IF NOT EXISTS
FOR (p:CH26Product) REQUIRE p.productId IS UNIQUE;
CREATE CONSTRAINT ch26_order_id IF NOT EXISTS
FOR (o:CH26Order) REQUIRE o.orderId IS UNIQUE;
CREATE INDEX ch26_order_day IF NOT EXISTS
FOR (o:CH26Order) ON (o.createdDay);

UNWIND range(1,6) AS i
CREATE (:Category:CH26Category {categoryId:'CH26-CAT-'+toString(i), chapter26:true});

UNWIND range(1,200) AS i
CREATE (:Customer:CH26Customer {
  customerId:'CH26-C-'+right('0000'+toString(i),4),
  segment: CASE WHEN i % 5 = 0 THEN 'business' ELSE 'consumer' END,
  chapter26:true
});

UNWIND range(1,300) AS i
CREATE (:Product:CH26Product {
  productId:'CH26-P-'+right('0000'+toString(i),4),
  name:'Atlas Product '+toString(i),
  price:toFloat(10 + (i % 90)),
  chapter26:true
});

MATCH (p:CH26Product), (cat:CH26Category)
WHERE toInteger(right(p.productId,4)) % 6 + 1 = toInteger(replace(cat.categoryId,'CH26-CAT-',''))
CREATE (p)-[:IN_CATEGORY {chapter26:true}]->(cat);

UNWIND range(1,2000) AS i
MATCH (c:CH26Customer {customerId:'CH26-C-'+right('0000'+toString(((i-1)%200)+1),4)})
CREATE (o:Order:CH26Order {
  orderId:'CH26-O-'+right('00000'+toString(i),5),
  createdDay:i % 120,
  status:CASE WHEN i % 9 = 0 THEN 'RETURNED' ELSE 'COMPLETE' END,
  total:toFloat(40 + (i % 260)),
  chapter26:true
})
CREATE (c)-[:PLACED {chapter26:true}]->(o);

MATCH (o:CH26Order)
WITH o, toInteger(right(o.orderId,5)) AS i
UNWIND range(0,2) AS j
WITH o, i, j,
     CASE WHEN j=0 AND i % 5 = 0 THEN 1
          ELSE ((i*17 + j*43) % 300)+1 END AS pno
MATCH (p:CH26Product {productId:'CH26-P-'+right('0000'+toString(pno),4)})
CREATE (o)-[:CONTAINS {quantity:1+(i+j)%3, chapter26:true}]->(p);

MATCH (c:CH26Customer)
WITH c, toInteger(right(c.customerId,4)) AS i
UNWIND range(0,5) AS j
WITH c, ((i*13 + j*29) % 300)+1 AS pno
MATCH (p:CH26Product {productId:'CH26-P-'+right('0000'+toString(pno),4)})
CREATE (c)-[:VIEWED_CH26 {chapter26:true}]->(p);

MERGE (:CH26HotCounter {id:'global', n:0, chapter26:true});

MATCH (c:CH26Customer) WITH count(c) AS customers
MATCH (p:CH26Product) WITH customers, count(p) AS products
MATCH (o:CH26Order) WITH customers, products, count(o) AS orders
MATCH (:CH26Order)-[r:CONTAINS]->(:CH26Product)
WITH customers, products, orders, count(r) AS contains
MATCH (:CH26Customer)-[v:VIEWED_CH26]->(:CH26Product)
RETURN customers, products, orders, contains, count(v) AS viewed;

// After seeding, run the Python harness from Lesson 1 with the same versioned file.
Capacity sweep harness
# pip install neo4j==6.3.0
from concurrent.futures import ThreadPoolExecutor, as_completed
from statistics import median
from time import perf_counter
from neo4j import GraphDatabase
import math, random

URI = "bolt://localhost:17687"
AUTH = ("neo4j", "atlasmart-course-2026")
DB = "neo4j"

QUERIES = {
    "product": ("""
        MATCH (p:CH26Product {productId:$id})<-[:CONTAINS]-(o:CH26Order)<-[:PLACED]-(c:CH26Customer)
        RETURN c.customerId AS customerId, o.orderId AS orderId
        ORDER BY o.orderId LIMIT 20
    """, {"id":"CH26-P-0001"}),
    "customer": ("""
        MATCH (c:CH26Customer {customerId:$id})-[:VIEWED_CH26]->(p:CH26Product)-[:IN_CATEGORY]->(cat:CH26Category)
        RETURN cat.categoryId AS categoryId, count(*) AS views
        ORDER BY views DESC, categoryId LIMIT 6
    """, {"id":"CH26-C-0042"}),
    "recent": ("""
        MATCH (o:CH26Order)
        WHERE o.createdDay >= $cutoff
        RETURN o.status AS status, count(*) AS orders, sum(o.total) AS revenue
        ORDER BY orders DESC
    """, {"cutoff":100}),
}

def pct(xs, p):
    ys=sorted(xs)
    if not ys: return float('nan')
    k=(len(ys)-1)*p/100
    lo=math.floor(k); hi=math.ceil(k)
    return ys[lo] if lo==hi else ys[lo]*(hi-k)+ys[hi]*(k-lo)

def one(driver, name):
    q, params = QUERIES[name]
    t0=perf_counter()
    try:
        records, summary, keys = driver.execute_query(q, parameters_=params, database_=DB)
        ms=(perf_counter()-t0)*1000
        return name, ms, True, len(records)
    except Exception as e:
        return name, (perf_counter()-t0)*1000, False, type(e).__name__

def run(concurrency=4, operations=400, seed=26):
    rnd=random.Random(seed)
    mix=["product"]*50+["customer"]*30+["recent"]*20
    lat=[]; failures=0; rows=0
    t0=perf_counter()
    with GraphDatabase.driver(
        URI, auth=AUTH,
        max_connection_pool_size=max(4, concurrency),
        connection_acquisition_timeout=5,
        max_transaction_retry_time=3,
    ) as driver:
        driver.verify_connectivity()
        with ThreadPoolExecutor(max_workers=concurrency) as ex:
            fs=[ex.submit(one, driver, rnd.choice(mix)) for _ in range(operations)]
            for f in as_completed(fs):
                _, ms, ok, detail=f.result(); lat.append(ms)
                if ok: rows += detail
                else: failures += 1
    elapsed=perf_counter()-t0
    print({
        "concurrency":concurrency,
        "operations":operations,
        "seconds":round(elapsed,3),
        "throughput_ops_s":round(operations/elapsed,2),
        "p50_ms":round(pct(lat,50),2),
        "p95_ms":round(pct(lat,95),2),
        "p99_ms":round(pct(lat,99),2),
        "failures":failures,
        "rows_materialized":rows,
    })

if __name__ == "__main__":
    for c in (1,2,4,8,16):
        run(concurrency=c)

3. Interpret the saturation curve

At low concurrency, throughput may rise as more CPU/I/O/network capacity is used. Eventually another request adds queueing rather than useful throughput. The signature is not one magic number: p95/p99 rise sharply, connection acquisition grows, CPU or I/O approaches its effective limit, throughput gains flatten, and errors/timeouts may appear. The chosen operating point should leave capacity headroom for skew, background work and failure—not sit exactly at the highest measured throughput.

Measurement template

Create a table with one row per concurrency level: concurrency, operations, elapsed seconds, ops/s, p50/p95/p99, failures, container CPU %, memory, disk read/write and client pool-wait. Fill it from your machine. This generated lesson intentionally does not pre-populate fake latency/QPS values.

4. Bounded failure injection: pause only the disposable container

Blast radius: only atlasmart-neo4j-perf on ports 17474/17687. Do not perform this against the continuity container, Aura, a cluster, or production. Start a longer harness run in one terminal. In a second terminal, pause the container briefly and unpause it. The exact client outcomes depend on in-flight transaction phase, driver timeout/retry configuration and OS/network behavior.

Controlled pause/unpause exercise
# Terminal B — while Terminal A is running a measurement interval.
docker pause atlasmart-neo4j-perf
Start-Sleep -Seconds 5
docker unpause atlasmart-neo4j-perf

# Verify the server returns and inspect logs; do not assume every failed call was uncommitted.
docker ps --filter "name=atlasmart-neo4j-perf"
docker logs --tail 120 atlasmart-neo4j-perf

For read-only benchmark calls, retry risk is simpler. For writes, a connection loss around commit can leave the client uncertain whether the server committed. Production retry logic therefore requires idempotent transaction design/business keys; “the driver retried” is not proof that a side effect occurred exactly once.

5. Add a deliberate hot-write experiment

The fixture contains one CH26HotCounter. Concurrently incrementing it creates an intentionally shared write point and can expose lock serialization. Run it only in the disposable lab, then compare with updates distributed across per-customer counters if that better matches the business model.

Hot-node write — isolated diagnostic
// Execute concurrently from several workers only in the disposable lab.
MATCH (h:CH26HotCounter {id:'global'})
SET h.n = h.n + 1
RETURN h.n;

// Inspect active transactions while the test runs.
SHOW TRANSACTIONS YIELD *;

The lesson is not “never use counters.” It is to recognize that a globally shared write entity imposes serialization/lock behavior. If the business invariant requires a single counter, correctness may outweigh throughput; if not, partition or derive the aggregate and reconcile it.

6. Regression gates: derive them from SLO + noise

A gate should be tied to a stable baseline and service objective. For this course lab, you may choose an exercise-only rule such as “fail the experiment if p95 regresses more than the measured repeat-run noise band or if any unexpected failures appear.” In a real service, quantify the noise through repeated runs and use an SLO-driven threshold. A fixed “10% slower = fail” rule copied between machines/workloads is vendor folklore.

Gate dimension Good question
Correctness Are result sets/invariants identical before comparing speed?
Latency Did p95/p99 move beyond baseline variance and SLO budget?
Throughput Did useful successful throughput change at the chosen offered load?
Errors Did timeouts/connection failures/retries increase?
Resources Did the change shift cost to CPU, heap, page cache, storage, network or client memory?
Failure recovery Does the application recover within the tested objective, and are ambiguous writes reconciled?

7. Deliberately wrong approach: change model, query, heap and CPU together

If a run gets faster after four simultaneous changes, AtlasMart cannot identify the causal factor, the cost tradeoff, or a safe rollback. It also cannot tell whether the result came from cache warmup.

Repair: freeze the baseline, change one factor, repeat runs, compare plans/resources/latency, revert, and rerun if necessary. When interactions matter, add a deliberately designed multi-factor experiment after single-factor effects are understood. Preserve all scripts/configs and raw measurements.

Production judgment and bridge to the capstone

A capacity statement is valid only inside its tested envelope. Document graph size/density/skew, query mix, concurrency, resource limits, pool/timeouts, versions, indexes, security/TLS, backup/background workload and failure conditions. Leave headroom for growth and single-component degradation; include cost per useful workload, not just maximum QPS. Separate Community lab findings from Enterprise/Aura topology/observability features. Keep rollback criteria for model/index/config changes and a restore/upgrade path from Chapters 16–18.

Chapter 27 is the production capstone: it combines graph fit, modeling, import, Cypher, search/vector/GraphRAG, GDS, security, backup/recovery, failure handling, observability and the evidence-driven performance envelope created here.

Check your understanding

  1. What makes a benchmark reproducible?
  2. How do you recognize saturation?
  3. Why can a write failure around commit be ambiguous?
  4. Why should regression gates include correctness?
  5. Why leave capacity headroom?
Review the answers

1. The environment, software, data, query mix/parameters, warmup, concurrency, resources and raw outputs are recorded so another run can recreate the envelope.

2. Throughput gains flatten while queueing/tail latency/resource utilization and possibly errors rise as offered concurrency increases.

3. The client can lose the connection after the server commits but before the acknowledgement is observed; retries therefore need idempotency/reconciliation.

4. A faster query/model is not an improvement if it changes the business result or violates invariants.

5. Production must absorb skew, background work, growth and degraded/failure conditions without operating continuously at the saturation knee.

Cleanup/reset
# PowerShell — delete only the disposable Chapter 26 performance lab.
docker rm -f atlasmart-neo4j-perf 2>$null
docker volume rm atlasmart-neo4j-perf-data atlasmart-neo4j-perf-logs 2>$null
# The continuity container atlasmart-neo4j is not touched.

Summary and next step

Run Controlled Benchmarks and Capacity Tests with Tail Latency, Saturation, Failure, and Regression Thresholds is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Define Requirements, Workloads, Availability/RPO/RTO, Data Sources, Security Boundaries, and Graph-Fit Criteria. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.