Chapter 26 · Performance Engineering: Data Model, Traversal Shape, Page Cache, Memory, and Workload Isolation
Infrastructure Tuning: Heap, Page Cache, Transaction Memory, CPU, Storage, Containers, and OS Limits
Divide a bounded memory and I/O budget deliberately among JVM heap, Neo4j page cache, transaction/query memory, native buffers, the OS/container, CPU, and storage while preserving failure headroom.
Learning outcomes
Separate JVM heap, Neo4j page cache, transaction/query memory, native/off-heap memory, network buffers and OS/container headroom.
Use explicit memory configuration and neo4j-admin memory recommendations as starting evidence, not universal sizing formulas.
Explain how CPU saturation, storage latency/IOPS, page faults, GC and file-descriptor limits create different symptoms.
Tune a disposable container one resource dimension at a time while respecting Docker Desktop/host/VM boundaries on Windows.
Define workload isolation and capacity headroom so analytical/batch traffic cannot consume the latency budget of transactional requests.
Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.
1. AtlasMart problem: increasing heap can make the database slower
An operator sees GC activity and doubles the JVM heap until the container consumes nearly all available memory. That leaves less memory for Neo4j page cache, native buffers and the environment hosting the container. If the working graph no longer fits well in page cache, store reads can move to disk and latency can worsen. Memory tuning is a budget-allocation problem, not a contest to maximize one region.
| Dimension | Chapter 26 reproducible assumption |
|---|---|
| Neo4j |
2026.07.1 Community. The continuity database remains
neo4j; performance experiments use a separate
disposable container named
atlasmart-neo4j-perf so tuning/failure tests
do not disturb earlier labs.
|
| Cypher | Cypher 25 examples. Planner/operator names and numeric PROFILE values are evidence to capture locally, not constants to memorize. |
| Java | Neo4j 2026 line with a supported Java 21/25 runtime as supplied/required by the chosen distribution. |
| Driver | Neo4j Python driver 6.3.x for the optional load harness; the driver object is shared, sessions/transactions are not shared between worker threads. |
| Auth/TLS |
User neo4j, password
atlasmart-course-2026. Loopback Bolt without
TLS only for the isolated disposable lab;
remote/production traffic should use validated TLS.
|
| Ports |
Performance container maps HTTP
17474→7474 and Bolt 17687→7687,
avoiding the continuity container on 7474/7687.
|
| Initial resource envelope |
Exercise baseline: Docker limit 2 GiB,
2 CPUs, explicit heap 512 MiB,
explicit page cache 512 MiB. These are lab
controls, not production recommendations.
|
| Dataset | 200 CH26 customers, 300 products, 2,000 orders, 6,000 CONTAINS relationships, 1,200 VIEWED_CH26 relationships, six categories, plus an isolated hot-counter node. Recount locally after setup. |
| Observability | Community labs rely on PROFILE, SHOW commands, Docker/OS counters, driver timing, container stats and logs. Neo4j metrics exporters and query.log are Enterprise surfaces and are optional, clearly labeled. |
| Evidence rule | This generated material does not execute your Docker host. Latencies, DB Hits, page-cache behavior, saturation points, GC, throughput, errors and recovery time must be measured locally; illustrative tables are labeled as templates. |
2. The memory budget has distinct mechanisms
| Region | What it primarily holds | What to observe / common mistake |
|---|---|---|
| JVM heap | Java objects, query/transaction structures, runtime metadata | GC pauses/occupancy and query memory. Oversizing can starve page cache/OS/native needs. |
| Neo4j page cache | Cached native store/index pages | Page hits/faults and storage I/O. It is separate from heap. |
| Transaction/query memory | Uncommitted state and intermediate query structures | SHOW TRANSACTIONS/query evidence; per-transaction/database/global limits can terminate pathological work safely. |
| Native/off-heap + network buffers | Direct/native allocations outside GC-managed heap | Container RSS/native growth; do not assume heap max equals total process memory. |
| OS/container headroom | Filesystem cache, vector-index OS memory, kernel/runtime, Docker/VM overhead | Swap/host pressure/container OOM. Leave explicit headroom. |
Current Neo4j exposes
server.memory.heap.initial_size,
server.memory.heap.max_size,
server.memory.pagecache.size, and
transaction-memory controls including
dbms.memory.transaction.total.max,
db.memory.transaction.total.max, and
db.memory.transaction.max. Exact defaults can be
calculated dynamically and can change; query
SHOW SETTINGS on the target release.
RETURN 1 AS cypherReachable;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;
SHOW SETTINGS YIELD name, value
WHERE name IN [
'server.memory.heap.initial_size',
'server.memory.heap.max_size',
'server.memory.pagecache.size',
'dbms.memory.transaction.total.max',
'db.memory.transaction.total.max',
'db.memory.transaction.max'
]
RETURN name, value ORDER BY name;
SHOW INDEXES YIELD name, state, type, entityType, labelsOrTypes, properties
WHERE name STARTS WITH 'ch26_'
RETURN name, state, type, entityType, labelsOrTypes, properties ORDER BY name;
# Run inside the disposable container. This is a recommendation, not a benchmark result.
docker exec atlasmart-neo4j-perf neo4j-admin server memory-recommendation --memory=2g
# Observe process/container resource use during a controlled run.
docker stats atlasmart-neo4j-perf
# Inspect filesystem/volume growth from the container shell if needed.
docker exec atlasmart-neo4j-perf sh -lc 'du -sh /data /logs; ulimit -n'
3. CPU, storage, and file descriptors are different bottlenecks
| Signal pattern | Plausible hypothesis | Evidence before remediation |
|---|---|---|
| CPU near allocated limit; throughput plateaus | Compute-bound query/operators or too much offered concurrency | PROFILE operator work, host/container CPU, runnable threads, same dataset at lower concurrency. |
| Low CPU but high latency + page faults/I/O wait | Working set misses page cache or storage latency dominates | Cold/warm comparison, disk latency/throughput, page-cache evidence where available. |
| Long GC pauses / heap pressure | Large intermediate rows, result materialization, transaction memory, insufficient/poorly balanced heap | GC logs, PROFILE memory, query shape, heap occupancy—not heap size alone. |
| Connection/open-file failures | File descriptor/socket exhaustion or pool misuse | OS/container limits, driver pool/session lifecycle, active connections. |
| Driver waits while server CPU idle | Pool acquisition/network/client bottleneck | Driver timing, pool limits, connection acquisition timeout, network RTT/result bytes. |
4. One-factor container experiments
Use a new disposable container/volume or restore the same fixture before each resource experiment. Keep the query mix and data identical. For example, compare the 2 GiB exercise baseline with a deliberately constrained 1.25 GiB container while keeping explicit heap/page cache small enough to leave native headroom. Do not run a configuration whose memory regions sum to the whole container limit.
# Remove only the disposable perf container; keep/reload a known fixture as documented.
docker rm -f atlasmart-neo4j-perf
docker run -d --name atlasmart-neo4j-perf `
--memory=1280m --cpus=1.5 `
-p 127.0.0.1:17474:7474 -p 127.0.0.1:17687:7687 `
-e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
-e NEO4J_server_memory_heap_initial__size=384M `
-e NEO4J_server_memory_heap_max__size=384M `
-e NEO4J_server_memory_pagecache_size=384M `
-v atlasmart-neo4j-perf-data:/data `
-v atlasmart-neo4j-perf-logs:/logs `
neo4j:2026.07.1
Record the configuration as part of the benchmark identity. On Windows Docker Desktop, the Linux container runs through Docker’s virtualization layer; distinguish the container limit from total Windows host resources and do not equate host Task Manager numbers one-to-one with Neo4j process counters.
5. Workload isolation is often safer than squeezing one instance
If long analytical traversals or GDS jobs compete with latency-sensitive order APIs, increasing every memory/thread limit can amplify contention. Isolation choices include scheduling batch jobs outside peak windows, using separate databases/instances where edition/architecture allows, constraining application concurrency, routing workloads by service, and using dedicated analytical infrastructure. In Enterprise/Aura, additional operational/topology controls may exist; the mandatory Community lab teaches the mechanism without claiming those paid capabilities.
6. Deliberately wrong approach: “set heap to almost all RAM”
Heap is only one consumer. Neo4j page cache, native/direct memory, transaction state, network buffers, the OS and other processes still require memory. Large heaps also change GC behavior. A container can be killed by its memory limit even when the configured Java heap is below that limit.
Repair: inventory all memory regions, use
neo4j-admin server memory-recommendation as an
initial estimate, set heap/page cache explicitly, leave measured
headroom, observe real workload/GC/page faults, and adjust one
region at a time. Verify both steady state and restart/failure
behavior.
Production judgment
Infrastructure tuning is subordinate to correctness and workload shape. A bad traversal can consume any amount of hardware. Keep indexes/constraints and plan evidence aligned with the data distribution; cap pathological transaction memory rather than allowing one tenant/query to destabilize the DBMS; budget CPU/I/O for backups, checkpoints, recovery and page-cache warmup; configure driver timeouts/pool limits so saturation fails predictably; preserve TLS/security overhead in realistic tests; and treat managed Aura resource abstractions differently from self-managed host tuning. Lesson 5 turns all these dimensions into a controlled capacity/regression experiment.
Check your understanding
- Why is heap not the same as total Neo4j memory?
- What does neo4j-admin memory-recommendation provide?
- Why can a cold page cache create high I/O?
- Why preserve container headroom?
- When is workload isolation preferable to more tuning?
Review the answers
1. Page cache, native/direct allocations, transaction state, buffers, JVM overhead and OS/container needs exist outside or alongside the configured heap.
2. A starting recommendation based on available memory/store/index observations; it still requires workload-specific testing and tuning.
3. Store/index pages are loaded on demand after startup, so page faults and disk reads are higher until the working set warms.
4. Heap/page cache are not the only allocations; native memory, buffers and runtime/OS needs can otherwise trigger pressure or OOM termination.
5. When competing workloads have incompatible latency/resource profiles and shared saturation creates unacceptable interference.
Summary and next step
Infrastructure Tuning: Heap, Page Cache, Transaction Memory, CPU, Storage, Containers, and OS Limits is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.
Next, continue to Run Controlled Benchmarks and Capacity Tests with Tail Latency, Saturation, Failure, and Regression Thresholds. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.
Authoritative references
- Neo4j current versions — Current database release and LTS baseline.
- Neo4j Operations Manual — Performance — Current performance topics and operational tuning surface.
- Memory configuration — Heap, page cache, transaction/native memory, OS headroom, and memory recommendation guidance.
- Disks, RAM and other tips — Page-cache warmup, storage and RAM behavior.
- Configuration settings — Authoritative current setting names and edition/dynamic boundaries.
- Docker configuration — Container configuration mapping and production configuration guidance.
- Cypher execution plans — EXPLAIN/PROFILE semantics and runtime evidence.
- Cypher operators in detail — Current operators, Rows, DB Hits, memory and plan behavior.
- Indexes for search performance — Current range/text/point/token index behavior and syntax.
- Cypher query tuning — Planner, statistics and query-tuning concepts.
- Neo4j Python driver performance — Driver-side result streaming, database selection and performance guidance.
- Neo4j Python driver API — Connection pool, timeout, retry and fetch-size configuration.
- Neo4j logging — Current debug/query/security/GC logging surfaces and edition boundaries.
- Neo4j metrics — Enterprise metrics surfaces and operational evidence.
- Transaction management — Transaction lifecycle and operational behavior.
- Java requirements — Current supported Java/runtime and platform requirements.
- Neo4j status codes — Classify transient/client/database failures rather than collapsing them into latency.