Chapter 26 · Performance Engineering: Data Model, Traversal Shape, Page Cache, Memory, and Workload Isolation

Infrastructure Tuning: Heap, Page Cache, Transaction Memory, CPU, Storage, Containers, and OS Limits

Divide a bounded memory and I/O budget deliberately among JVM heap, Neo4j page cache, transaction/query memory, native buffers, the OS/container, CPU, and storage while preserving failure headroom.

Advanced280–420 minutesHeap · page cache · transaction memory · I/ONeo4j 2026.07.1 · Community mandatoryCypher 25 · Python driver 6.3.x optionalDisposable perf container · 2 GiB/2 CPU baselineJava 21/25 · no APOC/GDS requiredLast reviewed: September 2026

Learning outcomes

01

Separate JVM heap, Neo4j page cache, transaction/query memory, native/off-heap memory, network buffers and OS/container headroom.

02

Use explicit memory configuration and neo4j-admin memory recommendations as starting evidence, not universal sizing formulas.

03

Explain how CPU saturation, storage latency/IOPS, page faults, GC and file-descriptor limits create different symptoms.

04

Tune a disposable container one resource dimension at a time while respecting Docker Desktop/host/VM boundaries on Windows.

05

Define workload isolation and capacity headroom so analytical/batch traffic cannot consume the latency budget of transactional requests.

Execution and safety note

Treat every command, query, configuration change, benchmark, security change, failure injection, and cleanup step in this lesson as scoped to the disposable AtlasMart course lab unless the text explicitly says otherwise. Verify the actual Neo4j, Cypher, driver, plugin/GDS, edition/tier, authentication, TLS, and deployment state before execution. Expected results describe invariants and evidence shapes; they are not fabricated claims that this generated lesson captured a live production run.

1. AtlasMart problem: increasing heap can make the database slower

An operator sees GC activity and doubles the JVM heap until the container consumes nearly all available memory. That leaves less memory for Neo4j page cache, native buffers and the environment hosting the container. If the working graph no longer fits well in page cache, store reads can move to disk and latency can worsen. Memory tuning is a budget-allocation problem, not a contest to maximize one region.

Dimension Chapter 26 reproducible assumption
Neo4j 2026.07.1 Community. The continuity database remains neo4j; performance experiments use a separate disposable container named atlasmart-neo4j-perf so tuning/failure tests do not disturb earlier labs.
Cypher Cypher 25 examples. Planner/operator names and numeric PROFILE values are evidence to capture locally, not constants to memorize.
Java Neo4j 2026 line with a supported Java 21/25 runtime as supplied/required by the chosen distribution.
Driver Neo4j Python driver 6.3.x for the optional load harness; the driver object is shared, sessions/transactions are not shared between worker threads.
Auth/TLS User neo4j, password atlasmart-course-2026. Loopback Bolt without TLS only for the isolated disposable lab; remote/production traffic should use validated TLS.
Ports Performance container maps HTTP 17474→7474 and Bolt 17687→7687, avoiding the continuity container on 7474/7687.
Initial resource envelope Exercise baseline: Docker limit 2 GiB, 2 CPUs, explicit heap 512 MiB, explicit page cache 512 MiB. These are lab controls, not production recommendations.
Dataset 200 CH26 customers, 300 products, 2,000 orders, 6,000 CONTAINS relationships, 1,200 VIEWED_CH26 relationships, six categories, plus an isolated hot-counter node. Recount locally after setup.
Observability Community labs rely on PROFILE, SHOW commands, Docker/OS counters, driver timing, container stats and logs. Neo4j metrics exporters and query.log are Enterprise surfaces and are optional, clearly labeled.
Evidence rule This generated material does not execute your Docker host. Latencies, DB Hits, page-cache behavior, saturation points, GC, throughput, errors and recovery time must be measured locally; illustrative tables are labeled as templates.

2. The memory budget has distinct mechanisms

Region What it primarily holds What to observe / common mistake
JVM heap Java objects, query/transaction structures, runtime metadata GC pauses/occupancy and query memory. Oversizing can starve page cache/OS/native needs.
Neo4j page cache Cached native store/index pages Page hits/faults and storage I/O. It is separate from heap.
Transaction/query memory Uncommitted state and intermediate query structures SHOW TRANSACTIONS/query evidence; per-transaction/database/global limits can terminate pathological work safely.
Native/off-heap + network buffers Direct/native allocations outside GC-managed heap Container RSS/native growth; do not assume heap max equals total process memory.
OS/container headroom Filesystem cache, vector-index OS memory, kernel/runtime, Docker/VM overhead Swap/host pressure/container OOM. Leave explicit headroom.

Current Neo4j exposes server.memory.heap.initial_size, server.memory.heap.max_size, server.memory.pagecache.size, and transaction-memory controls including dbms.memory.transaction.total.max, db.memory.transaction.total.max, and db.memory.transaction.max. Exact defaults can be calculated dynamically and can change; query SHOW SETTINGS on the target release.

Inspect the actual memory configuration
RETURN 1 AS cypherReachable;
CALL dbms.components() YIELD name, versions, edition
RETURN name, versions, edition;
SHOW SETTINGS YIELD name, value
WHERE name IN [
  'server.memory.heap.initial_size',
  'server.memory.heap.max_size',
  'server.memory.pagecache.size',
  'dbms.memory.transaction.total.max',
  'db.memory.transaction.total.max',
  'db.memory.transaction.max'
]
RETURN name, value ORDER BY name;
SHOW INDEXES YIELD name, state, type, entityType, labelsOrTypes, properties
WHERE name STARTS WITH 'ch26_'
RETURN name, state, type, entityType, labelsOrTypes, properties ORDER BY name;
Get a sizing starting point from neo4j-admin
# Run inside the disposable container. This is a recommendation, not a benchmark result.
docker exec atlasmart-neo4j-perf neo4j-admin server memory-recommendation --memory=2g

# Observe process/container resource use during a controlled run.
docker stats atlasmart-neo4j-perf

# Inspect filesystem/volume growth from the container shell if needed.
docker exec atlasmart-neo4j-perf sh -lc 'du -sh /data /logs; ulimit -n' 

3. CPU, storage, and file descriptors are different bottlenecks

Signal pattern Plausible hypothesis Evidence before remediation
CPU near allocated limit; throughput plateaus Compute-bound query/operators or too much offered concurrency PROFILE operator work, host/container CPU, runnable threads, same dataset at lower concurrency.
Low CPU but high latency + page faults/I/O wait Working set misses page cache or storage latency dominates Cold/warm comparison, disk latency/throughput, page-cache evidence where available.
Long GC pauses / heap pressure Large intermediate rows, result materialization, transaction memory, insufficient/poorly balanced heap GC logs, PROFILE memory, query shape, heap occupancy—not heap size alone.
Connection/open-file failures File descriptor/socket exhaustion or pool misuse OS/container limits, driver pool/session lifecycle, active connections.
Driver waits while server CPU idle Pool acquisition/network/client bottleneck Driver timing, pool limits, connection acquisition timeout, network RTT/result bytes.

4. One-factor container experiments

Use a new disposable container/volume or restore the same fixture before each resource experiment. Keep the query mix and data identical. For example, compare the 2 GiB exercise baseline with a deliberately constrained 1.25 GiB container while keeping explicit heap/page cache small enough to leave native headroom. Do not run a configuration whose memory regions sum to the whole container limit.

Example alternate envelope — intentionally a lab control, not a recommendation
# Remove only the disposable perf container; keep/reload a known fixture as documented.
docker rm -f atlasmart-neo4j-perf

docker run -d --name atlasmart-neo4j-perf `
  --memory=1280m --cpus=1.5 `
  -p 127.0.0.1:17474:7474 -p 127.0.0.1:17687:7687 `
  -e NEO4J_AUTH=neo4j/atlasmart-course-2026 `
  -e NEO4J_server_memory_heap_initial__size=384M `
  -e NEO4J_server_memory_heap_max__size=384M `
  -e NEO4J_server_memory_pagecache_size=384M `
  -v atlasmart-neo4j-perf-data:/data `
  -v atlasmart-neo4j-perf-logs:/logs `
  neo4j:2026.07.1

Record the configuration as part of the benchmark identity. On Windows Docker Desktop, the Linux container runs through Docker’s virtualization layer; distinguish the container limit from total Windows host resources and do not equate host Task Manager numbers one-to-one with Neo4j process counters.

5. Workload isolation is often safer than squeezing one instance

If long analytical traversals or GDS jobs compete with latency-sensitive order APIs, increasing every memory/thread limit can amplify contention. Isolation choices include scheduling batch jobs outside peak windows, using separate databases/instances where edition/architecture allows, constraining application concurrency, routing workloads by service, and using dedicated analytical infrastructure. In Enterprise/Aura, additional operational/topology controls may exist; the mandatory Community lab teaches the mechanism without claiming those paid capabilities.

6. Deliberately wrong approach: “set heap to almost all RAM”

Heap is only one consumer. Neo4j page cache, native/direct memory, transaction state, network buffers, the OS and other processes still require memory. Large heaps also change GC behavior. A container can be killed by its memory limit even when the configured Java heap is below that limit.

Repair: inventory all memory regions, use neo4j-admin server memory-recommendation as an initial estimate, set heap/page cache explicitly, leave measured headroom, observe real workload/GC/page faults, and adjust one region at a time. Verify both steady state and restart/failure behavior.

Production judgment

Infrastructure tuning is subordinate to correctness and workload shape. A bad traversal can consume any amount of hardware. Keep indexes/constraints and plan evidence aligned with the data distribution; cap pathological transaction memory rather than allowing one tenant/query to destabilize the DBMS; budget CPU/I/O for backups, checkpoints, recovery and page-cache warmup; configure driver timeouts/pool limits so saturation fails predictably; preserve TLS/security overhead in realistic tests; and treat managed Aura resource abstractions differently from self-managed host tuning. Lesson 5 turns all these dimensions into a controlled capacity/regression experiment.

Check your understanding

  1. Why is heap not the same as total Neo4j memory?
  2. What does neo4j-admin memory-recommendation provide?
  3. Why can a cold page cache create high I/O?
  4. Why preserve container headroom?
  5. When is workload isolation preferable to more tuning?
Review the answers

1. Page cache, native/direct allocations, transaction state, buffers, JVM overhead and OS/container needs exist outside or alongside the configured heap.

2. A starting recommendation based on available memory/store/index observations; it still requires workload-specific testing and tuning.

3. Store/index pages are loaded on demand after startup, so page faults and disk reads are higher until the working set warms.

4. Heap/page cache are not the only allocations; native memory, buffers and runtime/OS needs can otherwise trigger pressure or OOM termination.

5. When competing workloads have incompatible latency/resource profiles and shared saturation creates unacceptable interference.

Summary and next step

Infrastructure Tuning: Heap, Page Cache, Transaction Memory, CPU, Storage, Containers, and OS Limits is useful only when its assumptions and observed evidence stay attached to the decision. The examples above establish a reproducible mechanism and boundary; they do not turn one lab result into a universal production rule.

Next, continue to Run Controlled Benchmarks and Capacity Tests with Tail Latency, Saturation, Failure, and Regression Thresholds. Carry forward the verified assumptions, fixture state, version/edition boundaries, and measurements from this lesson instead of treating the next topic as an isolated recipe.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.