Chapter 18 · Monitoring and Troubleshooting: Metrics, Logs, Query Visibility, Memory, Store, and Capacity

Page Cache vs Heap, Transaction Memory, Query Memory, File Descriptors, CPU, Storage I/O, and Network

Separate Neo4j memory and host-resource planes precisely: heap, page cache, transaction/query memory, native buffers, file handles, CPU, disk, and network—and diagnose pressure from the correct evidence.

Advanced220–290 minutesMemory and host-resource labNeo4j 2026.07.1 Community mandatory · Enterprise observability optionalCypher 25 · SHOW TRANSACTIONS / PROFILE / SHOW SETTINGSJava 21/25 · Python driver 6.3Last reviewed: September 2026

AtlasMart sees a memory alert and a disk-latency spike at the same time. The fastest wrong response is to increase heap. That could actually reduce room for page cache and the operating system, increasing disk pressure. The correct question is: which memory/resource plane is constrained, and what work is consuming it?

Resource-budget mental model

Physical RAM is shared. Heap, page cache, native/direct buffers, Lucene/vector needs, the OS, filesystem cache and other processes all compete for a finite host budget.

Learning outcomes

01

Distinguish JVM heap, Neo4j page cache, transaction/query memory, and native/direct memory by mechanism.

02

Read current transaction memory/page counters and selected settings without assuming estimates are exact.

03

Correlate file handles, CPU, storage I/O and network with database and client evidence.

04

Diagnose a memory-heavy collection query safely and repair result/cardinality shape before tuning global memory.

05

Build a host memory/resource budget with workload-specific headroom instead of universal percentages.

Reproducible lab baseline

Chapter 18 baseline · reviewed 9 September 2026

The continuity environment remains Neo4j Community 2026.07.1, database neo4j, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic local credential neo4j / atlasmart-course-2026, Java 21 or 25, explicit CYPHER 25 where version-sensitive syntax matters, and Python driver 6.3. Current 5.26 LTS is 5.26.30. No runtime command is claimed to have been executed during lesson generation; expected outputs are fixture invariants or clearly labeled examples.

Monitoring edition boundary

The mandatory path stays free and local. Neo4j Enterprise provides the DBMS metrics framework (CSV/JMX/Prometheus/Graphite integrations), query.log, security.log, query CPU tracking, broader RBAC-based operational visibility, and additional cluster signals. Community still supports useful first-party evidence including SHOW TRANSACTIONS, SHOW SETTINGS on self-managed deployments, PROFILE, general/debug logs, optional JVM GC logging, neo4j-admin server memory-recommendation, driver timing, and operating-system/container counters. Do not fabricate Enterprise-only metrics in the Community lab.

Term Operational meaning
symptom What the user or SLO observes: high p99, timeout, error, stale response, blocked write, memory termination, disk pressure.
evidence A timestamped observation that can support or contradict a hypothesis: plan, transaction row, log entry, OS counter, size snapshot, driver timing.
metric A numeric time series such as CPU load, page faults, heap use, open file descriptors, or request latency; a metric is not a root cause by itself.
heap JVM-managed memory for Java objects and much transaction/query state; garbage collection reclaims unused heap objects.
page cache Neo4j-managed cache for graph-store pages and native indexes read from disk; it is distinct from JVM heap.
transaction/query memory Estimated memory used for uncommitted state, intermediate rows, collections, sort/aggregation/path work, and results; bounded by transaction-memory settings.
native/off-heap memory Memory allocated outside JVM heap, including direct/network buffers and other native structures.
page hit / page fault A hit finds a requested store page in page cache; a fault requires loading a page. Neither ratio alone explains end-to-end latency.
lock wait Time a transaction cannot proceed because another transaction owns a conflicting lock.
database hit A PROFILE plan counter for work against the graph store/index abstractions; it is not a literal physical-disk-read count.
tail latency High-percentile latency such as p95/p99; averages can remain healthy while a minority of requests violate the SLO.
capacity headroom Resources intentionally left unused so spikes, recovery, maintenance, compaction/index builds, and failures do not immediately saturate the system.
Assumption Pinned value / rule
server/current Neo4j 2026.07.1 Community mandatory; 5.26.30 LTS only for compatibility comparison
database neo4j; chapter fixture uses isolated Obs* labels and relationship types
driver Python driver 6.3; one maintained Driver, separate Session/Transaction scopes
TLS/auth loopback disposable lab may use bolt without TLS; remote production requires verified TLS and non-embedded secrets
plugins APOC/GDS are not required
metrics Enterprise metrics exporters are optional/reference-only; Community lab uses SHOW/PROFILE/logs/admin/OS/driver evidence
failure injection only isolated, reversible slow-query/lock/memory/cold-cache exercises; no production cache clearing or destructive disk pressure
measurement record p50/p95/p99, errors, queue/wait, memory, I/O/network and result size where the local tool exposes them; never invent timing values

1. Four memory planes that must not be collapsed

Plane Stores / serves Pressure symptoms First evidence
JVM heap Java objects, planner/runtime objects, much transaction/query state GC pressure, allocation failure, transaction termination, long pauses if poorly sized/workloaded heap settings, GC log, SHOW TRANSACTIONS memory fields
page cache graph store and native index pages page faults, disk reads, cold-start I/O, working-set misses page hit/fault fields, Enterprise page-cache metrics, store size vs page-cache budget
transaction/query memory uncommitted state, collections, sort/aggregation/path intermediates/results one/few queries consume large estimated heap or hit transaction limits currentQueryAllocatedBytes, estimatedUsedHeapMemory, transaction memory settings
native/direct/OS Bolt/network buffers, native structures, JVM overhead, OS/filesystem, Lucene/vector outside page cache process memory exceeds heap+pagecache estimate, handle/direct memory pressure, swapping/OOM allocatedDirectBytes where available, process/container memory, OS counters

2. Read the configured budget

Cypher 25 · inspect relevant settings
CYPHER 25
SHOW SETTINGS
YIELD name, value, isDynamic, isExplicitlySet
WHERE name IN [
  'server.memory.heap.initial_size',
  'server.memory.heap.max_size',
  'server.memory.pagecache.size',
  'dbms.memory.transaction.total.max',
  'db.memory.transaction.total.max',
  'db.memory.transaction.max',
  'db.tx_log.rotation.size',
  'db.tx_log.rotation.retention_policy'
]
RETURN name, value, isDynamic, isExplicitlySet
ORDER BY name;
Configuration is a budget, not a diagnosis

Even a “reasonable” heap/page-cache split can be wrong for a workload with huge transaction state, many databases, vector indexes, GDS projections, or other co-located processes. Measure actual pressure before changing it.

3. Observe a bounded memory-heavy query

The following query materializes a list of maps. It is intentionally inefficient for teaching: start at 50,000 rows in a disposable lab and increase only if your workstation has headroom. While it runs, inspect the same user’s transaction from another session.

Cypher 25 · controlled memory pressure
CYPHER 25
// Disposable lab only. Start smaller; increase gradually while observing transaction memory.
UNWIND range(1,50000) AS n
WITH collect({n:n, bucket:n % 100}) AS rows
RETURN size(rows) AS materializedRows,
       reduce(s=0, x IN rows | s + x.n) AS checksum;
Cypher 25 · observe estimated transaction/query memory
CYPHER 25
SHOW TRANSACTIONS
YIELD transactionId, username, database, currentQuery, currentQueryStatus,
      elapsedTime, waitTime, currentQueryElapsedTime, currentQueryWaitTime,
      activeLockCount, resourceInformation,
      currentQueryAllocatedBytes, allocatedDirectBytes, estimatedUsedHeapMemory,
      pageHits, pageFaults, currentQueryPageHits, currentQueryPageFaults,
      indexes, runtime, statusDetails
RETURN *
ORDER BY currentQueryElapsedTime DESC;
Evidence Interpretation
currentQueryAllocatedBytes rises query is allocating heap during execution
estimatedUsedHeapMemory rises transaction’s estimated heap footprint is growing; estimate is conservative, not exact process RSS
page faults stay low this particular workload is mostly allocation/materialization, not store-page reads
container memory rises + GC activity rises client/query allocation may be stressing heap; align timestamps before concluding

4. Repair the query before the server

If the business requirement is a checksum/count, do not collect 50,000 maps just to reduce them afterward. Stream rows through aggregation instead.

Cypher 25 · streaming aggregation shape
CYPHER 25
UNWIND range(1,50000) AS n
RETURN count(*) AS rowCount, sum(n) AS checksum;
Mechanism

Removing the large materialized list changes peak intermediate state. Verify with the same SHOW TRANSACTIONS/driver/container measurements before considering heap changes.

5. Page cache and disk are coupled

The page cache avoids disk access for frequently used store pages. After a restart it begins cold; page faults and block reads can be high until the active working set warms. If the store+native-index working set is much larger than the page-cache budget, sustained random traversal can continue faulting. But a high page-cache hit ratio cannot excuse high p99 caused by locks, CPU saturation, network, or huge result serialization.

Pattern Likely evidence Do not jump to
cold after restart fault spike + block reads, then decline “disk is broken”
working set > page cache ongoing faults under representative workload “heap should be larger”
large scan many DB hits/pages even with good hit ratio “cache ratio means query is cheap”
lock contention wait time high, page behavior normal “add page cache”

6. File descriptors / handles

Neo4j and Lucene open many files. Enterprise metrics expose current and maximum file-descriptor gauges on supported JVM/OS combinations. The free lab can inspect host/container process handles using platform tools. The question is not “is 5,000 high?” but “is usage approaching the configured OS limit, growing unexpectedly, or correlated with failures?”

Host evidence examples
# Portable container-level snapshot (Docker Desktop / Linux / macOS):
docker stats --no-stream atlasmart-neo4j

# Neo4j store/index sizing guidance from inside the official image:
docker exec atlasmart-neo4j neo4j-admin server memory-recommendation

# Linux host examples when Neo4j runs directly (not Docker Desktop VM):
# ps -o pid,%cpu,%mem,rss,vsz,etime -p <neo4j_pid>
# ls /proc/<neo4j_pid>/fd | wc -l
# iostat -xz 1
# ss -s

# Windows host examples for a directly running Neo4j process:
# Get-Process neo4j | Select-Object Id,CPU,WorkingSet64,PrivateMemorySize64,HandleCount
# Get-Counter '\\PhysicalDisk(*)\\Avg. Disk sec/Read','\\Network Interface(*)\\Bytes Total/sec'

7. CPU, storage I/O and network need denominator/context

Signal Useful question Common false conclusion
CPU 95% Is throughput near saturation? Is it Neo4j process, GC, import, indexing, or another process? 95% means “bad” regardless of completed work
disk latency/I/O wait Do page faults/checkpoints/log writes align with request latency? any disk activity means page cache is too small
network bytes Are large records/result payloads or retries driving traffic? network throughput alone proves network bottleneck
driver pool queue Is database fast but requests wait for connections? server must need more CPU

8. Memory budget worksheet

Budget area Evidence to collect Decision
OS + other processes host used/free/swap, container/runtime overhead reserve enough to avoid host pressure/swapping
heap GC/allocation, transaction estimates, concurrency size for live object/query workload, not store size
page cache store/native-index size, hit/fault behavior, active working set cache useful graph/index pages without starving OS/heap
Lucene/vector Lucene/vector index footprint and OS memory requirements budget separately from Neo4j page cache where documented
transaction headroom concurrent query memory, limits/terminations bound pathological work and preserve multi-query fairness

Check your understanding

  1. Why can increasing heap worsen I/O?
  2. What is page cache for?
  3. Why is estimatedUsedHeapMemory not process RSS?
  4. What should you do before raising transaction memory limits?
  5. Why monitor tail latency with resource saturation?
Review the answers

1. On a fixed-RAM host it can reduce page-cache/OS headroom, causing more page faults or OS pressure.

2. Caching graph-store pages and native indexes to avoid repeated disk access.

3. It estimates transaction-associated heap, not every JVM/native/OS allocation.

4. Fix query/cardinality shape, measure concurrency and prove legitimate work needs more memory without threatening system stability.

5. Throughput/averages can look healthy while queueing makes a minority of requests violate the SLO.

Production judgment

Decision surface Production questions
graph/workload fit Is latency dominated by graph fan-out, result cardinality, locks, indexes, client chatter, or infrastructure rather than raw store size?
correctness Can any proposed remediation change result semantics, transaction boundaries, retry behavior, or causal-read guarantees?
model/cardinality/degree Which labels/types are high-degree or skewed? Are a few hot nodes causing contention or explosive expansion?
latency What are p50/p95/p99 and timeout/error distributions at representative load, not one warm single-user query?
transactions/concurrency How long do transactions stay open? How much wait time and lock scope is normal, and which workflows create hot entities?
memory How are RAM budgets split among OS, heap, page cache, transaction/query memory, direct/native buffers, Lucene/vector needs, and GDS if present?
CPU/disk/network Which resource saturates first? Is disk latency/page-fault activity aligned with slowdown? Is client/network/serialization time dominant?
indexes/constraints Do plans use the intended access paths? What storage/write overhead and population/build headroom do indexes add?
driver Are pool acquisition, retry, timeout, fetch/result consumption and request correlation visible alongside server evidence?
security/tenant risk Do logs/metrics expose sensitive parameters or tenant data? Are monitoring privileges and endpoints restricted appropriately?
backup/recovery Is capacity monitoring independent from backup evidence, and is disk headroom sufficient for logs, dumps/backups and recovery operations?
observability What baseline defines alerts? Which signals are causal/proximate versus merely correlated? How long is telemetry retained?
testing/failure injection Can slow query, lock wait, memory pressure, cold start and network delay be reproduced in an isolated environment with reset steps?
version/edition/Aura Which fields/settings/logs/metrics exist on the exact server, Cypher version, edition and managed tier?
cost/migration What telemetry/storage overhead is acceptable, and how will capacity changes or schema/model remediations be rolled back?

Summary and next step

Resource planes are now separated. Lesson 3 joins that resource evidence to Cypher plans, running transactions, lock waits and application timing for a concrete slow-request diagnosis.

Authoritative references

  • Current Neo4j versions — Current Neo4j 2026.07.1 and 5.26.30 LTS release snapshot.
  • Monitoring overview — Current monitoring surfaces: logs, metrics, query/transaction management, connections, jobs, and reports.
  • Metrics — Enterprise metrics architecture and supported export surfaces.
  • Essential metrics — Server, Neo4j, cluster, and workload signals that require correlation rather than single-metric tuning.
  • Metrics reference — Current page-cache, JVM, CPU, file-descriptor, transaction and store metric names.
  • Logging — Current neo4j/debug/http/gc/query/security log behavior and edition boundaries.
  • Show and terminate transactions — Current SHOW TRANSACTIONS fields for locks, wait time, memory, page hits/faults, runtime, indexes, and status.
  • Manage queries — Current query visibility through transaction management rather than legacy dbms.listQueries().
  • Memory configuration — Heap, page cache, native memory, transaction limits, and memory-recommendation guidance.
  • Disks, RAM and other tips — Page-cache warmup, disk behavior, and operating-system considerations.
  • Performance — Performance topics spanning memory, disks, statistics, execution plans, Bolt threads, and filesystem tuning.
  • Show configuration settings — Current SHOW SETTINGS syntax and self-managed server scope.
  • Transaction logs — Transaction-log location, rotation, retention, pruning, and capacity implications.
  • Query tuning and PROFILE — Execution-plan evidence for rows, database hits, memory, access paths, and operator shape.
  • Python driver performance — Driver-side timing, database selection, result consumption, and application/database boundary considerations.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.