Chapter 18 · Monitoring and Troubleshooting: Metrics, Logs, Query Visibility, Memory, Store, and Capacity
Page Cache vs Heap, Transaction Memory, Query Memory, File Descriptors, CPU, Storage I/O, and Network
Separate Neo4j memory and host-resource planes precisely: heap, page cache, transaction/query memory, native buffers, file handles, CPU, disk, and network—and diagnose pressure from the correct evidence.
AtlasMart sees a memory alert and a disk-latency spike at the same time. The fastest wrong response is to increase heap. That could actually reduce room for page cache and the operating system, increasing disk pressure. The correct question is: which memory/resource plane is constrained, and what work is consuming it?
Physical RAM is shared. Heap, page cache, native/direct buffers, Lucene/vector needs, the OS, filesystem cache and other processes all compete for a finite host budget.
Learning outcomes
Distinguish JVM heap, Neo4j page cache, transaction/query memory, and native/direct memory by mechanism.
Read current transaction memory/page counters and selected settings without assuming estimates are exact.
Correlate file handles, CPU, storage I/O and network with database and client evidence.
Diagnose a memory-heavy collection query safely and repair result/cardinality shape before tuning global memory.
Build a host memory/resource budget with workload-specific headroom instead of universal percentages.
Reproducible lab baseline
The continuity environment remains
Neo4j Community 2026.07.1, database
neo4j, container atlasmart-neo4j,
loopback Bolt bolt://127.0.0.1:7687, synthetic
local credential neo4j / atlasmart-course-2026,
Java 21 or 25, explicit CYPHER 25 where
version-sensitive syntax matters, and Python driver 6.3.
Current 5.26 LTS is 5.26.30. No runtime command
is claimed to have been executed during lesson generation;
expected outputs are fixture invariants or clearly labeled
examples.
The mandatory path stays free and local. Neo4j Enterprise
provides the DBMS metrics framework
(CSV/JMX/Prometheus/Graphite integrations), query.log,
security.log, query CPU tracking, broader RBAC-based
operational visibility, and additional cluster signals.
Community still supports useful first-party evidence including
SHOW TRANSACTIONS, SHOW SETTINGS on
self-managed deployments, PROFILE, general/debug
logs, optional JVM GC logging,
neo4j-admin server memory-recommendation, driver
timing, and operating-system/container counters. Do not
fabricate Enterprise-only metrics in the Community lab.
| Term | Operational meaning |
|---|---|
| symptom | What the user or SLO observes: high p99, timeout, error, stale response, blocked write, memory termination, disk pressure. |
| evidence | A timestamped observation that can support or contradict a hypothesis: plan, transaction row, log entry, OS counter, size snapshot, driver timing. |
| metric | A numeric time series such as CPU load, page faults, heap use, open file descriptors, or request latency; a metric is not a root cause by itself. |
| heap | JVM-managed memory for Java objects and much transaction/query state; garbage collection reclaims unused heap objects. |
| page cache | Neo4j-managed cache for graph-store pages and native indexes read from disk; it is distinct from JVM heap. |
| transaction/query memory | Estimated memory used for uncommitted state, intermediate rows, collections, sort/aggregation/path work, and results; bounded by transaction-memory settings. |
| native/off-heap memory | Memory allocated outside JVM heap, including direct/network buffers and other native structures. |
| page hit / page fault | A hit finds a requested store page in page cache; a fault requires loading a page. Neither ratio alone explains end-to-end latency. |
| lock wait | Time a transaction cannot proceed because another transaction owns a conflicting lock. |
| database hit | A PROFILE plan counter for work against the graph store/index abstractions; it is not a literal physical-disk-read count. |
| tail latency | High-percentile latency such as p95/p99; averages can remain healthy while a minority of requests violate the SLO. |
| capacity headroom | Resources intentionally left unused so spikes, recovery, maintenance, compaction/index builds, and failures do not immediately saturate the system. |
| Assumption | Pinned value / rule |
|---|---|
| server/current | Neo4j 2026.07.1 Community mandatory; 5.26.30 LTS only for compatibility comparison |
| database | neo4j; chapter fixture uses isolated Obs* labels and relationship types |
| driver | Python driver 6.3; one maintained Driver, separate Session/Transaction scopes |
| TLS/auth | loopback disposable lab may use bolt without TLS; remote production requires verified TLS and non-embedded secrets |
| plugins | APOC/GDS are not required |
| metrics | Enterprise metrics exporters are optional/reference-only; Community lab uses SHOW/PROFILE/logs/admin/OS/driver evidence |
| failure injection | only isolated, reversible slow-query/lock/memory/cold-cache exercises; no production cache clearing or destructive disk pressure |
| measurement | record p50/p95/p99, errors, queue/wait, memory, I/O/network and result size where the local tool exposes them; never invent timing values |
1. Four memory planes that must not be collapsed
| Plane | Stores / serves | Pressure symptoms | First evidence |
|---|---|---|---|
| JVM heap | Java objects, planner/runtime objects, much transaction/query state | GC pressure, allocation failure, transaction termination, long pauses if poorly sized/workloaded | heap settings, GC log, SHOW TRANSACTIONS memory fields |
| page cache | graph store and native index pages | page faults, disk reads, cold-start I/O, working-set misses | page hit/fault fields, Enterprise page-cache metrics, store size vs page-cache budget |
| transaction/query memory | uncommitted state, collections, sort/aggregation/path intermediates/results | one/few queries consume large estimated heap or hit transaction limits | currentQueryAllocatedBytes, estimatedUsedHeapMemory, transaction memory settings |
| native/direct/OS | Bolt/network buffers, native structures, JVM overhead, OS/filesystem, Lucene/vector outside page cache | process memory exceeds heap+pagecache estimate, handle/direct memory pressure, swapping/OOM | allocatedDirectBytes where available, process/container memory, OS counters |
2. Read the configured budget
CYPHER 25
SHOW SETTINGS
YIELD name, value, isDynamic, isExplicitlySet
WHERE name IN [
'server.memory.heap.initial_size',
'server.memory.heap.max_size',
'server.memory.pagecache.size',
'dbms.memory.transaction.total.max',
'db.memory.transaction.total.max',
'db.memory.transaction.max',
'db.tx_log.rotation.size',
'db.tx_log.rotation.retention_policy'
]
RETURN name, value, isDynamic, isExplicitlySet
ORDER BY name;
Even a “reasonable” heap/page-cache split can be wrong for a workload with huge transaction state, many databases, vector indexes, GDS projections, or other co-located processes. Measure actual pressure before changing it.
3. Observe a bounded memory-heavy query
The following query materializes a list of maps. It is intentionally inefficient for teaching: start at 50,000 rows in a disposable lab and increase only if your workstation has headroom. While it runs, inspect the same user’s transaction from another session.
CYPHER 25
// Disposable lab only. Start smaller; increase gradually while observing transaction memory.
UNWIND range(1,50000) AS n
WITH collect({n:n, bucket:n % 100}) AS rows
RETURN size(rows) AS materializedRows,
reduce(s=0, x IN rows | s + x.n) AS checksum;
CYPHER 25
SHOW TRANSACTIONS
YIELD transactionId, username, database, currentQuery, currentQueryStatus,
elapsedTime, waitTime, currentQueryElapsedTime, currentQueryWaitTime,
activeLockCount, resourceInformation,
currentQueryAllocatedBytes, allocatedDirectBytes, estimatedUsedHeapMemory,
pageHits, pageFaults, currentQueryPageHits, currentQueryPageFaults,
indexes, runtime, statusDetails
RETURN *
ORDER BY currentQueryElapsedTime DESC;
| Evidence | Interpretation |
|---|---|
| currentQueryAllocatedBytes rises | query is allocating heap during execution |
| estimatedUsedHeapMemory rises | transaction’s estimated heap footprint is growing; estimate is conservative, not exact process RSS |
| page faults stay low | this particular workload is mostly allocation/materialization, not store-page reads |
| container memory rises + GC activity rises | client/query allocation may be stressing heap; align timestamps before concluding |
4. Repair the query before the server
If the business requirement is a checksum/count, do not collect 50,000 maps just to reduce them afterward. Stream rows through aggregation instead.
CYPHER 25
UNWIND range(1,50000) AS n
RETURN count(*) AS rowCount, sum(n) AS checksum;
Removing the large materialized list changes peak intermediate state. Verify with the same SHOW TRANSACTIONS/driver/container measurements before considering heap changes.
5. Page cache and disk are coupled
The page cache avoids disk access for frequently used store pages. After a restart it begins cold; page faults and block reads can be high until the active working set warms. If the store+native-index working set is much larger than the page-cache budget, sustained random traversal can continue faulting. But a high page-cache hit ratio cannot excuse high p99 caused by locks, CPU saturation, network, or huge result serialization.
| Pattern | Likely evidence | Do not jump to |
|---|---|---|
| cold after restart | fault spike + block reads, then decline | “disk is broken” |
| working set > page cache | ongoing faults under representative workload | “heap should be larger” |
| large scan | many DB hits/pages even with good hit ratio | “cache ratio means query is cheap” |
| lock contention | wait time high, page behavior normal | “add page cache” |
6. File descriptors / handles
Neo4j and Lucene open many files. Enterprise metrics expose current and maximum file-descriptor gauges on supported JVM/OS combinations. The free lab can inspect host/container process handles using platform tools. The question is not “is 5,000 high?” but “is usage approaching the configured OS limit, growing unexpectedly, or correlated with failures?”
# Portable container-level snapshot (Docker Desktop / Linux / macOS):
docker stats --no-stream atlasmart-neo4j
# Neo4j store/index sizing guidance from inside the official image:
docker exec atlasmart-neo4j neo4j-admin server memory-recommendation
# Linux host examples when Neo4j runs directly (not Docker Desktop VM):
# ps -o pid,%cpu,%mem,rss,vsz,etime -p <neo4j_pid>
# ls /proc/<neo4j_pid>/fd | wc -l
# iostat -xz 1
# ss -s
# Windows host examples for a directly running Neo4j process:
# Get-Process neo4j | Select-Object Id,CPU,WorkingSet64,PrivateMemorySize64,HandleCount
# Get-Counter '\\PhysicalDisk(*)\\Avg. Disk sec/Read','\\Network Interface(*)\\Bytes Total/sec'
7. CPU, storage I/O and network need denominator/context
| Signal | Useful question | Common false conclusion |
|---|---|---|
| CPU 95% | Is throughput near saturation? Is it Neo4j process, GC, import, indexing, or another process? | 95% means “bad” regardless of completed work |
| disk latency/I/O wait | Do page faults/checkpoints/log writes align with request latency? | any disk activity means page cache is too small |
| network bytes | Are large records/result payloads or retries driving traffic? | network throughput alone proves network bottleneck |
| driver pool queue | Is database fast but requests wait for connections? | server must need more CPU |
8. Memory budget worksheet
| Budget area | Evidence to collect | Decision |
|---|---|---|
| OS + other processes | host used/free/swap, container/runtime overhead | reserve enough to avoid host pressure/swapping |
| heap | GC/allocation, transaction estimates, concurrency | size for live object/query workload, not store size |
| page cache | store/native-index size, hit/fault behavior, active working set | cache useful graph/index pages without starving OS/heap |
| Lucene/vector | Lucene/vector index footprint and OS memory requirements | budget separately from Neo4j page cache where documented |
| transaction headroom | concurrent query memory, limits/terminations | bound pathological work and preserve multi-query fairness |
Check your understanding
- Why can increasing heap worsen I/O?
- What is page cache for?
- Why is estimatedUsedHeapMemory not process RSS?
- What should you do before raising transaction memory limits?
- Why monitor tail latency with resource saturation?
Review the answers
1. On a fixed-RAM host it can reduce page-cache/OS headroom, causing more page faults or OS pressure.
2. Caching graph-store pages and native indexes to avoid repeated disk access.
3. It estimates transaction-associated heap, not every JVM/native/OS allocation.
4. Fix query/cardinality shape, measure concurrency and prove legitimate work needs more memory without threatening system stability.
5. Throughput/averages can look healthy while queueing makes a minority of requests violate the SLO.
Production judgment
| Decision surface | Production questions |
|---|---|
| graph/workload fit | Is latency dominated by graph fan-out, result cardinality, locks, indexes, client chatter, or infrastructure rather than raw store size? |
| correctness | Can any proposed remediation change result semantics, transaction boundaries, retry behavior, or causal-read guarantees? |
| model/cardinality/degree | Which labels/types are high-degree or skewed? Are a few hot nodes causing contention or explosive expansion? |
| latency | What are p50/p95/p99 and timeout/error distributions at representative load, not one warm single-user query? |
| transactions/concurrency | How long do transactions stay open? How much wait time and lock scope is normal, and which workflows create hot entities? |
| memory | How are RAM budgets split among OS, heap, page cache, transaction/query memory, direct/native buffers, Lucene/vector needs, and GDS if present? |
| CPU/disk/network | Which resource saturates first? Is disk latency/page-fault activity aligned with slowdown? Is client/network/serialization time dominant? |
| indexes/constraints | Do plans use the intended access paths? What storage/write overhead and population/build headroom do indexes add? |
| driver | Are pool acquisition, retry, timeout, fetch/result consumption and request correlation visible alongside server evidence? |
| security/tenant risk | Do logs/metrics expose sensitive parameters or tenant data? Are monitoring privileges and endpoints restricted appropriately? |
| backup/recovery | Is capacity monitoring independent from backup evidence, and is disk headroom sufficient for logs, dumps/backups and recovery operations? |
| observability | What baseline defines alerts? Which signals are causal/proximate versus merely correlated? How long is telemetry retained? |
| testing/failure injection | Can slow query, lock wait, memory pressure, cold start and network delay be reproduced in an isolated environment with reset steps? |
| version/edition/Aura | Which fields/settings/logs/metrics exist on the exact server, Cypher version, edition and managed tier? |
| cost/migration | What telemetry/storage overhead is acceptable, and how will capacity changes or schema/model remediations be rolled back? |
Summary and next step
Resource planes are now separated. Lesson 3 joins that resource evidence to Cypher plans, running transactions, lock waits and application timing for a concrete slow-request diagnosis.
Authoritative references
- Current Neo4j versions — Current Neo4j 2026.07.1 and 5.26.30 LTS release snapshot.
- Monitoring overview — Current monitoring surfaces: logs, metrics, query/transaction management, connections, jobs, and reports.
- Metrics — Enterprise metrics architecture and supported export surfaces.
- Essential metrics — Server, Neo4j, cluster, and workload signals that require correlation rather than single-metric tuning.
- Metrics reference — Current page-cache, JVM, CPU, file-descriptor, transaction and store metric names.
- Logging — Current neo4j/debug/http/gc/query/security log behavior and edition boundaries.
- Show and terminate transactions — Current SHOW TRANSACTIONS fields for locks, wait time, memory, page hits/faults, runtime, indexes, and status.
- Manage queries — Current query visibility through transaction management rather than legacy dbms.listQueries().
- Memory configuration — Heap, page cache, native memory, transaction limits, and memory-recommendation guidance.
- Disks, RAM and other tips — Page-cache warmup, disk behavior, and operating-system considerations.
- Performance — Performance topics spanning memory, disks, statistics, execution plans, Bolt threads, and filesystem tuning.
- Show configuration settings — Current SHOW SETTINGS syntax and self-managed server scope.
- Transaction logs — Transaction-log location, rotation, retention, pruning, and capacity implications.
- Query tuning and PROFILE — Execution-plan evidence for rows, database hits, memory, access paths, and operator shape.
- Python driver performance — Driver-side timing, database selection, result consumption, and application/database boundary considerations.