Chapter 18 · Monitoring and Troubleshooting: Metrics, Logs, Query Visibility, Memory, Store, and Capacity

DBMS Metrics, Query/Transaction Listings, Logs, GC/JVM Evidence, Page Cache, and Operating-System Signals

Build an evidence map across Neo4j, JVM, logs, transactions, page cache, and operating-system signals so AtlasMart can distinguish a database problem from client, network, memory, or host pressure.

Advanced210–280 minutesObservability evidence-map labNeo4j 2026.07.1 Community mandatory · Enterprise observability optionalCypher 25 · SHOW TRANSACTIONS / PROFILE / SHOW SETTINGSJava 21/25 · Python driver 6.3Last reviewed: September 2026

AtlasMart receives “Neo4j is slow” tickets every afternoon. One engineer proposes a larger heap, another wants more indexes, and a third blames the network. None has yet shown whether the slow requests are running, waiting on locks, faulting pages, blocked in the driver pool, or spending time outside Neo4j. This lesson replaces opinion with an evidence plane: observe the transaction, plan, logs, JVM/memory, host, driver, and result together before changing anything.

Troubleshooting rule

A symptom narrows where to look; it does not identify the root cause. Capture synchronized evidence before tuning.

Learning outcomes

01

Map a latency or failure symptom to Neo4j transaction, plan, log, memory/JVM, OS and driver evidence.

02

Use SHOW TRANSACTIONS and SHOW SETTINGS without confusing unavailable/nullable fields with zero usage.

03

Separate Community-visible evidence from Enterprise metrics, query.log and security.log capabilities.

04

Interpret page hits/faults, GC, CPU, disk and network as correlated signals rather than standalone diagnoses.

05

Create a reproducible baseline that can be compared before and after one remediation.

Reproducible lab baseline

Chapter 18 baseline · reviewed 9 September 2026

The continuity environment remains Neo4j Community 2026.07.1, database neo4j, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic local credential neo4j / atlasmart-course-2026, Java 21 or 25, explicit CYPHER 25 where version-sensitive syntax matters, and Python driver 6.3. Current 5.26 LTS is 5.26.30. No runtime command is claimed to have been executed during lesson generation; expected outputs are fixture invariants or clearly labeled examples.

Monitoring edition boundary

The mandatory path stays free and local. Neo4j Enterprise provides the DBMS metrics framework (CSV/JMX/Prometheus/Graphite integrations), query.log, security.log, query CPU tracking, broader RBAC-based operational visibility, and additional cluster signals. Community still supports useful first-party evidence including SHOW TRANSACTIONS, SHOW SETTINGS on self-managed deployments, PROFILE, general/debug logs, optional JVM GC logging, neo4j-admin server memory-recommendation, driver timing, and operating-system/container counters. Do not fabricate Enterprise-only metrics in the Community lab.

Term Operational meaning
symptom What the user or SLO observes: high p99, timeout, error, stale response, blocked write, memory termination, disk pressure.
evidence A timestamped observation that can support or contradict a hypothesis: plan, transaction row, log entry, OS counter, size snapshot, driver timing.
metric A numeric time series such as CPU load, page faults, heap use, open file descriptors, or request latency; a metric is not a root cause by itself.
heap JVM-managed memory for Java objects and much transaction/query state; garbage collection reclaims unused heap objects.
page cache Neo4j-managed cache for graph-store pages and native indexes read from disk; it is distinct from JVM heap.
transaction/query memory Estimated memory used for uncommitted state, intermediate rows, collections, sort/aggregation/path work, and results; bounded by transaction-memory settings.
native/off-heap memory Memory allocated outside JVM heap, including direct/network buffers and other native structures.
page hit / page fault A hit finds a requested store page in page cache; a fault requires loading a page. Neither ratio alone explains end-to-end latency.
lock wait Time a transaction cannot proceed because another transaction owns a conflicting lock.
database hit A PROFILE plan counter for work against the graph store/index abstractions; it is not a literal physical-disk-read count.
tail latency High-percentile latency such as p95/p99; averages can remain healthy while a minority of requests violate the SLO.
capacity headroom Resources intentionally left unused so spikes, recovery, maintenance, compaction/index builds, and failures do not immediately saturate the system.
Assumption Pinned value / rule
server/current Neo4j 2026.07.1 Community mandatory; 5.26.30 LTS only for compatibility comparison
database neo4j; chapter fixture uses isolated Obs* labels and relationship types
driver Python driver 6.3; one maintained Driver, separate Session/Transaction scopes
TLS/auth loopback disposable lab may use bolt without TLS; remote production requires verified TLS and non-embedded secrets
plugins APOC/GDS are not required
metrics Enterprise metrics exporters are optional/reference-only; Community lab uses SHOW/PROFILE/logs/admin/OS/driver evidence
failure injection only isolated, reversible slow-query/lock/memory/cold-cache exercises; no production cache clearing or destructive disk pressure
measurement record p50/p95/p99, errors, queue/wait, memory, I/O/network and result size where the local tool exposes them; never invent timing values

1. Monitoring is a multi-plane model

Plane Question Community/local evidence Enterprise/additional evidence
application What did the user experience? request ID, p50/p95/p99, error class, result bytes same plus centralized tracing/APM if used
driver/Bolt Was time spent acquiring, sending, fetching, retrying? client timestamps, driver debug where enabled, pool/timeouts configured in app centralized telemetry correlation
Cypher/transaction Running, waiting, allocating, faulting, using which index? SHOW TRANSACTIONS, PROFILE, transaction status/wait/memory/page fields query.log and broader privilege-scoped visibility
DBMS/JVM Heap/GC/page cache/native pressure? SHOW SETTINGS, debug/gc logs, memory-recommendation CSV/JMX/Prometheus metrics and query CPU tracking
OS/container CPU, memory, block I/O, network, handles? docker stats, host OS counters same plus monitoring agent
store/capacity How fast are data, indexes and tx logs growing? memory-recommendation, file-size snapshots, tx-log config metrics/central retention and fleet views

2. Build a reversible observation fixture

The fixture is isolated from the course’s Customer/Product/Order graph. It creates 120 synthetic customers, 1,200 items, a small set of observed relationships, and one deliberately hot inventory node for later lock tests. The counts are intentionally modest: the goal is reproducibility, not to force a workstation into distress.

Cypher 25 · Chapter 18 observation fixture
CYPHER 25
CREATE CONSTRAINT obs_customer_id IF NOT EXISTS
FOR (c:ObsCustomer) REQUIRE c.obsId IS UNIQUE;
CREATE CONSTRAINT obs_item_id IF NOT EXISTS
FOR (i:ObsItem) REQUIRE i.obsId IS UNIQUE;
CREATE CONSTRAINT obs_inventory_id IF NOT EXISTS
FOR (x:ObsInventory) REQUIRE x.obsId IS UNIQUE;

UNWIND range(1,120) AS n
MERGE (c:ObsCustomer {obsId:'OC-' + toString(n)})
SET c.segment = n % 12,
    c.region = CASE n % 4 WHEN 0 THEN 'west' WHEN 1 THEN 'east' WHEN 2 THEN 'north' ELSE 'south' END;

UNWIND range(1,1200) AS n
MERGE (i:ObsItem {obsId:'OI-' + toString(n)})
SET i.segment = n % 12,
    i.price = 10.0 + (n % 100),
    i.active = n % 10 <> 0;

MATCH (c:ObsCustomer), (i:ObsItem)
WHERE toInteger(substring(c.obsId,3)) % 40 = 0
  AND toInteger(substring(i.obsId,3)) % 120 = 0
MERGE (c)-[:OBS_VIEWED]->(i);

MERGE (x:ObsInventory {obsId:'INV-HOT-18'})
SET x.available = 1000, x.reserved = 0, x.version = 0;
Expected invariant

After a clean setup: 120 ObsCustomer nodes, 1,200 ObsItem nodes, one ObsInventory node, three uniqueness constraints, and deterministic OBS_VIEWED relationships. Rerunning setup does not multiply nodes or relationships because identities use MERGE.

3. Read active transactions before guessing

SHOW TRANSACTIONS replaces the old “list queries” mental model because a query executes inside a transaction. Current output can expose current query text/status, elapsed and lock-wait time, lock counts and blocker metadata, estimated heap/direct memory, page hits/faults, runtime, indexes, and status details. Some values can be null when unavailable; CPU tracking in particular has edition/configuration requirements.

Cypher 25 · transaction evidence snapshot
CYPHER 25
SHOW TRANSACTIONS
YIELD transactionId, username, database, currentQuery, currentQueryStatus,
      elapsedTime, waitTime, currentQueryElapsedTime, currentQueryWaitTime,
      activeLockCount, resourceInformation,
      currentQueryAllocatedBytes, allocatedDirectBytes, estimatedUsedHeapMemory,
      pageHits, pageFaults, currentQueryPageHits, currentQueryPageFaults,
      indexes, runtime, statusDetails
RETURN *
ORDER BY currentQueryElapsedTime DESC;
Observation Mechanism it suggests What it does not prove
high waitTime / waiting status lock contention or blocked resource which business workflow should be redesigned
high currentQueryAllocatedBytes query has allocated substantial heap that heap size globally is too small
pageFaults increasing requested pages were not resident in page cache that disk is the only latency bottleneck
index list present query is using reported static label/type indexes that the chosen index is optimal or selective
elapsed high, wait low work may be CPU/DB hits/result/network/client dominated root cause without plan/host/client correlation

4. Inspect configuration, not folklore

Before saying “heap is 8 GB” or “page cache is automatic,” ask the server. SHOW SETTINGS reports current, startup and default values on the executing self-managed server. It is not available on Aura, and dynamic changes are an Enterprise capability even though settings can be listed for inspection.

Cypher 25 · memory and transaction-log settings
CYPHER 25
SHOW SETTINGS
YIELD name, value, isDynamic, isExplicitlySet
WHERE name IN [
  'server.memory.heap.initial_size',
  'server.memory.heap.max_size',
  'server.memory.pagecache.size',
  'dbms.memory.transaction.total.max',
  'db.memory.transaction.total.max',
  'db.memory.transaction.max',
  'db.tx_log.rotation.size',
  'db.tx_log.rotation.retention_policy'
]
RETURN name, value, isDynamic, isExplicitlySet
ORDER BY name;

5. Logs answer different questions

Log Default / edition boundary Use it for
neo4j.log general log, enabled startup/shutdown and general user-facing server information
debug.log enabled; current default structured JSON format support-grade diagnostics, exceptions and internal operational context
gc.log disabled until configured JVM GC pauses, allocation/collection behavior; correlate with latency
query.log Enterprise executed/slow-query evidence with configured threshold/mode; not a Community requirement
security.log Enterprise authentication/admin/authorization security events; do not treat as a query-performance log
http.log disabled by default HTTP API request evidence when that interface matters
Do not enable everything blindly

Verbose query/security/HTTP logging and high-cardinality telemetry have storage, CPU and privacy costs. Enable the evidence required by an explicit operational question and define retention/redaction.

6. JVM/GC and operating-system evidence

GC pauses only explain latency if their timestamps align with the slow window. Likewise, high CPU is useful only when paired with workload and queueing evidence. Start with a cross-platform Docker snapshot and the server’s own memory recommendation, then add host-native counters when Neo4j runs directly.

Shell / PowerShell notes · host and container evidence
# Portable container-level snapshot (Docker Desktop / Linux / macOS):
docker stats --no-stream atlasmart-neo4j

# Neo4j store/index sizing guidance from inside the official image:
docker exec atlasmart-neo4j neo4j-admin server memory-recommendation

# Linux host examples when Neo4j runs directly (not Docker Desktop VM):
# ps -o pid,%cpu,%mem,rss,vsz,etime -p <neo4j_pid>
# ls /proc/<neo4j_pid>/fd | wc -l
# iostat -xz 1
# ss -s

# Windows host examples for a directly running Neo4j process:
# Get-Process neo4j | Select-Object Id,CPU,WorkingSet64,PrivateMemorySize64,HandleCount
# Get-Counter '\\PhysicalDisk(*)\\Avg. Disk sec/Read','\\Network Interface(*)\\Bytes Total/sec'

7. Page cache is not “Neo4j memory” as a whole

Page cache caches graph-store pages and native indexes; heap holds Java objects and much query/transaction state; native memory includes direct buffers and other off-heap structures. A cold page cache after startup naturally produces more faults and disk reads until the active working set warms. That is not a reason to clear caches in production to make a benchmark “fair.”

Wrong approach: clear caches to fix latency

Clearing page/OS caches destroys useful state, causes artificial I/O, and changes the workload. In a disposable lab, a restart can demonstrate cold-start behavior; in production, compare naturally occurring cold/warm windows and keep the service state intact.

8. Baseline worksheet

Timestamp/request App p99/error Driver acquire/fetch Tx status/wait PROFILE rows/DB hits Heap/GC/page CPU/I/O/net Conclusion
baseline A record record record record record record no conclusion yet
incident B record record record record record record one or more hypotheses
after one change record same load record record record record record accept/reject hypothesis

Check your understanding

  1. Why is “high CPU” not a root cause?
  2. What does a database hit measure?
  3. Why can SHOW TRANSACTIONS fields be null?
  4. Which monitoring features are explicitly Enterprise in this chapter?
  5. What is the first remediation rule?
Review the answers

1. It says a resource is busy, not which query, workload shape, GC activity, client pattern, or plan produced that work.

2. Logical graph/index access work reported by PROFILE, not a one-to-one physical disk read.

3. Some measurements are unavailable, not enabled, or no current query exists; null must not be read as zero.

4. DBMS metrics exporters, query.log, security.log and some tracking/privilege capabilities; the Community path uses other evidence.

5. Change one causal dimension at a time and repeat the same measurements under comparable load.

Production judgment

Decision surface Production questions
graph/workload fit Is latency dominated by graph fan-out, result cardinality, locks, indexes, client chatter, or infrastructure rather than raw store size?
correctness Can any proposed remediation change result semantics, transaction boundaries, retry behavior, or causal-read guarantees?
model/cardinality/degree Which labels/types are high-degree or skewed? Are a few hot nodes causing contention or explosive expansion?
latency What are p50/p95/p99 and timeout/error distributions at representative load, not one warm single-user query?
transactions/concurrency How long do transactions stay open? How much wait time and lock scope is normal, and which workflows create hot entities?
memory How are RAM budgets split among OS, heap, page cache, transaction/query memory, direct/native buffers, Lucene/vector needs, and GDS if present?
CPU/disk/network Which resource saturates first? Is disk latency/page-fault activity aligned with slowdown? Is client/network/serialization time dominant?
indexes/constraints Do plans use the intended access paths? What storage/write overhead and population/build headroom do indexes add?
driver Are pool acquisition, retry, timeout, fetch/result consumption and request correlation visible alongside server evidence?
security/tenant risk Do logs/metrics expose sensitive parameters or tenant data? Are monitoring privileges and endpoints restricted appropriately?
backup/recovery Is capacity monitoring independent from backup evidence, and is disk headroom sufficient for logs, dumps/backups and recovery operations?
observability What baseline defines alerts? Which signals are causal/proximate versus merely correlated? How long is telemetry retained?
testing/failure injection Can slow query, lock wait, memory pressure, cold start and network delay be reproduced in an isolated environment with reset steps?
version/edition/Aura Which fields/settings/logs/metrics exist on the exact server, Cypher version, edition and managed tier?
cost/migration What telemetry/storage overhead is acceptable, and how will capacity changes or schema/model remediations be rolled back?

Summary and next step

You now have the evidence map. Lesson 2 separates the memory and host-resource planes precisely so AtlasMart stops treating every memory symptom as a heap-sizing problem.

Authoritative references

  • Current Neo4j versions — Current Neo4j 2026.07.1 and 5.26.30 LTS release snapshot.
  • Monitoring overview — Current monitoring surfaces: logs, metrics, query/transaction management, connections, jobs, and reports.
  • Metrics — Enterprise metrics architecture and supported export surfaces.
  • Essential metrics — Server, Neo4j, cluster, and workload signals that require correlation rather than single-metric tuning.
  • Metrics reference — Current page-cache, JVM, CPU, file-descriptor, transaction and store metric names.
  • Logging — Current neo4j/debug/http/gc/query/security log behavior and edition boundaries.
  • Show and terminate transactions — Current SHOW TRANSACTIONS fields for locks, wait time, memory, page hits/faults, runtime, indexes, and status.
  • Manage queries — Current query visibility through transaction management rather than legacy dbms.listQueries().
  • Memory configuration — Heap, page cache, native memory, transaction limits, and memory-recommendation guidance.
  • Disks, RAM and other tips — Page-cache warmup, disk behavior, and operating-system considerations.
  • Performance — Performance topics spanning memory, disks, statistics, execution plans, Bolt threads, and filesystem tuning.
  • Show configuration settings — Current SHOW SETTINGS syntax and self-managed server scope.
  • Transaction logs — Transaction-log location, rotation, retention, pruning, and capacity implications.
  • Query tuning and PROFILE — Execution-plan evidence for rows, database hits, memory, access paths, and operator shape.
  • Python driver performance — Driver-side timing, database selection, result consumption, and application/database boundary considerations.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.