Chapter 18 · Monitoring and Troubleshooting: Metrics, Logs, Query Visibility, Memory, Store, and Capacity

Slow Query Capture, PROFILE Evidence, Lock/Transaction Diagnostics, and Workload Correlation

Capture a slow or blocked AtlasMart request without guesswork, correlate SHOW TRANSACTIONS with PROFILE and driver timing, and prove whether the root cause is access path, fan-out, lock wait, or an external boundary.

Advanced230–310 minutesSlow-query and lock-diagnostics labNeo4j 2026.07.1 Community mandatory · Enterprise observability optionalCypher 25 · SHOW TRANSACTIONS / PROFILE / SHOW SETTINGSJava 21/25 · Python driver 6.3Last reviewed: September 2026

An AtlasMart API call takes 4 seconds. The application log only says /recommendations timeout. One query engineer sees a scan in an old plan; another sees a lock wait from a different request. The useful investigation must correlate the same request window across driver timing, live transaction state, and a representative plan.

Correlation rule

A plan explains the work a query shape can perform; SHOW TRANSACTIONS explains what a running transaction is doing now; client timing explains what the user waited for. Use all three.

Learning outcomes

01

Capture slow-query evidence in Community without depending on Enterprise query.log.

02

Use PROFILE rows/DB hits/operators alongside SHOW TRANSACTIONS runtime/index/wait fields.

03

Reproduce a safe lock wait and interpret resourceInformation, lock counts and wait durations.

04

Differentiate access-path/fan-out problems from contention and client/network/result-consumption time.

05

Validate a remediation by repeating representative sparse/dense parameter and concurrency cases.

Reproducible lab baseline

Chapter 18 baseline · reviewed 9 September 2026

The continuity environment remains Neo4j Community 2026.07.1, database neo4j, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic local credential neo4j / atlasmart-course-2026, Java 21 or 25, explicit CYPHER 25 where version-sensitive syntax matters, and Python driver 6.3. Current 5.26 LTS is 5.26.30. No runtime command is claimed to have been executed during lesson generation; expected outputs are fixture invariants or clearly labeled examples.

Monitoring edition boundary

The mandatory path stays free and local. Neo4j Enterprise provides the DBMS metrics framework (CSV/JMX/Prometheus/Graphite integrations), query.log, security.log, query CPU tracking, broader RBAC-based operational visibility, and additional cluster signals. Community still supports useful first-party evidence including SHOW TRANSACTIONS, SHOW SETTINGS on self-managed deployments, PROFILE, general/debug logs, optional JVM GC logging, neo4j-admin server memory-recommendation, driver timing, and operating-system/container counters. Do not fabricate Enterprise-only metrics in the Community lab.

Term Operational meaning
symptom What the user or SLO observes: high p99, timeout, error, stale response, blocked write, memory termination, disk pressure.
evidence A timestamped observation that can support or contradict a hypothesis: plan, transaction row, log entry, OS counter, size snapshot, driver timing.
metric A numeric time series such as CPU load, page faults, heap use, open file descriptors, or request latency; a metric is not a root cause by itself.
heap JVM-managed memory for Java objects and much transaction/query state; garbage collection reclaims unused heap objects.
page cache Neo4j-managed cache for graph-store pages and native indexes read from disk; it is distinct from JVM heap.
transaction/query memory Estimated memory used for uncommitted state, intermediate rows, collections, sort/aggregation/path work, and results; bounded by transaction-memory settings.
native/off-heap memory Memory allocated outside JVM heap, including direct/network buffers and other native structures.
page hit / page fault A hit finds a requested store page in page cache; a fault requires loading a page. Neither ratio alone explains end-to-end latency.
lock wait Time a transaction cannot proceed because another transaction owns a conflicting lock.
database hit A PROFILE plan counter for work against the graph store/index abstractions; it is not a literal physical-disk-read count.
tail latency High-percentile latency such as p95/p99; averages can remain healthy while a minority of requests violate the SLO.
capacity headroom Resources intentionally left unused so spikes, recovery, maintenance, compaction/index builds, and failures do not immediately saturate the system.
Assumption Pinned value / rule
server/current Neo4j 2026.07.1 Community mandatory; 5.26.30 LTS only for compatibility comparison
database neo4j; chapter fixture uses isolated Obs* labels and relationship types
driver Python driver 6.3; one maintained Driver, separate Session/Transaction scopes
TLS/auth loopback disposable lab may use bolt without TLS; remote production requires verified TLS and non-embedded secrets
plugins APOC/GDS are not required
metrics Enterprise metrics exporters are optional/reference-only; Community lab uses SHOW/PROFILE/logs/admin/OS/driver evidence
failure injection only isolated, reversible slow-query/lock/memory/cold-cache exercises; no production cache clearing or destructive disk pressure
measurement record p50/p95/p99, errors, queue/wait, memory, I/O/network and result size where the local tool exposes them; never invent timing values

1. Community and Enterprise slow-query capture differ

Need Community path Enterprise addition
what is running now SHOW TRANSACTIONS same, with fine-grained operational privileges
slow history application/driver request logs + reproducible test + general/debug context query.log threshold/mode with query metadata
operator work EXPLAIN/PROFILE same; runtime-specific columns can differ
CPU time per query OS/process correlation; CPU fields may be unavailable db.track_query_cpu_time and related evidence where enabled
fleet time series external OS/container metrics you collect Neo4j metrics exporters/NOM/JMX/CSV
Do not call an old API current

Current operations documentation directs query visibility through SHOW TRANSACTIONS; do not teach legacy dbms.listQueries() as the primary current mechanism.

2. Baseline a deliberately broad query

The bad query forms a broad two-stream match and filters by a non-indexed shared segment. Its final aggregation looks compact, but the work before that aggregation can be much larger than the four returned region rows. This is the same anti-pattern Chapter 5 warned about: a small result is not proof of a small intermediate cardinality.

Cypher 25 · deliberately broad analytical query
CYPHER 25
PROFILE
MATCH (c:ObsCustomer), (i:ObsItem)
WHERE i.segment = c.segment
RETURN c.region AS region, count(*) AS pairs
ORDER BY pairs DESC;
Plan evidence Question
estimated vs actual rows Where did cardinality diverge from planner assumptions?
scan/seek operators Was there a selective starting point or broad label scan?
DB hits Which operator performed most graph/index accesses?
memory Did sort/aggregation/materialization allocate notable memory?
result rows Did a tiny final result hide large upstream work?

3. Change query/model dimension, not random settings

For the chapter fixture, the observed relationship already represents the relevant event. Traversing it avoids comparing every customer/item pair and asks the graph the question it was modeled to answer.

Cypher 25 · relationship-shaped query
CYPHER 25
PROFILE
MATCH (c:ObsCustomer)-[:OBS_VIEWED]->(i:ObsItem)
WHERE i.active
RETURN c.region AS region,
       count(*) AS views,
       round(avg(i.price), 2) AS avgPrice
ORDER BY views DESC;
Causal evidence

If rows/DB hits and request latency fall under the same representative load, the access/model/query shape—not a heap tweak—caused the improvement. Preserve both plans/results in the regression record.

4. Reproduce a lock wait safely

The following Python exercise opens separate sessions/transactions. Writer A updates the single hot inventory node and intentionally holds the transaction for five seconds; Writer B then requests the same write lock. While A sleeps, use a third session with SHOW TRANSACTIONS. This is isolated, short, and fully resettable.

Python 6.3 · controlled hot-node lock wait
from neo4j import GraphDatabase
from threading import Thread, Event
import time

URI = "bolt://127.0.0.1:7687"
AUTH = ("neo4j", "atlasmart-course-2026")
driver = GraphDatabase.driver(URI, auth=AUTH)
first_has_lock = Event()

def writer_a():
    with driver.session(database="neo4j") as session:
        tx = session.begin_transaction()
        tx.run("MATCH (x:ObsInventory {obsId:$id}) "
               "SET x.reserved = x.reserved + 1, x.version = x.version + 1",
               id="INV-HOT-18").consume()
        first_has_lock.set()
        time.sleep(5)  # intentionally hold the write lock in this disposable lab
        tx.commit()

def writer_b():
    first_has_lock.wait()
    with driver.session(database="neo4j") as session:
        tx = session.begin_transaction()
        tx.run("MATCH (x:ObsInventory {obsId:$id}) "
               "SET x.reserved = x.reserved + 1, x.version = x.version + 1",
               id="INV-HOT-18").consume()
        tx.commit()

a=Thread(target=writer_a); b=Thread(target=writer_b)
a.start(); b.start()
# While A sleeps, use a third cypher-shell/session with SHOW TRANSACTIONS.
a.join(); b.join(); driver.close()
Cypher 25 · observer session
CYPHER 25
SHOW TRANSACTIONS
YIELD transactionId, username, database, currentQuery, currentQueryStatus,
      elapsedTime, waitTime, currentQueryElapsedTime, currentQueryWaitTime,
      activeLockCount, resourceInformation,
      currentQueryAllocatedBytes, allocatedDirectBytes, estimatedUsedHeapMemory,
      pageHits, pageFaults, currentQueryPageHits, currentQueryPageFaults,
      indexes, runtime, statusDetails
RETURN *
ORDER BY currentQueryElapsedTime DESC;
Evidence during test Expected mechanism
one transaction has active lock(s) writer A owns the conflicting entity lock
writer B status waiting / wait time increasing B is blocked on lock acquisition
resourceInformation populated blocked/locking transaction details may identify dependency
CPU low while wait grows latency is contention, not CPU execution
both commits; reserved/version each +2 total serialization preserved the invariant despite wait

5. Do not “fix” contention with more connections

Increasing driver pool/concurrency can make a hot-node queue longer. If every reservation updates one global counter node, topology and connection count do not remove serialization. Repair options include narrower transaction duration, sharded/partitioned counters when semantics allow, relationship/event modeling, per-SKU inventory nodes, or an application workflow that reduces contention. The correct change depends on the invariant.

Wrong approach: hide wait behind retries

Retries can be correct for transient failures, but they do not increase the capacity of one serialized hot key. Measure wait, retry and p99 together.

6. Separate server time from fetch/result/client time

Neo4j can finish producing records while the application still spends time fetching, decoding, serializing JSON, compressing, or sending a huge payload. Record at least request start, connection acquisition, query submission, first record, result consumption/summary, serialization end and response end. If the database segment is 50 ms and JSON/network is 900 ms, an index will not fix the user-visible bottleneck.

Timestamp Boundary
t0 request accepted application queue/work admission
t1 session/connection acquired driver pool wait
t2 query sent network + server starts
t3 first record planning/execution + initial network
t4 result consumed/summary database + fetch
t5 payload serialized client CPU/memory
t6 response complete downstream network/client boundary

7. Parameter skew changes the incident

A customer with two relationships and a customer with 200,000 relationships are the same Cypher text but not the same workload shape. Preserve sparse, median, dense/high-degree, empty and worst-business-case parameters in performance tests. Correlate actual rows/DB hits/result size and transaction memory for each.

Case What to record
empty/no match startup/access-path overhead, zero-result behavior
typical degree normal latency and rows
high degree fan-out, memory, DB hits, result size, p99
concurrent hot key wait/locks/retries
large payload database vs fetch/serialize/network timing

8. Acceptance test

A remediation is accepted only if the business result remains correct and the same evidence improves under representative load. Keep the before/after plan, request-ID timing, error rate, p50/p95/p99, transaction wait/memory fields and OS snapshot together.

Check your understanding

  1. Why can a tiny RETURN result still be expensive?
  2. What proves lock contention?
  3. Why not add driver connections to fix a hot node?
  4. What does PROFILE contribute beyond elapsed time?
  5. How do you know latency is outside Neo4j?
Review the answers

1. Aggregation/filtering can collapse a much larger intermediate row set.

2. Waiting status/wait duration plus lock/resource dependency evidence correlated with concurrent writers.

3. More concurrent contenders can increase queueing around the same serialized invariant.

4. Operator shape, rows, DB hits, memory and access-path evidence that explain where work occurred.

5. Client boundary timestamps show database/result consumption is small while pool, serialization or downstream network dominates.

Production judgment

Decision surface Production questions
graph/workload fit Is latency dominated by graph fan-out, result cardinality, locks, indexes, client chatter, or infrastructure rather than raw store size?
correctness Can any proposed remediation change result semantics, transaction boundaries, retry behavior, or causal-read guarantees?
model/cardinality/degree Which labels/types are high-degree or skewed? Are a few hot nodes causing contention or explosive expansion?
latency What are p50/p95/p99 and timeout/error distributions at representative load, not one warm single-user query?
transactions/concurrency How long do transactions stay open? How much wait time and lock scope is normal, and which workflows create hot entities?
memory How are RAM budgets split among OS, heap, page cache, transaction/query memory, direct/native buffers, Lucene/vector needs, and GDS if present?
CPU/disk/network Which resource saturates first? Is disk latency/page-fault activity aligned with slowdown? Is client/network/serialization time dominant?
indexes/constraints Do plans use the intended access paths? What storage/write overhead and population/build headroom do indexes add?
driver Are pool acquisition, retry, timeout, fetch/result consumption and request correlation visible alongside server evidence?
security/tenant risk Do logs/metrics expose sensitive parameters or tenant data? Are monitoring privileges and endpoints restricted appropriately?
backup/recovery Is capacity monitoring independent from backup evidence, and is disk headroom sufficient for logs, dumps/backups and recovery operations?
observability What baseline defines alerts? Which signals are causal/proximate versus merely correlated? How long is telemetry retained?
testing/failure injection Can slow query, lock wait, memory pressure, cold start and network delay be reproduced in an isolated environment with reset steps?
version/edition/Aura Which fields/settings/logs/metrics exist on the exact server, Cypher version, edition and managed tier?
cost/migration What telemetry/storage overhead is acceptable, and how will capacity changes or schema/model remediations be rolled back?

Summary and next step

Slow-query and lock diagnosis now have observable mechanisms. Lesson 4 shifts from one incident to the longer horizon: store, index, degree and transaction-log growth and the capacity headroom required to survive it.

Authoritative references

  • Current Neo4j versions — Current Neo4j 2026.07.1 and 5.26.30 LTS release snapshot.
  • Monitoring overview — Current monitoring surfaces: logs, metrics, query/transaction management, connections, jobs, and reports.
  • Metrics — Enterprise metrics architecture and supported export surfaces.
  • Essential metrics — Server, Neo4j, cluster, and workload signals that require correlation rather than single-metric tuning.
  • Metrics reference — Current page-cache, JVM, CPU, file-descriptor, transaction and store metric names.
  • Logging — Current neo4j/debug/http/gc/query/security log behavior and edition boundaries.
  • Show and terminate transactions — Current SHOW TRANSACTIONS fields for locks, wait time, memory, page hits/faults, runtime, indexes, and status.
  • Manage queries — Current query visibility through transaction management rather than legacy dbms.listQueries().
  • Memory configuration — Heap, page cache, native memory, transaction limits, and memory-recommendation guidance.
  • Disks, RAM and other tips — Page-cache warmup, disk behavior, and operating-system considerations.
  • Performance — Performance topics spanning memory, disks, statistics, execution plans, Bolt threads, and filesystem tuning.
  • Show configuration settings — Current SHOW SETTINGS syntax and self-managed server scope.
  • Transaction logs — Transaction-log location, rotation, retention, pruning, and capacity implications.
  • Query tuning and PROFILE — Execution-plan evidence for rows, database hits, memory, access paths, and operator shape.
  • Python driver performance — Driver-side timing, database selection, result consumption, and application/database boundary considerations.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.