Chapter 18 · Monitoring and Troubleshooting: Metrics, Logs, Query Visibility, Memory, Store, and Capacity
Slow Query Capture, PROFILE Evidence, Lock/Transaction Diagnostics, and Workload Correlation
Capture a slow or blocked AtlasMart request without guesswork, correlate SHOW TRANSACTIONS with PROFILE and driver timing, and prove whether the root cause is access path, fan-out, lock wait, or an external boundary.
An AtlasMart API call takes 4 seconds. The application log only
says /recommendations timeout. One query engineer
sees a scan in an old plan; another sees a lock wait from a
different request. The useful investigation must correlate
the same request window across driver timing,
live transaction state, and a representative plan.
A plan explains the work a query shape can perform; SHOW TRANSACTIONS explains what a running transaction is doing now; client timing explains what the user waited for. Use all three.
Learning outcomes
Capture slow-query evidence in Community without depending on Enterprise query.log.
Use PROFILE rows/DB hits/operators alongside SHOW TRANSACTIONS runtime/index/wait fields.
Reproduce a safe lock wait and interpret resourceInformation, lock counts and wait durations.
Differentiate access-path/fan-out problems from contention and client/network/result-consumption time.
Validate a remediation by repeating representative sparse/dense parameter and concurrency cases.
Reproducible lab baseline
The continuity environment remains
Neo4j Community 2026.07.1, database
neo4j, container atlasmart-neo4j,
loopback Bolt bolt://127.0.0.1:7687, synthetic
local credential neo4j / atlasmart-course-2026,
Java 21 or 25, explicit CYPHER 25 where
version-sensitive syntax matters, and Python driver 6.3.
Current 5.26 LTS is 5.26.30. No runtime command
is claimed to have been executed during lesson generation;
expected outputs are fixture invariants or clearly labeled
examples.
The mandatory path stays free and local. Neo4j Enterprise
provides the DBMS metrics framework
(CSV/JMX/Prometheus/Graphite integrations), query.log,
security.log, query CPU tracking, broader RBAC-based
operational visibility, and additional cluster signals.
Community still supports useful first-party evidence including
SHOW TRANSACTIONS, SHOW SETTINGS on
self-managed deployments, PROFILE, general/debug
logs, optional JVM GC logging,
neo4j-admin server memory-recommendation, driver
timing, and operating-system/container counters. Do not
fabricate Enterprise-only metrics in the Community lab.
| Term | Operational meaning |
|---|---|
| symptom | What the user or SLO observes: high p99, timeout, error, stale response, blocked write, memory termination, disk pressure. |
| evidence | A timestamped observation that can support or contradict a hypothesis: plan, transaction row, log entry, OS counter, size snapshot, driver timing. |
| metric | A numeric time series such as CPU load, page faults, heap use, open file descriptors, or request latency; a metric is not a root cause by itself. |
| heap | JVM-managed memory for Java objects and much transaction/query state; garbage collection reclaims unused heap objects. |
| page cache | Neo4j-managed cache for graph-store pages and native indexes read from disk; it is distinct from JVM heap. |
| transaction/query memory | Estimated memory used for uncommitted state, intermediate rows, collections, sort/aggregation/path work, and results; bounded by transaction-memory settings. |
| native/off-heap memory | Memory allocated outside JVM heap, including direct/network buffers and other native structures. |
| page hit / page fault | A hit finds a requested store page in page cache; a fault requires loading a page. Neither ratio alone explains end-to-end latency. |
| lock wait | Time a transaction cannot proceed because another transaction owns a conflicting lock. |
| database hit | A PROFILE plan counter for work against the graph store/index abstractions; it is not a literal physical-disk-read count. |
| tail latency | High-percentile latency such as p95/p99; averages can remain healthy while a minority of requests violate the SLO. |
| capacity headroom | Resources intentionally left unused so spikes, recovery, maintenance, compaction/index builds, and failures do not immediately saturate the system. |
| Assumption | Pinned value / rule |
|---|---|
| server/current | Neo4j 2026.07.1 Community mandatory; 5.26.30 LTS only for compatibility comparison |
| database | neo4j; chapter fixture uses isolated Obs* labels and relationship types |
| driver | Python driver 6.3; one maintained Driver, separate Session/Transaction scopes |
| TLS/auth | loopback disposable lab may use bolt without TLS; remote production requires verified TLS and non-embedded secrets |
| plugins | APOC/GDS are not required |
| metrics | Enterprise metrics exporters are optional/reference-only; Community lab uses SHOW/PROFILE/logs/admin/OS/driver evidence |
| failure injection | only isolated, reversible slow-query/lock/memory/cold-cache exercises; no production cache clearing or destructive disk pressure |
| measurement | record p50/p95/p99, errors, queue/wait, memory, I/O/network and result size where the local tool exposes them; never invent timing values |
1. Community and Enterprise slow-query capture differ
| Need | Community path | Enterprise addition |
|---|---|---|
| what is running now | SHOW TRANSACTIONS | same, with fine-grained operational privileges |
| slow history | application/driver request logs + reproducible test + general/debug context | query.log threshold/mode with query metadata |
| operator work | EXPLAIN/PROFILE | same; runtime-specific columns can differ |
| CPU time per query | OS/process correlation; CPU fields may be unavailable | db.track_query_cpu_time and related evidence where enabled |
| fleet time series | external OS/container metrics you collect | Neo4j metrics exporters/NOM/JMX/CSV |
Current operations documentation directs query visibility
through SHOW TRANSACTIONS; do not teach legacy
dbms.listQueries() as the primary current
mechanism.
2. Baseline a deliberately broad query
The bad query forms a broad two-stream match and filters by a non-indexed shared segment. Its final aggregation looks compact, but the work before that aggregation can be much larger than the four returned region rows. This is the same anti-pattern Chapter 5 warned about: a small result is not proof of a small intermediate cardinality.
CYPHER 25
PROFILE
MATCH (c:ObsCustomer), (i:ObsItem)
WHERE i.segment = c.segment
RETURN c.region AS region, count(*) AS pairs
ORDER BY pairs DESC;
| Plan evidence | Question |
|---|---|
| estimated vs actual rows | Where did cardinality diverge from planner assumptions? |
| scan/seek operators | Was there a selective starting point or broad label scan? |
| DB hits | Which operator performed most graph/index accesses? |
| memory | Did sort/aggregation/materialization allocate notable memory? |
| result rows | Did a tiny final result hide large upstream work? |
3. Change query/model dimension, not random settings
For the chapter fixture, the observed relationship already represents the relevant event. Traversing it avoids comparing every customer/item pair and asks the graph the question it was modeled to answer.
CYPHER 25
PROFILE
MATCH (c:ObsCustomer)-[:OBS_VIEWED]->(i:ObsItem)
WHERE i.active
RETURN c.region AS region,
count(*) AS views,
round(avg(i.price), 2) AS avgPrice
ORDER BY views DESC;
If rows/DB hits and request latency fall under the same representative load, the access/model/query shape—not a heap tweak—caused the improvement. Preserve both plans/results in the regression record.
4. Reproduce a lock wait safely
The following Python exercise opens separate
sessions/transactions. Writer A updates the single hot inventory
node and intentionally holds the transaction for five seconds;
Writer B then requests the same write lock. While A sleeps, use
a third session with SHOW TRANSACTIONS. This is
isolated, short, and fully resettable.
from neo4j import GraphDatabase
from threading import Thread, Event
import time
URI = "bolt://127.0.0.1:7687"
AUTH = ("neo4j", "atlasmart-course-2026")
driver = GraphDatabase.driver(URI, auth=AUTH)
first_has_lock = Event()
def writer_a():
with driver.session(database="neo4j") as session:
tx = session.begin_transaction()
tx.run("MATCH (x:ObsInventory {obsId:$id}) "
"SET x.reserved = x.reserved + 1, x.version = x.version + 1",
id="INV-HOT-18").consume()
first_has_lock.set()
time.sleep(5) # intentionally hold the write lock in this disposable lab
tx.commit()
def writer_b():
first_has_lock.wait()
with driver.session(database="neo4j") as session:
tx = session.begin_transaction()
tx.run("MATCH (x:ObsInventory {obsId:$id}) "
"SET x.reserved = x.reserved + 1, x.version = x.version + 1",
id="INV-HOT-18").consume()
tx.commit()
a=Thread(target=writer_a); b=Thread(target=writer_b)
a.start(); b.start()
# While A sleeps, use a third cypher-shell/session with SHOW TRANSACTIONS.
a.join(); b.join(); driver.close()
CYPHER 25
SHOW TRANSACTIONS
YIELD transactionId, username, database, currentQuery, currentQueryStatus,
elapsedTime, waitTime, currentQueryElapsedTime, currentQueryWaitTime,
activeLockCount, resourceInformation,
currentQueryAllocatedBytes, allocatedDirectBytes, estimatedUsedHeapMemory,
pageHits, pageFaults, currentQueryPageHits, currentQueryPageFaults,
indexes, runtime, statusDetails
RETURN *
ORDER BY currentQueryElapsedTime DESC;
| Evidence during test | Expected mechanism |
|---|---|
| one transaction has active lock(s) | writer A owns the conflicting entity lock |
| writer B status waiting / wait time increasing | B is blocked on lock acquisition |
| resourceInformation populated | blocked/locking transaction details may identify dependency |
| CPU low while wait grows | latency is contention, not CPU execution |
| both commits; reserved/version each +2 total | serialization preserved the invariant despite wait |
5. Do not “fix” contention with more connections
Increasing driver pool/concurrency can make a hot-node queue longer. If every reservation updates one global counter node, topology and connection count do not remove serialization. Repair options include narrower transaction duration, sharded/partitioned counters when semantics allow, relationship/event modeling, per-SKU inventory nodes, or an application workflow that reduces contention. The correct change depends on the invariant.
Retries can be correct for transient failures, but they do not increase the capacity of one serialized hot key. Measure wait, retry and p99 together.
6. Separate server time from fetch/result/client time
Neo4j can finish producing records while the application still spends time fetching, decoding, serializing JSON, compressing, or sending a huge payload. Record at least request start, connection acquisition, query submission, first record, result consumption/summary, serialization end and response end. If the database segment is 50 ms and JSON/network is 900 ms, an index will not fix the user-visible bottleneck.
| Timestamp | Boundary |
|---|---|
| t0 request accepted | application queue/work admission |
| t1 session/connection acquired | driver pool wait |
| t2 query sent | network + server starts |
| t3 first record | planning/execution + initial network |
| t4 result consumed/summary | database + fetch |
| t5 payload serialized | client CPU/memory |
| t6 response complete | downstream network/client boundary |
7. Parameter skew changes the incident
A customer with two relationships and a customer with 200,000 relationships are the same Cypher text but not the same workload shape. Preserve sparse, median, dense/high-degree, empty and worst-business-case parameters in performance tests. Correlate actual rows/DB hits/result size and transaction memory for each.
| Case | What to record |
|---|---|
| empty/no match | startup/access-path overhead, zero-result behavior |
| typical degree | normal latency and rows |
| high degree | fan-out, memory, DB hits, result size, p99 |
| concurrent hot key | wait/locks/retries |
| large payload | database vs fetch/serialize/network timing |
8. Acceptance test
A remediation is accepted only if the business result remains correct and the same evidence improves under representative load. Keep the before/after plan, request-ID timing, error rate, p50/p95/p99, transaction wait/memory fields and OS snapshot together.
Check your understanding
- Why can a tiny RETURN result still be expensive?
- What proves lock contention?
- Why not add driver connections to fix a hot node?
- What does PROFILE contribute beyond elapsed time?
- How do you know latency is outside Neo4j?
Review the answers
1. Aggregation/filtering can collapse a much larger intermediate row set.
2. Waiting status/wait duration plus lock/resource dependency evidence correlated with concurrent writers.
3. More concurrent contenders can increase queueing around the same serialized invariant.
4. Operator shape, rows, DB hits, memory and access-path evidence that explain where work occurred.
5. Client boundary timestamps show database/result consumption is small while pool, serialization or downstream network dominates.
Production judgment
| Decision surface | Production questions |
|---|---|
| graph/workload fit | Is latency dominated by graph fan-out, result cardinality, locks, indexes, client chatter, or infrastructure rather than raw store size? |
| correctness | Can any proposed remediation change result semantics, transaction boundaries, retry behavior, or causal-read guarantees? |
| model/cardinality/degree | Which labels/types are high-degree or skewed? Are a few hot nodes causing contention or explosive expansion? |
| latency | What are p50/p95/p99 and timeout/error distributions at representative load, not one warm single-user query? |
| transactions/concurrency | How long do transactions stay open? How much wait time and lock scope is normal, and which workflows create hot entities? |
| memory | How are RAM budgets split among OS, heap, page cache, transaction/query memory, direct/native buffers, Lucene/vector needs, and GDS if present? |
| CPU/disk/network | Which resource saturates first? Is disk latency/page-fault activity aligned with slowdown? Is client/network/serialization time dominant? |
| indexes/constraints | Do plans use the intended access paths? What storage/write overhead and population/build headroom do indexes add? |
| driver | Are pool acquisition, retry, timeout, fetch/result consumption and request correlation visible alongside server evidence? |
| security/tenant risk | Do logs/metrics expose sensitive parameters or tenant data? Are monitoring privileges and endpoints restricted appropriately? |
| backup/recovery | Is capacity monitoring independent from backup evidence, and is disk headroom sufficient for logs, dumps/backups and recovery operations? |
| observability | What baseline defines alerts? Which signals are causal/proximate versus merely correlated? How long is telemetry retained? |
| testing/failure injection | Can slow query, lock wait, memory pressure, cold start and network delay be reproduced in an isolated environment with reset steps? |
| version/edition/Aura | Which fields/settings/logs/metrics exist on the exact server, Cypher version, edition and managed tier? |
| cost/migration | What telemetry/storage overhead is acceptable, and how will capacity changes or schema/model remediations be rolled back? |
Summary and next step
Slow-query and lock diagnosis now have observable mechanisms. Lesson 4 shifts from one incident to the longer horizon: store, index, degree and transaction-log growth and the capacity headroom required to survive it.
Authoritative references
- Current Neo4j versions — Current Neo4j 2026.07.1 and 5.26.30 LTS release snapshot.
- Monitoring overview — Current monitoring surfaces: logs, metrics, query/transaction management, connections, jobs, and reports.
- Metrics — Enterprise metrics architecture and supported export surfaces.
- Essential metrics — Server, Neo4j, cluster, and workload signals that require correlation rather than single-metric tuning.
- Metrics reference — Current page-cache, JVM, CPU, file-descriptor, transaction and store metric names.
- Logging — Current neo4j/debug/http/gc/query/security log behavior and edition boundaries.
- Show and terminate transactions — Current SHOW TRANSACTIONS fields for locks, wait time, memory, page hits/faults, runtime, indexes, and status.
- Manage queries — Current query visibility through transaction management rather than legacy dbms.listQueries().
- Memory configuration — Heap, page cache, native memory, transaction limits, and memory-recommendation guidance.
- Disks, RAM and other tips — Page-cache warmup, disk behavior, and operating-system considerations.
- Performance — Performance topics spanning memory, disks, statistics, execution plans, Bolt threads, and filesystem tuning.
- Show configuration settings — Current SHOW SETTINGS syntax and self-managed server scope.
- Transaction logs — Transaction-log location, rotation, retention, pruning, and capacity implications.
- Query tuning and PROFILE — Execution-plan evidence for rows, database hits, memory, access paths, and operator shape.
- Python driver performance — Driver-side timing, database selection, result consumption, and application/database boundary considerations.