Chapter 18 · Monitoring and Troubleshooting: Metrics, Logs, Query Visibility, Memory, Store, and Capacity
Build a Troubleshooting Runbook from Symptom to Evidence to Hypothesis to Verified Remediation
Practice a repeatable symptom-to-evidence runbook on isolated AtlasMart slow-query, lock-contention, memory, and I/O cases, then verify each remediation with the same measurements that exposed the problem.
AtlasMart wants a troubleshooting document that an on-call engineer can use at 03:00 without guessing. A useful runbook is not a list of commands. It is a state machine: symptom → evidence → hypotheses → discriminating test → one remediation → verification → rollback/escalation. This lesson applies that sequence to slow query, lock wait, memory pressure and cold-cache/I/O symptoms.
Every remediation must have an acceptance test and a rollback/stop condition. If you cannot say what evidence would falsify your hypothesis, you are not troubleshooting—you are experimenting on production.
Learning outcomes
Use a consistent incident worksheet that ties application SLOs to database, JVM, host and driver evidence.
Run four isolated symptom drills without destructive production techniques or fabricated metrics.
Choose one remediation per hypothesis and verify it against the same before/after evidence.
Define alert thresholds from baselines/headroom rather than universal numbers copied from another system.
Close the lab with cleanup, invariant checks, regression artifacts and an escalation package.
Reproducible lab baseline
The continuity environment remains
Neo4j Community 2026.07.1, database
neo4j, container atlasmart-neo4j,
loopback Bolt bolt://127.0.0.1:7687, synthetic
local credential neo4j / atlasmart-course-2026,
Java 21 or 25, explicit CYPHER 25 where
version-sensitive syntax matters, and Python driver 6.3.
Current 5.26 LTS is 5.26.30. No runtime command
is claimed to have been executed during lesson generation;
expected outputs are fixture invariants or clearly labeled
examples.
The mandatory path stays free and local. Neo4j Enterprise
provides the DBMS metrics framework
(CSV/JMX/Prometheus/Graphite integrations), query.log,
security.log, query CPU tracking, broader RBAC-based
operational visibility, and additional cluster signals.
Community still supports useful first-party evidence including
SHOW TRANSACTIONS, SHOW SETTINGS on
self-managed deployments, PROFILE, general/debug
logs, optional JVM GC logging,
neo4j-admin server memory-recommendation, driver
timing, and operating-system/container counters. Do not
fabricate Enterprise-only metrics in the Community lab.
| Term | Operational meaning |
|---|---|
| symptom | What the user or SLO observes: high p99, timeout, error, stale response, blocked write, memory termination, disk pressure. |
| evidence | A timestamped observation that can support or contradict a hypothesis: plan, transaction row, log entry, OS counter, size snapshot, driver timing. |
| metric | A numeric time series such as CPU load, page faults, heap use, open file descriptors, or request latency; a metric is not a root cause by itself. |
| heap | JVM-managed memory for Java objects and much transaction/query state; garbage collection reclaims unused heap objects. |
| page cache | Neo4j-managed cache for graph-store pages and native indexes read from disk; it is distinct from JVM heap. |
| transaction/query memory | Estimated memory used for uncommitted state, intermediate rows, collections, sort/aggregation/path work, and results; bounded by transaction-memory settings. |
| native/off-heap memory | Memory allocated outside JVM heap, including direct/network buffers and other native structures. |
| page hit / page fault | A hit finds a requested store page in page cache; a fault requires loading a page. Neither ratio alone explains end-to-end latency. |
| lock wait | Time a transaction cannot proceed because another transaction owns a conflicting lock. |
| database hit | A PROFILE plan counter for work against the graph store/index abstractions; it is not a literal physical-disk-read count. |
| tail latency | High-percentile latency such as p95/p99; averages can remain healthy while a minority of requests violate the SLO. |
| capacity headroom | Resources intentionally left unused so spikes, recovery, maintenance, compaction/index builds, and failures do not immediately saturate the system. |
| Assumption | Pinned value / rule |
|---|---|
| server/current | Neo4j 2026.07.1 Community mandatory; 5.26.30 LTS only for compatibility comparison |
| database | neo4j; chapter fixture uses isolated Obs* labels and relationship types |
| driver | Python driver 6.3; one maintained Driver, separate Session/Transaction scopes |
| TLS/auth | loopback disposable lab may use bolt without TLS; remote production requires verified TLS and non-embedded secrets |
| plugins | APOC/GDS are not required |
| metrics | Enterprise metrics exporters are optional/reference-only; Community lab uses SHOW/PROFILE/logs/admin/OS/driver evidence |
| failure injection | only isolated, reversible slow-query/lock/memory/cold-cache exercises; no production cache clearing or destructive disk pressure |
| measurement | record p50/p95/p99, errors, queue/wait, memory, I/O/network and result size where the local tool exposes them; never invent timing values |
1. Runbook skeleton
| Stage | Required record | Gate before proceeding |
|---|---|---|
| symptom | request ID/time window/SLO impact/error class | confirm reproducibility/scope |
| evidence | SHOW TX, plan, logs, driver boundaries, OS/container, sizes | timestamps aligned |
| hypotheses | ranked mechanisms with supporting + contradicting evidence | at least one discriminating test |
| test | isolated change or observation | safe/reversible/bounded |
| remediation | one causal dimension changed | business semantics preserved |
| verification | same workload/evidence before vs after | acceptance criteria met |
| close/escalate | cleanup, artifacts, unresolved risks | no hidden temporary config or test data |
2. Drill A — broad query / cardinality symptom
Symptom: p99 rises for one analytical endpoint while lock wait remains low. Evidence: broad PROFILE has much larger intermediate rows/DB hits than returned rows. Hypothesis: query shape ignores explicit observed relationships. Remediation: use the relationship-shaped query, not a heap change.
CYPHER 25
PROFILE
MATCH (c:ObsCustomer), (i:ObsItem)
WHERE i.segment = c.segment
RETURN c.region AS region, count(*) AS pairs
ORDER BY pairs DESC;
CYPHER 25
PROFILE
MATCH (c:ObsCustomer)-[:OBS_VIEWED]->(i:ObsItem)
WHERE i.active
RETURN c.region AS region,
count(*) AS views,
round(avg(i.price), 2) AS avgPrice
ORDER BY views DESC;
| Verify | Acceptance |
|---|---|
| result correctness | same business definition; changed model/query does not silently change semantics |
| rows/DB hits | reduced at the expensive operator(s) |
| p95/p99 | improves under representative concurrency/parameters |
| memory/result bytes | no new payload/materialization regression |
3. Drill B — lock contention symptom
Symptom: writes queue even though CPU is not
saturated. Evidence: SHOW TRANSACTIONS shows
waiting state, rising wait time and lock dependency on
INV-HOT-18. Hypothesis: hot shared
entity serializes reservations.
Remediation options: shorten transaction,
partition invariant if business semantics permit, remodel to
per-SKU/event state, or reduce concurrent access. Do not simply
enlarge the pool.
from neo4j import GraphDatabase
from threading import Thread, Event
import time
URI = "bolt://127.0.0.1:7687"
AUTH = ("neo4j", "atlasmart-course-2026")
driver = GraphDatabase.driver(URI, auth=AUTH)
first_has_lock = Event()
def writer_a():
with driver.session(database="neo4j") as session:
tx = session.begin_transaction()
tx.run("MATCH (x:ObsInventory {obsId:$id}) "
"SET x.reserved = x.reserved + 1, x.version = x.version + 1",
id="INV-HOT-18").consume()
first_has_lock.set()
time.sleep(5) # intentionally hold the write lock in this disposable lab
tx.commit()
def writer_b():
first_has_lock.wait()
with driver.session(database="neo4j") as session:
tx = session.begin_transaction()
tx.run("MATCH (x:ObsInventory {obsId:$id}) "
"SET x.reserved = x.reserved + 1, x.version = x.version + 1",
id="INV-HOT-18").consume()
tx.commit()
a=Thread(target=writer_a); b=Thread(target=writer_b)
a.start(); b.start()
# While A sleeps, use a third cypher-shell/session with SHOW TRANSACTIONS.
a.join(); b.join(); driver.close()
CYPHER 25
SHOW TRANSACTIONS
YIELD transactionId, username, database, currentQuery, currentQueryStatus,
elapsedTime, waitTime, currentQueryElapsedTime, currentQueryWaitTime,
activeLockCount, resourceInformation,
currentQueryAllocatedBytes, allocatedDirectBytes, estimatedUsedHeapMemory,
pageHits, pageFaults, currentQueryPageHits, currentQueryPageFaults,
indexes, runtime, statusDetails
RETURN *
ORDER BY currentQueryElapsedTime DESC;
4. Drill C — query-memory symptom
Symptom: one analytical request allocates heavily or reaches transaction memory limits. Evidence: current query memory rises while page-fault evidence is not dominant. Hypothesis: unnecessary collection materialization. Remediation: stream aggregation or bound/chunk the workload before changing global transaction limits.
CYPHER 25
// Disposable lab only. Start smaller; increase gradually while observing transaction memory.
UNWIND range(1,50000) AS n
WITH collect({n:n, bucket:n % 100}) AS rows
RETURN size(rows) AS materializedRows,
reduce(s=0, x IN rows | s + x.n) AS checksum;
CYPHER 25
UNWIND range(1,50000) AS n
RETURN count(*) AS rowCount, sum(n) AS checksum;
5. Drill D — cold-start page-cache / I/O symptom
On the disposable course container only, a planned restart naturally makes the page cache cold. Run the same read workload after restart and again after warmup; record page-fault/block-I/O and latency evolution. Do not clear production OS/page caches to manufacture this state.
docker restart atlasmart-neo4j
# Wait for normal readiness using your existing course health check.
# Run the same bounded read benchmark three times and record:
# request p50/p95/p99, SHOW TRANSACTIONS page hits/faults while active,
# and docker stats Block I/O / CPU / memory snapshots.
A cold-start comparison can demonstrate cache warmup effects on this fixture. It does not establish a universal warmup duration or justify manually clearing caches in production.
6. Alert thresholds come from baselines and headroom
| Alert family | Bad threshold design | Evidence-based design |
|---|---|---|
| CPU | alert at one universal percent | baseline sustained saturation + queueing/p99 + workload context |
| page faults | alert on any fault | rate change from normal working set correlated with disk/p99 |
| heap/GC | alert on high used heap snapshot | GC pause/frequency/allocation trend + OOM/transaction terminations + headroom |
| tx wait | alert on one short wait | wait distribution by workflow plus p99/SLO impact |
| disk | alert only at 95% full | growth rate + backup/log/index-build headroom + time-to-exhaustion |
| driver pool | alert only when errors start | acquisition wait/queue trend before timeout saturation |
7. Escalation package
If the incident survives local diagnosis, preserve enough evidence for another engineer or vendor support to reproduce the reasoning without leaking secrets.
| Artifact | Include | Exclude/redact |
|---|---|---|
| incident timeline | absolute timestamps, request IDs, deploy/config changes | credentials/tokens |
| query evidence | parameterized query shape, EXPLAIN/PROFILE, representative parameter class | PII/raw secrets |
| transactions | status/wait/memory/page/index/runtime fields | sensitive query parameters if present |
| logs | relevant debug/GC/query/security excerpts by timestamp/edition | unnecessary user data |
| environment | Neo4j/Java/driver/plugin versions, edition, topology, memory settings | passwords/private keys |
| resource history | CPU/memory/I/O/network/store/log growth around incident | unrelated tenant data |
8. Cleanup and proof of reset
CYPHER 25
MATCH (n)
WHERE n:ObsCustomer OR n:ObsItem OR n:ObsInventory
DETACH DELETE n;
DROP CONSTRAINT obs_customer_id IF EXISTS;
DROP CONSTRAINT obs_item_id IF EXISTS;
DROP CONSTRAINT obs_inventory_id IF EXISTS;
CYPHER 25
MATCH (n)
WHERE n:ObsCustomer OR n:ObsItem OR n:ObsInventory
RETURN count(n) AS remainingObsNodes;
SHOW CONSTRAINTS
YIELD name
WHERE name STARTS WITH 'obs_'
RETURN name;
remainingObsNodes = 0 and no obs_* constraints remain. Chapter 18 does not change global logging/metrics/memory settings in the mandatory path, so there is no hidden configuration to revert.
9. Final runbook acceptance
Check your understanding
- What sequence should an on-call engineer follow?
- Why change one dimension at a time?
- Why are universal alert thresholds weak?
- What is a safe cold-cache test?
- What closes the chapter lab?
Review the answers
1. Symptom → synchronized evidence → ranked hypotheses → discriminating test → one remediation → same-evidence verification → cleanup/escalation.
2. Otherwise improvement/regression cannot be causally attributed and rollback becomes ambiguous.
3. Different workloads, hardware, graph shapes and tiers have different normal ranges; thresholds need baselines, SLOs and capacity headroom.
4. A planned restart of an isolated disposable lab followed by repeated bounded reads; never production cache clearing as a tuning ritual.
5. Fixture/constraint cleanup, invariant checks, before/after evidence retained, and no temporary monitoring/config changes left behind.
Production judgment
| Decision surface | Production questions |
|---|---|
| graph/workload fit | Is latency dominated by graph fan-out, result cardinality, locks, indexes, client chatter, or infrastructure rather than raw store size? |
| correctness | Can any proposed remediation change result semantics, transaction boundaries, retry behavior, or causal-read guarantees? |
| model/cardinality/degree | Which labels/types are high-degree or skewed? Are a few hot nodes causing contention or explosive expansion? |
| latency | What are p50/p95/p99 and timeout/error distributions at representative load, not one warm single-user query? |
| transactions/concurrency | How long do transactions stay open? How much wait time and lock scope is normal, and which workflows create hot entities? |
| memory | How are RAM budgets split among OS, heap, page cache, transaction/query memory, direct/native buffers, Lucene/vector needs, and GDS if present? |
| CPU/disk/network | Which resource saturates first? Is disk latency/page-fault activity aligned with slowdown? Is client/network/serialization time dominant? |
| indexes/constraints | Do plans use the intended access paths? What storage/write overhead and population/build headroom do indexes add? |
| driver | Are pool acquisition, retry, timeout, fetch/result consumption and request correlation visible alongside server evidence? |
| security/tenant risk | Do logs/metrics expose sensitive parameters or tenant data? Are monitoring privileges and endpoints restricted appropriately? |
| backup/recovery | Is capacity monitoring independent from backup evidence, and is disk headroom sufficient for logs, dumps/backups and recovery operations? |
| observability | What baseline defines alerts? Which signals are causal/proximate versus merely correlated? How long is telemetry retained? |
| testing/failure injection | Can slow query, lock wait, memory pressure, cold start and network delay be reproduced in an isolated environment with reset steps? |
| version/edition/Aura | Which fields/settings/logs/metrics exist on the exact server, Cypher version, edition and managed tier? |
| cost/migration | What telemetry/storage overhead is acceptable, and how will capacity changes or schema/model remediations be rolled back? |
Summary and next step
Chapter 18 turns observability into an operational method rather than a dashboard catalog. You can now separate query work, locks, memory, page cache, host resources, result/client time and capacity growth with evidence. Chapter 19 uses that discipline when Neo4j data is split across multiple and composite databases, where domain boundaries add new routing, security, resource and operational consequences.
Authoritative references
- Current Neo4j versions — Current Neo4j 2026.07.1 and 5.26.30 LTS release snapshot.
- Monitoring overview — Current monitoring surfaces: logs, metrics, query/transaction management, connections, jobs, and reports.
- Metrics — Enterprise metrics architecture and supported export surfaces.
- Essential metrics — Server, Neo4j, cluster, and workload signals that require correlation rather than single-metric tuning.
- Metrics reference — Current page-cache, JVM, CPU, file-descriptor, transaction and store metric names.
- Logging — Current neo4j/debug/http/gc/query/security log behavior and edition boundaries.
- Show and terminate transactions — Current SHOW TRANSACTIONS fields for locks, wait time, memory, page hits/faults, runtime, indexes, and status.
- Manage queries — Current query visibility through transaction management rather than legacy dbms.listQueries().
- Memory configuration — Heap, page cache, native memory, transaction limits, and memory-recommendation guidance.
- Disks, RAM and other tips — Page-cache warmup, disk behavior, and operating-system considerations.
- Performance — Performance topics spanning memory, disks, statistics, execution plans, Bolt threads, and filesystem tuning.
- Show configuration settings — Current SHOW SETTINGS syntax and self-managed server scope.
- Transaction logs — Transaction-log location, rotation, retention, pruning, and capacity implications.
- Query tuning and PROFILE — Execution-plan evidence for rows, database hits, memory, access paths, and operator shape.
- Python driver performance — Driver-side timing, database selection, result consumption, and application/database boundary considerations.