Chapter 18 · Monitoring and Troubleshooting: Metrics, Logs, Query Visibility, Memory, Store, and Capacity

Build a Troubleshooting Runbook from Symptom to Evidence to Hypothesis to Verified Remediation

Practice a repeatable symptom-to-evidence runbook on isolated AtlasMart slow-query, lock-contention, memory, and I/O cases, then verify each remediation with the same measurements that exposed the problem.

Advanced250–340 minutesTroubleshooting runbook labNeo4j 2026.07.1 Community mandatory · Enterprise observability optionalCypher 25 · SHOW TRANSACTIONS / PROFILE / SHOW SETTINGSJava 21/25 · Python driver 6.3Last reviewed: September 2026

AtlasMart wants a troubleshooting document that an on-call engineer can use at 03:00 without guessing. A useful runbook is not a list of commands. It is a state machine: symptom → evidence → hypotheses → discriminating test → one remediation → verification → rollback/escalation. This lesson applies that sequence to slow query, lock wait, memory pressure and cold-cache/I/O symptoms.

Operational rule

Every remediation must have an acceptance test and a rollback/stop condition. If you cannot say what evidence would falsify your hypothesis, you are not troubleshooting—you are experimenting on production.

Learning outcomes

01

Use a consistent incident worksheet that ties application SLOs to database, JVM, host and driver evidence.

02

Run four isolated symptom drills without destructive production techniques or fabricated metrics.

03

Choose one remediation per hypothesis and verify it against the same before/after evidence.

04

Define alert thresholds from baselines/headroom rather than universal numbers copied from another system.

05

Close the lab with cleanup, invariant checks, regression artifacts and an escalation package.

Reproducible lab baseline

Chapter 18 baseline · reviewed 9 September 2026

The continuity environment remains Neo4j Community 2026.07.1, database neo4j, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic local credential neo4j / atlasmart-course-2026, Java 21 or 25, explicit CYPHER 25 where version-sensitive syntax matters, and Python driver 6.3. Current 5.26 LTS is 5.26.30. No runtime command is claimed to have been executed during lesson generation; expected outputs are fixture invariants or clearly labeled examples.

Monitoring edition boundary

The mandatory path stays free and local. Neo4j Enterprise provides the DBMS metrics framework (CSV/JMX/Prometheus/Graphite integrations), query.log, security.log, query CPU tracking, broader RBAC-based operational visibility, and additional cluster signals. Community still supports useful first-party evidence including SHOW TRANSACTIONS, SHOW SETTINGS on self-managed deployments, PROFILE, general/debug logs, optional JVM GC logging, neo4j-admin server memory-recommendation, driver timing, and operating-system/container counters. Do not fabricate Enterprise-only metrics in the Community lab.

Term Operational meaning
symptom What the user or SLO observes: high p99, timeout, error, stale response, blocked write, memory termination, disk pressure.
evidence A timestamped observation that can support or contradict a hypothesis: plan, transaction row, log entry, OS counter, size snapshot, driver timing.
metric A numeric time series such as CPU load, page faults, heap use, open file descriptors, or request latency; a metric is not a root cause by itself.
heap JVM-managed memory for Java objects and much transaction/query state; garbage collection reclaims unused heap objects.
page cache Neo4j-managed cache for graph-store pages and native indexes read from disk; it is distinct from JVM heap.
transaction/query memory Estimated memory used for uncommitted state, intermediate rows, collections, sort/aggregation/path work, and results; bounded by transaction-memory settings.
native/off-heap memory Memory allocated outside JVM heap, including direct/network buffers and other native structures.
page hit / page fault A hit finds a requested store page in page cache; a fault requires loading a page. Neither ratio alone explains end-to-end latency.
lock wait Time a transaction cannot proceed because another transaction owns a conflicting lock.
database hit A PROFILE plan counter for work against the graph store/index abstractions; it is not a literal physical-disk-read count.
tail latency High-percentile latency such as p95/p99; averages can remain healthy while a minority of requests violate the SLO.
capacity headroom Resources intentionally left unused so spikes, recovery, maintenance, compaction/index builds, and failures do not immediately saturate the system.
Assumption Pinned value / rule
server/current Neo4j 2026.07.1 Community mandatory; 5.26.30 LTS only for compatibility comparison
database neo4j; chapter fixture uses isolated Obs* labels and relationship types
driver Python driver 6.3; one maintained Driver, separate Session/Transaction scopes
TLS/auth loopback disposable lab may use bolt without TLS; remote production requires verified TLS and non-embedded secrets
plugins APOC/GDS are not required
metrics Enterprise metrics exporters are optional/reference-only; Community lab uses SHOW/PROFILE/logs/admin/OS/driver evidence
failure injection only isolated, reversible slow-query/lock/memory/cold-cache exercises; no production cache clearing or destructive disk pressure
measurement record p50/p95/p99, errors, queue/wait, memory, I/O/network and result size where the local tool exposes them; never invent timing values

1. Runbook skeleton

Stage Required record Gate before proceeding
symptom request ID/time window/SLO impact/error class confirm reproducibility/scope
evidence SHOW TX, plan, logs, driver boundaries, OS/container, sizes timestamps aligned
hypotheses ranked mechanisms with supporting + contradicting evidence at least one discriminating test
test isolated change or observation safe/reversible/bounded
remediation one causal dimension changed business semantics preserved
verification same workload/evidence before vs after acceptance criteria met
close/escalate cleanup, artifacts, unresolved risks no hidden temporary config or test data

2. Drill A — broad query / cardinality symptom

Symptom: p99 rises for one analytical endpoint while lock wait remains low. Evidence: broad PROFILE has much larger intermediate rows/DB hits than returned rows. Hypothesis: query shape ignores explicit observed relationships. Remediation: use the relationship-shaped query, not a heap change.

Baseline
CYPHER 25
PROFILE
MATCH (c:ObsCustomer), (i:ObsItem)
WHERE i.segment = c.segment
RETURN c.region AS region, count(*) AS pairs
ORDER BY pairs DESC;
Candidate remediation
CYPHER 25
PROFILE
MATCH (c:ObsCustomer)-[:OBS_VIEWED]->(i:ObsItem)
WHERE i.active
RETURN c.region AS region,
       count(*) AS views,
       round(avg(i.price), 2) AS avgPrice
ORDER BY views DESC;
Verify Acceptance
result correctness same business definition; changed model/query does not silently change semantics
rows/DB hits reduced at the expensive operator(s)
p95/p99 improves under representative concurrency/parameters
memory/result bytes no new payload/materialization regression

3. Drill B — lock contention symptom

Symptom: writes queue even though CPU is not saturated. Evidence: SHOW TRANSACTIONS shows waiting state, rising wait time and lock dependency on INV-HOT-18. Hypothesis: hot shared entity serializes reservations. Remediation options: shorten transaction, partition invariant if business semantics permit, remodel to per-SKU/event state, or reduce concurrent access. Do not simply enlarge the pool.

Controlled lock producer
from neo4j import GraphDatabase
from threading import Thread, Event
import time

URI = "bolt://127.0.0.1:7687"
AUTH = ("neo4j", "atlasmart-course-2026")
driver = GraphDatabase.driver(URI, auth=AUTH)
first_has_lock = Event()

def writer_a():
    with driver.session(database="neo4j") as session:
        tx = session.begin_transaction()
        tx.run("MATCH (x:ObsInventory {obsId:$id}) "
               "SET x.reserved = x.reserved + 1, x.version = x.version + 1",
               id="INV-HOT-18").consume()
        first_has_lock.set()
        time.sleep(5)  # intentionally hold the write lock in this disposable lab
        tx.commit()

def writer_b():
    first_has_lock.wait()
    with driver.session(database="neo4j") as session:
        tx = session.begin_transaction()
        tx.run("MATCH (x:ObsInventory {obsId:$id}) "
               "SET x.reserved = x.reserved + 1, x.version = x.version + 1",
               id="INV-HOT-18").consume()
        tx.commit()

a=Thread(target=writer_a); b=Thread(target=writer_b)
a.start(); b.start()
# While A sleeps, use a third cypher-shell/session with SHOW TRANSACTIONS.
a.join(); b.join(); driver.close()
Observer
CYPHER 25
SHOW TRANSACTIONS
YIELD transactionId, username, database, currentQuery, currentQueryStatus,
      elapsedTime, waitTime, currentQueryElapsedTime, currentQueryWaitTime,
      activeLockCount, resourceInformation,
      currentQueryAllocatedBytes, allocatedDirectBytes, estimatedUsedHeapMemory,
      pageHits, pageFaults, currentQueryPageHits, currentQueryPageFaults,
      indexes, runtime, statusDetails
RETURN *
ORDER BY currentQueryElapsedTime DESC;

4. Drill C — query-memory symptom

Symptom: one analytical request allocates heavily or reaches transaction memory limits. Evidence: current query memory rises while page-fault evidence is not dominant. Hypothesis: unnecessary collection materialization. Remediation: stream aggregation or bound/chunk the workload before changing global transaction limits.

Bad teaching shape
CYPHER 25
// Disposable lab only. Start smaller; increase gradually while observing transaction memory.
UNWIND range(1,50000) AS n
WITH collect({n:n, bucket:n % 100}) AS rows
RETURN size(rows) AS materializedRows,
       reduce(s=0, x IN rows | s + x.n) AS checksum;
Safer streaming shape
CYPHER 25
UNWIND range(1,50000) AS n
RETURN count(*) AS rowCount, sum(n) AS checksum;

5. Drill D — cold-start page-cache / I/O symptom

On the disposable course container only, a planned restart naturally makes the page cache cold. Run the same read workload after restart and again after warmup; record page-fault/block-I/O and latency evolution. Do not clear production OS/page caches to manufacture this state.

Disposable lab only · planned cold-start observation
docker restart atlasmart-neo4j
# Wait for normal readiness using your existing course health check.
# Run the same bounded read benchmark three times and record:
# request p50/p95/p99, SHOW TRANSACTIONS page hits/faults while active,
# and docker stats Block I/O / CPU / memory snapshots.
What this proves

A cold-start comparison can demonstrate cache warmup effects on this fixture. It does not establish a universal warmup duration or justify manually clearing caches in production.

6. Alert thresholds come from baselines and headroom

Alert family Bad threshold design Evidence-based design
CPU alert at one universal percent baseline sustained saturation + queueing/p99 + workload context
page faults alert on any fault rate change from normal working set correlated with disk/p99
heap/GC alert on high used heap snapshot GC pause/frequency/allocation trend + OOM/transaction terminations + headroom
tx wait alert on one short wait wait distribution by workflow plus p99/SLO impact
disk alert only at 95% full growth rate + backup/log/index-build headroom + time-to-exhaustion
driver pool alert only when errors start acquisition wait/queue trend before timeout saturation

7. Escalation package

If the incident survives local diagnosis, preserve enough evidence for another engineer or vendor support to reproduce the reasoning without leaking secrets.

Artifact Include Exclude/redact
incident timeline absolute timestamps, request IDs, deploy/config changes credentials/tokens
query evidence parameterized query shape, EXPLAIN/PROFILE, representative parameter class PII/raw secrets
transactions status/wait/memory/page/index/runtime fields sensitive query parameters if present
logs relevant debug/GC/query/security excerpts by timestamp/edition unnecessary user data
environment Neo4j/Java/driver/plugin versions, edition, topology, memory settings passwords/private keys
resource history CPU/memory/I/O/network/store/log growth around incident unrelated tenant data

8. Cleanup and proof of reset

Cypher 25 · remove Chapter 18 fixture
CYPHER 25
MATCH (n)
WHERE n:ObsCustomer OR n:ObsItem OR n:ObsInventory
DETACH DELETE n;
DROP CONSTRAINT obs_customer_id IF EXISTS;
DROP CONSTRAINT obs_item_id IF EXISTS;
DROP CONSTRAINT obs_inventory_id IF EXISTS;
Cypher 25 · verify cleanup
CYPHER 25
MATCH (n)
WHERE n:ObsCustomer OR n:ObsItem OR n:ObsInventory
RETURN count(n) AS remainingObsNodes;
SHOW CONSTRAINTS
YIELD name
WHERE name STARTS WITH 'obs_'
RETURN name;
Expected reset

remainingObsNodes = 0 and no obs_* constraints remain. Chapter 18 does not change global logging/metrics/memory settings in the mandatory path, so there is no hidden configuration to revert.

9. Final runbook acceptance

Check your understanding

  1. What sequence should an on-call engineer follow?
  2. Why change one dimension at a time?
  3. Why are universal alert thresholds weak?
  4. What is a safe cold-cache test?
  5. What closes the chapter lab?
Review the answers

1. Symptom → synchronized evidence → ranked hypotheses → discriminating test → one remediation → same-evidence verification → cleanup/escalation.

2. Otherwise improvement/regression cannot be causally attributed and rollback becomes ambiguous.

3. Different workloads, hardware, graph shapes and tiers have different normal ranges; thresholds need baselines, SLOs and capacity headroom.

4. A planned restart of an isolated disposable lab followed by repeated bounded reads; never production cache clearing as a tuning ritual.

5. Fixture/constraint cleanup, invariant checks, before/after evidence retained, and no temporary monitoring/config changes left behind.

Production judgment

Decision surface Production questions
graph/workload fit Is latency dominated by graph fan-out, result cardinality, locks, indexes, client chatter, or infrastructure rather than raw store size?
correctness Can any proposed remediation change result semantics, transaction boundaries, retry behavior, or causal-read guarantees?
model/cardinality/degree Which labels/types are high-degree or skewed? Are a few hot nodes causing contention or explosive expansion?
latency What are p50/p95/p99 and timeout/error distributions at representative load, not one warm single-user query?
transactions/concurrency How long do transactions stay open? How much wait time and lock scope is normal, and which workflows create hot entities?
memory How are RAM budgets split among OS, heap, page cache, transaction/query memory, direct/native buffers, Lucene/vector needs, and GDS if present?
CPU/disk/network Which resource saturates first? Is disk latency/page-fault activity aligned with slowdown? Is client/network/serialization time dominant?
indexes/constraints Do plans use the intended access paths? What storage/write overhead and population/build headroom do indexes add?
driver Are pool acquisition, retry, timeout, fetch/result consumption and request correlation visible alongside server evidence?
security/tenant risk Do logs/metrics expose sensitive parameters or tenant data? Are monitoring privileges and endpoints restricted appropriately?
backup/recovery Is capacity monitoring independent from backup evidence, and is disk headroom sufficient for logs, dumps/backups and recovery operations?
observability What baseline defines alerts? Which signals are causal/proximate versus merely correlated? How long is telemetry retained?
testing/failure injection Can slow query, lock wait, memory pressure, cold start and network delay be reproduced in an isolated environment with reset steps?
version/edition/Aura Which fields/settings/logs/metrics exist on the exact server, Cypher version, edition and managed tier?
cost/migration What telemetry/storage overhead is acceptable, and how will capacity changes or schema/model remediations be rolled back?

Summary and next step

Chapter 18 turns observability into an operational method rather than a dashboard catalog. You can now separate query work, locks, memory, page cache, host resources, result/client time and capacity growth with evidence. Chapter 19 uses that discipline when Neo4j data is split across multiple and composite databases, where domain boundaries add new routing, security, resource and operational consequences.

Authoritative references

  • Current Neo4j versions — Current Neo4j 2026.07.1 and 5.26.30 LTS release snapshot.
  • Monitoring overview — Current monitoring surfaces: logs, metrics, query/transaction management, connections, jobs, and reports.
  • Metrics — Enterprise metrics architecture and supported export surfaces.
  • Essential metrics — Server, Neo4j, cluster, and workload signals that require correlation rather than single-metric tuning.
  • Metrics reference — Current page-cache, JVM, CPU, file-descriptor, transaction and store metric names.
  • Logging — Current neo4j/debug/http/gc/query/security log behavior and edition boundaries.
  • Show and terminate transactions — Current SHOW TRANSACTIONS fields for locks, wait time, memory, page hits/faults, runtime, indexes, and status.
  • Manage queries — Current query visibility through transaction management rather than legacy dbms.listQueries().
  • Memory configuration — Heap, page cache, native memory, transaction limits, and memory-recommendation guidance.
  • Disks, RAM and other tips — Page-cache warmup, disk behavior, and operating-system considerations.
  • Performance — Performance topics spanning memory, disks, statistics, execution plans, Bolt threads, and filesystem tuning.
  • Show configuration settings — Current SHOW SETTINGS syntax and self-managed server scope.
  • Transaction logs — Transaction-log location, rotation, retention, pruning, and capacity implications.
  • Query tuning and PROFILE — Execution-plan evidence for rows, database hits, memory, access paths, and operator shape.
  • Python driver performance — Driver-side timing, database selection, result consumption, and application/database boundary considerations.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.