Chapter 18 · Monitoring and Troubleshooting: Metrics, Logs, Query Visibility, Memory, Store, and Capacity
Store Growth, Relationship Density, Index Size, Transaction Log Growth, and Capacity Forecasting
Turn store, index, relationship-degree, and transaction-log measurements into capacity evidence and forecasts instead of extrapolating from node counts or one snapshot.
AtlasMart’s graph is healthy today, but data volume grows 4% each week, one category hub’s degree doubles every campaign, and transaction logs occasionally consume unexpected disk. Capacity planning cannot be “count nodes and buy 2× RAM.” Graph shape, indexes, log retention, write rate, working set, backups and failure/maintenance headroom all consume different resources.
Forecast from measured deltas of the resources that actually saturate—store, indexes, tx logs, page working set, IOPS, CPU, network, memory and recovery headroom—not from one vanity count.
Learning outcomes
Measure data/native-index and Lucene/vector footprint separately enough to reason about memory and disk budgets.
Measure relationship degree/skew rather than using average node count as a traversal-cost proxy.
Explain transaction-log rotation/retention and why manual deletion is unsupported.
Create timestamped growth snapshots and a cautious forecast with explicit uncertainty/headroom.
Connect capacity decisions to backups, index builds, cluster recovery, maintenance and failure headroom.
Reproducible lab baseline
The continuity environment remains
Neo4j Community 2026.07.1, database
neo4j, container atlasmart-neo4j,
loopback Bolt bolt://127.0.0.1:7687, synthetic
local credential neo4j / atlasmart-course-2026,
Java 21 or 25, explicit CYPHER 25 where
version-sensitive syntax matters, and Python driver 6.3.
Current 5.26 LTS is 5.26.30. No runtime command
is claimed to have been executed during lesson generation;
expected outputs are fixture invariants or clearly labeled
examples.
The mandatory path stays free and local. Neo4j Enterprise
provides the DBMS metrics framework
(CSV/JMX/Prometheus/Graphite integrations), query.log,
security.log, query CPU tracking, broader RBAC-based
operational visibility, and additional cluster signals.
Community still supports useful first-party evidence including
SHOW TRANSACTIONS, SHOW SETTINGS on
self-managed deployments, PROFILE, general/debug
logs, optional JVM GC logging,
neo4j-admin server memory-recommendation, driver
timing, and operating-system/container counters. Do not
fabricate Enterprise-only metrics in the Community lab.
| Term | Operational meaning |
|---|---|
| symptom | What the user or SLO observes: high p99, timeout, error, stale response, blocked write, memory termination, disk pressure. |
| evidence | A timestamped observation that can support or contradict a hypothesis: plan, transaction row, log entry, OS counter, size snapshot, driver timing. |
| metric | A numeric time series such as CPU load, page faults, heap use, open file descriptors, or request latency; a metric is not a root cause by itself. |
| heap | JVM-managed memory for Java objects and much transaction/query state; garbage collection reclaims unused heap objects. |
| page cache | Neo4j-managed cache for graph-store pages and native indexes read from disk; it is distinct from JVM heap. |
| transaction/query memory | Estimated memory used for uncommitted state, intermediate rows, collections, sort/aggregation/path work, and results; bounded by transaction-memory settings. |
| native/off-heap memory | Memory allocated outside JVM heap, including direct/network buffers and other native structures. |
| page hit / page fault | A hit finds a requested store page in page cache; a fault requires loading a page. Neither ratio alone explains end-to-end latency. |
| lock wait | Time a transaction cannot proceed because another transaction owns a conflicting lock. |
| database hit | A PROFILE plan counter for work against the graph store/index abstractions; it is not a literal physical-disk-read count. |
| tail latency | High-percentile latency such as p95/p99; averages can remain healthy while a minority of requests violate the SLO. |
| capacity headroom | Resources intentionally left unused so spikes, recovery, maintenance, compaction/index builds, and failures do not immediately saturate the system. |
| Assumption | Pinned value / rule |
|---|---|
| server/current | Neo4j 2026.07.1 Community mandatory; 5.26.30 LTS only for compatibility comparison |
| database | neo4j; chapter fixture uses isolated Obs* labels and relationship types |
| driver | Python driver 6.3; one maintained Driver, separate Session/Transaction scopes |
| TLS/auth | loopback disposable lab may use bolt without TLS; remote production requires verified TLS and non-embedded secrets |
| plugins | APOC/GDS are not required |
| metrics | Enterprise metrics exporters are optional/reference-only; Community lab uses SHOW/PROFILE/logs/admin/OS/driver evidence |
| failure injection | only isolated, reversible slow-query/lock/memory/cold-cache exercises; no production cache clearing or destructive disk pressure |
| measurement | record p50/p95/p99, errors, queue/wait, memory, I/O/network and result size where the local tool exposes them; never invent timing values |
1. Store size is multiple stores
Neo4j persists node/relationship/property data, native indexes,
Lucene-backed indexes, transaction logs and other database
files.
neo4j-admin server memory-recommendation reports
useful aggregate data/native-index and Lucene-index sizes for
memory planning. File-system snapshots add disk evidence. Do not
infer page-cache requirement from “database folder size” without
knowing which files are served by page cache versus
OS/Lucene/vector mechanisms.
docker exec atlasmart-neo4j neo4j-admin server memory-recommendation
# Optional inside-container filesystem snapshots; paths depend on distribution/container mapping:
docker exec atlasmart-neo4j sh -lc 'du -sh /data/databases/neo4j /data/transactions/neo4j 2>/dev/null || true'
2. Degree distribution predicts traversal pressure better than average count
A graph with one million evenly connected nodes behaves differently from one where one influencer/product/category has five million incident relationships. Measure degree distribution on labels/types relevant to critical queries.
CYPHER 25
MATCH (c:ObsCustomer)
OPTIONAL MATCH (c)-[r:OBS_VIEWED]->()
WITH c, count(r) AS degree
RETURN count(*) AS customers,
min(degree) AS minDegree,
percentileCont(degree, 0.50) AS p50Degree,
percentileCont(degree, 0.95) AS p95Degree,
max(degree) AS maxDegree;
A stable node count with rising degree can increase traversal rows, lock hotspots, page working set and result volume. Track the workload-relevant graph shape, not just total nodes.
3. Index footprint is a write/storage tradeoff
Every index can accelerate supported predicates but adds storage, population/build work and write maintenance. Chapter 9 already established evidence-based index selection. Capacity planning extends that discipline: keep index inventory, store footprint and build headroom with the query evidence that justifies each index.
CYPHER 25
SHOW INDEXES
YIELD name, type, entityType, labelsOrTypes, properties, state, populationPercent
RETURN * ORDER BY type, name;
| Index evidence | Capacity question |
|---|---|
| ONLINE + justified by plans | reserve steady-state storage/write maintenance |
| POPULATING | is there temporary CPU/I/O/storage headroom for build? |
| redundant/unused by workload evidence | can it be removed after a safe review? |
| full-text/vector/Lucene-backed | what OS/native/Lucene memory/storage budget applies beyond page cache? |
4. Transaction logs are recovery/replication data, not monitoring logs
Transaction logs record database changes. They are separate from
debug/query/security logs. Current configuration includes
rotation size and retention policy; the documented default
retention policy is 2 days 2G, but production must
choose retention based on backup/recovery/cluster needs and disk
capacity. Neo4j explicitly does not support manually deleting
transaction-log files.
CYPHER 25
SHOW SETTINGS
YIELD name, value
WHERE name IN ['db.tx_log.rotation.size','db.tx_log.rotation.retention_policy']
RETURN name, value ORDER BY name;
A burst of writes can increase transaction-log generation even when net store size barely changes. Alert on log-area growth and free space, not only database-store size.
5. Build timestamped snapshots
| Snapshot field | Source | Why retain history |
|---|---|---|
| data + native index MiB | memory-recommendation / file size | page-cache/storage trend |
| Lucene/vector index MiB | memory-recommendation / index inventory/files | OS/native/storage budget trend |
| transaction log MiB | filesystem + retention config | write/recovery disk trend |
| nodes/rels by critical label/type | Cypher counts | domain volume trend |
| p50/p95/max degree | Cypher degree distribution | traversal/hotspot trend |
| p50/p95/p99 latency + throughput | application benchmark | capacity outcome, not just resource use |
| CPU/memory/I/O/network | host/container metrics | identify saturation dimension |
6. Forecast cautiously
A simple linear forecast is useful as an alerting/planning tool only when its assumptions are explicit. Campaigns, imports, index builds and retention changes break linearity. Use it to ask “when should we revisit capacity?” rather than to promise an exhaustion date.
from statistics import mean
from datetime import date, timedelta
# Replace with measured daily MiB snapshots from your environment.
points = [
(date(2026, 9, 1), 820),
(date(2026, 9, 2), 836),
(date(2026, 9, 3), 851),
(date(2026, 9, 4), 870),
(date(2026, 9, 5), 888),
]
deltas = [points[i][1]-points[i-1][1] for i in range(1,len(points))]
daily_growth = mean(deltas)
current = points[-1][1]
capacity = 2048
headroom = capacity-current
print({"avg_daily_growth_mib": round(daily_growth,2),
"current_mib": current,
"headroom_mib": headroom,
"linear_days_to_capacity": round(headroom/daily_growth,1)})
# This is a planning model, not a forecast guarantee; growth is rarely perfectly linear.
The code ships with illustrative values so the arithmetic is reproducible. Replace them with your own daily snapshots before making any capacity decision.
7. Headroom includes recovery and maintenance
| Event | Temporary resource need |
|---|---|
| index population/rebuild | CPU, I/O, disk, page/OS memory |
| backup/dump/recovery | disk throughput, space for artifact/restore target, CPU |
| cluster store copy/catch-up | network + disk + CPU on source/target; remaining members serve traffic |
| schema/data migration | write amplification, tx logs, query memory, rollback copy |
| failure of one host | survivors absorb workload; no headroom means failover can become saturation |
8. Capacity acceptance questions
Check your understanding
- Why is node count a weak capacity metric by itself?
- Why track transaction-log size separately from store growth?
- Can you manually delete old tx-log files to free space?
- Why keep degree percentiles?
- Why reserve headroom beyond normal peak?
Review the answers
1. It omits relationship degree/skew, properties, indexes, working set, write rate, logs and query shape.
2. Write volume/retention can grow tx logs rapidly even if net graph size changes little.
3. No; use supported retention/pruning configuration and recovery planning.
4. High-degree/skewed nodes can dominate traversal cost and contention while averages stay modest.
5. Recovery, maintenance, index builds and host failures add temporary resource demand exactly when the system is already stressed.
Production judgment
| Decision surface | Production questions |
|---|---|
| graph/workload fit | Is latency dominated by graph fan-out, result cardinality, locks, indexes, client chatter, or infrastructure rather than raw store size? |
| correctness | Can any proposed remediation change result semantics, transaction boundaries, retry behavior, or causal-read guarantees? |
| model/cardinality/degree | Which labels/types are high-degree or skewed? Are a few hot nodes causing contention or explosive expansion? |
| latency | What are p50/p95/p99 and timeout/error distributions at representative load, not one warm single-user query? |
| transactions/concurrency | How long do transactions stay open? How much wait time and lock scope is normal, and which workflows create hot entities? |
| memory | How are RAM budgets split among OS, heap, page cache, transaction/query memory, direct/native buffers, Lucene/vector needs, and GDS if present? |
| CPU/disk/network | Which resource saturates first? Is disk latency/page-fault activity aligned with slowdown? Is client/network/serialization time dominant? |
| indexes/constraints | Do plans use the intended access paths? What storage/write overhead and population/build headroom do indexes add? |
| driver | Are pool acquisition, retry, timeout, fetch/result consumption and request correlation visible alongside server evidence? |
| security/tenant risk | Do logs/metrics expose sensitive parameters or tenant data? Are monitoring privileges and endpoints restricted appropriately? |
| backup/recovery | Is capacity monitoring independent from backup evidence, and is disk headroom sufficient for logs, dumps/backups and recovery operations? |
| observability | What baseline defines alerts? Which signals are causal/proximate versus merely correlated? How long is telemetry retained? |
| testing/failure injection | Can slow query, lock wait, memory pressure, cold start and network delay be reproduced in an isolated environment with reset steps? |
| version/edition/Aura | Which fields/settings/logs/metrics exist on the exact server, Cypher version, edition and managed tier? |
| cost/migration | What telemetry/storage overhead is acceptable, and how will capacity changes or schema/model remediations be rolled back? |
Summary and next step
You can now connect graph/store/log growth to resource headroom. Lesson 5 turns the chapter into an operational runbook and tests whether each diagnosis/remediation survives evidence.
Authoritative references
- Current Neo4j versions — Current Neo4j 2026.07.1 and 5.26.30 LTS release snapshot.
- Monitoring overview — Current monitoring surfaces: logs, metrics, query/transaction management, connections, jobs, and reports.
- Metrics — Enterprise metrics architecture and supported export surfaces.
- Essential metrics — Server, Neo4j, cluster, and workload signals that require correlation rather than single-metric tuning.
- Metrics reference — Current page-cache, JVM, CPU, file-descriptor, transaction and store metric names.
- Logging — Current neo4j/debug/http/gc/query/security log behavior and edition boundaries.
- Show and terminate transactions — Current SHOW TRANSACTIONS fields for locks, wait time, memory, page hits/faults, runtime, indexes, and status.
- Manage queries — Current query visibility through transaction management rather than legacy dbms.listQueries().
- Memory configuration — Heap, page cache, native memory, transaction limits, and memory-recommendation guidance.
- Disks, RAM and other tips — Page-cache warmup, disk behavior, and operating-system considerations.
- Performance — Performance topics spanning memory, disks, statistics, execution plans, Bolt threads, and filesystem tuning.
- Show configuration settings — Current SHOW SETTINGS syntax and self-managed server scope.
- Transaction logs — Transaction-log location, rotation, retention, pruning, and capacity implications.
- Query tuning and PROFILE — Execution-plan evidence for rows, database hits, memory, access paths, and operator shape.
- Python driver performance — Driver-side timing, database selection, result consumption, and application/database boundary considerations.