Chapter 18 · Monitoring and Troubleshooting: Metrics, Logs, Query Visibility, Memory, Store, and Capacity

Store Growth, Relationship Density, Index Size, Transaction Log Growth, and Capacity Forecasting

Turn store, index, relationship-degree, and transaction-log measurements into capacity evidence and forecasts instead of extrapolating from node counts or one snapshot.

Advanced220–300 minutesCapacity-growth forecasting labNeo4j 2026.07.1 Community mandatory · Enterprise observability optionalCypher 25 · SHOW TRANSACTIONS / PROFILE / SHOW SETTINGSJava 21/25 · Python driver 6.3Last reviewed: September 2026

AtlasMart’s graph is healthy today, but data volume grows 4% each week, one category hub’s degree doubles every campaign, and transaction logs occasionally consume unexpected disk. Capacity planning cannot be “count nodes and buy 2× RAM.” Graph shape, indexes, log retention, write rate, working set, backups and failure/maintenance headroom all consume different resources.

Capacity rule

Forecast from measured deltas of the resources that actually saturate—store, indexes, tx logs, page working set, IOPS, CPU, network, memory and recovery headroom—not from one vanity count.

Learning outcomes

01

Measure data/native-index and Lucene/vector footprint separately enough to reason about memory and disk budgets.

02

Measure relationship degree/skew rather than using average node count as a traversal-cost proxy.

03

Explain transaction-log rotation/retention and why manual deletion is unsupported.

04

Create timestamped growth snapshots and a cautious forecast with explicit uncertainty/headroom.

05

Connect capacity decisions to backups, index builds, cluster recovery, maintenance and failure headroom.

Reproducible lab baseline

Chapter 18 baseline · reviewed 9 September 2026

The continuity environment remains Neo4j Community 2026.07.1, database neo4j, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic local credential neo4j / atlasmart-course-2026, Java 21 or 25, explicit CYPHER 25 where version-sensitive syntax matters, and Python driver 6.3. Current 5.26 LTS is 5.26.30. No runtime command is claimed to have been executed during lesson generation; expected outputs are fixture invariants or clearly labeled examples.

Monitoring edition boundary

The mandatory path stays free and local. Neo4j Enterprise provides the DBMS metrics framework (CSV/JMX/Prometheus/Graphite integrations), query.log, security.log, query CPU tracking, broader RBAC-based operational visibility, and additional cluster signals. Community still supports useful first-party evidence including SHOW TRANSACTIONS, SHOW SETTINGS on self-managed deployments, PROFILE, general/debug logs, optional JVM GC logging, neo4j-admin server memory-recommendation, driver timing, and operating-system/container counters. Do not fabricate Enterprise-only metrics in the Community lab.

Term Operational meaning
symptom What the user or SLO observes: high p99, timeout, error, stale response, blocked write, memory termination, disk pressure.
evidence A timestamped observation that can support or contradict a hypothesis: plan, transaction row, log entry, OS counter, size snapshot, driver timing.
metric A numeric time series such as CPU load, page faults, heap use, open file descriptors, or request latency; a metric is not a root cause by itself.
heap JVM-managed memory for Java objects and much transaction/query state; garbage collection reclaims unused heap objects.
page cache Neo4j-managed cache for graph-store pages and native indexes read from disk; it is distinct from JVM heap.
transaction/query memory Estimated memory used for uncommitted state, intermediate rows, collections, sort/aggregation/path work, and results; bounded by transaction-memory settings.
native/off-heap memory Memory allocated outside JVM heap, including direct/network buffers and other native structures.
page hit / page fault A hit finds a requested store page in page cache; a fault requires loading a page. Neither ratio alone explains end-to-end latency.
lock wait Time a transaction cannot proceed because another transaction owns a conflicting lock.
database hit A PROFILE plan counter for work against the graph store/index abstractions; it is not a literal physical-disk-read count.
tail latency High-percentile latency such as p95/p99; averages can remain healthy while a minority of requests violate the SLO.
capacity headroom Resources intentionally left unused so spikes, recovery, maintenance, compaction/index builds, and failures do not immediately saturate the system.
Assumption Pinned value / rule
server/current Neo4j 2026.07.1 Community mandatory; 5.26.30 LTS only for compatibility comparison
database neo4j; chapter fixture uses isolated Obs* labels and relationship types
driver Python driver 6.3; one maintained Driver, separate Session/Transaction scopes
TLS/auth loopback disposable lab may use bolt without TLS; remote production requires verified TLS and non-embedded secrets
plugins APOC/GDS are not required
metrics Enterprise metrics exporters are optional/reference-only; Community lab uses SHOW/PROFILE/logs/admin/OS/driver evidence
failure injection only isolated, reversible slow-query/lock/memory/cold-cache exercises; no production cache clearing or destructive disk pressure
measurement record p50/p95/p99, errors, queue/wait, memory, I/O/network and result size where the local tool exposes them; never invent timing values

1. Store size is multiple stores

Neo4j persists node/relationship/property data, native indexes, Lucene-backed indexes, transaction logs and other database files. neo4j-admin server memory-recommendation reports useful aggregate data/native-index and Lucene-index sizes for memory planning. File-system snapshots add disk evidence. Do not infer page-cache requirement from “database folder size” without knowing which files are served by page cache versus OS/Lucene/vector mechanisms.

Shell · size-oriented evidence
docker exec atlasmart-neo4j neo4j-admin server memory-recommendation
# Optional inside-container filesystem snapshots; paths depend on distribution/container mapping:
docker exec atlasmart-neo4j sh -lc 'du -sh /data/databases/neo4j /data/transactions/neo4j 2>/dev/null || true'

2. Degree distribution predicts traversal pressure better than average count

A graph with one million evenly connected nodes behaves differently from one where one influencer/product/category has five million incident relationships. Measure degree distribution on labels/types relevant to critical queries.

Cypher 25 · degree distribution for the chapter fixture
CYPHER 25
MATCH (c:ObsCustomer)
OPTIONAL MATCH (c)-[r:OBS_VIEWED]->()
WITH c, count(r) AS degree
RETURN count(*) AS customers,
       min(degree) AS minDegree,
       percentileCont(degree, 0.50) AS p50Degree,
       percentileCont(degree, 0.95) AS p95Degree,
       max(degree) AS maxDegree;
Edge case

A stable node count with rising degree can increase traversal rows, lock hotspots, page working set and result volume. Track the workload-relevant graph shape, not just total nodes.

3. Index footprint is a write/storage tradeoff

Every index can accelerate supported predicates but adds storage, population/build work and write maintenance. Chapter 9 already established evidence-based index selection. Capacity planning extends that discipline: keep index inventory, store footprint and build headroom with the query evidence that justifies each index.

Cypher 25 · index inventory
CYPHER 25
SHOW INDEXES
YIELD name, type, entityType, labelsOrTypes, properties, state, populationPercent
RETURN * ORDER BY type, name;
Index evidence Capacity question
ONLINE + justified by plans reserve steady-state storage/write maintenance
POPULATING is there temporary CPU/I/O/storage headroom for build?
redundant/unused by workload evidence can it be removed after a safe review?
full-text/vector/Lucene-backed what OS/native/Lucene memory/storage budget applies beyond page cache?

4. Transaction logs are recovery/replication data, not monitoring logs

Transaction logs record database changes. They are separate from debug/query/security logs. Current configuration includes rotation size and retention policy; the documented default retention policy is 2 days 2G, but production must choose retention based on backup/recovery/cluster needs and disk capacity. Neo4j explicitly does not support manually deleting transaction-log files.

Cypher 25 · inspect tx-log policy
CYPHER 25
SHOW SETTINGS
YIELD name, value
WHERE name IN ['db.tx_log.rotation.size','db.tx_log.rotation.retention_policy']
RETURN name, value ORDER BY name;
Capacity consequence

A burst of writes can increase transaction-log generation even when net store size barely changes. Alert on log-area growth and free space, not only database-store size.

5. Build timestamped snapshots

Snapshot field Source Why retain history
data + native index MiB memory-recommendation / file size page-cache/storage trend
Lucene/vector index MiB memory-recommendation / index inventory/files OS/native/storage budget trend
transaction log MiB filesystem + retention config write/recovery disk trend
nodes/rels by critical label/type Cypher counts domain volume trend
p50/p95/max degree Cypher degree distribution traversal/hotspot trend
p50/p95/p99 latency + throughput application benchmark capacity outcome, not just resource use
CPU/memory/I/O/network host/container metrics identify saturation dimension

6. Forecast cautiously

A simple linear forecast is useful as an alerting/planning tool only when its assumptions are explicit. Campaigns, imports, index builds and retention changes break linearity. Use it to ask “when should we revisit capacity?” rather than to promise an exhaustion date.

Python · simple measured-delta forecast
from statistics import mean
from datetime import date, timedelta

# Replace with measured daily MiB snapshots from your environment.
points = [
    (date(2026, 9, 1), 820),
    (date(2026, 9, 2), 836),
    (date(2026, 9, 3), 851),
    (date(2026, 9, 4), 870),
    (date(2026, 9, 5), 888),
]
deltas = [points[i][1]-points[i-1][1] for i in range(1,len(points))]
daily_growth = mean(deltas)
current = points[-1][1]
capacity = 2048
headroom = capacity-current
print({"avg_daily_growth_mib": round(daily_growth,2),
       "current_mib": current,
       "headroom_mib": headroom,
       "linear_days_to_capacity": round(headroom/daily_growth,1)})
# This is a planning model, not a forecast guarantee; growth is rarely perfectly linear.
Do not forecast from synthetic example numbers

The code ships with illustrative values so the arithmetic is reproducible. Replace them with your own daily snapshots before making any capacity decision.

7. Headroom includes recovery and maintenance

Event Temporary resource need
index population/rebuild CPU, I/O, disk, page/OS memory
backup/dump/recovery disk throughput, space for artifact/restore target, CPU
cluster store copy/catch-up network + disk + CPU on source/target; remaining members serve traffic
schema/data migration write amplification, tx logs, query memory, rollback copy
failure of one host survivors absorb workload; no headroom means failover can become saturation

8. Capacity acceptance questions

Check your understanding

  1. Why is node count a weak capacity metric by itself?
  2. Why track transaction-log size separately from store growth?
  3. Can you manually delete old tx-log files to free space?
  4. Why keep degree percentiles?
  5. Why reserve headroom beyond normal peak?
Review the answers

1. It omits relationship degree/skew, properties, indexes, working set, write rate, logs and query shape.

2. Write volume/retention can grow tx logs rapidly even if net graph size changes little.

3. No; use supported retention/pruning configuration and recovery planning.

4. High-degree/skewed nodes can dominate traversal cost and contention while averages stay modest.

5. Recovery, maintenance, index builds and host failures add temporary resource demand exactly when the system is already stressed.

Production judgment

Decision surface Production questions
graph/workload fit Is latency dominated by graph fan-out, result cardinality, locks, indexes, client chatter, or infrastructure rather than raw store size?
correctness Can any proposed remediation change result semantics, transaction boundaries, retry behavior, or causal-read guarantees?
model/cardinality/degree Which labels/types are high-degree or skewed? Are a few hot nodes causing contention or explosive expansion?
latency What are p50/p95/p99 and timeout/error distributions at representative load, not one warm single-user query?
transactions/concurrency How long do transactions stay open? How much wait time and lock scope is normal, and which workflows create hot entities?
memory How are RAM budgets split among OS, heap, page cache, transaction/query memory, direct/native buffers, Lucene/vector needs, and GDS if present?
CPU/disk/network Which resource saturates first? Is disk latency/page-fault activity aligned with slowdown? Is client/network/serialization time dominant?
indexes/constraints Do plans use the intended access paths? What storage/write overhead and population/build headroom do indexes add?
driver Are pool acquisition, retry, timeout, fetch/result consumption and request correlation visible alongside server evidence?
security/tenant risk Do logs/metrics expose sensitive parameters or tenant data? Are monitoring privileges and endpoints restricted appropriately?
backup/recovery Is capacity monitoring independent from backup evidence, and is disk headroom sufficient for logs, dumps/backups and recovery operations?
observability What baseline defines alerts? Which signals are causal/proximate versus merely correlated? How long is telemetry retained?
testing/failure injection Can slow query, lock wait, memory pressure, cold start and network delay be reproduced in an isolated environment with reset steps?
version/edition/Aura Which fields/settings/logs/metrics exist on the exact server, Cypher version, edition and managed tier?
cost/migration What telemetry/storage overhead is acceptable, and how will capacity changes or schema/model remediations be rolled back?

Summary and next step

You can now connect graph/store/log growth to resource headroom. Lesson 5 turns the chapter into an operational runbook and tests whether each diagnosis/remediation survives evidence.

Authoritative references

  • Current Neo4j versions — Current Neo4j 2026.07.1 and 5.26.30 LTS release snapshot.
  • Monitoring overview — Current monitoring surfaces: logs, metrics, query/transaction management, connections, jobs, and reports.
  • Metrics — Enterprise metrics architecture and supported export surfaces.
  • Essential metrics — Server, Neo4j, cluster, and workload signals that require correlation rather than single-metric tuning.
  • Metrics reference — Current page-cache, JVM, CPU, file-descriptor, transaction and store metric names.
  • Logging — Current neo4j/debug/http/gc/query/security log behavior and edition boundaries.
  • Show and terminate transactions — Current SHOW TRANSACTIONS fields for locks, wait time, memory, page hits/faults, runtime, indexes, and status.
  • Manage queries — Current query visibility through transaction management rather than legacy dbms.listQueries().
  • Memory configuration — Heap, page cache, native memory, transaction limits, and memory-recommendation guidance.
  • Disks, RAM and other tips — Page-cache warmup, disk behavior, and operating-system considerations.
  • Performance — Performance topics spanning memory, disks, statistics, execution plans, Bolt threads, and filesystem tuning.
  • Show configuration settings — Current SHOW SETTINGS syntax and self-managed server scope.
  • Transaction logs — Transaction-log location, rotation, retention, pruning, and capacity implications.
  • Query tuning and PROFILE — Execution-plan evidence for rows, database hits, memory, access paths, and operator shape.
  • Python driver performance — Driver-side timing, database selection, result consumption, and application/database boundary considerations.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.