Chapter 01 · Graph Database Foundations, Neo4j Editions, Deployment Choices, and Lab Setup

Databases, DBMS Processes, Bolt, HTTP, Drivers, Configuration, Storage Locations, and Transaction Logs

Look behind a successful query without treating internal files as APIs: separate databases, connectors, config, stores, logs, plugins, and transaction durability evidence.

Intermediate105–125 minutesRuntime and storage-boundary labNeo4j 2026.07.1 Community · Docker baselineLast reviewed: September 2026

Learning outcomes

AtlasMart can now connect to Neo4j, but a production operator needs to know what lies behind that successful query. This lesson draws the boundaries between the DBMS process, the neo4j database, the system database, Bolt and HTTP connectors, drivers, configuration inputs, persistent stores, logs, plugins, and transaction logs. The goal is observability without treating internal files as an application API.

01

Distinguish the DBMS process from an individual database and explain the special system database at an introductory level.

02

Explain Bolt, HTTP, Browser, Cypher Shell, and official drivers as separate client/connector layers.

03

Locate configuration, data, logs, plugins, and transaction logs in the Docker baseline while recognizing distribution-specific path differences.

04

Explain why transaction logs are durability/recovery machinery and why direct file mutation is unsafe.

05

Correlate a Cypher write with client evidence, server logs, mounted storage, and transaction-log state without claiming filesystem internals are stable APIs.

Boundary rule

Application code talks to Neo4j through supported protocols/APIs. Operators inspect and manage files with documented Neo4j tooling. A file visible under /data is not permission to edit, copy selectively, truncate, or delete it while the server runs.

DBMS process versus database

The DBMS is the running Neo4j management/runtime process. A database is a transactional data-management domain hosted by that DBMS. A standard current installation includes the system database for DBMS metadata/security/administration and a default user database named neo4j. Community supports one user database plus system; Enterprise extends multi-database administration.

A transaction cannot be casually treated as spanning arbitrary databases. Application code should specify the intended database rather than relying forever on a default. The driver example in Lesson 3 used database_="neo4j" deliberately. That makes routing and intent explicit and avoids surprises if a deployment’s default/home database changes.

Cypher · DBMS and user-context evidence that works in the baseline
CALL dbms.components() YIELD name, versions, edition RETURN name, versions, edition;SHOW CURRENT USER;MATCH (n) RETURN count(n) AS nodesInCurrentDatabase;

Enterprise learners may additionally inspect SHOW DATABASES; Community learners should not make that command a mandatory check.

Connectors: Browser is not Bolt and Bolt is not a driver

Neo4j exposes network connectors. The current defaults include HTTP on 7474, HTTPS on 7473 when enabled, and Bolt on 7687. Browser is a web client typically loaded through HTTP/HTTPS; it then connects to the database. Cypher Shell and official language drivers speak Bolt. A driver adds language-level connection pooling, sessions/transactions, routing, retries and result APIs around Bolt—it is not just a socket.

Layer Chapter 01 example What it owns
HTTP connector 127.0.0.1:7474 Web/HTTP endpoint; unencrypted in this local baseline
Bolt connector 127.0.0.1:7687 Binary database protocol used by Cypher Shell/drivers
Browser/Query Web UI Interactive query authoring/visualization; not a production app runtime
Cypher Shell CLI client Scriptable/interactive Cypher over Bolt
Python Driver 6.3.0 Application library Pools, transactions, query execution, routing/retry behavior and result mapping
Cypher · inspect selected connector settings
SHOW SETTINGS 'server.bolt.enabled', 'server.bolt.listen_address', 'server.bolt.advertised_address', 'server.http.enabled', 'server.http.listen_address' YIELD name, value RETURN name, value ORDER BY name;

Listen and advertised addresses solve different problems. A listener says where a server accepts traffic; an advertised address is what clients/routing may be told to use. Container NAT, DNS, clusters, proxies, and cloud networking can make them differ. The lab’s host-loopback mapping is Docker-level evidence in addition to Neo4j’s internal connector settings.

Configuration provenance: file, environment, and runtime evidence

For archive/server distributions, neo4j.conf is the principal configuration file. Docker can also derive settings from environment variables and mounted configuration. The safest operational practice is to know which source generated the effective setting and then confirm the runtime value. Editing a file without restarting or without understanding precedence does not prove the active DBMS changed.

shell · compare configuration input with runtime-oriented evidence
docker exec atlasmart-neo4j sh -lc 'echo "$NEO4J_HOME"; sed -n "1,80p" "$NEO4J_HOME/conf/neo4j.conf"'docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p atlasmart-course-2026 "SHOW SETTINGS 'db.query.default_language', 'server.bolt.enabled', 'server.http.enabled' YIELD name, value RETURN name, value ORDER BY name;"docker inspect atlasmart-neo4j --format '{{json .Config.Env}}' 

Do not paste a production docker inspect environment dump into tickets if it may contain secrets. This course environment contains only a disposable password. Production credentials should come from a secrets mechanism rather than plaintext Docker environment variables when practical.

Storage locations are operational evidence, not public data structures

Neo4j’s documented default locations separate configuration, data, logs, plugins, metrics/products and other runtime directories. In Docker, official mount points include /data, /logs, /conf, and /plugins. This chapter mounts only data and logs because no plugin or custom configuration is required yet.

Surface Docker baseline Why it matters Safe rule
Configuration $NEO4J_HOME/conf/neo4j.conf Startup/default settings Version-control desired config; verify effective settings
Graph/data stores /data Persistent database state Use supported DBMS/admin tools; do not hand-edit files
Transaction logs /data/transactions/<database> by default Recovery and change durability history Never delete/truncate to “save space”; configure retention correctly
Server logs /logs Startup, warnings, failures, operational evidence Collect/rotate with documented settings and protect sensitive content
Plugins /plugins Optional extension JARs Pin compatible plugin versions and permissions; none required in Chapter 01
shell · read-only directory evidence
docker exec atlasmart-neo4j sh -lc 'find /data -maxdepth 2 -type d -print | sort | head -80'docker exec atlasmart-neo4j sh -lc 'find /logs -maxdepth 1 -type f -printf "%f %s bytes\n" 2>/dev/null | sort'docker exec atlasmart-neo4j sh -lc 'find /data/transactions -maxdepth 2 -type f -printf "%p %s bytes\n" 2>/dev/null | head -40' 

Exact internal file names, store format, rotation, and directory contents can change by version and edition. The lesson observes them only to connect operations to documented concepts. An application must never depend on a particular store filename.

Transaction logs connect committed writes to recovery

Neo4j records transactional changes in transaction logs. They support recovery and other operational mechanisms. A successful transaction means the DBMS committed according to its transactional rules; it does not mean a separate off-host backup exists. Transaction logs are not a replacement for a backup policy, and replication—where available—is also not a backup.

Create one clearly identifiable write, then compare transaction-log directory metadata before and after. File sizes can be buffered/rotated and therefore should not be treated as a one-write-equals-N-bytes formula.

shell · list log state before the write
docker exec atlasmart-neo4j sh -lc 'find /data/transactions/neo4j -maxdepth 1 -type f -printf "%f %s bytes %TY-%Tm-%TdT%TH:%TM:%TS\n" 2>/dev/null | sort' 
Cypher · one committed AtlasMart write
CYPHER 25MERGE (p:Product {productId:'P-2001'})SET p.name='Smart Shelf Sensor',    p.category='Store IoT',    p.updatedAt=datetime('2026-09-09T09:00:00Z')RETURN p.productId, p.name, p.updatedAt;
shell · list transaction-log state again, read-only
docker exec atlasmart-neo4j sh -lc 'find /data/transactions/neo4j -maxdepth 1 -type f -printf "%f %s bytes %TY-%Tm-%TdT%TH:%TM:%TS\n" 2>/dev/null | sort' 

The accepted evidence is that the graph query returns the committed state and the transaction-log area exists/changes over normal operation. Do not claim that one timestamp or byte delta uniquely identifies that transaction. Background activity, checkpoints, rotation, and buffering make that inference unsound.

What Neo4j is doing during the write

The driver/Cypher Shell sends the statement over Bolt. Neo4j executes it in a transaction, obtains the necessary locks/concurrency control, evaluates the MERGE match/create semantics, updates graph state and relevant indexes/constraints, records transactional changes, and commits or rolls back as a unit. MERGE is not a universal SQL-style UPSERT: its match pattern and constraints determine what can match/create, and concurrent writers require a deliberate uniqueness model. Later chapters analyze these details.

Memory also has distinct surfaces: JVM heap for runtime objects/query execution, page cache for store pages, native/off-heap allocations, and—when used later—Graph Data Science memory. Increasing one pool blindly can starve another. Chapter 01 records the vocabulary; performance engineering later measures it.

Deliberately wrong approach: manipulate files because they are visible

A common local-lab mistake is to find a large file under /data/transactions or a store file under /data/databases and delete it manually. That bypasses Neo4j’s recovery, retention, checkpoint, and consistency mechanisms. The same anti-pattern applies to copying a subset of live files and calling it a backup.

The repair is to manage transaction-log retention through supported configuration and to use supported dump/backup/restore tooling appropriate to the edition. If a lab needs destructive file experiments in a later internals chapter, it must operate on a stopped disposable copy with an explicit reset path.

Production judgment

For self-managed production, document NEO4J_HOME/NEO4J_CONF strategy, persistent-volume ownership, data/log separation, filesystem permissions, secrets, TLS, connector exposure, log collection, transaction-log retention, backup paths, disk-headroom alarms, and upgrade-compatible layouts. Windows, packages, archives, containers, and orchestrators place files differently; avoid runbooks that assume Docker paths everywhere.

For Aura, most of these server-file details disappear from the customer control surface, but the application still must use the correct database, secure URI, driver settings, query timeouts, credentials, and service-supported recovery/observability mechanisms.

Check your understanding

  1. What is the difference between the Neo4j DBMS and the neo4j database?
  2. Why can Browser load over HTTP while the Python driver still fails?
  3. Why compare neo4j.conf/env input with SHOW SETTINGS output?
  4. What do transaction logs provide that an application query does not?
  5. Why is direct mutation of /data unsafe even in a graph database that exposes its storage volume?
Review the answers

1. The DBMS is the running management/runtime process that hosts database domains. neo4j is the default user database inside it; system holds DBMS administration metadata.

2. Browser/HTTP reachability and Bolt/driver connectivity are different connectors/protocols. Bolt can have independent URI, listener, TLS, authentication, or routing problems.

3. Inputs may be overridden or may require restart/new database creation. Runtime evidence tells you what the server actually exposes now.

4. They persist transactional change history used by Neo4j recovery/operational mechanisms. They are DBMS state, not a client response and not an off-host backup.

5. File formats, checkpoints, indexes, transaction state and recovery invariants are coordinated by Neo4j. Hand edits bypass those invariants and can corrupt or invalidate the store.

Summary and next step

You can now locate the boundary between supported client interfaces and implementation storage, distinguish DBMS from database, explain HTTP/Bolt/driver roles, inspect effective configuration, and connect a committed write to transaction-log state without over-interpreting internal files.

Next, turn these pieces into a repeatable course lab with a richer AtlasMart fixture, administrative-account evidence, metrics, failure checks, and safe reset procedures.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.