Chapter 01 · Graph Database Foundations, Neo4j Editions, Deployment Choices, and Lab Setup
Databases, DBMS Processes, Bolt, HTTP, Drivers, Configuration, Storage Locations, and Transaction Logs
Look behind a successful query without treating internal files as APIs: separate databases, connectors, config, stores, logs, plugins, and transaction durability evidence.
Learning outcomes
AtlasMart can now connect to Neo4j, but a production operator
needs to know what lies behind that successful query. This
lesson draws the boundaries between the DBMS process, the
neo4j database, the system database,
Bolt and HTTP connectors, drivers, configuration inputs,
persistent stores, logs, plugins, and transaction logs. The goal
is observability without treating internal files as an
application API.
Distinguish the DBMS process from an individual database and explain the special system database at an introductory level.
Explain Bolt, HTTP, Browser, Cypher Shell, and official drivers as separate client/connector layers.
Locate configuration, data, logs, plugins, and transaction logs in the Docker baseline while recognizing distribution-specific path differences.
Explain why transaction logs are durability/recovery machinery and why direct file mutation is unsafe.
Correlate a Cypher write with client evidence, server logs, mounted storage, and transaction-log state without claiming filesystem internals are stable APIs.
Application code talks to Neo4j through supported
protocols/APIs. Operators inspect and manage files with
documented Neo4j tooling. A file visible under
/data is not permission to edit, copy
selectively, truncate, or delete it while the server runs.
DBMS process versus database
The DBMS is the running Neo4j
management/runtime process. A database is a
transactional data-management domain hosted by that DBMS. A
standard current installation includes the
system database for DBMS
metadata/security/administration and a default user database
named neo4j. Community supports one user database
plus system; Enterprise extends multi-database
administration.
A transaction cannot be casually treated as spanning arbitrary
databases. Application code should specify the intended database
rather than relying forever on a default. The driver example in
Lesson 3 used database_="neo4j" deliberately. That
makes routing and intent explicit and avoids surprises if a
deployment’s default/home database changes.
CALL dbms.components() YIELD name, versions, edition RETURN name, versions, edition;SHOW CURRENT USER;MATCH (n) RETURN count(n) AS nodesInCurrentDatabase;
Enterprise learners may additionally inspect
SHOW DATABASES; Community learners should not make
that command a mandatory check.
Connectors: Browser is not Bolt and Bolt is not a driver
Neo4j exposes network connectors. The current defaults include
HTTP on 7474, HTTPS on 7473 when
enabled, and Bolt on 7687. Browser is a web client
typically loaded through HTTP/HTTPS; it then connects to the
database. Cypher Shell and official language drivers speak Bolt.
A driver adds language-level connection
pooling, sessions/transactions, routing, retries and result APIs
around Bolt—it is not just a socket.
| Layer | Chapter 01 example | What it owns |
|---|---|---|
| HTTP connector | 127.0.0.1:7474 | Web/HTTP endpoint; unencrypted in this local baseline |
| Bolt connector | 127.0.0.1:7687 | Binary database protocol used by Cypher Shell/drivers |
| Browser/Query | Web UI | Interactive query authoring/visualization; not a production app runtime |
| Cypher Shell | CLI client | Scriptable/interactive Cypher over Bolt |
| Python Driver 6.3.0 | Application library | Pools, transactions, query execution, routing/retry behavior and result mapping |
SHOW SETTINGS 'server.bolt.enabled', 'server.bolt.listen_address', 'server.bolt.advertised_address', 'server.http.enabled', 'server.http.listen_address' YIELD name, value RETURN name, value ORDER BY name;
Listen and advertised addresses solve different problems. A listener says where a server accepts traffic; an advertised address is what clients/routing may be told to use. Container NAT, DNS, clusters, proxies, and cloud networking can make them differ. The lab’s host-loopback mapping is Docker-level evidence in addition to Neo4j’s internal connector settings.
Configuration provenance: file, environment, and runtime evidence
For archive/server distributions, neo4j.conf is the
principal configuration file. Docker can also derive settings
from environment variables and mounted configuration. The safest
operational practice is to know which source generated the
effective setting and then confirm the runtime value. Editing a
file without restarting or without understanding precedence does
not prove the active DBMS changed.
docker exec atlasmart-neo4j sh -lc 'echo "$NEO4J_HOME"; sed -n "1,80p" "$NEO4J_HOME/conf/neo4j.conf"'docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p atlasmart-course-2026 "SHOW SETTINGS 'db.query.default_language', 'server.bolt.enabled', 'server.http.enabled' YIELD name, value RETURN name, value ORDER BY name;"docker inspect atlasmart-neo4j --format '{{json .Config.Env}}'
Do not paste a production
docker inspect environment dump into tickets if it
may contain secrets. This course environment contains only a
disposable password. Production credentials should come from a
secrets mechanism rather than plaintext Docker environment
variables when practical.
Storage locations are operational evidence, not public data structures
Neo4j’s documented default locations separate configuration,
data, logs, plugins, metrics/products and other runtime
directories. In Docker, official mount points include
/data, /logs, /conf, and
/plugins. This chapter mounts only data and logs
because no plugin or custom configuration is required yet.
| Surface | Docker baseline | Why it matters | Safe rule |
|---|---|---|---|
| Configuration | $NEO4J_HOME/conf/neo4j.conf |
Startup/default settings | Version-control desired config; verify effective settings |
| Graph/data stores | /data |
Persistent database state | Use supported DBMS/admin tools; do not hand-edit files |
| Transaction logs |
/data/transactions/<database> by
default
|
Recovery and change durability history | Never delete/truncate to “save space”; configure retention correctly |
| Server logs | /logs |
Startup, warnings, failures, operational evidence | Collect/rotate with documented settings and protect sensitive content |
| Plugins | /plugins |
Optional extension JARs | Pin compatible plugin versions and permissions; none required in Chapter 01 |
docker exec atlasmart-neo4j sh -lc 'find /data -maxdepth 2 -type d -print | sort | head -80'docker exec atlasmart-neo4j sh -lc 'find /logs -maxdepth 1 -type f -printf "%f %s bytes\n" 2>/dev/null | sort'docker exec atlasmart-neo4j sh -lc 'find /data/transactions -maxdepth 2 -type f -printf "%p %s bytes\n" 2>/dev/null | head -40'
Exact internal file names, store format, rotation, and directory contents can change by version and edition. The lesson observes them only to connect operations to documented concepts. An application must never depend on a particular store filename.
Transaction logs connect committed writes to recovery
Neo4j records transactional changes in transaction logs. They support recovery and other operational mechanisms. A successful transaction means the DBMS committed according to its transactional rules; it does not mean a separate off-host backup exists. Transaction logs are not a replacement for a backup policy, and replication—where available—is also not a backup.
Create one clearly identifiable write, then compare transaction-log directory metadata before and after. File sizes can be buffered/rotated and therefore should not be treated as a one-write-equals-N-bytes formula.
docker exec atlasmart-neo4j sh -lc 'find /data/transactions/neo4j -maxdepth 1 -type f -printf "%f %s bytes %TY-%Tm-%TdT%TH:%TM:%TS\n" 2>/dev/null | sort'
CYPHER 25MERGE (p:Product {productId:'P-2001'})SET p.name='Smart Shelf Sensor', p.category='Store IoT', p.updatedAt=datetime('2026-09-09T09:00:00Z')RETURN p.productId, p.name, p.updatedAt;
docker exec atlasmart-neo4j sh -lc 'find /data/transactions/neo4j -maxdepth 1 -type f -printf "%f %s bytes %TY-%Tm-%TdT%TH:%TM:%TS\n" 2>/dev/null | sort'
The accepted evidence is that the graph query returns the committed state and the transaction-log area exists/changes over normal operation. Do not claim that one timestamp or byte delta uniquely identifies that transaction. Background activity, checkpoints, rotation, and buffering make that inference unsound.
What Neo4j is doing during the write
The driver/Cypher Shell sends the statement over Bolt. Neo4j
executes it in a transaction, obtains the necessary
locks/concurrency control, evaluates the
MERGE match/create semantics, updates graph state
and relevant indexes/constraints, records transactional changes,
and commits or rolls back as a unit. MERGE is not a
universal SQL-style UPSERT: its match pattern and constraints
determine what can match/create, and concurrent writers require
a deliberate uniqueness model. Later chapters analyze these
details.
Memory also has distinct surfaces: JVM heap for runtime objects/query execution, page cache for store pages, native/off-heap allocations, and—when used later—Graph Data Science memory. Increasing one pool blindly can starve another. Chapter 01 records the vocabulary; performance engineering later measures it.
Deliberately wrong approach: manipulate files because they are visible
A common local-lab mistake is to find a large file under
/data/transactions or a store file under
/data/databases and delete it manually. That
bypasses Neo4j’s recovery, retention, checkpoint, and
consistency mechanisms. The same anti-pattern applies to copying
a subset of live files and calling it a backup.
The repair is to manage transaction-log retention through supported configuration and to use supported dump/backup/restore tooling appropriate to the edition. If a lab needs destructive file experiments in a later internals chapter, it must operate on a stopped disposable copy with an explicit reset path.
Production judgment
For self-managed production, document
NEO4J_HOME/NEO4J_CONF strategy,
persistent-volume ownership, data/log separation, filesystem
permissions, secrets, TLS, connector exposure, log collection,
transaction-log retention, backup paths, disk-headroom alarms,
and upgrade-compatible layouts. Windows, packages, archives,
containers, and orchestrators place files differently; avoid
runbooks that assume Docker paths everywhere.
For Aura, most of these server-file details disappear from the customer control surface, but the application still must use the correct database, secure URI, driver settings, query timeouts, credentials, and service-supported recovery/observability mechanisms.
Check your understanding
-
What is the difference between the Neo4j DBMS and the
neo4jdatabase? - Why can Browser load over HTTP while the Python driver still fails?
- Why compare neo4j.conf/env input with SHOW SETTINGS output?
- What do transaction logs provide that an application query does not?
-
Why is direct mutation of
/dataunsafe even in a graph database that exposes its storage volume?
Review the answers
1. The DBMS is the running
management/runtime process that hosts database domains.
neo4j is the default user database inside it;
system holds DBMS administration metadata.
2. Browser/HTTP reachability and Bolt/driver connectivity are different connectors/protocols. Bolt can have independent URI, listener, TLS, authentication, or routing problems.
3. Inputs may be overridden or may require restart/new database creation. Runtime evidence tells you what the server actually exposes now.
4. They persist transactional change history used by Neo4j recovery/operational mechanisms. They are DBMS state, not a client response and not an off-host backup.
5. File formats, checkpoints, indexes, transaction state and recovery invariants are coordinated by Neo4j. Hand edits bypass those invariants and can corrupt or invalidate the store.
Summary and next step
You can now locate the boundary between supported client interfaces and implementation storage, distinguish DBMS from database, explain HTTP/Bolt/driver roles, inspect effective configuration, and connect a committed write to transaction-log state without over-interpreting internal files.
Next, turn these pieces into a repeatable course lab with a richer AtlasMart fixture, administrative-account evidence, metrics, failure checks, and safe reset procedures.
Authoritative references
- Current Neo4j versions — Official current-release and LTS patch snapshot.
- Neo4j Operations Manual — Authoritative self-managed operational documentation for the current release.
- Neo4j system requirements — Supported operating systems and JVM requirements.
- Cypher Manual — Current Cypher 25 reference and language semantics.
- Configure the Cypher default version — Cypher 5 versus Cypher 25 default and override behavior.
- Neo4j in Docker — Official image tags, ports, editions, and Docker starting point.
- Default file locations — Distribution-specific config/data/log/plugin/product locations.
- Neo4j transaction logs — Transaction-log purpose and management.
- Neo4j ports — Connector and administrative port reference.
- Configure network connectors — Listen/advertised address and protocol configuration.
- Docker configuration — Environment/configuration behavior for official images.
- Docker volumes — Official mount points and persistence behavior.