Chapter 01 · Graph Database Foundations, Neo4j Editions, Deployment Choices, and Lab Setup
Graph Databases vs Relational, Document, and Search Systems: Connectivity, Traversal, and Workload Fit
Build the graph mental model from workload evidence: connected patterns, first-class relationships, transaction/query boundaries, and observable Neo4j runtime state.
Learning outcomes
AtlasMart already uses relational tables for orders, a document store for flexible catalog payloads, and a search engine for product discovery. A new fraud-and-recommendation initiative asks questions such as “which customers share devices, addresses, payment instruments, and products with this suspicious account within three relationship hops?” The architecture decision is not whether graphs look intuitive on a whiteboard. It is whether connected access patterns are important enough that representing and traversing relationships as first-class data produces a simpler, more predictable system than repeatedly reconstructing those connections elsewhere.
Explain the labeled property graph model and distinguish a graph database from a generic visualization or network data structure.
Compare relational, document, search, and graph systems by the work they make cheap or explicit rather than by slogans such as “NoSQL scales.”
Trace a small Cypher read from client connection through transaction, planning, operators, indexes/page cache/store access, and returned rows.
Define node, relationship, label, relationship type, property, pattern, Cypher, transaction, index, execution plan, driver, Bolt, and DBMS before using them operationally.
Prove which Neo4j server and Cypher configuration answered a request using observable runtime evidence.
The mandatory lab pins Neo4j Community 2026.07.1.
Neo4j also maintains the 5.26.30 LTS line. For
the current 2026 line, supported JVMs are Java 21 and 25 on
supported platforms. Neo4j 2026.02+ distribution configuration
explicitly sets
db.query.default_language=CYPHER_25 for new
deployments, but existing deployments can retain Cypher 5
defaults. Every lab below asks the server what it is actually
using instead of inferring behavior from the calendar version.
These commands were authored against current official Neo4j documentation and pinned versions, but this generation environment does not provide a running Docker daemon or Neo4j server, so the chapter does not pretend that example output was captured here. Expected output is described by invariant and shape. The lab binds HTTP and Bolt to host loopback, uses course-specific volumes and a synthetic graph, and deliberately avoids production credentials or data.
The property graph is about connected facts
Neo4j stores a labeled property graph. A node represents an entity or event-like thing. A node can have one or more labels, which classify it for modeling, querying, constraints, and indexes. A relationship connects exactly two nodes and has a stored direction and one relationship type. Nodes and relationships can carry properties: named values such as a customer ID, product name, amount, timestamp, or confidence score.
The critical difference from “an application that happens to
contain IDs” is that the relationship itself is first-class
database state. AtlasMart can store
(Customer)-[:PLACED]->(Order),
(Order)-[:CONTAINS]->(Product), and
(Customer)-[:USED_DEVICE]->(Device). A traversal
can then follow those stored relationships. Direction is part of
relationship identity, although Cypher may intentionally match a
relationship without specifying direction when the query means
either direction. Labels and types describe structure; they do
not by themselves enforce identity, uniqueness, or
every business invariant. Later chapters add constraints
deliberately.
| Term | Precise meaning in this chapter | Common misconception |
|---|---|---|
| DBMS | The Neo4j database-management process and administrative boundary | A single graph visualization tab |
| Database | A transactional data-management domain inside the DBMS | A directory that applications may mutate directly |
| Node | A graph entity with zero or more labels and properties | A JSON document that must contain all related data |
| Relationship | A directed, typed connection between two nodes, optionally with properties | A foreign-key value that has no independent semantics |
| Pattern | A shape described in Cypher and matched against graph data | A procedural loop that manually walks every node |
| Cypher | Neo4j’s declarative graph query language | A generic name for every graph API |
| Transaction | An ACID unit in which reads/writes are committed or rolled back | A request retry policy in the driver |
| Index | An access structure that can help find starting entities/properties | A structure that makes every traversal constant-time |
| Execution plan | The operators Neo4j chooses to execute a Cypher statement | The textual query itself |
| Driver / Bolt | The client library and Neo4j binary protocol used to communicate with the DBMS | A guarantee that every retry is safe |
Graph, relational, document, and search systems optimize different work
A relational database is usually strongest when normalized records, joins, constraints, set-oriented processing, and transactional invariants are the center of the workload. A document database is often attractive when an aggregate can be loaded and changed as a self-contained document with flexible nested shape. A search engine is optimized for retrieval over indexed text, terms, scoring, faceting, and related search primitives. A graph database earns its place when relationships and multi-step connected patterns are core application questions, not merely when data can be drawn as circles and arrows.
These systems are not mutually exclusive. AtlasMart may keep order-of-record accounting in a relational database, catalog text search in a search engine, and use Neo4j for fraud neighborhoods, recommendation paths, supply-chain dependency exploration, or identity resolution. The architecture question is the ownership boundary: which system is authoritative for which fact, how data is synchronized, what stale data is acceptable, and how failure or rollback works.
| Workload question | Relational | Document | Search | Neo4j / graph |
|---|---|---|---|---|
| Join two well-indexed business tables with strong constraints | Usually natural | May duplicate/embed | Not its primary purpose | Possible, but graph value may be low |
| Load one aggregate-shaped catalog object | Possible | Usually natural | Useful as indexed projection | Possible, but avoid turning properties into a giant document |
| Rank products by tokenized text relevance | Possible with extensions | Possible with extensions | Usually natural | Use full-text/vector features when graph context adds value |
| Find customers connected through devices, addresses, orders, and products across several hops | Possible, but joins may become cumbersome as depth/path choices grow | Usually requires application-side reference chasing or duplication | Search can retrieve candidates but is not a general relationship traversal engine | Often natural when the modeled relationships are first-class and selectively traversed |
| Enforce a global financial invariant | Often strong | Depends on product/model | Usually not the system of record | Requires careful transaction/model scope; graph shape does not automatically solve invariants |
If AtlasMart only needs
customer_id -> orders and a relational index
already answers it predictably, moving the data to Neo4j can
add another datastore without adding value. Graph value
generally increases when connected structure, path selection,
changing depth, relationship semantics, or graph algorithms
are themselves part of the requirement.
Follow one Cypher request through the system
Consider a service asking for products bought by customers who
used the same device as customer C-1001. The
application creates or reuses a Neo4j driver.
The driver manages connection pooling and speaks
Bolt to Neo4j. A query executes inside a
transaction boundary. Neo4j parses the Cypher statement, builds
or reuses a plan where appropriate, chooses operators, locates
useful starting data through indexes or other access operators,
traverses relationships, filters rows, and returns records to
the client.
This is a declarative pipeline, not an instruction to scan every node. The query planner estimates cardinalities and costs; the execution plan is the chosen operator tree/pipeline. The page cache is Neo4j-managed memory used to cache store pages, while the JVM heap serves other runtime structures. A successful query proves that the server completed that transaction/request; it does not prove the query plan is scalable, the graph is correctly constrained, a backup exists, or retries are safe.
CYPHER 25MATCH (seed:Customer {customerId: $customerId})-[:USED_DEVICE]->(d:Device)<-[:USED_DEVICE]-(peer:Customer)MATCH (peer)-[:PLACED]->(:Order)-[:CONTAINS]->(p:Product)RETURN DISTINCT p.productId, p.nameORDER BY p.productIdLIMIT 20;
The parameter $customerId separates data from query
text. The relationship types constrain which edges can be
followed. DISTINCT controls duplicate projected
rows; it does not change the graph. LIMIT bounds
returned rows but does not, by itself, guarantee that the
preceding search space is cheap. Later chapters use
EXPLAIN and PROFILE to prove what the
planner and runtime actually did.
Minimal observable lab: prove which server answered
The mandatory path uses the official Community image because it
is free/local and reproducible. Docker Desktop on Windows/macOS
or Docker Engine on Linux can run the same single-line commands.
Host mappings are bound to 127.0.0.1 so this first
lab does not intentionally expose an unencrypted HTTP endpoint
or Bolt listener to the LAN. The password below is a disposable
course credential; never reuse it outside this sandbox.
docker volume create atlasmart-neo4j-datadocker volume create atlasmart-neo4j-logsdocker run -d --name atlasmart-neo4j -p 127.0.0.1:7474:7474 -p 127.0.0.1:7687:7687 -e NEO4J_AUTH=neo4j/atlasmart-course-2026 -v atlasmart-neo4j-data:/data -v atlasmart-neo4j-logs:/logs neo4j:2026.07.1docker logs --tail 120 atlasmart-neo4j
Wait until the logs indicate the server is ready before diagnosing a failed client connection. Then inspect the process and version surfaces separately:
docker ps --filter name=atlasmart-neo4jdocker exec atlasmart-neo4j java -versiondocker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p atlasmart-course-2026 "CALL dbms.components() YIELD name, versions, edition RETURN name, versions, edition;"docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p atlasmart-course-2026 "SHOW SETTINGS 'db.query.default_language' YIELD name, value RETURN name, value;"docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p atlasmart-course-2026 "SHOW CURRENT USER;"docker exec atlasmart-neo4j sh -lc 'printf "NEO4J_HOME=%s\n" "$NEO4J_HOME"; grep -E "^[[:space:]]*db.query.default_language" "$NEO4J_HOME/conf/neo4j.conf" || true'
For this pinned image, the expected invariant is a Community
server in the 2026.07.1 line, a compatible Java
21/25 runtime supplied by the image, and a distributed
configuration that explicitly selects Cypher 25 for new
databases. Exact JVM vendor/build strings can differ. If a
setting is absent in a different deployment, remember that the
semantic default and a distributed configuration file’s explicit
value are different facts.
Create the smallest AtlasMart graph
CYPHER 25 CREATE CONSTRAINT customer_id IF NOT EXISTS FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;CYPHER 25 CREATE CONSTRAINT product_id IF NOT EXISTS FOR (p:Product) REQUIRE p.productId IS UNIQUE;CYPHER 25 MERGE (c:Customer {customerId:'C-1001'}) SET c.name='Mina Rahimi' MERGE (p:Product {productId:'P-1001'}) SET p.name='Trail Camera', p.category='Cameras', p.price=129.90 MERGE (c)-[:VIEWED {at:datetime('2026-09-09T08:00:00Z')}]->(p);CYPHER 25 MATCH (c:Customer)-[r:VIEWED]->(p:Product) RETURN c.customerId, type(r), r.at, p.productId, p.name;
Run the four statements in Browser/Query or pass each statement
separately to cypher-shell. The result should
contain one customer-to-product relationship. That proves graph
state and query execution. It does not prove that
VIEWED is the correct production model, that the
graph is large-scale performant, or that the current Community
deployment is highly available.
Controlled failures make the boundary visible
docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p definitely-wrong "RETURN 1;"docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:9999 -u neo4j -p atlasmart-course-2026 "RETURN 1;"
The first failure exercises authentication; the second exercises reachability/endpoint selection. Neither is evidence of a bad Cypher query. Separating connection, authentication, authorization, query, and data-model failures is a production habit worth establishing immediately.
Deliberately wrong approach: choose the technology from the diagram
A practitioner may see an AtlasMart entity-relationship diagram, notice many lines, and conclude “this is a graph, therefore use Neo4j.” That skips the workload. A normalized relational schema is already a graph in the mathematical sense of connected records; what matters is how frequently the application must navigate those connections, how dynamic the path shapes are, what invariants must be enforced, and what operational platform the team can support.
A second failure is operational: using
neo4j:latest for the course and assuming every
learner receives the same Cypher semantics. The moving tag can
change the server release, defaults, deprecations, plugin
compatibility, and JVM. Repair both errors by writing an
access-pattern decision record and pinning the server and driver
versions. Then verify them at runtime.
Production judgment and migration boundaries
For AtlasMart, a credible graph decision names the target questions, expected graph size, node/relationship degree distribution, write rate, path depth, selectivity of starting predicates, consistency requirements, transaction scope, latency objectives, security boundaries, and data ownership. High-degree “supernodes” are not automatically wrong, but they change traversal fan-out and must be measured. An index can make finding the starting customer cheap while a poorly bounded traversal after that starting point remains expensive.
Neo4j Community is appropriate for this learning lab and single-instance use cases, but it does not prove clustered availability, online backup, or enterprise RBAC. Aura changes the operational-responsibility boundary; Enterprise changes available self-managed capabilities; Infinigraph changes the scale architecture further. Those choices come next. If AtlasMart migrates an existing workload, plan dual-read/write or backfill validation, data reconciliation, cutover criteria, and rollback. A graph model that cannot be verified against the source of truth is not production-ready merely because queries are expressive.
Check your understanding
- Why is “our data has relationships” insufficient justification for adopting Neo4j?
- What is first-class about a Neo4j relationship that a foreign-key scalar alone does not represent?
-
What does a successful
CALL dbms.components()prove, and what does it not prove? - Why can an indexed starting node still lead to a slow graph query?
-
Why does pinning
neo4j:{SERVER}matter for a reproducible course?
Review the answers
1. Almost every business dataset has relationships. The decision depends on whether connected traversal/pattern workloads, relationship semantics, and the graph operational model materially improve the target system compared with existing stores.
2. A Neo4j relationship is stored as a typed, directed graph entity connecting two nodes and can carry its own properties. It participates directly in pattern matching and traversal.
3. It identifies the responding DBMS component/version/edition. It does not prove query scalability, backups, clustering, authorization design, data correctness, or a suitable production topology.
4. An index can cheaply locate the start but cannot eliminate fan-out created by following many matching relationships or paths. Cardinality through the whole plan matters.
5. It freezes the server baseline. A moving tag can silently change defaults, Cypher behavior, security fixes, plugins, JVM compatibility, and expected output.
Summary and next step
A Neo4j graph is a transactional property graph, not a generic replacement for relational, document, or search systems. You should now be able to state what relationships make first-class, trace a request through the DBMS, prove the exact server/Cypher baseline, and reject workloads whose connected-access needs do not justify the operational trade.
Next, compare Community, Enterprise, Aura, Desktop’s Developer license, and Infinigraph without treating them as interchangeable packaging choices.
Authoritative references
- Current Neo4j versions — Official current-release and LTS patch snapshot.
- Neo4j Operations Manual — Authoritative self-managed operational documentation for the current release.
- Neo4j system requirements — Supported operating systems and JVM requirements.
- Cypher Manual — Current Cypher 25 reference and language semantics.
- Configure the Cypher default version — Cypher 5 versus Cypher 25 default and override behavior.
- Neo4j in Docker — Official image tags, ports, editions, and Docker starting point.
- Cypher and Neo4j editions — Edition-aware Cypher capability boundaries.
- Neo4j ports — Official Bolt/HTTP/HTTPS and administrative port reference.
- Built-in procedures — Current procedure availability including dbms.components().