Chapter 01 · Graph Database Foundations, Neo4j Editions, Deployment Choices, and Lab Setup

Graph Databases vs Relational, Document, and Search Systems: Connectivity, Traversal, and Workload Fit

Build the graph mental model from workload evidence: connected patterns, first-class relationships, transaction/query boundaries, and observable Neo4j runtime state.

Beginner → Intermediate110–130 minutesPinned single-instance evidence labNeo4j 2026.07.1 Community · Cypher 25Last reviewed: September 2026

Learning outcomes

AtlasMart already uses relational tables for orders, a document store for flexible catalog payloads, and a search engine for product discovery. A new fraud-and-recommendation initiative asks questions such as “which customers share devices, addresses, payment instruments, and products with this suspicious account within three relationship hops?” The architecture decision is not whether graphs look intuitive on a whiteboard. It is whether connected access patterns are important enough that representing and traversing relationships as first-class data produces a simpler, more predictable system than repeatedly reconstructing those connections elsewhere.

01

Explain the labeled property graph model and distinguish a graph database from a generic visualization or network data structure.

02

Compare relational, document, search, and graph systems by the work they make cheap or explicit rather than by slogans such as “NoSQL scales.”

03

Trace a small Cypher read from client connection through transaction, planning, operators, indexes/page cache/store access, and returned rows.

04

Define node, relationship, label, relationship type, property, pattern, Cypher, transaction, index, execution plan, driver, Bolt, and DBMS before using them operationally.

05

Prove which Neo4j server and Cypher configuration answered a request using observable runtime evidence.

Version baseline reviewed 9 September 2026

The mandatory lab pins Neo4j Community 2026.07.1. Neo4j also maintains the 5.26.30 LTS line. For the current 2026 line, supported JVMs are Java 21 and 25 on supported platforms. Neo4j 2026.02+ distribution configuration explicitly sets db.query.default_language=CYPHER_25 for new deployments, but existing deployments can retain Cypher 5 defaults. Every lab below asks the server what it is actually using instead of inferring behavior from the calendar version.

Execution and safety note

These commands were authored against current official Neo4j documentation and pinned versions, but this generation environment does not provide a running Docker daemon or Neo4j server, so the chapter does not pretend that example output was captured here. Expected output is described by invariant and shape. The lab binds HTTP and Bolt to host loopback, uses course-specific volumes and a synthetic graph, and deliberately avoids production credentials or data.

The property graph is about connected facts

Neo4j stores a labeled property graph. A node represents an entity or event-like thing. A node can have one or more labels, which classify it for modeling, querying, constraints, and indexes. A relationship connects exactly two nodes and has a stored direction and one relationship type. Nodes and relationships can carry properties: named values such as a customer ID, product name, amount, timestamp, or confidence score.

The critical difference from “an application that happens to contain IDs” is that the relationship itself is first-class database state. AtlasMart can store (Customer)-[:PLACED]->(Order), (Order)-[:CONTAINS]->(Product), and (Customer)-[:USED_DEVICE]->(Device). A traversal can then follow those stored relationships. Direction is part of relationship identity, although Cypher may intentionally match a relationship without specifying direction when the query means either direction. Labels and types describe structure; they do not by themselves enforce identity, uniqueness, or every business invariant. Later chapters add constraints deliberately.

Term Precise meaning in this chapter Common misconception
DBMS The Neo4j database-management process and administrative boundary A single graph visualization tab
Database A transactional data-management domain inside the DBMS A directory that applications may mutate directly
Node A graph entity with zero or more labels and properties A JSON document that must contain all related data
Relationship A directed, typed connection between two nodes, optionally with properties A foreign-key value that has no independent semantics
Pattern A shape described in Cypher and matched against graph data A procedural loop that manually walks every node
Cypher Neo4j’s declarative graph query language A generic name for every graph API
Transaction An ACID unit in which reads/writes are committed or rolled back A request retry policy in the driver
Index An access structure that can help find starting entities/properties A structure that makes every traversal constant-time
Execution plan The operators Neo4j chooses to execute a Cypher statement The textual query itself
Driver / Bolt The client library and Neo4j binary protocol used to communicate with the DBMS A guarantee that every retry is safe

Graph, relational, document, and search systems optimize different work

A relational database is usually strongest when normalized records, joins, constraints, set-oriented processing, and transactional invariants are the center of the workload. A document database is often attractive when an aggregate can be loaded and changed as a self-contained document with flexible nested shape. A search engine is optimized for retrieval over indexed text, terms, scoring, faceting, and related search primitives. A graph database earns its place when relationships and multi-step connected patterns are core application questions, not merely when data can be drawn as circles and arrows.

These systems are not mutually exclusive. AtlasMart may keep order-of-record accounting in a relational database, catalog text search in a search engine, and use Neo4j for fraud neighborhoods, recommendation paths, supply-chain dependency exploration, or identity resolution. The architecture question is the ownership boundary: which system is authoritative for which fact, how data is synchronized, what stale data is acceptable, and how failure or rollback works.

Workload question Relational Document Search Neo4j / graph
Join two well-indexed business tables with strong constraints Usually natural May duplicate/embed Not its primary purpose Possible, but graph value may be low
Load one aggregate-shaped catalog object Possible Usually natural Useful as indexed projection Possible, but avoid turning properties into a giant document
Rank products by tokenized text relevance Possible with extensions Possible with extensions Usually natural Use full-text/vector features when graph context adds value
Find customers connected through devices, addresses, orders, and products across several hops Possible, but joins may become cumbersome as depth/path choices grow Usually requires application-side reference chasing or duplication Search can retrieve candidates but is not a general relationship traversal engine Often natural when the modeled relationships are first-class and selectively traversed
Enforce a global financial invariant Often strong Depends on product/model Usually not the system of record Requires careful transaction/model scope; graph shape does not automatically solve invariants
Boundary case: one hop is not automatically a graph problem.

If AtlasMart only needs customer_id -> orders and a relational index already answers it predictably, moving the data to Neo4j can add another datastore without adding value. Graph value generally increases when connected structure, path selection, changing depth, relationship semantics, or graph algorithms are themselves part of the requirement.

Follow one Cypher request through the system

Consider a service asking for products bought by customers who used the same device as customer C-1001. The application creates or reuses a Neo4j driver. The driver manages connection pooling and speaks Bolt to Neo4j. A query executes inside a transaction boundary. Neo4j parses the Cypher statement, builds or reuses a plan where appropriate, chooses operators, locates useful starting data through indexes or other access operators, traverses relationships, filters rows, and returns records to the client.

This is a declarative pipeline, not an instruction to scan every node. The query planner estimates cardinalities and costs; the execution plan is the chosen operator tree/pipeline. The page cache is Neo4j-managed memory used to cache store pages, while the JVM heap serves other runtime structures. A successful query proves that the server completed that transaction/request; it does not prove the query plan is scalable, the graph is correctly constrained, a backup exists, or retries are safe.

Cypher · connected pattern expressed declaratively
CYPHER 25MATCH (seed:Customer {customerId: $customerId})-[:USED_DEVICE]->(d:Device)<-[:USED_DEVICE]-(peer:Customer)MATCH (peer)-[:PLACED]->(:Order)-[:CONTAINS]->(p:Product)RETURN DISTINCT p.productId, p.nameORDER BY p.productIdLIMIT 20;

The parameter $customerId separates data from query text. The relationship types constrain which edges can be followed. DISTINCT controls duplicate projected rows; it does not change the graph. LIMIT bounds returned rows but does not, by itself, guarantee that the preceding search space is cheap. Later chapters use EXPLAIN and PROFILE to prove what the planner and runtime actually did.

Minimal observable lab: prove which server answered

The mandatory path uses the official Community image because it is free/local and reproducible. Docker Desktop on Windows/macOS or Docker Engine on Linux can run the same single-line commands. Host mappings are bound to 127.0.0.1 so this first lab does not intentionally expose an unencrypted HTTP endpoint or Bolt listener to the LAN. The password below is a disposable course credential; never reuse it outside this sandbox.

shell · create persistent course volumes and start pinned Community
docker volume create atlasmart-neo4j-datadocker volume create atlasmart-neo4j-logsdocker run -d --name atlasmart-neo4j -p 127.0.0.1:7474:7474 -p 127.0.0.1:7687:7687 -e NEO4J_AUTH=neo4j/atlasmart-course-2026 -v atlasmart-neo4j-data:/data -v atlasmart-neo4j-logs:/logs neo4j:2026.07.1docker logs --tail 120 atlasmart-neo4j

Wait until the logs indicate the server is ready before diagnosing a failed client connection. Then inspect the process and version surfaces separately:

shell · DBMS, JVM, Cypher default, user and connector evidence
docker ps --filter name=atlasmart-neo4jdocker exec atlasmart-neo4j java -versiondocker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p atlasmart-course-2026 "CALL dbms.components() YIELD name, versions, edition RETURN name, versions, edition;"docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p atlasmart-course-2026 "SHOW SETTINGS 'db.query.default_language' YIELD name, value RETURN name, value;"docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p atlasmart-course-2026 "SHOW CURRENT USER;"docker exec atlasmart-neo4j sh -lc 'printf "NEO4J_HOME=%s\n" "$NEO4J_HOME"; grep -E "^[[:space:]]*db.query.default_language" "$NEO4J_HOME/conf/neo4j.conf" || true' 

For this pinned image, the expected invariant is a Community server in the 2026.07.1 line, a compatible Java 21/25 runtime supplied by the image, and a distributed configuration that explicitly selects Cypher 25 for new databases. Exact JVM vendor/build strings can differ. If a setting is absent in a different deployment, remember that the semantic default and a distributed configuration file’s explicit value are different facts.

Create the smallest AtlasMart graph

Cypher · identity constraints and a tiny connected fixture
CYPHER 25 CREATE CONSTRAINT customer_id IF NOT EXISTS FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;CYPHER 25 CREATE CONSTRAINT product_id IF NOT EXISTS FOR (p:Product) REQUIRE p.productId IS UNIQUE;CYPHER 25 MERGE (c:Customer {customerId:'C-1001'}) SET c.name='Mina Rahimi' MERGE (p:Product {productId:'P-1001'}) SET p.name='Trail Camera', p.category='Cameras', p.price=129.90 MERGE (c)-[:VIEWED {at:datetime('2026-09-09T08:00:00Z')}]->(p);CYPHER 25 MATCH (c:Customer)-[r:VIEWED]->(p:Product) RETURN c.customerId, type(r), r.at, p.productId, p.name;

Run the four statements in Browser/Query or pass each statement separately to cypher-shell. The result should contain one customer-to-product relationship. That proves graph state and query execution. It does not prove that VIEWED is the correct production model, that the graph is large-scale performant, or that the current Community deployment is highly available.

Controlled failures make the boundary visible

shell · bad authentication and wrong endpoint, expected to fail
docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:7687 -u neo4j -p definitely-wrong "RETURN 1;"docker exec atlasmart-neo4j cypher-shell -a neo4j://localhost:9999 -u neo4j -p atlasmart-course-2026 "RETURN 1;"

The first failure exercises authentication; the second exercises reachability/endpoint selection. Neither is evidence of a bad Cypher query. Separating connection, authentication, authorization, query, and data-model failures is a production habit worth establishing immediately.

Deliberately wrong approach: choose the technology from the diagram

A practitioner may see an AtlasMart entity-relationship diagram, notice many lines, and conclude “this is a graph, therefore use Neo4j.” That skips the workload. A normalized relational schema is already a graph in the mathematical sense of connected records; what matters is how frequently the application must navigate those connections, how dynamic the path shapes are, what invariants must be enforced, and what operational platform the team can support.

A second failure is operational: using neo4j:latest for the course and assuming every learner receives the same Cypher semantics. The moving tag can change the server release, defaults, deprecations, plugin compatibility, and JVM. Repair both errors by writing an access-pattern decision record and pinning the server and driver versions. Then verify them at runtime.

Production judgment and migration boundaries

For AtlasMart, a credible graph decision names the target questions, expected graph size, node/relationship degree distribution, write rate, path depth, selectivity of starting predicates, consistency requirements, transaction scope, latency objectives, security boundaries, and data ownership. High-degree “supernodes” are not automatically wrong, but they change traversal fan-out and must be measured. An index can make finding the starting customer cheap while a poorly bounded traversal after that starting point remains expensive.

Neo4j Community is appropriate for this learning lab and single-instance use cases, but it does not prove clustered availability, online backup, or enterprise RBAC. Aura changes the operational-responsibility boundary; Enterprise changes available self-managed capabilities; Infinigraph changes the scale architecture further. Those choices come next. If AtlasMart migrates an existing workload, plan dual-read/write or backfill validation, data reconciliation, cutover criteria, and rollback. A graph model that cannot be verified against the source of truth is not production-ready merely because queries are expressive.

Check your understanding

  1. Why is “our data has relationships” insufficient justification for adopting Neo4j?
  2. What is first-class about a Neo4j relationship that a foreign-key scalar alone does not represent?
  3. What does a successful CALL dbms.components() prove, and what does it not prove?
  4. Why can an indexed starting node still lead to a slow graph query?
  5. Why does pinning neo4j:{SERVER} matter for a reproducible course?
Review the answers

1. Almost every business dataset has relationships. The decision depends on whether connected traversal/pattern workloads, relationship semantics, and the graph operational model materially improve the target system compared with existing stores.

2. A Neo4j relationship is stored as a typed, directed graph entity connecting two nodes and can carry its own properties. It participates directly in pattern matching and traversal.

3. It identifies the responding DBMS component/version/edition. It does not prove query scalability, backups, clustering, authorization design, data correctness, or a suitable production topology.

4. An index can cheaply locate the start but cannot eliminate fan-out created by following many matching relationships or paths. Cardinality through the whole plan matters.

5. It freezes the server baseline. A moving tag can silently change defaults, Cypher behavior, security fixes, plugins, JVM compatibility, and expected output.

Summary and next step

A Neo4j graph is a transactional property graph, not a generic replacement for relational, document, or search systems. You should now be able to state what relationships make first-class, trace a request through the DBMS, prove the exact server/Cypher baseline, and reject workloads whose connected-access needs do not justify the operational trade.

Next, compare Community, Enterprise, Aura, Desktop’s Developer license, and Infinigraph without treating them as interchangeable packaging choices.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.