Chapter 02 · Property Graph Model: Nodes, Relationships, Labels, Properties, Types, and Schema Discipline
Translate a Relational/Document Domain into a Property Graph and Defend Every Modeling Decision
Prove the model by translation: turn source records into a connected graph only where identity, lifecycle, traversal, and workload evidence justify it.
Learning outcomes
Modeling discipline is proven by translation, not vocabulary quizzes. AtlasMart will take a small relational/document source shape—customers, orders, order lines, catalog products, categories, suppliers, stores, inventory, and shipments—and decide which facts become nodes, relationships, or properties. Every decision must point back to identity, cardinality, traversal questions, lifecycle, integrity, and migration cost.
Translate tables/documents into a property graph without copying source storage structure table-for-table.
Choose node versus property versus relationship using identity, lifecycle, cardinality, and query patterns.
Preserve stable source keys while preventing opaque foreign-key strings from becoming the only connection representation.
Validate graph cardinalities, missing properties, relationship endpoints, and duplicate identity after import.
Defend where Neo4j adds value and where relational/document/search systems should remain authoritative.
Continue Chapter 01's free/local baseline: Neo4j Community
2026.07.1, database neo4j, explicit
CYPHER 25 in version-sensitive examples, local
container atlasmart-neo4j, Bolt
127.0.0.1:7687, HTTP 127.0.0.1:7474,
and disposable password atlasmart-course-2026.
Existing AtlasMart identifiers use Community-supported
uniqueness constraints. This chapter adds
:Category, :Supplier,
:Store, and :Shipment identities and
makes graph-model contracts observable.
The generation environment does not provide a running Docker daemon or Neo4j server, so commands are documentation-checked but not presented as captured output. Community supports the uniqueness constraints used in the mandatory lab. Current property-existence, property-type, and key constraints are Enterprise-only; when those stronger contracts are discussed, the Community path uses validation queries plus application/import checks rather than fabricating successful Community DDL.
Start with questions, not boxes and arrows
AtlasMart's source systems expose relational records such as
customers, orders,
order_lines, products,
suppliers, stores,
inventory, and shipments, plus
document-oriented catalog attributes. A weak graph migration
creates one node per table row and one relationship per foreign
key without asking whether those elements matter to connected
queries. A strong migration begins with questions.
| AtlasMart question | Graph implication | Keep elsewhere too? |
|---|---|---|
| Which products are in the same categories and share suppliers? |
Product-IN_CATEGORY-Category and
Product-SUPPLIED_BY-Supplier
|
Catalog/document/search may remain authoritative for rich product content. |
| Which customers are connected to orders containing products supplied by the same supplier? | Customer→Order→Product→Supplier traversal | Financial/order system may remain relational source of record. |
| Where is a product stocked and what is current on-hand? |
Product-STOCKED_AT-Store with relationship
properties
|
Inventory service may remain operational authority; graph may be a projection. |
| What shipment fulfilled an order and where is it going? | Order→Shipment plus destination/customer/location connection | Logistics source can remain authoritative; graph captures connected context. |
| Render one product detail page | Graph may not add value to the basic aggregate read | Document/search/catalog service may be simpler. |
Classify source fields by graph role
Use four questions for each source field: does it have
independent identity; does it need independent lifecycle; will
queries traverse to/from it; does the value describe an entity
or a connection? Product.name is a property.
Category becomes a node because AtlasMart traverses
categories, categories have their own identity/name hierarchy
potential, and many products connect to one category.
order_lines.quantity becomes a
CONTAINS.quantity relationship property because it
describes a specific order-product pair.
Not every relational join table becomes a node. A pure two-party association with a few connection-specific properties often maps naturally to a relationship. But if the join row has its own ID, approvals, documents, state transitions, temporal history, or connects more than two parties, an intermediate node can be more faithful.
| Source concept | Graph choice | Reason |
|---|---|---|
| Customer | Node :Customer |
Independent identity and many connected questions |
| Product | Node :Product |
Independent identity; connected to orders/categories/suppliers/stores |
| Category | Node :Category |
Shared classification with traversal/hierarchy potential |
| Order line | Relationship CONTAINS with quantity |
Fact about an Order–Product pair in this simplified model |
| Inventory row | Relationship STOCKED_AT with onHand |
Fact about Product–Store pair |
| Shipment | Node :Shipment |
Independent tracking/lifecycle/status |
| Product description blob | Property or external document/search projection | No independent graph identity unless chunk/entity retrieval later requires it |
Build the target AtlasMart graph deterministically
CYPHER 25 CREATE CONSTRAINT customer_id IF NOT EXISTS FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;CYPHER 25 CREATE CONSTRAINT product_id IF NOT EXISTS FOR (p:Product) REQUIRE p.productId IS UNIQUE;CYPHER 25 CREATE CONSTRAINT order_id IF NOT EXISTS FOR (o:Order) REQUIRE o.orderId IS UNIQUE;CYPHER 25 CREATE CONSTRAINT category_id IF NOT EXISTS FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;CYPHER 25 CREATE CONSTRAINT supplier_id IF NOT EXISTS FOR (s:Supplier) REQUIRE s.supplierId IS UNIQUE;CYPHER 25 CREATE CONSTRAINT store_id IF NOT EXISTS FOR (s:Store) REQUIRE s.storeId IS UNIQUE;CYPHER 25 CREATE CONSTRAINT shipment_id IF NOT EXISTS FOR (s:Shipment) REQUIRE s.shipmentId IS UNIQUE;
CYPHER 25MERGE (c:Customer {customerId:'C-1001'}) SET c.name='Mina Rahimi'MERGE (o:Order {orderId:'O-5001'}) SET o.orderedAt=datetime('2026-09-08T16:30:00Z'), o.status='PAID'MERGE (p:Product {productId:'P-1001'}) SET p.name='Trail Camera', p.price=129.90MERGE (cat:Category {categoryId:'CAT-CAMERAS'}) SET cat.name='Cameras'MERGE (sup:Supplier {supplierId:'SUP-1001'}) SET sup.name='Northwind Optics'MERGE (st:Store {storeId:'ST-TEH-01'}) SET st.name='AtlasMart Central', st.location=point({latitude:35.6892,longitude:51.3890})MERGE (sh:Shipment {shipmentId:'SH-7001'}) SET sh.status='IN_TRANSIT', sh.dispatchedAt=datetime('2026-09-09T09:15:00Z')MERGE (c)-[:PLACED]->(o)MERGE (o)-[line:CONTAINS]->(p) SET line.quantity=1, line.unitPrice=129.90MERGE (p)-[:IN_CATEGORY]->(cat)MERGE (p)-[supply:SUPPLIED_BY]->(sup) SET supply.preferred=trueMERGE (p)-[stock:STOCKED_AT]->(st) SET stock.onHand=14, stock.observedAt=datetime('2026-09-09T10:00:00Z')MERGE (o)-[:FULFILLED_BY]->(sh)MERGE (sh)-[:SHIPPED_TO]->(c);
The FULFILLED_BY type is introduced deliberately in
Chapter 02 because Shipment needs a semantic link to Order; it
is not a generic RELATED_TO edge.
SHIPPED_TO points to Customer in this compact
fixture. A richer logistics model might instead connect Shipment
to a :Location or immutable delivery-address node
so address-at-order-time is not confused with the customer's
current profile address.
Prove the graph answers connected questions
CYPHER 25MATCH (c:Customer {customerId:'C-1001'})-[:PLACED]->(o:Order)-[line:CONTAINS]->(p:Product)MATCH (p)-[:IN_CATEGORY]->(cat:Category)MATCH (p)-[:SUPPLIED_BY]->(sup:Supplier)OPTIONAL MATCH (p)-[stock:STOCKED_AT]->(st:Store)OPTIONAL MATCH (o)-[:FULFILLED_BY]->(sh:Shipment)RETURN o.orderId, p.productId, line.quantity, cat.name AS category, sup.name AS supplier, collect({storeId:st.storeId,onHand:stock.onHand}) AS inventory, sh.shipmentId, sh.status;
The query is valuable because the connected structure is the
question. It would be poor evidence for choosing Neo4j if
AtlasMart only needed
Product.productId → Product.name. Graph adoption
should be defended at workload boundaries, not by demonstrating
that Cypher can reproduce a relational join.
Deliberately wrong translation: one node per source row
A mechanical migration might create
:OrderLine nodes for every relational order-line
row, :InventoryRow nodes for every inventory row,
and :ProductCategoryJoin nodes for every
product-category row even when those rows have no independent
identity or lifecycle. That graph is not automatically
incorrect, but it can add extra hops and complexity without
modeling value.
CYPHER 25MATCH (o:Order)-[line:CONTAINS]->(p:Product)RETURN o.orderId, p.productId, line.quantity, line.unitPrice;CYPHER 25MATCH (p:Product)-[stock:STOCKED_AT]->(st:Store)RETURN p.productId, st.storeId, stock.onHand, stock.observedAt;
Keep an intermediate node when the source row genuinely has entity semantics—for example, an order line that receives its own fulfillment status, returns, tax allocations, serial-number assignments, or references from other systems. The decision is query/lifecycle-driven, not dogmatic “relationships are always better.”
Integrity audit after translation
CYPHER 25MATCH (n)WHERE any(label IN labels(n) WHERE label IN ['Customer','Product','Order','Category','Supplier','Store','Shipment'])RETURN labels(n) AS labels, count(*) AS nodes ORDER BY labels;CYPHER 25MATCH ()-[r]->()RETURN type(r) AS relationshipType, count(*) AS relationships ORDER BY relationshipType;CYPHER 25MATCH (o:Order)OPTIONAL MATCH (c:Customer)-[:PLACED]->(o)WITH o, count(c) AS placingCustomersWHERE placingCustomers <> 1RETURN o.orderId, placingCustomers;CYPHER 25MATCH (p:Product)WHERE p.productId IS NULL OR p.name IS NULL OR (p.price IS NOT NULL AND NOT (p.price IS :: INTEGER | FLOAT))RETURN p.productId, p.name, p.price, valueType(p.price) AS priceType;
Add endpoint validations when the graph can be written by
multiple producers. For example, find any
SUPPLIED_BY relationship whose start node lacks
:Product or end node lacks :Supplier.
Labels/types classify but do not by themselves stop a writer
from creating a semantically wrong edge between other nodes.
CYPHER 25MATCH (a)-[r:SUPPLIED_BY]->(b)WHERE NOT a:Product OR NOT b:SupplierRETURN elementId(r) AS relationshipElementId, labels(a) AS startLabels, labels(b) AS endLabels;CYPHER 25MATCH (a)-[r:STOCKED_AT]->(b)WHERE NOT a:Product OR NOT b:StoreRETURN elementId(r) AS relationshipElementId, labels(a) AS startLabels, labels(b) AS endLabels;
Migration and rollback judgment
Do not replace an authoritative relational or document system merely because the graph projection works. A staged production migration needs source-of-truth ownership, backfill checkpoints, dual-write or CDC strategy, replay/idempotency rules, reconciliation counts, cutover criteria, and rollback. The graph may initially be a read-optimized connected projection. If it later becomes authoritative for selected facts, that is a separate architecture decision with transaction, availability, backup, and operational consequences.
For AtlasMart, Chapter 02's model now has explicit nodes where identity/lifecycle/traversal justify them, relationship properties where facts belong to a pair, stable domain IDs, temporal types, and validations for cardinality and property drift. Chapter 03 can teach MATCH/RETURN/WHERE against a graph whose semantics are known rather than accidental.
Check your understanding
- Why should a relational join table not automatically become a graph node?
- Why is Category a node in this AtlasMart model instead of only Product.category text?
- What justifies quantity as a CONTAINS relationship property?
- What does the endpoint validation for SUPPLIED_BY catch?
- Why might Neo4j remain a projection rather than the source of truth after this chapter?
Review the answers
1. Graph shape should follow identity, lifecycle, cardinality, and query semantics. A pure association may be better represented directly as a relationship.
2. Categories have stable identity, are shared across products, and may participate in traversals/hierarchies; those needs exceed a duplicated scalar string.
3. Quantity describes a specific Order–Product pair, not the order or product independently.
4. Relationships of the right type attached to semantically wrong endpoint labels, which labels/types alone do not automatically forbid.
5. The original systems may still own transactional/business facts; proving connected-query value does not by itself justify changing authority, operational responsibility, migration risk, or rollback architecture.
Summary and bridge to Chapter 03
Chapter 02 turned Neo4j's property graph into an explicit
AtlasMart contract: stable identities, meaningful
directions/types, legal property values, null semantics,
cardinality validation, relationship-local facts, temporal
validity, and a defensible source-to-graph translation. Chapter
03 now focuses on reading this known graph precisely with
MATCH, RETURN, WHERE,
parameters, ordering, and pagination.
Authoritative references
- Current Neo4j versions — Official current-release and 5.26 LTS patch snapshot.
- Cypher Manual — Current Cypher language reference.
- Property, structural, and constructed values — Current property/storage type rules, including lists and maps.
- Working with null — Current missing-property and three-valued null semantics.
- Create constraints — Constraint syntax, integrity semantics, backing indexes, and edition boundaries.
- Scalar functions: elementId() — Current elementId() guarantees and warning against durable application identity.
- MATCH — Directed and undirected relationship-pattern matching semantics.
- CREATE — Node and relationship creation semantics.
- Naming rules and recommendations — Current naming recommendations and case sensitivity.
- Basic Cypher queries — Official data-model and basic pattern-query examples used to ground source-to-graph translation.
- MERGE — Match-or-create behavior and the role of constraints for identity under concurrent writes.