Chapter 07 · Graph Data Modeling: Aggregates, Relationships, Hyperedges, Hierarchies, and Temporal Graphs
Review a Poor Graph Model and Refactor It for Traversal Clarity, Integrity, and Query Performance
Refactor a poor graph model using workload, invariant, cardinality and temporal evidence, then validate the migration path before production cutover.
Learning outcomes
The capstone begins with a deliberately poor supply/order model: source-table-shaped nodes, generic links, string foreign keys, duplicated current state and no history. Learners must defend each refactoring with a query, invariant or lifecycle requirement.
Identify table-for-table graph translation, generic-edge, opaque-FK and current-state-only anti-patterns.
Refactor toward stable domain nodes, explicit relationships and intermediate facts.
Compare traversal cardinality and query readability before and after refactoring.
Validate identity, DAG and temporal invariants with deterministic queries.
Plan a production migration with dual-read/write, backfill, verification and rollback boundaries.
The mandatory lab continues the accepted course baseline:
Neo4j Community 2026.07.1, database
neo4j, explicit CYPHER 25 for
version-sensitive examples, authentication enabled, no
mandatory APOC/GDS plugin, and stable AtlasMart domain
identifiers from Chapters 01–06. Neo4j
5.26.30 remains the LTS comparison line. Modeling
examples use only Community-compatible graph and uniqueness
features.
This generation environment does not run Neo4j or Docker.
Commands were checked against current official documentation
but were not executed here. Expected outputs are deterministic
fixture invariants, not fabricated captures. All
destructive/refactoring steps are scoped to Chapter 07
identifiers or labTag='ch07'; never replace them
with unconstrained production matches.
Re-establish the Chapter 07 modeling fixture
The lab deliberately creates a small supply-chain slice whose questions require explicit semantics: two suppliers, one product, one store, dated supply agreements, and a component hierarchy. Stable domain IDs remain the durable identity; internal element IDs are not used as business keys.
CYPHER 25CREATE CONSTRAINT supplier_id IF NOT EXISTS FOR (s:Supplier) REQUIRE s.supplierId IS UNIQUE;CREATE CONSTRAINT store_id IF NOT EXISTS FOR (s:Store) REQUIRE s.storeId IS UNIQUE;CREATE CONSTRAINT product_id IF NOT EXISTS FOR (p:Product) REQUIRE p.productId IS UNIQUE;CREATE CONSTRAINT category_id IF NOT EXISTS FOR (c:Category) REQUIRE c.categoryId IS UNIQUE;CREATE CONSTRAINT agreement_id IF NOT EXISTS FOR (a:SupplyAgreement) REQUIRE a.agreementId IS UNIQUE;CREATE CONSTRAINT component_id IF NOT EXISTS FOR (c:Component) REQUIRE c.componentId IS UNIQUE;
CYPHER 25MERGE (sup1:Supplier {supplierId:'S-7001'}) SET sup1.name='Northstar Components', sup1.labTag='ch07'MERGE (sup2:Supplier {supplierId:'S-7002'}) SET sup2.name='BlueRiver Plastics', sup2.labTag='ch07'MERGE (store:Store {storeId:'ST-7001'}) SET store.name='AtlasMart Central', store.labTag='ch07'MERGE (prod:Product {productId:'P-7001'}) SET prod.name='Trail Camera Kit', prod.labTag='ch07'MERGE (cat:Category {categoryId:'CAT-7001'}) SET cat.name='Outdoor Imaging', cat.labTag='ch07'MERGE (prod)-[:IN_CATEGORY]->(cat)MERGE (a1:SupplyAgreement {agreementId:'SA-7001'})SET a1.validFrom=date('2026-01-01'), a1.validTo=date('2026-07-01'), a1.recordedFrom=datetime('2026-01-02T09:00:00Z'), a1.recordedTo=datetime('9999-12-31T00:00:00Z'), a1.unitPrice=84.0, a1.currency='USD', a1.labTag='ch07'MERGE (sup1)-[:PARTY_TO]->(a1)MERGE (a1)-[:SUPPLIES]->(prod)MERGE (a1)-[:DELIVERS_TO]->(store)MERGE (a2:SupplyAgreement {agreementId:'SA-7002'})SET a2.validFrom=date('2026-07-01'), a2.validTo=date('2027-01-01'), a2.recordedFrom=datetime('2026-06-15T10:00:00Z'), a2.recordedTo=datetime('9999-12-31T00:00:00Z'), a2.unitPrice=79.0, a2.currency='USD', a2.labTag='ch07'MERGE (sup2)-[:PARTY_TO]->(a2)MERGE (a2)-[:SUPPLIES]->(prod)MERGE (a2)-[:DELIVERS_TO]->(store);
CYPHER 25MERGE (kit:Component {componentId:'CMP-KIT'}) SET kit.name='Trail Camera Kit', kit.labTag='ch07'MERGE (cam:Component {componentId:'CMP-CAM'}) SET cam.name='Camera Module', cam.labTag='ch07'MERGE (case:Component {componentId:'CMP-CASE'}) SET case.name='Weather Case', case.labTag='ch07'MERGE (lens:Component {componentId:'CMP-LENS'}) SET lens.name='Lens Assembly', lens.labTag='ch07'MERGE (kit)-[:CONTAINS_COMPONENT {quantity:1}]->(cam)MERGE (kit)-[:CONTAINS_COMPONENT {quantity:1}]->(case)MERGE (cam)-[:CONTAINS_COMPONENT {quantity:1}]->(lens);
Use SHOW CONSTRAINTS and targeted counts before
refactoring. A successful query proves only the fixture state,
not that the model scales to production degree distributions or
workload volume.
1. Construct a deliberately poor model safely
CYPHER 25CREATE (x:LegacySupplierRow {rowId:'LS-1', supplier_name:'Northstar Components', labTag:'ch07'})CREATE (y:LegacyProductRow {rowId:'LP-1', product_name:'Trail Camera Kit', supplier_row_id:'LS-1', store_row_id:'LST-1', valid_from:'2026-01-01', labTag:'ch07'})CREATE (z:LegacyStoreRow {rowId:'LST-1', store_name:'AtlasMart Central', labTag:'ch07'})CREATE (x)-[:RELATED_TO {kind:'supplier-product'}]->(y)CREATE (y)-[:RELATED_TO {kind:'product-store'}]->(z);
The problem is not that these nodes are illegal. It is that queries must decode source schema artifacts rather than traverse domain semantics, date strings are untyped, and the supply fact has no durable identity.
2. Refactor by question inventory
| Question | Poor model cost | Refactored path |
|---|---|---|
| Who supplies this product to this store? | Decode RELATED_TO.kind and string IDs | Supplier→PARTY_TO→SupplyAgreement→SUPPLIES/Product + DELIVERS_TO/Store |
| What was valid on date D? | Parse string current fields | Predicate native validFrom/validTo |
| What components build the kit? | No hierarchy semantics | Component→CONTAINS_COMPONENT DAG |
| Can this fact be corrected/audited? | No fact identity/history | SupplyAgreement ID + recorded interval |
The explicit model is not automatically faster for every query;
it is semantically clearer and can be indexed/optimized around
stable anchors. Performance claims require
PROFILE on representative graph sizes and degree
distributions.
3. Compare before/after traversals and cardinality
CYPHER 25MATCH (s:Supplier)-[:PARTY_TO]->(a:SupplyAgreement)-[:SUPPLIES]->(p:Product {productId:$productId}), (a)-[:DELIVERS_TO]->(st:Store {storeId:$storeId})WHERE a.validFrom <= date($asOf) AND date($asOf) < a.validToRETURN s.supplierId,a.agreementId,a.unitPrice;
CYPHER 25PROFILEMATCH (s:Supplier)-[:PARTY_TO]->(a:SupplyAgreement)-[:SUPPLIES]->(p:Product {productId:'P-7001'}), (a)-[:DELIVERS_TO]->(st:Store {storeId:'ST-7001'})WHERE a.validFrom <= date('2026-05-01') AND date('2026-05-01') < a.validToRETURN s.supplierId,a.agreementId;
Record actual rows and database hits in your environment. This lesson does not invent plan numbers. If the graph becomes dense, selective anchors and appropriate indexes/constraints become increasingly important.
4. Validation gate after refactoring
CYPHER 25// stable-key duplicates should return no rowsMATCH (a:SupplyAgreement) WHERE a.labTag='ch07'WITH a.agreementId AS id, count(*) AS copies WHERE copies <> 1RETURN id,copies;// BOM cycles should be zeroMATCH p=(c:Component)-[:CONTAINS_COMPONENT*1..10]->(c)RETURN count(p) AS bomCycles;// expected supply snapshotsUNWIND [date('2026-05-01'),date('2026-07-01'),date('2026-12-31')] AS dMATCH (a:SupplyAgreement)-[:SUPPLIES]->(:Product {productId:'P-7001'})WHERE a.validFrom <= d AND d < a.validToRETURN d, count(a) AS agreements ORDER BY d;
For this fixture, the duplicate-key query should return no rows,
bomCycles should be zero, and each listed snapshot
should return one agreement. Those are fixture invariants, not
universal production rules.
5. Migration and production judgment
Refactoring a production graph is a data migration, not a diagram edit. Define stable mapping from old identities to new identities; backfill in bounded batches; dual-write or capture changes during migration; compare old/new read results; observe degree distributions, latency and error rates; and maintain a rollback window until the new model is proven. Shared concepts such as Customer, Product, Supplier and Category need governance so independent teams do not create incompatible synonyms or generic relationships.
| Decision surface | Production question |
|---|---|
| Semantic clarity | Can a reviewer infer relationship meaning from type and endpoints? |
| Degree distribution | Which nodes become hubs and how does fan-out affect tail latency? |
| Write complexity | Does reification/temporal history multiply write work acceptably? |
| Retention | How long must historical versions remain queryable? |
| Migration | Can old/new models coexist long enough for verification and rollback? |
| Governance | Who owns shared labels, relationship types and identifiers? |
Capstone lab and cleanup
Run the legacy fixture, answer the four business questions, then
answer them against the refactored model. Record query text,
result counts and PROFILE evidence. Run the
invariant suite, remove the deliberate overlap scratch agreement
if present, and delete only labTag='ch07' data when
you are ready to reset.
CYPHER 25MATCH (n) WHERE n.labTag='ch07' DETACH DELETE n;
Check your understanding
- Why is copying every relational table to a node label often weak graph modeling?
- What evidence should justify a refactor?
- Why can denormalization/precomputation be legitimate?
- What must a temporal migration preserve?
- When is the new model ready for cutover?
Review the answers
1. It preserves source storage shape instead of modeling domain identity and traversal questions.
2. Clearer semantics plus invariant, cardinality, query-result and representative plan/performance evidence.
3. If a measured hot query benefits and the write/update consistency cost is explicit and controlled.
4. Business-valid history, recorded/audit history if required, and stable mapping between old and new identities.
5. After backfill/change capture, old/new result comparison, invariant checks, representative workload tests and a defined rollback strategy.
Summary and next step
Chapter 07 turns graph modeling into an evidence-driven design discipline: questions define identity and semantics; degree predicts traversal cost; intermediate facts represent richer relationships; hierarchy rules define valid paths; and temporal requirements preserve history. Chapter 08 applies this model to importing data at different scales while validating the resulting graph.
Authoritative references
- Current Neo4j versions — Release/LTS snapshot used for the chapter baseline.
- Cypher Manual — patterns — Current property-graph pattern semantics.
- Temporal values — Native temporal value types that can be stored on nodes and relationships.
- Data modeling — Official modeling workflow and domain-to-graph guidance.
- PROFILE and EXPLAIN — Plan evidence for representative query shapes.
- LOAD CSV — Bridge to Chapter 08 import workflows.