Build an AtlasMart property graph from explicit identities, typed relationships, labels, properties, constraints, and observable adjacency state.
Nodes, Relationships/Edges, Labels, Properties, and Graph Schema Choices
Model relationship-heavy AtlasMart fraud data as a property graph without treating “schema-light” as “schema-free,” then prove how identity and edge choices change traversal correctness.
Learning outcomes
Model relationship-heavy AtlasMart fraud data as a property graph without treating “schema-light” as “schema-free,” then prove how identity and edge choices change traversal correctness.
Define nodes, labels/types, directed relationships/edges, and properties precisely.
Separate stable business identity from a database-internal node identifier.
Explain how edge modeling changes navigation, integrity checks, and relationship metadata.
Use uniqueness/validation contracts so schema flexibility does not become silent data corruption.
Mandatory work uses Python 3.13+ standard library only on a single local process. No graph database, Docker image, cloud account, paid feature, network manipulation, or destructive failure injection is required. Neo4j 2026.07.1 is an optional current implementation reference; Community Edition is GPLv3. Product-specific clustering, sharding, security, and enterprise features are not assumed by the lab.
1. Start with the business question, not a diagram aesthetic
AtlasMart fraud analysts do not begin with “which database is fashionable?” They ask relationship questions: which customers share a payment instrument, which orders arrived from the same device or IP address, and which accounts sit within two hops of a previously confirmed fraud case? A graph is a collection of vertices and edges. In a property graph, vertices are usually called nodes; connections are relationships or edges; nodes may have one or more labels (or analogous types); relationships have a type/direction; and both nodes and relationships may carry named properties. ISO/IEC 39075:2024 standardizes a property-graph data model and the Graph Query Language (GQL), but individual products still differ in storage, constraints, query planning, distribution, and operational semantics.
2. Identity is part of the model
A graph can be syntactically valid and still be semantically
wrong. AtlasMart's customer ID c17 is a business
identity. If an import creates two different nodes that both
claim to represent c17, half of the customer's
devices may attach to one node and half of the payment history
to the other. A traversal then returns a false negative even
though every edge exists. Production graph models therefore need
explicit identity policy: tenant scope, normalization,
uniqueness, merge rules, deletion semantics, and a migration
path when identifiers change. Indexes help locate an anchor
node; uniqueness constraints or equivalent validation protect
identity. Neither is optional merely because graph models can
evolve flexibly.
3. AtlasMart lab: make the mechanism observable
Save the following as lesson1_property_graph.py and
run it with python lesson1_property_graph.py. The
program has no dependencies and mutates no external state.
from collections import defaultdict
nodes = {}
adj = defaultdict(list)
def add_node(node_id, labels, **props):
if node_id in nodes:
raise ValueError(f"duplicate node identity: {node_id}")
nodes[node_id] = {"labels": set(labels), "props": props}
def add_edge(src, rel_type, dst, **props):
if src not in nodes or dst not in nodes:
raise KeyError("edge endpoint missing")
edge = {"type": rel_type, "to": dst, "props": props}
adj[src].append(edge)
add_node("cust:c17", ["Customer"], tenant="t1", name="Mina")
add_node("card:901", ["PaymentInstrument"], tenant="t1", fingerprint="fp-901")
add_node("dev:d3", ["Device"], tenant="t1", trusted=False)
add_node("order:o55", ["Order"], tenant="t1", amount=480)
add_node("ip:203.0.113.8", ["IPAddress"], tenant="t1")
add_edge("cust:c17", "USES_CARD", "card:901", first_seen="2026-08-01")
add_edge("cust:c17", "USES_DEVICE", "dev:d3", first_seen="2026-08-25")
add_edge("order:o55", "PLACED_BY", "cust:c17")
add_edge("order:o55", "FROM_IP", "ip:203.0.113.8")
add_edge("order:o55", "PAID_WITH", "card:901", authorized=True)
print("labels(customer):", sorted(nodes["cust:c17"]["labels"]))
print("customer adjacency:", [(e["type"], e["to"]) for e in adj["cust:c17"]])
print("order adjacency:", [(e["type"], e["to"]) for e in adj["order:o55"]])
# Deliberately wrong: business identity is not enforced.
bad_nodes = [
{"internal": 1, "customer_id": "c17", "name": "Mina"},
{"internal": 2, "customer_id": "c17", "name": "Mina A."},
]
print("bad duplicate business ids:", [n["customer_id"] for n in bad_nodes])
print("fraud neighborhood split across internal nodes:", len(bad_nodes))
# Repair: enforce a stable tenant-scoped business identity before adding edges.
identity_registry = set()
def claim_customer(tenant, customer_id):
key = (tenant, customer_id)
if key in identity_registry:
return False
identity_registry.add(key)
return True
print("claim t1/c17 first:", claim_customer("t1", "c17"))
print("claim t1/c17 duplicate:", claim_customer("t1", "c17"))
print("claim t2/c17 distinct tenant:", claim_customer("t2", "c17"))
Expected evidence: Customer c17 has two outgoing relationship types, order o55 connects to the customer, IP address, and payment instrument, the deliberately duplicated business ID splits one logical identity into two graph nodes, and the repaired tenant-scoped identity registry rejects the duplicate while allowing the same external ID in another tenant.
4. A relationship is data, not decorative syntax
Consider
Customer-[:USES_CARD]->PaymentInstrument. If the
association has facts such as first_seen,
verification source, confidence, or role, those facts naturally
belong to the relationship because they describe the connection.
Storing only card_id as a customer property loses
that first-class relationship identity and makes multi-card or
historical associations awkward. The opposite mistake is turning
every scalar attribute into a node. A product price, a
customer's display name, or an order status usually remains a
property unless the domain needs independent identity,
relationships, lifecycle, or shared reference semantics for that
fact.
5. Schema-light is not schema-free
Graph databases often permit nodes with different labels and property sets, but applications still impose a schema through labels, relationship types, property names/types, constraints, indexes, authorization rules, ingestion contracts, and query expectations. Neo4j's current documentation, for example, exposes indexes and multiple constraint types; that is implementation evidence that flexible graph structure does not eliminate schema governance. AtlasMart should version ingestion mappings, reject impossible edge endpoints, enforce tenant ownership, and test that migrations keep old and new readers compatible.
6. What the observable state proves—and does not prove
The lab prints labels, adjacency lists, and identity-claim results. That proves that a particular in-memory model has one stable customer identity and explicit typed edges. It does not prove durability, transactional isolation, replica consistency, authorization, index behavior, or Neo4j performance. Those are separate contracts. In a real graph deployment, add query plans, constraint status, transaction logs, replication/backup evidence, and tenant-aware authorization tests to the same verification mindset.
7. Production judgment
Choose a graph representation when relationships and multi-hop navigation are central to the workload, not merely because the domain contains relationships. Define which facts are authoritative, whether the graph itself is the system of record or a derived projection, what consistency is needed when edges are created/deleted, how tenant boundaries are enforced on both nodes and edges, and how identity reconciliation works. Relationship-heavy writes can increase contention around popular nodes; cross-region or sharded deployments add ownership and traversal costs; and backups must restore both topology and constraints. If a simple foreign key plus one indexed join serves the dominant workload more clearly, the relational model may remain the better operational choice.
Wrong approach: let imports invent duplicate graph identities
A team bulk-loads customers using generated internal IDs and
assumes matching customer_id properties will be
“close enough.” The result is two nodes for the same customer.
Fraud edges are split, so a two-hop search misses connections.
The safe repair is to define a tenant-scoped identity key,
validate uniqueness before/while loading, reconcile duplicates
with domain rules, redirect edges, and verify the resulting
neighborhood—not to blindly merge nodes that merely share a
display name.
Verification, cleanup, and production checklist
Verification is the program output plus the conceptual checks
below. Cleanup is simply deleting the local
lesson1_property_graph.py file; the simulation
creates no sockets, services, databases, containers,
credentials, or persistent data. In production, additionally
verify tenant authorization on traversals, identity/constraint
health, degree and path-cardinality distributions, p95/p99
latency, cache and remote-hop behavior where applicable,
backup/restore or projection rebuild, software/security
advisories, and edition/license constraints before adopting
product-specific features.
Check your understanding
- What makes a property graph different from a plain key-value map?
- Why is a stable business identity important?
- Does schema-light mean no schema?
- When should a fact become an edge?
- What does an index typically help with first?
Review the answers
1. It represents first-class nodes and relationships, with labels/types and properties, so topology itself is queryable data.
2. Because duplicate logical entities fragment the neighborhood and can produce false negatives or contradictory updates.
3. No. Labels/types, property contracts, constraints, indexes, authorization rules, and application expectations still form a schema.
4. When the association has independent meaning, multiplicity, lifecycle, properties, or is a central traversal target.
5. Locating an anchor node or relationship by a property; adjacency traversal after anchoring is a separate mechanism.
References
Foundational statements use the published property-graph standard where appropriate; implementation-sensitive examples use current official documentation and are labeled as examples rather than universal graph guarantees.
- ISO/IEC 39075:2024 — Database languages — GQL — published property-graph data structures and operations.
- Neo4j — What is a graph database — current property-graph terminology for nodes, labels, relationships, properties, indexes, and constraints.
- Neo4j Cypher Manual — Create constraints — current implementation reference for uniqueness and other graph constraints.
- Neo4j release notes — Neo4j 2026.07.1 is the current release as of the August 2026 review.
- Neo4j editions and licensing — Community Edition is GPLv3; enterprise capabilities are separate and not required by this lesson.