Build an AtlasMart property graph from explicit identities, typed relationships, labels, properties, constraints, and observable adjacency state.

Nodes, Relationships/Edges, Labels, Properties, and Graph Schema Choices

Model relationship-heavy AtlasMart fraud data as a property graph without treating “schema-light” as “schema-free,” then prove how identity and edge choices change traversal correctness.

Intermediate95–120 minutesProperty graph + identity labPython 3.13+ · standard libraryNeo4j 2026.07.1 optional referenceLast reviewed: August 2026

Learning outcomes

Model relationship-heavy AtlasMart fraud data as a property graph without treating “schema-light” as “schema-free,” then prove how identity and edge choices change traversal correctness.

01

Define nodes, labels/types, directed relationships/edges, and properties precisely.

02

Separate stable business identity from a database-internal node identifier.

03

Explain how edge modeling changes navigation, integrity checks, and relationship metadata.

04

Use uniqueness/validation contracts so schema flexibility does not become silent data corruption.

Implementation snapshot

Mandatory work uses Python 3.13+ standard library only on a single local process. No graph database, Docker image, cloud account, paid feature, network manipulation, or destructive failure injection is required. Neo4j 2026.07.1 is an optional current implementation reference; Community Edition is GPLv3. Product-specific clustering, sharding, security, and enterprise features are not assumed by the lab.

1. Start with the business question, not a diagram aesthetic

AtlasMart fraud analysts do not begin with “which database is fashionable?” They ask relationship questions: which customers share a payment instrument, which orders arrived from the same device or IP address, and which accounts sit within two hops of a previously confirmed fraud case? A graph is a collection of vertices and edges. In a property graph, vertices are usually called nodes; connections are relationships or edges; nodes may have one or more labels (or analogous types); relationships have a type/direction; and both nodes and relationships may carry named properties. ISO/IEC 39075:2024 standardizes a property-graph data model and the Graph Query Language (GQL), but individual products still differ in storage, constraints, query planning, distribution, and operational semantics.

2. Identity is part of the model

A graph can be syntactically valid and still be semantically wrong. AtlasMart's customer ID c17 is a business identity. If an import creates two different nodes that both claim to represent c17, half of the customer's devices may attach to one node and half of the payment history to the other. A traversal then returns a false negative even though every edge exists. Production graph models therefore need explicit identity policy: tenant scope, normalization, uniqueness, merge rules, deletion semantics, and a migration path when identifiers change. Indexes help locate an anchor node; uniqueness constraints or equivalent validation protect identity. Neither is optional merely because graph models can evolve flexibly.

3. AtlasMart lab: make the mechanism observable

Save the following as lesson1_property_graph.py and run it with python lesson1_property_graph.py. The program has no dependencies and mutates no external state.

python · AtlasMart deterministic simulation
from collections import defaultdict

nodes = {}
adj = defaultdict(list)

def add_node(node_id, labels, **props):
    if node_id in nodes:
        raise ValueError(f"duplicate node identity: {node_id}")
    nodes[node_id] = {"labels": set(labels), "props": props}

def add_edge(src, rel_type, dst, **props):
    if src not in nodes or dst not in nodes:
        raise KeyError("edge endpoint missing")
    edge = {"type": rel_type, "to": dst, "props": props}
    adj[src].append(edge)

add_node("cust:c17", ["Customer"], tenant="t1", name="Mina")
add_node("card:901", ["PaymentInstrument"], tenant="t1", fingerprint="fp-901")
add_node("dev:d3", ["Device"], tenant="t1", trusted=False)
add_node("order:o55", ["Order"], tenant="t1", amount=480)
add_node("ip:203.0.113.8", ["IPAddress"], tenant="t1")

add_edge("cust:c17", "USES_CARD", "card:901", first_seen="2026-08-01")
add_edge("cust:c17", "USES_DEVICE", "dev:d3", first_seen="2026-08-25")
add_edge("order:o55", "PLACED_BY", "cust:c17")
add_edge("order:o55", "FROM_IP", "ip:203.0.113.8")
add_edge("order:o55", "PAID_WITH", "card:901", authorized=True)

print("labels(customer):", sorted(nodes["cust:c17"]["labels"]))
print("customer adjacency:", [(e["type"], e["to"]) for e in adj["cust:c17"]])
print("order adjacency:", [(e["type"], e["to"]) for e in adj["order:o55"]])

# Deliberately wrong: business identity is not enforced.
bad_nodes = [
    {"internal": 1, "customer_id": "c17", "name": "Mina"},
    {"internal": 2, "customer_id": "c17", "name": "Mina A."},
]
print("bad duplicate business ids:", [n["customer_id"] for n in bad_nodes])
print("fraud neighborhood split across internal nodes:", len(bad_nodes))

# Repair: enforce a stable tenant-scoped business identity before adding edges.
identity_registry = set()
def claim_customer(tenant, customer_id):
    key = (tenant, customer_id)
    if key in identity_registry:
        return False
    identity_registry.add(key)
    return True

print("claim t1/c17 first:", claim_customer("t1", "c17"))
print("claim t1/c17 duplicate:", claim_customer("t1", "c17"))
print("claim t2/c17 distinct tenant:", claim_customer("t2", "c17"))
Expected evidence

Expected evidence: Customer c17 has two outgoing relationship types, order o55 connects to the customer, IP address, and payment instrument, the deliberately duplicated business ID splits one logical identity into two graph nodes, and the repaired tenant-scoped identity registry rejects the duplicate while allowing the same external ID in another tenant.

4. A relationship is data, not decorative syntax

Consider Customer-[:USES_CARD]->PaymentInstrument. If the association has facts such as first_seen, verification source, confidence, or role, those facts naturally belong to the relationship because they describe the connection. Storing only card_id as a customer property loses that first-class relationship identity and makes multi-card or historical associations awkward. The opposite mistake is turning every scalar attribute into a node. A product price, a customer's display name, or an order status usually remains a property unless the domain needs independent identity, relationships, lifecycle, or shared reference semantics for that fact.

5. Schema-light is not schema-free

Graph databases often permit nodes with different labels and property sets, but applications still impose a schema through labels, relationship types, property names/types, constraints, indexes, authorization rules, ingestion contracts, and query expectations. Neo4j's current documentation, for example, exposes indexes and multiple constraint types; that is implementation evidence that flexible graph structure does not eliminate schema governance. AtlasMart should version ingestion mappings, reject impossible edge endpoints, enforce tenant ownership, and test that migrations keep old and new readers compatible.

6. What the observable state proves—and does not prove

The lab prints labels, adjacency lists, and identity-claim results. That proves that a particular in-memory model has one stable customer identity and explicit typed edges. It does not prove durability, transactional isolation, replica consistency, authorization, index behavior, or Neo4j performance. Those are separate contracts. In a real graph deployment, add query plans, constraint status, transaction logs, replication/backup evidence, and tenant-aware authorization tests to the same verification mindset.

7. Production judgment

Choose a graph representation when relationships and multi-hop navigation are central to the workload, not merely because the domain contains relationships. Define which facts are authoritative, whether the graph itself is the system of record or a derived projection, what consistency is needed when edges are created/deleted, how tenant boundaries are enforced on both nodes and edges, and how identity reconciliation works. Relationship-heavy writes can increase contention around popular nodes; cross-region or sharded deployments add ownership and traversal costs; and backups must restore both topology and constraints. If a simple foreign key plus one indexed join serves the dominant workload more clearly, the relational model may remain the better operational choice.

Wrong approach: let imports invent duplicate graph identities

Failure injection / diagnosis

A team bulk-loads customers using generated internal IDs and assumes matching customer_id properties will be “close enough.” The result is two nodes for the same customer. Fraud edges are split, so a two-hop search misses connections. The safe repair is to define a tenant-scoped identity key, validate uniqueness before/while loading, reconcile duplicates with domain rules, redirect edges, and verify the resulting neighborhood—not to blindly merge nodes that merely share a display name.

Verification, cleanup, and production checklist

Verification is the program output plus the conceptual checks below. Cleanup is simply deleting the local lesson1_property_graph.py file; the simulation creates no sockets, services, databases, containers, credentials, or persistent data. In production, additionally verify tenant authorization on traversals, identity/constraint health, degree and path-cardinality distributions, p95/p99 latency, cache and remote-hop behavior where applicable, backup/restore or projection rebuild, software/security advisories, and edition/license constraints before adopting product-specific features.

Check your understanding

  1. What makes a property graph different from a plain key-value map?
  2. Why is a stable business identity important?
  3. Does schema-light mean no schema?
  4. When should a fact become an edge?
  5. What does an index typically help with first?
Review the answers

1. It represents first-class nodes and relationships, with labels/types and properties, so topology itself is queryable data.

2. Because duplicate logical entities fragment the neighborhood and can produce false negatives or contradictory updates.

3. No. Labels/types, property contracts, constraints, indexes, authorization rules, and application expectations still form a schema.

4. When the association has independent meaning, multiplicity, lifecycle, properties, or is a central traversal target.

5. Locating an anchor node or relationship by a property; adjacency traversal after anchoring is a separate mechanism.

References

Foundational statements use the published property-graph standard where appropriate; implementation-sensitive examples use current official documentation and are labeled as examples rather than universal graph guarantees.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.