Decide when AtlasMart relationship-centric queries justify graph storage and when relational, search, key-value, or analytical systems remain simpler.
Use Cases: Identity, Fraud, Recommendations, Networks, Knowledge Graphs, and Dependency Analysis
Close the graph chapter with a workload-fit matrix, explicit system-of-record boundaries, derived-graph repair, and a failure case that exposes the cost of treating every relationship as a graph-database problem.
Learning outcomes
Close the graph chapter with a workload-fit matrix, explicit system-of-record boundaries, derived-graph repair, and a failure case that exposes the cost of treating every relationship as a graph-database problem.
Identify the relationship-centric query that justifies graph storage for six common domains.
Name a simpler alternative when the workload is predominantly point lookup, search, or aggregation.
Separate authoritative transactional data from a derived graph projection when appropriate.
Define freshness, reconciliation, security, and recovery evidence for a production graph.
Mandatory work uses Python 3.13+ standard library only on a single local process. No graph database, Docker image, cloud account, paid feature, network manipulation, or destructive failure injection is required. Neo4j 2026.07.1 is an optional current implementation reference; Community Edition is GPLv3. Product-specific clustering, sharding, security, and enterprise features are not assumed by the lab.
1. Graph fit starts with a relationship-centric question
A domain name is not a technology requirement. “Fraud,” “recommendations,” and “knowledge graph” do not automatically require a graph database. The useful discriminator is the query shape. Fraud becomes graph-shaped when an investigator needs multi-hop shared-device/card/IP neighborhoods with relationship provenance. Identity resolution becomes graph-shaped when evidence links many aliases/identifiers and confidence depends on paths. Dependency analysis becomes graph-shaped when engineers ask reachability and blast-radius questions over a changing service/package topology.
2. Evaluate each common use case against a simpler baseline
For fraud rings, graph storage earns its cost when ring structure and multi-hop shared resources matter; isolated risk scoring over transaction rows may fit relational/stream processing. For identity resolution, graph paths can combine aliases and evidence; a simple customer-master lookup may not need them. For recommendations, user-item neighborhoods can be valuable, while a fixed top-products list or embedding search may be simpler elsewhere. Networks and dependency graphs naturally expose reachability; static trees can live in ordinary relational structures. Knowledge graphs are useful when typed entities and semantic relationships are the product; keyword retrieval belongs to search, and heavy global aggregation may belong to analytical storage.
3. AtlasMart lab: make the mechanism observable
Save the following as lesson5_graph_fit.py and run
it with python lesson5_graph_fit.py. The program
has no dependencies and mutates no external state.
from dataclasses import dataclass
@dataclass(frozen=True)
class Workload:
name: str
multi_hop: int
changing_relationships: int
edge_properties: int
path_constraints: int
point_lookup: int
full_text: int
heavy_aggregate: int
workloads = [
Workload("fraud-ring", 3,3,3,3,1,0,1),
Workload("identity-resolution", 3,3,2,3,1,0,1),
Workload("recommendation-neighborhood", 3,2,2,2,1,1,1),
Workload("dependency-impact", 3,2,2,3,1,0,0),
Workload("order-by-id", 0,0,0,0,3,0,0),
Workload("catalog-keyword-search", 0,1,0,0,1,3,1),
Workload("daily-revenue-report", 0,0,0,0,0,0,3),
]
def graph_score(w):
return 3*w.multi_hop + 2*w.changing_relationships + 2*w.edge_properties + 2*w.path_constraints - 2*w.point_lookup - 2*w.full_text - 2*w.heavy_aggregate
for w in workloads:
score = graph_score(w)
verdict = "graph candidate" if score >= 12 else "other model likely simpler"
print(w.name, score, verdict)
# Source-of-truth vs derived fraud graph.
authoritative_events = [
("evt-1", "cust:c1", "USES_CARD", "card:p9"),
("evt-2", "cust:c2", "USES_CARD", "card:p9"),
("evt-3", "cust:c2", "USES_DEVICE", "dev:d1"),
]
derived_edges = set()
for event in authoritative_events[:2]: # simulate lost dual write / consumer outage
_,a,t,b = event
derived_edges.add((a,t,b))
print("derived before reconciliation:", sorted(derived_edges))
missing = {(a,t,b) for _,a,t,b in authoritative_events} - derived_edges
print("missing derived edges:", sorted(missing))
derived_edges |= missing
print("derived after reconciliation:", sorted(derived_edges))
Expected evidence: relationship-centric workloads such as fraud-ring, identity-resolution, and dependency-impact score as graph candidates in the transparent teaching heuristic, while order-by-ID, full-text catalog search, and daily revenue aggregation point elsewhere. The deliberately incomplete derived fraud graph is missing one authoritative edge until reconciliation restores it.
4. System of record versus derived graph is an architecture decision
AtlasMart's orders and payments already have strong transactional invariants. Replacing that ledger with a graph solely to support fraud traversal may create unnecessary migration and operational risk. A common polyglot design keeps the authoritative event/order state in its transactional store and builds a graph projection from Change Data Capture (CDC), an outbox, or an event stream. That graph can then be optimized for traversal. The cost is eventual synchronization: every edge needs a source event, idempotent application, replayability, freshness Service-Level Objectives (SLOs), and reconciliation when events are missed.
5. Wrong derived views fail silently unless you reconcile
A naïve dual write performs the payment transaction and separately updates the graph. If the graph write fails after the payment commits, the system of record is correct but the graph is stale. Fraud queries may then miss a connection—the worst kind of “available but wrong” result. The lab simulates exactly that missing edge. The repair is to make synchronization replayable from an authoritative log/outbox and periodically compare expected relationship counts/keys or domain-specific invariants. A graph projection must be rebuildable or its backup/recovery posture must match its authority level.
6. Security is topology-sensitive
Graph access control must consider paths, not only rows. A user permitted to see a customer node might infer restricted merchants, employees, or other tenants through edges. Store tenant ownership or visibility metadata where enforcement can evaluate it, test path traversal across shared infrastructure entities, redact sensitive relationship properties, and audit broad neighborhood/path queries. Identity and fraud graphs often contain especially sensitive linkage data, so data minimization and retention policies matter as much as traversal performance.
7. Production judgment and bridge to search
Choose a graph database when changing relationships and multi-hop/path queries are first-class workload requirements and when your team can operate the necessary indexes, backups, monitoring, authorization, and—if distributed—placement/failure semantics. Keep point lookups in key-value stores, strong transactional ledgers in relational/document systems when they fit, and textual relevance in search infrastructure. A polyglot architecture is justified only when ownership and synchronization are explicit. Chapter 14 continues with another derived-access-path family: secondary indexes, inverted indexes, and vector retrieval.
Wrong approach: make the graph authoritative just because fraud queries use it
AtlasMart moves the authoritative payment ledger into a graph without a workload or invariant reason, then also dual-writes a relational reporting store. The migration increases operational surface and creates another synchronization direction. A safer design keeps the transactional authority where its invariants are already well served, projects the relationship data needed by fraud queries through a replayable event path, measures graph freshness, and proves reconciliation/restore before relying on it for decisions.
Verification, cleanup, and production checklist
Verification is the program output plus the conceptual checks
below. Cleanup is simply deleting the local
lesson5_graph_fit.py file; the simulation creates
no sockets, services, databases, containers, credentials, or
persistent data. In production, additionally verify tenant
authorization on traversals, identity/constraint health, degree
and path-cardinality distributions, p95/p99 latency, cache and
remote-hop behavior where applicable, backup/restore or
projection rebuild, software/security advisories, and
edition/license constraints before adopting product-specific
features.
Check your understanding
- What is a better graph-selection criterion than the domain label “fraud”?
- When is a relational system simpler for relationship data?
- Why might a fraud graph be derived rather than authoritative?
- What failure does naïve dual writing create?
- What proves a derived graph is operationally trustworthy?
Review the answers
1. Whether the important queries repeatedly traverse changing relationships, paths, neighborhoods, and edge properties.
2. When queries are bounded/simple joins, transactions dominate, and arbitrary multi-hop traversal is not a first-class requirement.
3. The order/payment store may already own strong business invariants; the graph can optimize traversal while remaining rebuildable.
4. One store can commit while the other update fails, leaving a silently inconsistent derived graph.
5. Measured freshness, idempotent replay, reconciliation evidence, restore/rebuild tests, and authorization/path-leak tests.
References
Foundational statements use the published property-graph standard where appropriate; implementation-sensitive examples use current official documentation and are labeled as examples rather than universal graph guarantees.
- Neo4j — What is a graph database — current relationship-first property-graph overview.
- Neo4j Cypher Manual — Variable-length paths — current path-query semantics and path-explosion caution.
- ISO/IEC 39075:2024 — GQL — standard property-graph model/query context.
- Neo4j release notes — current Neo4j 2026.07.1 snapshot.
- Neo4j editions and licensing — Community Edition GPLv3; enterprise clustering/sharding capabilities are edition-specific and not required in the lab.