Decide when AtlasMart relationship-centric queries justify graph storage and when relational, search, key-value, or analytical systems remain simpler.

Use Cases: Identity, Fraud, Recommendations, Networks, Knowledge Graphs, and Dependency Analysis

Close the graph chapter with a workload-fit matrix, explicit system-of-record boundaries, derived-graph repair, and a failure case that exposes the cost of treating every relationship as a graph-database problem.

Intermediate95–120 minutesGraph-fit + derived-view repair labPython 3.13+ · standard libraryNeo4j 2026.07.1 optional referenceLast reviewed: August 2026

Learning outcomes

Close the graph chapter with a workload-fit matrix, explicit system-of-record boundaries, derived-graph repair, and a failure case that exposes the cost of treating every relationship as a graph-database problem.

01

Identify the relationship-centric query that justifies graph storage for six common domains.

02

Name a simpler alternative when the workload is predominantly point lookup, search, or aggregation.

03

Separate authoritative transactional data from a derived graph projection when appropriate.

04

Define freshness, reconciliation, security, and recovery evidence for a production graph.

Implementation snapshot

Mandatory work uses Python 3.13+ standard library only on a single local process. No graph database, Docker image, cloud account, paid feature, network manipulation, or destructive failure injection is required. Neo4j 2026.07.1 is an optional current implementation reference; Community Edition is GPLv3. Product-specific clustering, sharding, security, and enterprise features are not assumed by the lab.

1. Graph fit starts with a relationship-centric question

A domain name is not a technology requirement. “Fraud,” “recommendations,” and “knowledge graph” do not automatically require a graph database. The useful discriminator is the query shape. Fraud becomes graph-shaped when an investigator needs multi-hop shared-device/card/IP neighborhoods with relationship provenance. Identity resolution becomes graph-shaped when evidence links many aliases/identifiers and confidence depends on paths. Dependency analysis becomes graph-shaped when engineers ask reachability and blast-radius questions over a changing service/package topology.

2. Evaluate each common use case against a simpler baseline

For fraud rings, graph storage earns its cost when ring structure and multi-hop shared resources matter; isolated risk scoring over transaction rows may fit relational/stream processing. For identity resolution, graph paths can combine aliases and evidence; a simple customer-master lookup may not need them. For recommendations, user-item neighborhoods can be valuable, while a fixed top-products list or embedding search may be simpler elsewhere. Networks and dependency graphs naturally expose reachability; static trees can live in ordinary relational structures. Knowledge graphs are useful when typed entities and semantic relationships are the product; keyword retrieval belongs to search, and heavy global aggregation may belong to analytical storage.

3. AtlasMart lab: make the mechanism observable

Save the following as lesson5_graph_fit.py and run it with python lesson5_graph_fit.py. The program has no dependencies and mutates no external state.

python · AtlasMart deterministic simulation
from dataclasses import dataclass

@dataclass(frozen=True)
class Workload:
    name: str
    multi_hop: int
    changing_relationships: int
    edge_properties: int
    path_constraints: int
    point_lookup: int
    full_text: int
    heavy_aggregate: int

workloads = [
    Workload("fraud-ring", 3,3,3,3,1,0,1),
    Workload("identity-resolution", 3,3,2,3,1,0,1),
    Workload("recommendation-neighborhood", 3,2,2,2,1,1,1),
    Workload("dependency-impact", 3,2,2,3,1,0,0),
    Workload("order-by-id", 0,0,0,0,3,0,0),
    Workload("catalog-keyword-search", 0,1,0,0,1,3,1),
    Workload("daily-revenue-report", 0,0,0,0,0,0,3),
]

def graph_score(w):
    return 3*w.multi_hop + 2*w.changing_relationships + 2*w.edge_properties + 2*w.path_constraints - 2*w.point_lookup - 2*w.full_text - 2*w.heavy_aggregate

for w in workloads:
    score = graph_score(w)
    verdict = "graph candidate" if score >= 12 else "other model likely simpler"
    print(w.name, score, verdict)

# Source-of-truth vs derived fraud graph.
authoritative_events = [
    ("evt-1", "cust:c1", "USES_CARD", "card:p9"),
    ("evt-2", "cust:c2", "USES_CARD", "card:p9"),
    ("evt-3", "cust:c2", "USES_DEVICE", "dev:d1"),
]
derived_edges = set()
for event in authoritative_events[:2]:  # simulate lost dual write / consumer outage
    _,a,t,b = event
    derived_edges.add((a,t,b))
print("derived before reconciliation:", sorted(derived_edges))
missing = {(a,t,b) for _,a,t,b in authoritative_events} - derived_edges
print("missing derived edges:", sorted(missing))
derived_edges |= missing
print("derived after reconciliation:", sorted(derived_edges))
Expected evidence

Expected evidence: relationship-centric workloads such as fraud-ring, identity-resolution, and dependency-impact score as graph candidates in the transparent teaching heuristic, while order-by-ID, full-text catalog search, and daily revenue aggregation point elsewhere. The deliberately incomplete derived fraud graph is missing one authoritative edge until reconciliation restores it.

4. System of record versus derived graph is an architecture decision

AtlasMart's orders and payments already have strong transactional invariants. Replacing that ledger with a graph solely to support fraud traversal may create unnecessary migration and operational risk. A common polyglot design keeps the authoritative event/order state in its transactional store and builds a graph projection from Change Data Capture (CDC), an outbox, or an event stream. That graph can then be optimized for traversal. The cost is eventual synchronization: every edge needs a source event, idempotent application, replayability, freshness Service-Level Objectives (SLOs), and reconciliation when events are missed.

5. Wrong derived views fail silently unless you reconcile

A naïve dual write performs the payment transaction and separately updates the graph. If the graph write fails after the payment commits, the system of record is correct but the graph is stale. Fraud queries may then miss a connection—the worst kind of “available but wrong” result. The lab simulates exactly that missing edge. The repair is to make synchronization replayable from an authoritative log/outbox and periodically compare expected relationship counts/keys or domain-specific invariants. A graph projection must be rebuildable or its backup/recovery posture must match its authority level.

6. Security is topology-sensitive

Graph access control must consider paths, not only rows. A user permitted to see a customer node might infer restricted merchants, employees, or other tenants through edges. Store tenant ownership or visibility metadata where enforcement can evaluate it, test path traversal across shared infrastructure entities, redact sensitive relationship properties, and audit broad neighborhood/path queries. Identity and fraud graphs often contain especially sensitive linkage data, so data minimization and retention policies matter as much as traversal performance.

7. Production judgment and bridge to search

Choose a graph database when changing relationships and multi-hop/path queries are first-class workload requirements and when your team can operate the necessary indexes, backups, monitoring, authorization, and—if distributed—placement/failure semantics. Keep point lookups in key-value stores, strong transactional ledgers in relational/document systems when they fit, and textual relevance in search infrastructure. A polyglot architecture is justified only when ownership and synchronization are explicit. Chapter 14 continues with another derived-access-path family: secondary indexes, inverted indexes, and vector retrieval.

Wrong approach: make the graph authoritative just because fraud queries use it

Failure injection / diagnosis

AtlasMart moves the authoritative payment ledger into a graph without a workload or invariant reason, then also dual-writes a relational reporting store. The migration increases operational surface and creates another synchronization direction. A safer design keeps the transactional authority where its invariants are already well served, projects the relationship data needed by fraud queries through a replayable event path, measures graph freshness, and proves reconciliation/restore before relying on it for decisions.

Verification, cleanup, and production checklist

Verification is the program output plus the conceptual checks below. Cleanup is simply deleting the local lesson5_graph_fit.py file; the simulation creates no sockets, services, databases, containers, credentials, or persistent data. In production, additionally verify tenant authorization on traversals, identity/constraint health, degree and path-cardinality distributions, p95/p99 latency, cache and remote-hop behavior where applicable, backup/restore or projection rebuild, software/security advisories, and edition/license constraints before adopting product-specific features.

Check your understanding

  1. What is a better graph-selection criterion than the domain label “fraud”?
  2. When is a relational system simpler for relationship data?
  3. Why might a fraud graph be derived rather than authoritative?
  4. What failure does naïve dual writing create?
  5. What proves a derived graph is operationally trustworthy?
Review the answers

1. Whether the important queries repeatedly traverse changing relationships, paths, neighborhoods, and edge properties.

2. When queries are bounded/simple joins, transactions dominate, and arbitrary multi-hop traversal is not a first-class requirement.

3. The order/payment store may already own strong business invariants; the graph can optimize traversal while remaining rebuildable.

4. One store can commit while the other update fails, leaving a silently inconsistent derived graph.

5. Measured freshness, idempotent replay, reconciliation evidence, restore/rebuild tests, and authorization/path-leak tests.

References

Foundational statements use the published property-graph standard where appropriate; implementation-sensitive examples use current official documentation and are labeled as examples rather than universal graph guarantees.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.