Chapter 19 · Composite Databases, Multiple Databases, Federation, and Data-Domain Boundaries

Federated Queries, Data Locality, Cross-Graph Joins/Traversal Constraints, and Latency Tradeoffs

Write and reason about federated Cypher across AtlasMart domains, quantify locality and network costs, use proxy identities instead of impossible cross-graph relationships, and respect the one-write-graph-per-transaction rule.

Advanced240–330 minutesFederated query and locality labNeo4j 2026.07.1 · Community multi-instance mandatory · Enterprise composite optionalCypher 25 · USE / CALL / SHOW DATABASES / SHOW ALIASESJava 21/25 · Python driver 6.3Last reviewed: September 2026

An AtlasMart order-details endpoint now needs the Order graph plus Customer and Catalog facts. In a single graph this is one traversal; in a federated design it is a sequence of constituent reads and value joins. The danger is to celebrate a single logical Cypher query while hiding three physical execution sites, network round trips, result transfer and failure dependencies. This lesson makes those costs explicit.

Federation rule

Move selective predicates and aggregation to the owning constituent before returning rows across a boundary. Federation should exchange the smallest stable identity/value set that answers the business question.

Learning outcomes

01

Structure federated queries so each graph performs selective local work before cross-graph value joins.

02

Explain why relationships and graph traversals remain local while proxy IDs cross the federation boundary.

03

Measure local/remote constituent and application-side timings instead of attributing all latency to Neo4j generically.

04

Demonstrate the one-write-constituent rule and design safe workflows for cross-domain state changes.

05

Choose between federation, local read models, duplication, batching and domain recomposition based on workload evidence.

Reproducible Chapter 19 lab baseline

Chapter 19 baseline · reviewed 9 September 2026

Current self-managed Neo4j is 2026.07.1; current 5.26 LTS is 5.26.30. The course remains on Java 21 or 25, explicit CYPHER 25 for version-sensitive examples, and Python driver 6.3. The mandatory chapter lab uses three disposable Neo4j Community 2026.07.1 instances because Community can host exactly one standard database per DBMS. No runtime output or latency value in these lessons is claimed to have been executed during generation; deterministic expected rows are fixture invariants and timings must be measured by the learner.

Edition and platform boundary

Self-managed Community Edition can have exactly one standard database. Self-managed Enterprise Edition can have multiple standard databases; CREATE DATABASE and composite-database administration are Enterprise features and are not available on Aura. Composite databases are Enterprise-only and explicitly unavailable on Aura. Therefore the free learning path uses separate Community DBMS instances and application-side federation; optional Enterprise commands are labeled and must not be mistaken for Community or Aura behavior.

Term Mechanism-first meaning
standard database A physical Neo4j database that contains one graph in Neo4j 2026.07; it is an execution context and transaction domain.
DBMS A Neo4j database-management process/deployment that hosts the system database plus the standard databases allowed by its edition.
system database Built-in metadata/security database used for database, alias, server and access administration; it does not contain AtlasMart domain graph data.
multiple databases Several standard databases managed by one Enterprise DBMS; separation is stronger than labels but still shares DBMS/server resources and operations.
composite database Enterprise logical execution/federation context containing aliases to constituent graphs; it stores no graph data independently.
constituent A local or remote standard database exposed inside a composite through a namespaced alias.
local alias Alias whose target standard database is in the same DBMS.
remote alias Alias whose target is another Neo4j DBMS over a driver connection and whose authentication/security is governed at that remote boundary.
federated query One Cypher query whose graph-specific subqueries read from more than one constituent.
location transparency The caller uses a logical constituent name while the alias determines whether its target is local or remote; latency/failure locality is not magically erased.
proxy node A deliberately duplicated identity-only node used to join facts across disjoint graphs because Neo4j relationships cannot span graphs.
transaction domain The set of graph updates that can commit atomically together. A standard database is one transaction domain; a composite permits multi-graph reads but updates only one constituent per transaction.
tenant boundary A technical/operational separation choice for tenant data. Database separation does not automatically provide CPU/memory/noisy-neighbor isolation, billing isolation, or legal compliance.
domain ownership The team/system accountable for a fact’s schema, invariants, writes, recovery and lifecycle—not merely the graph where a convenient copy exists.
Community instance Purpose HTTP Bolt Container Volume
catalog Product/catalog source of truth 7574 7767 atlasmart-ch19-catalog atlasmart-ch19-catalog-data
orders Order facts plus CustomerRef/ProductRef proxies 7575 7768 atlasmart-ch19-orders atlasmart-ch19-orders-data
customers Customer source of truth 7576 7769 atlasmart-ch19-customers atlasmart-ch19-customers-data
Assumption Pinned value / rule
deployment Three isolated local Community containers on one workstation; this simulates domain separation, not a composite database
database Each Community DBMS uses its single standard database named neo4j
auth Synthetic lab-only neo4j / atlasmart-course-2026 credential; never use it outside the disposable lab
TLS Loopback lab uses bolt:// for simplicity; remote production aliases/drivers require verified TLS and credential governance
plugins No APOC or GDS required
indexes Uniqueness constraints on domain IDs only; no cross-database constraint exists
graph size Tiny deterministic fixture: 3 products, 3 customers, 3 orders, 4 line items/proxy references
failure injection Stop one disposable instance or add an application-side artificial delay; no destructive network or disk fault is required
Enterprise option Commands are examples for a licensed self-managed Enterprise environment and are not executed by the free path
Aura Composite databases and self-managed CREATE DATABASE are not taught as Aura capabilities

1. Compare three query shapes

Shape Requests / graph scopes Strength Failure/cost risk
N+1 application federation Orders once, then one Customer/Product call per row simple to write round-trip explosion and pool/network overhead
batched application federation Orders once, Customer once, Catalog with IN $ids once free/portable and explicit timing application owns join/correlation and partial failures
Enterprise composite federation one logical Cypher query with graph-specific CALL/USE scopes central logical query and alias abstraction still performs local/remote constituent work; Enterprise only

2. Data locality starts with selectivity

The Orders constituent should find one order and return only its customer/product keys and commercial line facts. Customer should fetch one customer by constrained ID. Catalog should batch the product IDs. Returning every product to the root scope and filtering there would convert an index-selective local operation into a network/result-volume problem.

Python · free batched federation (reuse from Lesson 1)
from neo4j import GraphDatabase
from time import perf_counter

AUTH = ("neo4j", "atlasmart-course-2026")
drivers = {
    "orders": GraphDatabase.driver("bolt://127.0.0.1:7768", auth=AUTH),
    "customers": GraphDatabase.driver("bolt://127.0.0.1:7769", auth=AUTH),
    "catalog": GraphDatabase.driver("bolt://127.0.0.1:7767", auth=AUTH),
}

def run(domain, query, **params):
    t0 = perf_counter()
    with drivers[domain].session(database="neo4j") as s:
        rows = [r.data() for r in s.run(query, **params)]
    return rows, perf_counter() - t0

order_rows, t_orders = run("orders", """
MATCH (o:Order {orderId:$orderId})-[:PLACED_BY]->(cr:CustomerRef)
MATCH (o)-[li:CONTAINS]->(pr:ProductRef)
RETURN o.orderId AS orderId, o.status AS status,
       cr.customerId AS customerId, pr.productId AS productId,
       li.quantity AS quantity, li.unitPrice AS unitPrice
ORDER BY productId
""", orderId="O-1901")

customer_id = order_rows[0]["customerId"]
product_ids = [r["productId"] for r in order_rows]
customer_rows, t_customers = run("customers", """
MATCH (c:Customer {customerId:$id})
RETURN c.customerId AS customerId, c.name AS customerName, c.tier AS tier
""", id=customer_id)
product_rows, t_catalog = run("catalog", """
MATCH (p:Product) WHERE p.productId IN $ids
RETURN p.productId AS productId, p.name AS productName, p.price AS currentPrice
""", ids=product_ids)

products = {p["productId"]: p for p in product_rows}
result = {
    "orderId": order_rows[0]["orderId"],
    "status": order_rows[0]["status"],
    "customer": customer_rows[0],
    "lines": [{**r, **products[r["productId"]]} for r in order_rows],
}
print(result)
print({"orders_s": t_orders, "customers_s": t_customers, "catalog_s": t_catalog})
for d in drivers.values(): d.close()
Measure these separately

Record Orders query time, Customer query time, Catalog query time, total request time, number of rows from each domain and serialized result bytes. A slow total with fast constituent queries can still be client/network/serialization overhead.

3. Proxy identity replaces impossible cross-graph relationships

Need Bad cross-graph idea Correct federated model
Order → Customer relationship from Orders node to Customer node in another database Order → CustomerRef locally; join customerId to Customers domain
Order line → Product relationship to remote Product Order → ProductRef locally; join productId to Catalog
current product metadata copy mutable Product name/category everywhere resolve by stable productId or maintain an intentional read model with freshness contract
historical charged price look up today’s Catalog price store line.unitPrice in Orders because Orders owns the historical commercial fact

4. One logical query still has physical boundaries

In the optional composite query, each CALL scope executes against a selected graph. Root-level scalar/value operations combine returned data. If a constituent is remote, those operations incur remote connectivity and can fail independently. Location transparency should make topology evolvable, not make operators blind to topology.

Cypher 25 · composite query for O-1901
CYPHER 25
CALL {
  USE atlasmart.orders
  MATCH (o:Order {orderId:$orderId})-[:PLACED_BY]->(cr:CustomerRef)
  MATCH (o)-[li:CONTAINS]->(pr:ProductRef)
  RETURN o.orderId AS orderId, o.status AS status,
         cr.customerId AS customerId, pr.productId AS productId,
         li.quantity AS quantity, li.unitPrice AS unitPrice
}
CALL {
  USE atlasmart.customers
  WITH customerId
  MATCH (c:Customer {customerId:customerId})
  RETURN c.name AS customerName, c.tier AS tier
}
CALL {
  USE atlasmart.catalog
  WITH productId
  MATCH (p:Product {productId:productId})
  RETURN p.name AS productName, p.price AS currentPrice
}
RETURN orderId, status, customerId, customerName, tier,
       productId, productName, quantity, unitPrice, currentPrice
ORDER BY productId;
Evidence What to compare
local constituent same-DBMS query execution and returned rows
remote constituent network/TLS/auth plus remote query execution and transfer
root coordination rows crossing scopes, join cardinality, final projection
driver request end-to-end p50/p95/p99, retries/timeouts, result consumption

5. Edge case: multi-graph writes are not distributed ACID

Cypher 25 · intentionally invalid two-constituent write
CYPHER 25
// OPTIONAL Enterprise composite example: this must fail because one transaction writes two constituents.
CALL {
  USE atlasmart.orders
  CREATE (:AuditMarker {id:'two-write-test-orders'})
  RETURN 1 AS x
}
CALL {
  USE atlasmart.catalog
  WITH x
  CREATE (:AuditMarker {id:'two-write-test-catalog'})
  RETURN 1 AS y
}
RETURN x, y;
// Expected mechanism: writing to more than one database per transaction is not allowed.
Expected failure mechanism

The transaction attempts to create data in Orders and Catalog. Current composite rules reject writing to more than one database in one transaction. This is a correctness boundary, not a performance limitation.

Business requirement Safer design
Create order and decrement inventory atomically put the invariant in one owning transactional domain if strict atomicity is mandatory
Order accepted then notify Catalog/Inventory commit Orders; publish an outbox/event with idempotent consumer and reconciliation
Two services must eventually agree use stable operation IDs, retries, idempotency, compensating states and reconciliation
Client timeout after write treat outcome as potentially ambiguous; query by idempotency/operation ID before blindly repeating side effects

6. Failure injection: one domain unavailable

Docker · disposable failure injection
# Keep Orders and Customers running; stop only Catalog.
docker stop atlasmart-ch19-catalog
# Run the free federation script: Orders/Customer may succeed; Catalog lookup must fail/timeout per driver settings.
# Record error class and elapsed time. Then restore Catalog.
docker start atlasmart-ch19-catalog
# Re-run and verify the O-1901 result is complete again.
What this proves

A federated endpoint inherits dependency failure modes. A single logical business response can fail because one remote/local constituent is unavailable. It does not prove that every request should fail—your API may define partial/degraded semantics, but that must be an explicit product contract.

7. When not to federate at request time

Signal Consider
same high-volume cross-domain join on every request materialized/read projection owned by the consuming domain with documented freshness
remote domain has high tail latency cache/read model or co-location if correctness permits
cross-domain result cardinality explodes push filtering/aggregation down; redesign split or precompute
strict multi-domain atomicity required redefine aggregate/ownership boundary instead of adding retries around impossible semantics
rare analyst query with tolerant latency federation may be a good fit because operational coupling is acceptable

Check your understanding

  1. Why should filtering occur inside each constituent?
  2. Can a returned node object become an edge endpoint in another graph?
  3. What is wrong with measuring only total endpoint latency?
  4. What happens if a transaction writes two constituents?
  5. When is a local read model preferable?
Review the answers

1. To reduce graph work and especially rows/bytes crossing the federation boundary.

2. No. Relationships are graph-local; cross-domain joins use stable values/proxy identities.

3. It cannot distinguish constituent execution, network/TLS, driver/pool, serialization and application join time.

4. Current composite semantics reject it; multi-graph updates are not ordinary distributed ACID.

5. When the dominant workload repeatedly pays expensive/fragile federation and bounded-staleness duplication is acceptable and governed.

Production judgment

Decision surface Production questions
graph/workload fit Does domain separation reduce ownership/capacity coupling, or does it turn the dominant traversal into repeated remote joins?
correctness Which invariants remain atomic inside one graph, and which become asynchronous/application-coordinated across domains?
model/cardinality/degree Which high-degree relationships must remain local for traversal cost and invariant enforcement? Which references are safe as proxy keys?
latency How much p95/p99 is local graph work versus remote constituent/network/application join time? What happens during remote tail spikes?
transactions/concurrency Which writes must commit together? Composite transactions may read many graphs but update only one constituent; plan compensating/outbox workflows elsewhere.
memory/resources Do multiple databases share a DBMS resource envelope? Do separate instances need independent memory/page-cache/process budgets and capacity headroom?
CPU/disk/network Does federation move filtering to constituents or ship excessive rows across network boundaries? Are region/zone egress and serialization material?
indexes/constraints Are stable IDs constrained and indexed independently on each owning/proxy graph? No cross-graph relationship or uniqueness constraint exists.
driver Are database/alias names explicit, drivers long-lived, timeouts/retries bounded, and per-domain timings correlated?
security/tenant risk Are ACCESS/MATCH/procedure rights granted on every required constituent? How are remote credentials/OIDC forwarding, TLS and tenant boundaries governed?
backup/recovery Can every constituent be restored and reconciled independently? A composite alias layer is not a backup of its target stores.
observability Can operators attribute latency/errors to the root composite scope and each local/remote constituent without hiding network waits?
testing/failure injection Have one-domain-down, stale proxy, version mismatch, denied constituent, slow remote graph and ambiguous cross-service write cases been tested safely?
version/edition/Aura Is the design pinned to self-managed Enterprise composite support? What is the fallback for Community/Aura or mixed-version constituents?
licensing/cost/migration Does federation justify Enterprise/ops/network cost, and can the split be rolled back or re-partitioned without breaking identity contracts?

Summary and next step

Federation is now a measurable execution plan across data domains, not a magical graph traversal. Lesson 4 turns from mechanics to governance: who owns identity, reference data, schema changes, access, backups and compatibility as those graphs evolve independently.

Authoritative references

  • Current Neo4j versions — Current self-managed release 2026.07.1 and 5.26.30 LTS snapshot.
  • Database administration — Current database/transaction-domain model and the one-standard-database Community versus multi-database Enterprise boundary.
  • Create standard databases — Enterprise-only self-managed CREATE DATABASE semantics and system-database administration.
  • Show databases — SHOW DATABASES fields including type, role, writer, status, aliases and constituents.
  • Composite database concepts — Enterprise-only, unavailable-on-Aura composite semantics, local/remote constituents, compatibility, transactions and proxy-node federation.
  • Create composite databases — CREATE COMPOSITE DATABASE and current default Cypher language behavior.
  • Query composite databases — USE/CALL graph selection, graph functions, update restrictions, root-scope limits and runtime behavior.
  • Composite aliases — Local and remote constituent aliases, SHOW ALIASES evidence, namespace rules and alias limitations.
  • Standard aliases — Local/remote alias behavior, credentials, access visibility and current OIDC-forwarding option for remote aliases.
  • Composite RBAC — Access must be granted to the composite and constituents; remote constituents enforce remote-user RBAC.
  • Composite tutorial — Official federation/sharding example and proxy-node model.
  • Database alias command syntax — Current SHOW/CREATE/ALTER alias command forms.
  • Cypher USE clause — Current graph-selection clause semantics for composite/federated queries.
  • Python driver manual — Official driver lifecycle, sessions, parameters, database selection and application-side orchestration.
  • Backup and restore — Operational reminder that constituent data stores—not a composite alias layer—are the recovery units.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.