Chapter 19 · Composite Databases, Multiple Databases, Federation, and Data-Domain Boundaries
Multiple Databases: Isolation, Administration, Security, Resource Considerations, and Tenant Boundaries
Decide when AtlasMart domains deserve separate databases or separate DBMS instances, and understand exactly what isolation, security, resource, transaction, backup, and tenant boundaries those choices do and do not provide.
AtlasMart has grown from one graph into three accountable domains: Catalog owns product descriptions and current prices; Customer owns profile/tier data; Orders owns commercial facts such as order status and the unit price charged at purchase time. A proposal says “put each domain in a database and we get tenant isolation for free.” That statement mixes at least five different boundaries: data files, transactions, security, resources, and failure/recovery. This lesson separates them before any federation feature is introduced.
Separate a domain because ownership, invariants, lifecycle, capacity, security or recovery require it—not merely because labels look different.
Learning outcomes
Distinguish one standard database, multiple Enterprise databases, separate Community DBMS instances, and composite databases.
Explain what database separation guarantees for transactions/data and what it does not guarantee for shared CPU, memory, network or tenant SLOs.
Design AtlasMart Catalog, Orders and Customer ownership with stable cross-domain proxy IDs.
Observe the free multi-instance design through deterministic fixture counts and explicit application-side federation.
Evaluate security, backup, resource and migration consequences before using database boundaries as tenant boundaries.
Reproducible Chapter 19 lab baseline
Current self-managed Neo4j is 2026.07.1; current
5.26 LTS is 5.26.30. The course remains on Java
21 or 25, explicit CYPHER 25 for
version-sensitive examples, and Python driver 6.3. The
mandatory chapter lab uses three disposable Neo4j Community
2026.07.1 instances because Community can host exactly one
standard database per DBMS. No runtime output or latency value
in these lessons is claimed to have been executed during
generation; deterministic expected rows are fixture invariants
and timings must be measured by the learner.
Self-managed
Community Edition can have exactly one standard
database. Self-managed
Enterprise Edition can have multiple standard
databases; CREATE DATABASE and composite-database
administration are Enterprise features and are not available
on Aura. Composite databases are Enterprise-only and
explicitly unavailable on Aura. Therefore the free learning
path uses separate Community DBMS instances and
application-side federation; optional Enterprise commands are
labeled and must not be mistaken for Community or Aura
behavior.
| Term | Mechanism-first meaning |
|---|---|
| standard database | A physical Neo4j database that contains one graph in Neo4j 2026.07; it is an execution context and transaction domain. |
| DBMS | A Neo4j database-management process/deployment that hosts the system database plus the standard databases allowed by its edition. |
| system database | Built-in metadata/security database used for database, alias, server and access administration; it does not contain AtlasMart domain graph data. |
| multiple databases | Several standard databases managed by one Enterprise DBMS; separation is stronger than labels but still shares DBMS/server resources and operations. |
| composite database | Enterprise logical execution/federation context containing aliases to constituent graphs; it stores no graph data independently. |
| constituent | A local or remote standard database exposed inside a composite through a namespaced alias. |
| local alias | Alias whose target standard database is in the same DBMS. |
| remote alias | Alias whose target is another Neo4j DBMS over a driver connection and whose authentication/security is governed at that remote boundary. |
| federated query | One Cypher query whose graph-specific subqueries read from more than one constituent. |
| location transparency | The caller uses a logical constituent name while the alias determines whether its target is local or remote; latency/failure locality is not magically erased. |
| proxy node | A deliberately duplicated identity-only node used to join facts across disjoint graphs because Neo4j relationships cannot span graphs. |
| transaction domain | The set of graph updates that can commit atomically together. A standard database is one transaction domain; a composite permits multi-graph reads but updates only one constituent per transaction. |
| tenant boundary | A technical/operational separation choice for tenant data. Database separation does not automatically provide CPU/memory/noisy-neighbor isolation, billing isolation, or legal compliance. |
| domain ownership | The team/system accountable for a fact’s schema, invariants, writes, recovery and lifecycle—not merely the graph where a convenient copy exists. |
| Community instance | Purpose | HTTP | Bolt | Container | Volume |
|---|---|---|---|---|---|
| catalog | Product/catalog source of truth | 7574 | 7767 | atlasmart-ch19-catalog | atlasmart-ch19-catalog-data |
| orders | Order facts plus CustomerRef/ProductRef proxies | 7575 | 7768 | atlasmart-ch19-orders | atlasmart-ch19-orders-data |
| customers | Customer source of truth | 7576 | 7769 | atlasmart-ch19-customers | atlasmart-ch19-customers-data |
| Assumption | Pinned value / rule |
|---|---|
| deployment | Three isolated local Community containers on one workstation; this simulates domain separation, not a composite database |
| database | Each Community DBMS uses its single standard database named neo4j |
| auth | Synthetic lab-only neo4j / atlasmart-course-2026 credential; never use it outside the disposable lab |
| TLS | Loopback lab uses bolt:// for simplicity; remote production aliases/drivers require verified TLS and credential governance |
| plugins | No APOC or GDS required |
| indexes | Uniqueness constraints on domain IDs only; no cross-database constraint exists |
| graph size | Tiny deterministic fixture: 3 products, 3 customers, 3 orders, 4 line items/proxy references |
| failure injection | Stop one disposable instance or add an application-side artificial delay; no destructive network or disk fault is required |
| Enterprise option | Commands are examples for a licensed self-managed Enterprise environment and are not executed by the free path |
| Aura | Composite databases and self-managed CREATE DATABASE are not taught as Aura capabilities |
1. The first edition fact changes the lab design
| Deployment choice | Standard databases available | Composite? | What Chapter 19 uses it for |
|---|---|---|---|
| Community DBMS | exactly one | No | Mandatory free path: three separate DBMS instances |
| Self-managed Enterprise DBMS | multiple | Yes | Optional real multi-database/composite commands |
| Aura | managed instance semantics; self-managed CREATE DATABASE is unavailable | Composite unavailable | Discuss managed-domain alternatives; do not present composites as Aura behavior |
Three labels in one Community database are not three database/transaction/security boundaries. Three separate Community instances are three DBMS/process boundaries, but they add independent operations and cannot provide a single atomic Neo4j transaction across them.
2. Database isolation has several dimensions
| Dimension | Separate standard databases in one Enterprise DBMS | Separate DBMS instances |
|---|---|---|
| graph/store files | separate database stores | separate stores/process directories/volumes |
| transaction domain | a normal transaction is bound to one standard database | also separate; application coordination required |
| schema/indexes | defined per database | defined per instance/database |
| security | Enterprise privileges can target databases/graphs | security configured independently per DBMS |
| page cache / server resources | not a hard noisy-neighbor boundary; databases consume shared server resources | stronger process/host budgeting possible, but same host can still contend |
| failure/restart | database-level operations can be scoped, but DBMS/server failure can affect several databases | one process can fail independently; infrastructure may still be shared |
| backup/recovery | backup/restore remains per data database | backup/restore per instance/database |
| operational cost | fewer processes, richer admin/topology | more ports, secrets, upgrades, monitoring and orchestration |
3. Start the free multi-instance laboratory
The ports deliberately avoid the course baseline container on
7474/7687. Each Chapter 19 container runs one Community standard
database named neo4j. The Docker network is for
optional container-to-container exercises; the Python lab uses
loopback Bolt ports from the host.
# PowerShell / Bash-compatible Docker commands (line continuations omitted intentionally)
docker network create atlasmart-ch19-net
docker run -d --name atlasmart-ch19-catalog --network atlasmart-ch19-net \
-p 7574:7474 -p 7767:7687 \
-e NEO4J_AUTH=neo4j/atlasmart-course-2026 \
-v atlasmart-ch19-catalog-data:/data neo4j:2026.07.1
docker run -d --name atlasmart-ch19-orders --network atlasmart-ch19-net \
-p 7575:7474 -p 7768:7687 \
-e NEO4J_AUTH=neo4j/atlasmart-course-2026 \
-v atlasmart-ch19-orders-data:/data neo4j:2026.07.1
docker run -d --name atlasmart-ch19-customers --network atlasmart-ch19-net \
-p 7576:7474 -p 7769:7687 \
-e NEO4J_AUTH=neo4j/atlasmart-course-2026 \
-v atlasmart-ch19-customers-data:/data neo4j:2026.07.1
# Verify process/port state; availability timing varies by machine.
docker ps --filter name=atlasmart-ch19-
docker ps should list the three named containers with the pinned host ports after startup. Server readiness timing depends on the machine; do not hard-code a startup duration.
4. Seed facts where their owner lives
Catalog owns Product; Customers owns
Customer; Orders owns Order plus
identity-only CustomerRef and
ProductRef proxies. The order line stores
unitPrice because that is an immutable commercial
fact at purchase time; Catalog’s currentPrice can
later change without rewriting history.
CYPHER 25
CREATE CONSTRAINT ch19_product_id IF NOT EXISTS
FOR (p:Product) REQUIRE p.productId IS UNIQUE;
UNWIND [
{id:'P-1901', name:'Trail Camera', category:'Cameras', price:129.90},
{id:'P-1902', name:'Weather Case', category:'Accessories', price:39.50},
{id:'P-1903', name:'Sensor Hub', category:'Store IoT', price:249.00}
] AS row
MERGE (p:Product {productId:row.id})
SET p.name=row.name, p.category=row.category, p.price=row.price, p.labTag='ch19';
CYPHER 25
CREATE CONSTRAINT ch19_customer_id IF NOT EXISTS
FOR (c:Customer) REQUIRE c.customerId IS UNIQUE;
UNWIND [
{id:'C-1901', name:'Ada Lovelace', tier:'gold'},
{id:'C-1902', name:'Grace Hopper', tier:'silver'},
{id:'C-1903', name:'Katherine Johnson', tier:'gold'}
] AS row
MERGE (c:Customer {customerId:row.id})
SET c.name=row.name, c.tier=row.tier, c.labTag='ch19';
CYPHER 25
CREATE CONSTRAINT ch19_order_id IF NOT EXISTS
FOR (o:Order) REQUIRE o.orderId IS UNIQUE;
CREATE CONSTRAINT ch19_customer_ref_id IF NOT EXISTS
FOR (c:CustomerRef) REQUIRE c.customerId IS UNIQUE;
CREATE CONSTRAINT ch19_product_ref_id IF NOT EXISTS
FOR (p:ProductRef) REQUIRE p.productId IS UNIQUE;
UNWIND [
{orderId:'O-1901', customerId:'C-1901', status:'PAID'},
{orderId:'O-1902', customerId:'C-1902', status:'SHIPPED'},
{orderId:'O-1903', customerId:'C-1903', status:'PAID'}
] AS row
MERGE (o:Order {orderId:row.orderId})
SET o.status=row.status, o.labTag='ch19'
MERGE (c:CustomerRef {customerId:row.customerId})
SET c.labTag='ch19'
MERGE (o)-[:PLACED_BY]->(c);
UNWIND [
{orderId:'O-1901', productId:'P-1901', quantity:1, unitPrice:129.90},
{orderId:'O-1901', productId:'P-1902', quantity:2, unitPrice:39.50},
{orderId:'O-1902', productId:'P-1903', quantity:1, unitPrice:249.00},
{orderId:'O-1903', productId:'P-1901', quantity:1, unitPrice:129.90}
] AS row
MATCH (o:Order {orderId:row.orderId})
MERGE (p:ProductRef {productId:row.productId})
SET p.labTag='ch19'
MERGE (o)-[r:CONTAINS]->(p)
SET r.quantity=row.quantity, r.unitPrice=row.unitPrice;
CYPHER 25
// Run on catalog
MATCH (p:Product {labTag:'ch19'}) RETURN count(p) AS products;
// Expected: 3
// Run on customers
MATCH (c:Customer {labTag:'ch19'}) RETURN count(c) AS customers;
// Expected: 3
// Run on orders
MATCH (o:Order {labTag:'ch19'}) RETURN count(o) AS orders;
MATCH (:Order {labTag:'ch19'})-[r:CONTAINS]->(:ProductRef) RETURN count(r) AS lines;
// Expected: orders=3, lines=4
5. A proxy is an identity contract, not a cross-graph edge
| Orders graph element | Meaning | Owned elsewhere? |
|---|---|---|
| Order O-1901 | commercial aggregate and status | No — Orders owns it |
| CustomerRef C-1901 | foreign identity/proxy only | Customer profile is owned by Customers |
| ProductRef P-1901 | foreign identity/proxy only | Product descriptive/current-price data is owned by Catalog |
| CONTAINS.unitPrice | price actually charged | Orders owns historical fact even though Catalog owns current price |
Duplicating all descriptive fields into Orders creates competing sources of truth and ambiguous update ownership. Duplicate only what the Orders domain intentionally owns (for example historical price) and keep proxy identity narrow.
6. Simulate federation explicitly in application code
The free path queries Orders first, extracts stable keys, then fetches Customer and Product data in bounded batches. The program prints measured per-domain timings from the learner’s machine; no timing threshold is prescribed. This is intentionally not called a composite database.
from neo4j import GraphDatabase
from time import perf_counter
AUTH = ("neo4j", "atlasmart-course-2026")
drivers = {
"orders": GraphDatabase.driver("bolt://127.0.0.1:7768", auth=AUTH),
"customers": GraphDatabase.driver("bolt://127.0.0.1:7769", auth=AUTH),
"catalog": GraphDatabase.driver("bolt://127.0.0.1:7767", auth=AUTH),
}
def run(domain, query, **params):
t0 = perf_counter()
with drivers[domain].session(database="neo4j") as s:
rows = [r.data() for r in s.run(query, **params)]
return rows, perf_counter() - t0
order_rows, t_orders = run("orders", """
MATCH (o:Order {orderId:$orderId})-[:PLACED_BY]->(cr:CustomerRef)
MATCH (o)-[li:CONTAINS]->(pr:ProductRef)
RETURN o.orderId AS orderId, o.status AS status,
cr.customerId AS customerId, pr.productId AS productId,
li.quantity AS quantity, li.unitPrice AS unitPrice
ORDER BY productId
""", orderId="O-1901")
customer_id = order_rows[0]["customerId"]
product_ids = [r["productId"] for r in order_rows]
customer_rows, t_customers = run("customers", """
MATCH (c:Customer {customerId:$id})
RETURN c.customerId AS customerId, c.name AS customerName, c.tier AS tier
""", id=customer_id)
product_rows, t_catalog = run("catalog", """
MATCH (p:Product) WHERE p.productId IN $ids
RETURN p.productId AS productId, p.name AS productName, p.price AS currentPrice
""", ids=product_ids)
products = {p["productId"]: p for p in product_rows}
result = {
"orderId": order_rows[0]["orderId"],
"status": order_rows[0]["status"],
"customer": customer_rows[0],
"lines": [{**r, **products[r["productId"]]} for r in order_rows],
}
print(result)
print({"orders_s": t_orders, "customers_s": t_customers, "catalog_s": t_catalog})
for d in drivers.values(): d.close()
For O-1901: customer C-1901 / Ada Lovelace; products P-1901 Trail Camera and P-1902 Weather Case. Exact timing values are environment-dependent and must be measured locally.
7. Tenant boundary judgment
| Question | If yes, prefer… | Reason |
|---|---|---|
| Must writes across two domains be atomic? | keep the invariant in one owning graph or redesign aggregate boundary | separate databases/instances cannot offer an ordinary one-graph transaction across both |
| Must one tenant survive another tenant’s pathological CPU/memory workload? | separate resource-isolated deployment/tier, not merely another database | database is not automatically a hard resource quota |
| Different backup/RPO/security lifecycle? | stronger database/instance boundary may be justified | operational ownership is a real boundary driver |
| Dominant query traverses both domains on every request? | reconsider split, co-locate read model, or budget federation explicitly | remote joins add latency/failure coupling |
8. Cleanup is part of the boundary
# Remove disposable Chapter 19 Community instances and their data volumes.
docker rm -f atlasmart-ch19-catalog atlasmart-ch19-orders atlasmart-ch19-customers
docker volume rm atlasmart-ch19-catalog-data atlasmart-ch19-orders-data atlasmart-ch19-customers-data
docker network rm atlasmart-ch19-net
Lessons 2–5 reuse these instances. Run cleanup after Lesson 5 unless you intentionally want to preserve the disposable fixture.
Check your understanding
- How many standard databases can current Community host?
- Does a database boundary automatically guarantee CPU or memory isolation?
- Why does Orders store ProductRef rather than a relationship to Catalog Product?
- Why may Orders own unitPrice while Catalog owns currentPrice?
- What proves the free lab is not a composite database?
Review the answers
1. Exactly one standard database per Community installation/DBMS.
2. No. Multiple databases in one DBMS share server resources; separate processes/hosts can provide stronger resource controls but still need explicit quotas/capacity planning.
3. Relationships cannot span databases/graphs; the proxy key represents identity across the boundary.
4. They are different facts with different invariants: charged historical price versus mutable catalog price.
5. It uses three Community DBMS instances and explicit application queries/joins; no composite execution context or USE-based constituent query exists.
Production judgment
| Decision surface | Production questions |
|---|---|
| graph/workload fit | Does domain separation reduce ownership/capacity coupling, or does it turn the dominant traversal into repeated remote joins? |
| correctness | Which invariants remain atomic inside one graph, and which become asynchronous/application-coordinated across domains? |
| model/cardinality/degree | Which high-degree relationships must remain local for traversal cost and invariant enforcement? Which references are safe as proxy keys? |
| latency | How much p95/p99 is local graph work versus remote constituent/network/application join time? What happens during remote tail spikes? |
| transactions/concurrency | Which writes must commit together? Composite transactions may read many graphs but update only one constituent; plan compensating/outbox workflows elsewhere. |
| memory/resources | Do multiple databases share a DBMS resource envelope? Do separate instances need independent memory/page-cache/process budgets and capacity headroom? |
| CPU/disk/network | Does federation move filtering to constituents or ship excessive rows across network boundaries? Are region/zone egress and serialization material? |
| indexes/constraints | Are stable IDs constrained and indexed independently on each owning/proxy graph? No cross-graph relationship or uniqueness constraint exists. |
| driver | Are database/alias names explicit, drivers long-lived, timeouts/retries bounded, and per-domain timings correlated? |
| security/tenant risk | Are ACCESS/MATCH/procedure rights granted on every required constituent? How are remote credentials/OIDC forwarding, TLS and tenant boundaries governed? |
| backup/recovery | Can every constituent be restored and reconciled independently? A composite alias layer is not a backup of its target stores. |
| observability | Can operators attribute latency/errors to the root composite scope and each local/remote constituent without hiding network waits? |
| testing/failure injection | Have one-domain-down, stale proxy, version mismatch, denied constituent, slow remote graph and ambiguous cross-service write cases been tested safely? |
| version/edition/Aura | Is the design pinned to self-managed Enterprise composite support? What is the fallback for Community/Aura or mixed-version constituents? |
| licensing/cost/migration | Does federation justify Enterprise/ops/network cost, and can the split be rolled back or re-partitioned without breaking identity contracts? |
Summary and next step
You now have honest domain boundaries and a free reproducible multi-instance model. Lesson 2 adds the Enterprise composite layer and shows exactly what it contributes: logical constituent naming and one query execution context—not a new place where AtlasMart data is stored.
Authoritative references
- Current Neo4j versions — Current self-managed release 2026.07.1 and 5.26.30 LTS snapshot.
- Database administration — Current database/transaction-domain model and the one-standard-database Community versus multi-database Enterprise boundary.
- Create standard databases — Enterprise-only self-managed CREATE DATABASE semantics and system-database administration.
- Show databases — SHOW DATABASES fields including type, role, writer, status, aliases and constituents.
- Composite database concepts — Enterprise-only, unavailable-on-Aura composite semantics, local/remote constituents, compatibility, transactions and proxy-node federation.
- Create composite databases — CREATE COMPOSITE DATABASE and current default Cypher language behavior.
- Query composite databases — USE/CALL graph selection, graph functions, update restrictions, root-scope limits and runtime behavior.
- Composite aliases — Local and remote constituent aliases, SHOW ALIASES evidence, namespace rules and alias limitations.
- Standard aliases — Local/remote alias behavior, credentials, access visibility and current OIDC-forwarding option for remote aliases.
- Composite RBAC — Access must be granted to the composite and constituents; remote constituents enforce remote-user RBAC.
- Composite tutorial — Official federation/sharding example and proxy-node model.
- Database alias command syntax — Current SHOW/CREATE/ALTER alias command forms.
- Cypher USE clause — Current graph-selection clause semantics for composite/federated queries.
- Python driver manual — Official driver lifecycle, sessions, parameters, database selection and application-side orchestration.
- Backup and restore — Operational reminder that constituent data stores—not a composite alias layer—are the recovery units.