Chapter 17 · Clustering, High Availability, Routing, Consensus, and Failure Domains
Server Discovery, Routing, Leader Changes, Read Scaling, and Causal Consistency Expectations
Follow an AtlasMart request from discovery through a routing table to the current writer or a read endpoint, then use bookmarks deliberately when read-after-write causality matters across sessions.
AtlasMart’s API is now pointed at a cluster. The team asks whether clients should pin the writer hostname, load-balance Bolt connections themselves, or send all reads to the nearest secondary. All three shortcuts can fail. A routing-aware official driver starts from one or more reachable addresses, obtains the database’s current routing table, sends writes to the writer, sends eligible reads to reader endpoints, and refreshes topology as leadership or membership changes.
Learning outcomes
Separate server discovery from client routing and explain why both exist.
Interpret writers/readers/routers and routing-table TTL without hard-coding a writer host.
Use routing-aware driver access modes for AtlasMart reads and writes.
Explain secondary lag and use bookmarks only when a causal read dependency exists.
Diagnose leader changes and routing failures from server, database, driver and application evidence.
Lab baseline and scope
The continuity environment remains
Neo4j Community 2026.07.1, database
neo4j, explicit CYPHER 25 where
language selection matters, container
atlasmart-neo4j, loopback Bolt
bolt://127.0.0.1:7687, synthetic local credential
neo4j / atlasmart-course-2026, Java 21 or 25, and
Python driver 6.3. Current 5.26 LTS is 5.26.30.
Community is intentionally retained for the mandatory path,
but it cannot run a Neo4j DBMS cluster; cluster output is
therefore never fabricated.
Neo4j DBMS clustering, multi-database topology management,
fine-grained server management, and the cluster administrative
evidence used in this chapter are Enterprise Edition
capabilities. The mandatory path is a deterministic local
simulation/trace plus application-level routing/bookmark
exercises that do not pretend Community is clustered. The
optional licensed lab assumes Neo4j Enterprise
2026.07.1 on isolated test servers and uses
commands documented for the current release. Aura manages
topology for you, so self-managed server allocation,
discovery, seeding, and failure-domain commands do not map
one-for-one to Aura operations.
AtlasMart remains the same customer/order/product graph used
throughout earlier chapters. This chapter changes deployment
topology, not business identity. The optional licensed topology
uses a separate database name atlasmart on
disposable test infrastructure so cluster administration cannot
damage the single-server continuity database.
| Term | Precise meaning in current Neo4j clustering |
|---|---|
| server | A Neo4j DBMS process/machine that can host allocations for multiple databases. A server is not permanently “a primary” or “a secondary”. |
| database topology | The set of primary and secondary allocations for one database. Each standard database has its own topology. |
| primary | A database copy that participates in fault-tolerant write processing and is eligible to become that database’s writer/leader. |
| writer / leader | Exactly one eligible primary at a time orders writes for a database. Different databases can have different leaders on the same physical cluster. |
| secondary | An asynchronously replicated database copy used primarily for read scaling; it does not vote in that database’s write quorum and can lag. |
| quorum / simple majority |
The minimum number of primaries needed to
acknowledge/advance fault-tolerant writes. For
N primaries, majority is
floor(N/2)+1.
|
| routing table | A driver-consumable list of routers/readers/writers for a database, cached for a TTL and refreshed when topology/leadership changes. |
| bookmark | A causal token used to ensure later work does not execute before the represented committed state is available. It is not global serializability. |
| failure domain | Infrastructure expected to fail together: host, rack, availability zone, data center, region, network path, or power domain. |
| Assumption | Pinned value / rule |
|---|---|
| current server | Neo4j 2026.07.1; 5.26.30 is the current LTS comparison line |
| Cypher | Cypher 25 for current administrative examples; Cypher 5 is the frozen compatibility line |
| mandatory path | Community/local deterministic simulation and trace; no successful Enterprise output is claimed |
| optional cluster | Enterprise 2026.07.1; at least three servers for primary quorum; fourth server used when demonstrating a secondary |
| driver |
Python driver 6.3 with neo4j:// routing URI
for a cluster; bolt:// is direct to one
address
|
| auth/TLS | synthetic credentials only; production/remote cluster transport requires properly verified TLS and intra-cluster encryption as appropriate |
| plugins | APOC/GDS not required for clustering lessons |
| measurement | all failover/lag/latency/alert values must be measured on the learner’s licensed environment; trace outputs are explicitly labeled deterministic simulations |
1. Discovery answers “who is in the cluster?”
A server joining a cluster needs contact points so it can discover topology. Current Neo4j supports resolver strategies including explicit lists, DNS and Kubernetes services. The discovery service continuously exchanges topology information after bootstrap. Discovery endpoints are cluster-internal membership plumbing; an application should not use that list as an improvised read/write load balancer.
# Current LIST discovery shape on each licensed server.
server.default_advertised_address=server01.example.test # change per server
dbms.cluster.discovery.resolver_type=LIST
dbms.cluster.endpoints=server01.example.test:6000,server02.example.test:6000,server03.example.test:6000,server04.example.test:6000
# DNS and Kubernetes resolver types are alternatives when appropriate.
# Discovery finds cluster members; it is not the same thing as driver routing.
The old discovery service v1 was removed in Neo4j 2025.01.
Current 2026.07 examples must use the current
discovery-service settings; do not revive removed
server.discovery.* configuration from old
tutorials.
2. Routing answers “where should this transaction go?”
The routing table is database-specific and contains endpoint capabilities: writers, readers and routers. There is usually one writer. Readers can include secondaries and, depending on routing configuration and topology, non-writer primaries. Routers are endpoints that can answer routing discovery. The table has a TTL, so the driver caches it for a bounded period and also refreshes when failures/topology changes invalidate its view.
:use system
CALL dbms.routing.getRoutingTable({}, 'atlasmart')
YIELD ttl, servers
RETURN ttl, servers;
| Routing role | Application interpretation | Do not assume |
|---|---|---|
| WRITE | current writer endpoint(s) for the database | that the same host remains writer forever |
| READ | eligible endpoints for read-mode work | that every reader has identical replication freshness without causal coordination |
| ROUTE | endpoints that can supply routing information | that a router must host the database allocation used for every query |
| TTL | how long callers may cache this view | that the table is immutable until TTL expiry after a known failure |
3. neo4j:// and access mode express routing intent
For a cluster, use a routing-aware neo4j:// URI
with the official driver. bolt:// is a direct
connection to one address and is useful for deliberate
direct-server operations, but it does not provide the same
client-side routing behavior. In the Python driver,
execute_query() defaults toward writers; read-only
calls should explicitly request read routing when appropriate.
from neo4j import GraphDatabase, RoutingControl
URI = "neo4j://server01.example.test:7687" # routing-aware scheme
AUTH = ("atlasmart_app", "<synthetic-test-secret>")
with GraphDatabase.driver(URI, auth=AUTH) as driver:
driver.verify_connectivity()
# Writes are routed to the current writer.
records, summary, keys = driver.execute_query(
"""
MERGE (m:HaMarker {markerId: $id})
SET m.updatedAt = datetime()
RETURN m.markerId AS id
""",
id="HA-TRACE-17",
database_="atlasmart",
)
# Explicitly request read routing for read-only work.
records, summary, keys = driver.execute_query(
"MATCH (m:HaMarker {markerId:$id}) RETURN m.markerId AS id",
id="HA-TRACE-17",
database_="atlasmart",
routing_=RoutingControl.READ,
)
Read/write routing is not authorization. A transaction routed in read mode should contain read-only work, but access control still comes from Neo4j privileges and application policy—not from treating routing mode as a permission system.
4. Leadership changes are normal, not exceptional architecture
If the writer stops or becomes unsuitable, another primary can be elected as long as quorum exists. The database writer can therefore move between servers. A robust driver refreshes routing and managed transactions may retry classified transient failures. Your code must not cache “server02 is the writer” in configuration or DNS.
| Time | Cluster event | Correct client behavior | Evidence to correlate |
|---|---|---|---|
| t0 | server01 is writer | write route contains server01 | SHOW DATABASE, routing table, driver debug |
| t1 | server01 lost | in-flight request may fail/turn ambiguous | transaction status/error code, server logs |
| t2 | server02 elected | routing view refreshes | writer=true moves to server02 |
| t3 | managed retry if safe | transaction re-executes only for retryable classification | driver retry log + idempotent business invariant |
| t4 | server01 returns | it catches up; it need not regain leadership | replication/catch-up status |
Pinning a writer hostname can turn normal leader movement into an application outage. Repair: routing-aware URI, current official driver, explicit database, bounded retry policy, idempotent transaction work and telemetry that shows routing refresh/retry events.
5. Read scaling introduces a freshness decision
Secondaries replicate asynchronously, so they can lag. That is often acceptable for product browse analytics, but not necessarily for “place order, immediately fetch the order from another session.” A bookmark represents a committed causal state. Passing or managing the bookmark causes later work to wait until that state is available on the chosen server. This is causal coordination, not a promise that all concurrent global operations are serialized.
from neo4j import GraphDatabase, RoutingControl
URI = "neo4j://server01.example.test:7687"
AUTH = ("atlasmart_app", "<synthetic-test-secret>")
with GraphDatabase.driver(URI, auth=AUTH) as driver:
# execute_query uses the driver's bookmark manager by default.
driver.execute_query(
"MERGE (:HaMarker {markerId:$id})",
id="HA-CAUSAL-17",
database_="atlasmart",
)
# Causally chained: this read waits as needed for the represented state.
rows, summary, keys = driver.execute_query(
"MATCH (m:HaMarker {markerId:$id}) RETURN m.markerId AS id",
id="HA-CAUSAL-17",
database_="atlasmart",
routing_=RoutingControl.READ,
)
| AtlasMart read | Can tolerate lag? | Bookmark/route decision |
|---|---|---|
| catalog browse count | often yes | read routing; bookmark may be unnecessary |
| immediate confirmation after order write | usually no | causally chain to the write/bookmark |
| fraud workflow depending on two prior writes | depends on workflow | coordinate bookmarks or group the dependent work transactionally |
| offline dashboard | usually yes within stated freshness SLO | read route to scale; measure replication lag |
6. Deterministic routing trace
Without an Enterprise cluster, practice the state machine instead of inventing output. Start with the following trace and answer where each request should go.
| Step | Writer set | Reader set | Request | Expected decision |
|---|---|---|---|---|
| 1 | {P1} | {P2,P3,S1} | write order O-9001 | P1 |
| 2 | {P1} | {P2,P3,S1} | read stale-tolerant catalog | one eligible reader |
| 3 | {} | {P2,P3,S1} | write during election | wait/fail/retry; no valid writer yet |
| 4 | {P2} | {P1,P3,S1} | retry idempotent order write | P2 after refreshed route |
| 5 | {P2} | {P1,P3,S1} | causal read requiring O-9001 | chosen reader must first satisfy bookmark |
It validates routing and causal-dependency reasoning. It does not measure election time, routing TTL refresh timing, secondary lag, or driver retry latency. Those are optional licensed-lab measurements.
7. Optional licensed evidence capture
:use system
SHOW DATABASE atlasmart YIELD address, role, writer, currentStatus, replicationLag RETURN * ORDER BY address;
CALL dbms.routing.getRoutingTable({}, 'atlasmart') YIELD ttl, servers RETURN ttl, servers;
| Capture before/after writer loss | Why |
|---|---|
| SHOW DATABASE rows | proves which allocation is writer and which are healthy |
| routing table | proves client-visible writer/reader endpoints |
| driver logs/retry counts | proves client reaction rather than only server state |
| application request ID / marker ID | ties a business operation to retries and final state |
| bookmark/read result | proves causal acceptance condition for the specific flow |
8. Verification checklist
- Discovery settings are not used as application routing logic.
- The application uses a routing-aware URI and one maintained driver.
- Reads explicitly request read routing when stale-tolerant/read-only.
- Bookmarks are applied only where causality requires them and their latency cost is measured.
- Leader change evidence correlates database state, routing table, driver retry and business invariant.
Check your understanding
- What does cluster discovery solve?
- What does the routing table solve?
- Should the application pin a writer hostname?
- Do bookmarks make every cluster read globally serializable?
- Why might a secondary read wait when a bookmark is supplied?
Review the answers
1. How servers find and continuously learn about other cluster members/topology.
2. Where a client can route reader, writer and routing requests for a specific database.
3. No. Leadership may change; use routing-aware drivers.
4. No. They coordinate causal dependencies represented by the bookmark.
5. The selected allocation may need to catch up to at least the bookmarked committed state.
Production judgment
| Decision surface | Questions that must be answered before production |
|---|---|
| graph/workload fit | Which writes need fault tolerance? How much read scaling is needed? Does graph fan-out make reader CPU/page-cache the bottleneck rather than topology? |
| correctness | Which operations need causal read-after-write? Which externally visible side effects require idempotency/reconciliation after an ambiguous client outcome? |
| topology/cardinality | How many primaries satisfy failure tolerance? Are secondaries justified by read demand? Are dense/hot entities producing lock contention that topology will not solve? |
| latency | What are write p50/p95/p99 with quorum across actual zones? What read tail latency is added by lag/bookmark waiting/routing refresh? |
| transactions/concurrency | How do leader changes, deadlocks, retries and long transactions interact? Is transaction work idempotent under managed retries? |
| memory/CPU/disk/network | Can remaining servers absorb a failed member? Is there space/bandwidth for store copy and catch-up while serving traffic? |
| indexes/constraints | Are access paths and uniqueness contracts identical/ONLINE across allocations after seeding or recovery? |
| driver | Is one maintained driver reused? Are routing URI, database selection, read/write modes, pool sizes, acquisition timeouts, max retry time and bookmarks intentional? |
| security | Are Bolt, HTTPS and intra-cluster links encrypted as required? Are server-management privileges and certificates isolated from application credentials? |
| backup/recovery | Replication is not backup. Where are off-host/immutable backups and when was an isolated restore last proven? |
| observability | Do alerts cover writer absence, unavailable primaries, replication lag, store copy, routing failures, retry spikes, queueing and capacity headroom? |
| testing/failure injection | Have host loss, reader loss, writer loss, majority loss, partition, slow network, maintenance and ambiguous commit outcomes been tested safely? |
| version/edition | Are server/Java/driver/APOC/GDS/Cypher/discovery settings compatible with the exact release and license? |
| Aura/self-managed | Which topology choices are operator-owned self-managed responsibilities versus managed-service controls/SLOs? |
| cost/migration | What is the cost of extra primaries/secondaries/zones and cross-zone traffic, and how will topology changes or rollback be staged? |
Summary and next step
Routing makes topology changes survivable to clients, but only if failure domains and quorum are sound. Lesson 3 moves from individual server loss to zones, regions and network partitions.
Authoritative references
- Current Neo4j versions — Current Neo4j 2026.07.1, 5.26.30 LTS, Cypher Shell and GDS release snapshot.
- Clustering architecture — Current server/database decoupling, primaries, secondaries, majority writes, and causal consistency.
- Deploy a basic cluster — Current discovery/bootstrap configuration and TOPOLOGY examples.
- Cluster server discovery — LIST/DNS/Kubernetes discovery and the current discovery-service model.
- Leadership, routing, and load balancing — Per-database Raft leadership, routing tables, readers/writers/routers, and routing policy.
- Managing databases in a cluster — Primary/secondary allocation and topology changes.
- Managing servers in a cluster — Server enablement, constraints, deallocation, removal, and operational lifecycle.
- Server management command syntax — SHOW/ENABLE/ALTER/DEALLOCATE/REALLOCATE/DROP server commands.
- Show databases — Role, writer, allocation counts, replication lag, and database-state evidence.
- Monitor databases in a cluster — SHOW DATABASES-based per-allocation monitoring and status interpretation.
- Resilient multi-data-center cluster — Failure-domain placement, latency/fault-tolerance tradeoffs, and recommended/anti-pattern layouts.
- Cluster disaster recovery — Recovery when allocations or whole failure domains are lost.
- Built-in procedures — Current routing, cordon, deallocation and related procedure names/deprecations.
- Python driver transactions/routing — Read/write routing and explicit database selection in the official driver.
- Python driver bookmarks — Bookmark propagation and causal-consistency coordination across sessions.