Chapter 17 · Clustering, High Availability, Routing, Consensus, and Failure Domains
Failure Domains, Zones/Regions, Network Partitions, Quorum Loss, and Split-Brain Avoidance
Design AtlasMart placement around real failure domains and quorum mathematics, reason through asymmetric network partitions, and understand why majority consensus prevents two independent writers from safely committing.
AtlasMart expands into multiple availability zones. A diagram with three server icons in different boxes looks resilient, but appearance is not enough. You need to ask which failures are correlated, where the primary majority remains after each failure, what cross-zone round trips cost every write, and what a network partition does when both sides are still running.
Learning outcomes
Translate host/zone/data-center/network failures into reachable-primary counts and write-quorum outcomes.
Explain how majority consensus prevents two partitioned sides from both safely committing independent write histories.
Compare 3-primary and 5-primary placement patterns across failure domains and maintenance scenarios.
Separate secondary read resilience from primary write fault tolerance and account for asynchronous lag.
Define failure-injection acceptance criteria that measure availability, correctness, alerts, latency and recovery rather than process survival.
Lab baseline and scope
The continuity environment remains
Neo4j Community 2026.07.1, database
neo4j, explicit CYPHER 25 where
language selection matters, container
atlasmart-neo4j, loopback Bolt
bolt://127.0.0.1:7687, synthetic local credential
neo4j / atlasmart-course-2026, Java 21 or 25, and
Python driver 6.3. Current 5.26 LTS is 5.26.30.
Community is intentionally retained for the mandatory path,
but it cannot run a Neo4j DBMS cluster; cluster output is
therefore never fabricated.
Neo4j DBMS clustering, multi-database topology management,
fine-grained server management, and the cluster administrative
evidence used in this chapter are Enterprise Edition
capabilities. The mandatory path is a deterministic local
simulation/trace plus application-level routing/bookmark
exercises that do not pretend Community is clustered. The
optional licensed lab assumes Neo4j Enterprise
2026.07.1 on isolated test servers and uses
commands documented for the current release. Aura manages
topology for you, so self-managed server allocation,
discovery, seeding, and failure-domain commands do not map
one-for-one to Aura operations.
AtlasMart remains the same customer/order/product graph used
throughout earlier chapters. This chapter changes deployment
topology, not business identity. The optional licensed topology
uses a separate database name atlasmart on
disposable test infrastructure so cluster administration cannot
damage the single-server continuity database.
| Term | Precise meaning in current Neo4j clustering |
|---|---|
| server | A Neo4j DBMS process/machine that can host allocations for multiple databases. A server is not permanently “a primary” or “a secondary”. |
| database topology | The set of primary and secondary allocations for one database. Each standard database has its own topology. |
| primary | A database copy that participates in fault-tolerant write processing and is eligible to become that database’s writer/leader. |
| writer / leader | Exactly one eligible primary at a time orders writes for a database. Different databases can have different leaders on the same physical cluster. |
| secondary | An asynchronously replicated database copy used primarily for read scaling; it does not vote in that database’s write quorum and can lag. |
| quorum / simple majority |
The minimum number of primaries needed to
acknowledge/advance fault-tolerant writes. For
N primaries, majority is
floor(N/2)+1.
|
| routing table | A driver-consumable list of routers/readers/writers for a database, cached for a TTL and refreshed when topology/leadership changes. |
| bookmark | A causal token used to ensure later work does not execute before the represented committed state is available. It is not global serializability. |
| failure domain | Infrastructure expected to fail together: host, rack, availability zone, data center, region, network path, or power domain. |
| Assumption | Pinned value / rule |
|---|---|
| current server | Neo4j 2026.07.1; 5.26.30 is the current LTS comparison line |
| Cypher | Cypher 25 for current administrative examples; Cypher 5 is the frozen compatibility line |
| mandatory path | Community/local deterministic simulation and trace; no successful Enterprise output is claimed |
| optional cluster | Enterprise 2026.07.1; at least three servers for primary quorum; fourth server used when demonstrating a secondary |
| driver |
Python driver 6.3 with neo4j:// routing URI
for a cluster; bolt:// is direct to one
address
|
| auth/TLS | synthetic credentials only; production/remote cluster transport requires properly verified TLS and intra-cluster encryption as appropriate |
| plugins | APOC/GDS not required for clustering lessons |
| measurement | all failover/lag/latency/alert values must be measured on the learner’s licensed environment; trace outputs are explicitly labeled deterministic simulations |
1. A failure domain is a shared cause, not a drawing box
A failure domain is anything likely to fail together: a host, hypervisor, rack switch, power feed, availability zone, data center or regional network. Putting three primaries on three virtual machines does not provide three independent failure domains if all VMs share one physical host or one power/network dependency. The topology review must map causes to allocations.
| Placement | Independent against | Still correlated against |
|---|---|---|
| 3 VMs / one host | individual process crash | host, storage, power, top-of-rack failure |
| 3 hosts / one rack | host failure | rack power/network and site failure |
| 3 zones / one region | single-zone outage | regional control-plane/network disasters |
| 3 regions | regional outage | global identity/DNS/application dependencies; cross-region write latency |
2. Majority consensus is the split-brain safety rule
Consider a network partition. Both sides may have live Neo4j processes, but only a side that can communicate with a majority of the database’s primaries can continue fault-tolerant commits. In a 3-primary 2|1 partition, the side with two can retain quorum; the single-primary side cannot independently become a safe writer. That sacrifices availability on the minority side to preserve one committed history.
from math import floor
# Each tuple: (layout name, primaries total, primaries reachable by side A, side B)
partitions = [
("3P split 2|1", 3, 2, 1),
("5P split 3|2", 5, 3, 2),
("3P split 1|1 with one down", 3, 1, 1),
]
for name, total, a, b in partitions:
majority = floor(total/2)+1
print(name, {
"majority": majority,
"side_A_can_commit": a >= majority,
"side_B_can_commit": b >= majority,
})
The script does not model real Raft election timeouts, partially delayed packets, disk stalls, membership changes, retries or client timeouts. It only proves the majority condition each partition side can satisfy.
3. Three primaries: tolerate one arbitrary primary loss
With three primaries, majority is two. This is compact and often appropriate when the required failure tolerance is one primary. But maintenance consumes the same headroom as a failure: if one primary is deliberately offline, a second failure removes quorum. Therefore a maintenance runbook must state the temporary reduction in fault tolerance.
| State | Live primaries | Can write? | Fault headroom |
|---|---|---|---|
| healthy 3P | 3 | yes | one arbitrary primary can fail |
| one primary maintenance | 2 | yes | zero additional primary failures tolerated |
| one maintenance + another failure | 1 | no | quorum lost |
4. Five primaries: more fault headroom, more coordination
Five primaries need a majority of three and can tolerate two arbitrary primary failures. That can be useful for combinations of maintenance and failure, or carefully designed multi-data-center placement. It also means more copies, more catch-up work and potentially more expensive quorum paths. “Five is safer” is incomplete unless the placement and network latency actually satisfy the intended failure model.
| Failure model | 3 primaries | 5 primaries |
|---|---|---|
| lose 1 primary | write remains available | write remains available |
| lose 2 primaries | write unavailable | write remains available |
| maintenance 1 + failure 1 | write unavailable | write remains available |
| cross-zone quorum latency | depends on placement | can involve more/longer paths; measure |
5. Multi-data-center placement: availability versus latency
Neo4j’s current guidance emphasizes that primary placement across distant data centers increases write latency because the writer must obtain quorum confirmation. Geo-distributing primaries can tolerate site loss when a majority remains, but every transactional write pays for the topology’s network path. A design must therefore define both fault-tolerance objective and latency SLO.
| Pattern | Strength | Risk / tradeoff |
|---|---|---|
| all primaries one DC + remote secondaries | fast local writes; remote read resilience | loss of primary DC removes write availability until recovery/recreation |
| 3 primaries across 3 DCs | can lose one DC and preserve 2-of-3 quorum | cross-DC write coordination on every commit |
| 5 primaries 2+2+1 across 3 DCs | can tolerate any one DC in the documented recommended shape | higher cost/coordination and capacity requirement |
| balanced 2-DC primary split | looks symmetric | loss/partition of either DC can lose majority; documentation strongly warns against this pattern |
6. Secondaries help reads, not partition quorum
If a remote site contains only secondaries, local clients may continue to serve stale-tolerant reads when the primary site is unreachable, subject to routing and the local database state. They cannot form a new write majority. Promoting or recreating databases after disaster is an explicit recovery operation with possible data-loss considerations; it is not the normal meaning of a secondary.
“If the primary data center fails, we will just write to the secondary.” That bypasses the consensus model. Repair: design sufficient primary placement for the required automatic write availability, or document a disaster-recovery/recreation procedure with RPO implications.
7. Optional licensed partition acceptance matrix
Use safe, isolated failure injection. Prefer stopping an Enterprise test process/container or disconnecting an isolated test network segment. Do not change the host firewall or production routes for a course exercise. Before each injection capture topology, current writer, baseline request success, backup status and reset procedure.
| Injection | Expected safety invariant | Availability signal | Recovery signal |
|---|---|---|---|
| stop secondary | no data divergence | writes continue; read capacity reduced | secondary catches up after restart |
| stop one of 3 primaries | single committed history | writes continue after/without leader change | primary returns/catches up |
| partition 2P from 1P | only 2P side can maintain majority | minority write requests fail/unavailable | partition heals and minority catches up |
| lose 2 of 3 primaries | no unsafe forced writer | writes stop | restore quorum/recreate per documented recovery |
8. Failure-domain verification checklist
- Every primary is mapped to an independent host/zone/site cause.
- For each planned failure, surviving primaries are counted and compared with majority.
- Write latency is measured across the real quorum path.
- Secondaries are not counted as write voters.
- Maintenance states are included in failure combinations.
- Backups remain independent from the cluster and restore-tested.
Check your understanding
- Why does the minority side of a partition stop writing?
- What does three primaries tolerate?
- Why can maintenance reduce HA?
- Do remote secondaries solve primary-site write loss?
- What must accompany cross-region primaries?
Review the answers
1. Because it cannot satisfy the primary majority required to safely advance the committed history.
2. One arbitrary primary failure while retaining a 2-of-3 majority.
3. A deliberately unavailable primary consumes the same quorum headroom as a failure.
4. No. They can help read scaling/resilience but do not form write quorum.
5. Measured write-latency impact, network/failure-domain analysis, capacity and recovery planning.
Production judgment
| Decision surface | Questions that must be answered before production |
|---|---|
| graph/workload fit | Which writes need fault tolerance? How much read scaling is needed? Does graph fan-out make reader CPU/page-cache the bottleneck rather than topology? |
| correctness | Which operations need causal read-after-write? Which externally visible side effects require idempotency/reconciliation after an ambiguous client outcome? |
| topology/cardinality | How many primaries satisfy failure tolerance? Are secondaries justified by read demand? Are dense/hot entities producing lock contention that topology will not solve? |
| latency | What are write p50/p95/p99 with quorum across actual zones? What read tail latency is added by lag/bookmark waiting/routing refresh? |
| transactions/concurrency | How do leader changes, deadlocks, retries and long transactions interact? Is transaction work idempotent under managed retries? |
| memory/CPU/disk/network | Can remaining servers absorb a failed member? Is there space/bandwidth for store copy and catch-up while serving traffic? |
| indexes/constraints | Are access paths and uniqueness contracts identical/ONLINE across allocations after seeding or recovery? |
| driver | Is one maintained driver reused? Are routing URI, database selection, read/write modes, pool sizes, acquisition timeouts, max retry time and bookmarks intentional? |
| security | Are Bolt, HTTPS and intra-cluster links encrypted as required? Are server-management privileges and certificates isolated from application credentials? |
| backup/recovery | Replication is not backup. Where are off-host/immutable backups and when was an isolated restore last proven? |
| observability | Do alerts cover writer absence, unavailable primaries, replication lag, store copy, routing failures, retry spikes, queueing and capacity headroom? |
| testing/failure injection | Have host loss, reader loss, writer loss, majority loss, partition, slow network, maintenance and ambiguous commit outcomes been tested safely? |
| version/edition | Are server/Java/driver/APOC/GDS/Cypher/discovery settings compatible with the exact release and license? |
| Aura/self-managed | Which topology choices are operator-owned self-managed responsibilities versus managed-service controls/SLOs? |
| cost/migration | What is the cost of extra primaries/secondaries/zones and cross-zone traffic, and how will topology changes or rollback be staged? |
Summary and next step
Quorum defines safety under failure, but day-two operations can accidentally consume that safety margin. Lesson 4 turns to seeding, membership changes, cordoning, deallocation, reallocation and maintenance headroom.
Authoritative references
- Current Neo4j versions — Current Neo4j 2026.07.1, 5.26.30 LTS, Cypher Shell and GDS release snapshot.
- Clustering architecture — Current server/database decoupling, primaries, secondaries, majority writes, and causal consistency.
- Deploy a basic cluster — Current discovery/bootstrap configuration and TOPOLOGY examples.
- Cluster server discovery — LIST/DNS/Kubernetes discovery and the current discovery-service model.
- Leadership, routing, and load balancing — Per-database Raft leadership, routing tables, readers/writers/routers, and routing policy.
- Managing databases in a cluster — Primary/secondary allocation and topology changes.
- Managing servers in a cluster — Server enablement, constraints, deallocation, removal, and operational lifecycle.
- Server management command syntax — SHOW/ENABLE/ALTER/DEALLOCATE/REALLOCATE/DROP server commands.
- Show databases — Role, writer, allocation counts, replication lag, and database-state evidence.
- Monitor databases in a cluster — SHOW DATABASES-based per-allocation monitoring and status interpretation.
- Resilient multi-data-center cluster — Failure-domain placement, latency/fault-tolerance tradeoffs, and recommended/anti-pattern layouts.
- Cluster disaster recovery — Recovery when allocations or whole failure domains are lost.
- Built-in procedures — Current routing, cordon, deallocation and related procedure names/deprecations.
- Python driver transactions/routing — Read/write routing and explicit database selection in the official driver.
- Python driver bookmarks — Bookmark propagation and causal-consistency coordination across sessions.