Chapter 17 · Clustering, High Availability, Routing, Consensus, and Failure Domains

Failure Domains, Zones/Regions, Network Partitions, Quorum Loss, and Split-Brain Avoidance

Design AtlasMart placement around real failure domains and quorum mathematics, reason through asymmetric network partitions, and understand why majority consensus prevents two independent writers from safely committing.

Advanced220–290 minutesFailure-domain partition labNeo4j 2026.07.1 Community mandatory trace · Enterprise cluster optionalCypher 25 · current primary/secondary topology terminologyJava 21/25 · Python driver 6.3Last reviewed: September 2026

AtlasMart expands into multiple availability zones. A diagram with three server icons in different boxes looks resilient, but appearance is not enough. You need to ask which failures are correlated, where the primary majority remains after each failure, what cross-zone round trips cost every write, and what a network partition does when both sides are still running.

Learning outcomes

01

Translate host/zone/data-center/network failures into reachable-primary counts and write-quorum outcomes.

02

Explain how majority consensus prevents two partitioned sides from both safely committing independent write histories.

03

Compare 3-primary and 5-primary placement patterns across failure domains and maintenance scenarios.

04

Separate secondary read resilience from primary write fault tolerance and account for asynchronous lag.

05

Define failure-injection acceptance criteria that measure availability, correctness, alerts, latency and recovery rather than process survival.

Lab baseline and scope

Chapter 17 baseline · reviewed 9 September 2026

The continuity environment remains Neo4j Community 2026.07.1, database neo4j, explicit CYPHER 25 where language selection matters, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic local credential neo4j / atlasmart-course-2026, Java 21 or 25, and Python driver 6.3. Current 5.26 LTS is 5.26.30. Community is intentionally retained for the mandatory path, but it cannot run a Neo4j DBMS cluster; cluster output is therefore never fabricated.

Edition and evidence boundary

Neo4j DBMS clustering, multi-database topology management, fine-grained server management, and the cluster administrative evidence used in this chapter are Enterprise Edition capabilities. The mandatory path is a deterministic local simulation/trace plus application-level routing/bookmark exercises that do not pretend Community is clustered. The optional licensed lab assumes Neo4j Enterprise 2026.07.1 on isolated test servers and uses commands documented for the current release. Aura manages topology for you, so self-managed server allocation, discovery, seeding, and failure-domain commands do not map one-for-one to Aura operations.

AtlasMart remains the same customer/order/product graph used throughout earlier chapters. This chapter changes deployment topology, not business identity. The optional licensed topology uses a separate database name atlasmart on disposable test infrastructure so cluster administration cannot damage the single-server continuity database.

Term Precise meaning in current Neo4j clustering
server A Neo4j DBMS process/machine that can host allocations for multiple databases. A server is not permanently “a primary” or “a secondary”.
database topology The set of primary and secondary allocations for one database. Each standard database has its own topology.
primary A database copy that participates in fault-tolerant write processing and is eligible to become that database’s writer/leader.
writer / leader Exactly one eligible primary at a time orders writes for a database. Different databases can have different leaders on the same physical cluster.
secondary An asynchronously replicated database copy used primarily for read scaling; it does not vote in that database’s write quorum and can lag.
quorum / simple majority The minimum number of primaries needed to acknowledge/advance fault-tolerant writes. For N primaries, majority is floor(N/2)+1.
routing table A driver-consumable list of routers/readers/writers for a database, cached for a TTL and refreshed when topology/leadership changes.
bookmark A causal token used to ensure later work does not execute before the represented committed state is available. It is not global serializability.
failure domain Infrastructure expected to fail together: host, rack, availability zone, data center, region, network path, or power domain.
Assumption Pinned value / rule
current server Neo4j 2026.07.1; 5.26.30 is the current LTS comparison line
Cypher Cypher 25 for current administrative examples; Cypher 5 is the frozen compatibility line
mandatory path Community/local deterministic simulation and trace; no successful Enterprise output is claimed
optional cluster Enterprise 2026.07.1; at least three servers for primary quorum; fourth server used when demonstrating a secondary
driver Python driver 6.3 with neo4j:// routing URI for a cluster; bolt:// is direct to one address
auth/TLS synthetic credentials only; production/remote cluster transport requires properly verified TLS and intra-cluster encryption as appropriate
plugins APOC/GDS not required for clustering lessons
measurement all failover/lag/latency/alert values must be measured on the learner’s licensed environment; trace outputs are explicitly labeled deterministic simulations

1. A failure domain is a shared cause, not a drawing box

A failure domain is anything likely to fail together: a host, hypervisor, rack switch, power feed, availability zone, data center or regional network. Putting three primaries on three virtual machines does not provide three independent failure domains if all VMs share one physical host or one power/network dependency. The topology review must map causes to allocations.

Placement Independent against Still correlated against
3 VMs / one host individual process crash host, storage, power, top-of-rack failure
3 hosts / one rack host failure rack power/network and site failure
3 zones / one region single-zone outage regional control-plane/network disasters
3 regions regional outage global identity/DNS/application dependencies; cross-region write latency

2. Majority consensus is the split-brain safety rule

Consider a network partition. Both sides may have live Neo4j processes, but only a side that can communicate with a majority of the database’s primaries can continue fault-tolerant commits. In a 3-primary 2|1 partition, the side with two can retain quorum; the single-primary side cannot independently become a safe writer. That sacrifices availability on the minority side to preserve one committed history.

Python · deterministic network-partition trace
from math import floor

# Each tuple: (layout name, primaries total, primaries reachable by side A, side B)
partitions = [
    ("3P split 2|1", 3, 2, 1),
    ("5P split 3|2", 5, 3, 2),
    ("3P split 1|1 with one down", 3, 1, 1),
]

for name, total, a, b in partitions:
    majority = floor(total/2)+1
    print(name, {
        "majority": majority,
        "side_A_can_commit": a >= majority,
        "side_B_can_commit": b >= majority,
    })
What this does not prove

The script does not model real Raft election timeouts, partially delayed packets, disk stalls, membership changes, retries or client timeouts. It only proves the majority condition each partition side can satisfy.

3. Three primaries: tolerate one arbitrary primary loss

With three primaries, majority is two. This is compact and often appropriate when the required failure tolerance is one primary. But maintenance consumes the same headroom as a failure: if one primary is deliberately offline, a second failure removes quorum. Therefore a maintenance runbook must state the temporary reduction in fault tolerance.

State Live primaries Can write? Fault headroom
healthy 3P 3 yes one arbitrary primary can fail
one primary maintenance 2 yes zero additional primary failures tolerated
one maintenance + another failure 1 no quorum lost

4. Five primaries: more fault headroom, more coordination

Five primaries need a majority of three and can tolerate two arbitrary primary failures. That can be useful for combinations of maintenance and failure, or carefully designed multi-data-center placement. It also means more copies, more catch-up work and potentially more expensive quorum paths. “Five is safer” is incomplete unless the placement and network latency actually satisfy the intended failure model.

Failure model 3 primaries 5 primaries
lose 1 primary write remains available write remains available
lose 2 primaries write unavailable write remains available
maintenance 1 + failure 1 write unavailable write remains available
cross-zone quorum latency depends on placement can involve more/longer paths; measure

5. Multi-data-center placement: availability versus latency

Neo4j’s current guidance emphasizes that primary placement across distant data centers increases write latency because the writer must obtain quorum confirmation. Geo-distributing primaries can tolerate site loss when a majority remains, but every transactional write pays for the topology’s network path. A design must therefore define both fault-tolerance objective and latency SLO.

Pattern Strength Risk / tradeoff
all primaries one DC + remote secondaries fast local writes; remote read resilience loss of primary DC removes write availability until recovery/recreation
3 primaries across 3 DCs can lose one DC and preserve 2-of-3 quorum cross-DC write coordination on every commit
5 primaries 2+2+1 across 3 DCs can tolerate any one DC in the documented recommended shape higher cost/coordination and capacity requirement
balanced 2-DC primary split looks symmetric loss/partition of either DC can lose majority; documentation strongly warns against this pattern

6. Secondaries help reads, not partition quorum

If a remote site contains only secondaries, local clients may continue to serve stale-tolerant reads when the primary site is unreachable, subject to routing and the local database state. They cannot form a new write majority. Promoting or recreating databases after disaster is an explicit recovery operation with possible data-loss considerations; it is not the normal meaning of a secondary.

Wrong approach

“If the primary data center fails, we will just write to the secondary.” That bypasses the consensus model. Repair: design sufficient primary placement for the required automatic write availability, or document a disaster-recovery/recreation procedure with RPO implications.

7. Optional licensed partition acceptance matrix

Use safe, isolated failure injection. Prefer stopping an Enterprise test process/container or disconnecting an isolated test network segment. Do not change the host firewall or production routes for a course exercise. Before each injection capture topology, current writer, baseline request success, backup status and reset procedure.

Injection Expected safety invariant Availability signal Recovery signal
stop secondary no data divergence writes continue; read capacity reduced secondary catches up after restart
stop one of 3 primaries single committed history writes continue after/without leader change primary returns/catches up
partition 2P from 1P only 2P side can maintain majority minority write requests fail/unavailable partition heals and minority catches up
lose 2 of 3 primaries no unsafe forced writer writes stop restore quorum/recreate per documented recovery

8. Failure-domain verification checklist

  • Every primary is mapped to an independent host/zone/site cause.
  • For each planned failure, surviving primaries are counted and compared with majority.
  • Write latency is measured across the real quorum path.
  • Secondaries are not counted as write voters.
  • Maintenance states are included in failure combinations.
  • Backups remain independent from the cluster and restore-tested.

Check your understanding

  1. Why does the minority side of a partition stop writing?
  2. What does three primaries tolerate?
  3. Why can maintenance reduce HA?
  4. Do remote secondaries solve primary-site write loss?
  5. What must accompany cross-region primaries?
Review the answers

1. Because it cannot satisfy the primary majority required to safely advance the committed history.

2. One arbitrary primary failure while retaining a 2-of-3 majority.

3. A deliberately unavailable primary consumes the same quorum headroom as a failure.

4. No. They can help read scaling/resilience but do not form write quorum.

5. Measured write-latency impact, network/failure-domain analysis, capacity and recovery planning.

Production judgment

Decision surface Questions that must be answered before production
graph/workload fit Which writes need fault tolerance? How much read scaling is needed? Does graph fan-out make reader CPU/page-cache the bottleneck rather than topology?
correctness Which operations need causal read-after-write? Which externally visible side effects require idempotency/reconciliation after an ambiguous client outcome?
topology/cardinality How many primaries satisfy failure tolerance? Are secondaries justified by read demand? Are dense/hot entities producing lock contention that topology will not solve?
latency What are write p50/p95/p99 with quorum across actual zones? What read tail latency is added by lag/bookmark waiting/routing refresh?
transactions/concurrency How do leader changes, deadlocks, retries and long transactions interact? Is transaction work idempotent under managed retries?
memory/CPU/disk/network Can remaining servers absorb a failed member? Is there space/bandwidth for store copy and catch-up while serving traffic?
indexes/constraints Are access paths and uniqueness contracts identical/ONLINE across allocations after seeding or recovery?
driver Is one maintained driver reused? Are routing URI, database selection, read/write modes, pool sizes, acquisition timeouts, max retry time and bookmarks intentional?
security Are Bolt, HTTPS and intra-cluster links encrypted as required? Are server-management privileges and certificates isolated from application credentials?
backup/recovery Replication is not backup. Where are off-host/immutable backups and when was an isolated restore last proven?
observability Do alerts cover writer absence, unavailable primaries, replication lag, store copy, routing failures, retry spikes, queueing and capacity headroom?
testing/failure injection Have host loss, reader loss, writer loss, majority loss, partition, slow network, maintenance and ambiguous commit outcomes been tested safely?
version/edition Are server/Java/driver/APOC/GDS/Cypher/discovery settings compatible with the exact release and license?
Aura/self-managed Which topology choices are operator-owned self-managed responsibilities versus managed-service controls/SLOs?
cost/migration What is the cost of extra primaries/secondaries/zones and cross-zone traffic, and how will topology changes or rollback be staged?

Summary and next step

Quorum defines safety under failure, but day-two operations can accidentally consume that safety margin. Lesson 4 turns to seeding, membership changes, cordoning, deallocation, reallocation and maintenance headroom.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.