Chapter 17 · Clustering, High Availability, Routing, Consensus, and Failure Domains
Cluster Roles, Primaries/Secondaries or Current Topology Concepts, Consensus, Quorum, and Write Availability
Build the correct mental model for current Neo4j clustering: databases own topologies, primaries form the write-consensus group, secondaries scale reads asynchronously, and a majority—not copy count alone—determines write availability.
AtlasMart is preparing a flash-sale launch. A single Neo4j process can preserve ACID correctness, but one host failure would make the graph unavailable. The platform team asks for “three replicas plus a read replica.” That phrase sounds simple but hides the central design question: which copies are allowed to participate in committing writes, how many acknowledgements are required, and what happens when one copy disappears? Current Neo4j answers those questions per database, not per server.
High availability is not “more copies.” It is a protocol plus a topology. Start with the database’s primary count, majority requirement, failure domains and client routing, then decide whether secondaries add useful read capacity.
Learning outcomes
Distinguish servers from per-database primary/secondary allocations and identify the elected writer.
Compute majority/quorum for one, three and five primaries and predict write availability after failures.
Explain synchronous primary acknowledgement versus asynchronous secondary replication without confusing either with backup.
Use SHOW DATABASES/SHOW SERVERS evidence in an optional licensed cluster instead of inferring health from process count.
Choose a topology by explicit fault-tolerance and latency requirements rather than copying a fixed server count.
Lab baseline and scope
The continuity environment remains
Neo4j Community 2026.07.1, database
neo4j, explicit CYPHER 25 where
language selection matters, container
atlasmart-neo4j, loopback Bolt
bolt://127.0.0.1:7687, synthetic local credential
neo4j / atlasmart-course-2026, Java 21 or 25, and
Python driver 6.3. Current 5.26 LTS is 5.26.30.
Community is intentionally retained for the mandatory path,
but it cannot run a Neo4j DBMS cluster; cluster output is
therefore never fabricated.
Neo4j DBMS clustering, multi-database topology management,
fine-grained server management, and the cluster administrative
evidence used in this chapter are Enterprise Edition
capabilities. The mandatory path is a deterministic local
simulation/trace plus application-level routing/bookmark
exercises that do not pretend Community is clustered. The
optional licensed lab assumes Neo4j Enterprise
2026.07.1 on isolated test servers and uses
commands documented for the current release. Aura manages
topology for you, so self-managed server allocation,
discovery, seeding, and failure-domain commands do not map
one-for-one to Aura operations.
AtlasMart remains the same customer/order/product graph used
throughout earlier chapters. This chapter changes deployment
topology, not business identity. The optional licensed topology
uses a separate database name atlasmart on
disposable test infrastructure so cluster administration cannot
damage the single-server continuity database.
| Term | Precise meaning in current Neo4j clustering |
|---|---|
| server | A Neo4j DBMS process/machine that can host allocations for multiple databases. A server is not permanently “a primary” or “a secondary”. |
| database topology | The set of primary and secondary allocations for one database. Each standard database has its own topology. |
| primary | A database copy that participates in fault-tolerant write processing and is eligible to become that database’s writer/leader. |
| writer / leader | Exactly one eligible primary at a time orders writes for a database. Different databases can have different leaders on the same physical cluster. |
| secondary | An asynchronously replicated database copy used primarily for read scaling; it does not vote in that database’s write quorum and can lag. |
| quorum / simple majority |
The minimum number of primaries needed to
acknowledge/advance fault-tolerant writes. For
N primaries, majority is
floor(N/2)+1.
|
| routing table | A driver-consumable list of routers/readers/writers for a database, cached for a TTL and refreshed when topology/leadership changes. |
| bookmark | A causal token used to ensure later work does not execute before the represented committed state is available. It is not global serializability. |
| failure domain | Infrastructure expected to fail together: host, rack, availability zone, data center, region, network path, or power domain. |
| Assumption | Pinned value / rule |
|---|---|
| current server | Neo4j 2026.07.1; 5.26.30 is the current LTS comparison line |
| Cypher | Cypher 25 for current administrative examples; Cypher 5 is the frozen compatibility line |
| mandatory path | Community/local deterministic simulation and trace; no successful Enterprise output is claimed |
| optional cluster | Enterprise 2026.07.1; at least three servers for primary quorum; fourth server used when demonstrating a secondary |
| driver |
Python driver 6.3 with neo4j:// routing URI
for a cluster; bolt:// is direct to one
address
|
| auth/TLS | synthetic credentials only; production/remote cluster transport requires properly verified TLS and intra-cluster encryption as appropriate |
| plugins | APOC/GDS not required for clustering lessons |
| measurement | all failover/lag/latency/alert values must be measured on the learner’s licensed environment; trace outputs are explicitly labeled deterministic simulations |
1. The topology belongs to the database
In modern Neo4j, a cluster is a pool of servers that can host
allocations for many databases. The labels
primary and secondary describe
a copy of one database. Server 01 might host the writer primary
for atlasmart, a follower primary for another
database, and no copy of a third. This is why old mental models
that call machines “Core servers” or “read replicas” permanently
are misleading for current releases.
| Server | atlasmart allocation | another database allocation | What this proves |
|---|---|---|---|
| server01 | primary · writer | primary · follower | role is database-specific |
| server02 | primary · follower | secondary | one server can host different roles |
| server03 | primary · follower | primary · writer | different databases elect different writers |
| server04 | secondary | none | secondary can serve read scale without joining atlasmart write quorum |
2. One writer, multiple voting primaries
Neo4j uses Raft for the write-consensus group of each database. One primary is elected leader/writer. It orders transactions and synchronously replicates the corresponding log entries to enough primaries for the fault-tolerance contract. A client does not choose an arbitrary primary and create an independent write history. Routing directs writes to the current writer, and if leadership changes the routing view must converge on the new writer.
“Leader” and “follower” are Raft roles among primary allocations. “Primary” and “secondary” are the durable topology modes you model operationally. Do not write application code that assumes one named host is permanently the writer.
3. Quorum mathematics determines write availability
For N primaries, a simple majority is
floor(N/2)+1. A three-primary database therefore
needs two live primaries to continue fault-tolerant writes; a
five-primary database needs three. A single-primary database
needs its only primary: there is no automatic write-failover
tolerance if that allocation fails.
| Primaries | Majority | Arbitrary primary failures tolerated while still writing | Operational note |
|---|---|---|---|
| 1 | 1 | 0 | lowest coordination cost; no primary failure tolerance |
| 3 | 2 | 1 | common HA baseline; one primary can be lost |
| 5 | 3 | 2 | more failure/maintenance headroom, but more resource/network coordination |
| 7 | 4 | 3 | possible but not automatically “better”; latency/cost/operational complexity rise |
from math import floor
SCENARIOS = [
(1, 0), (1, 1),
(3, 0), (3, 1), (3, 2),
(5, 1), (5, 2), (5, 3),
]
for primaries, failed in SCENARIOS:
majority = floor(primaries / 2) + 1
alive = primaries - failed
print({
"primaries": primaries,
"failed": failed,
"alive": alive,
"majority": majority,
"write_available": alive >= majority,
})
The script is a local arithmetic simulation, not a Neo4j
cluster test. For 3 primaries / 1 failed it
reports write_available=True; for 3 / 2 failed it
reports False. It proves majority arithmetic, not election
timing, network behavior, recovery time or client retry
success.
4. Primaries and secondaries solve different problems
A secondary receives database changes asynchronously. That is
useful for read scaling and can reduce read load on primaries,
but it does not vote in write consensus. Therefore
1 primary + 2 secondaries gives three copies yet
still loses write availability when the single primary is
unavailable. This is the chapter’s first deliberately misleading
architecture.
| Topology | Copies | Write quorum | Primary failure tolerance | Read-scale potential |
|---|---|---|---|---|
| 1P + 2S | 3 | 1 of 1 | 0 | high relative to one server |
| 3P + 0S | 3 | 2 of 3 | 1 | reads can use configured eligible primaries |
| 3P + 1S | 4 | 2 of 3 | 1 | adds an asynchronous read target |
| 5P + 0S | 5 | 3 of 5 | 2 | higher write-fault tolerance; coordination cost grows |
“I have three copies, therefore I can lose any one copy and still write.” This is false for 1P+2S. Repair the reasoning by counting primaries, computing the majority, then separately evaluating secondary read capacity and lag.
5. Optional licensed cluster setup and evidence
The following is intentionally separated from the mandatory Community path. It assumes a licensed, isolated Enterprise 2026.07.1 test cluster. Do not copy the addresses to a public network; production cluster links need deliberate TLS/intra-cluster security, DNS and firewall design.
# OPTIONAL LICENSED LAB — illustrative per-server neo4j.conf fragments.
# Use distinct resolvable addresses in an isolated Enterprise 2026.07.1 test environment.
# On server01, server02, server03 (system primaries):
server.default_listen_address=0.0.0.0
server.default_advertised_address=server01.example.test # change per server
dbms.cluster.discovery.resolver_type=LIST
dbms.cluster.endpoints=server01.example.test:6000,server02.example.test:6000,server03.example.test:6000,server04.example.test:6000
server.cluster.system_database_mode=PRIMARY
initial.dbms.default_primaries_count=3
initial.dbms.default_secondaries_count=0
# On server04 use its own advertised address. It may join with NONE modeConstraint
# and later host a secondary allocation for atlasmart.
:use system
CREATE DATABASE atlasmart TOPOLOGY 3 PRIMARIES 1 SECONDARY;
SHOW DATABASE atlasmart YIELD
address, role, writer, currentStatus,
currentPrimariesCount, currentSecondariesCount,
requestedPrimariesCount, requestedSecondariesCount,
replicationLag
RETURN * ORDER BY role, address;
Wait until every requested allocation is online. Exactly one
row for atlasmart should have
writer=true; three rows should have
role=primary and one role=secondary.
Record actual addresses, statuses and replicationLag. These
are acceptance criteria, not captured output from this
generation environment.
6. Controlled failure reasoning before touching a server
Before stopping anything, predict the outcome. In a 3P+1S topology: losing the secondary should not change write quorum; losing one primary should retain a 2-of-3 majority; losing a second primary should remove the majority and therefore write availability. Reads may still be possible from surviving allocations, but “readable” is not the same as “healthy for writes.”
| Injected condition | Expected write state | What must be observed |
|---|---|---|
| secondary stopped | writes remain available | writer unchanged or routable; secondary row unavailable/unknown; replication lag catches up after return |
| writer primary stopped | new writer election if 2 primaries remain | old writer unavailable, another primary becomes writer, driver refresh/retry behavior measured |
| one follower primary stopped | writes remain available | 2 primaries including writer remain; fault-tolerance headroom now zero |
| two primaries stopped | writes unavailable | no simple majority; alert must fire; do not force two independent writers |
7. Replication is not backup
All cluster copies can faithfully replicate a bad write, an accidental delete, compromised credentials, or application corruption. They can also share a common infrastructure disaster. Chapter 16’s off-host/immutable backup and isolated restore evidence therefore remains mandatory even after HA is introduced.
HA addresses selected infrastructure failures and service continuity. Backup addresses recoverability from destructive state changes and disasters. A production design needs both; neither substitutes for the other.
8. AtlasMart verification checklist
- Can you name the writer by database rather than by permanent server?
- Can you calculate the majority from requested/current primary counts?
- Can you explain why secondaries do not increase write quorum?
-
Can you prove topology with
SHOW DATABASE ... YIELD *rather than process count? - Can you identify the backup evidence that remains independent of all cluster copies?
Check your understanding
- Do servers have permanent primary/secondary roles?
- How many primaries must be alive for a three-primary database to write?
- Does a secondary vote in the write quorum?
- Can three copies always tolerate any one copy loss for writes?
- Why is a cluster not a backup?
Review the answers
1. No. Primary/secondary describe database allocations; a server can host different roles for different databases subject to its constraints.
2. A simple majority: two.
3. No. It is asynchronously replicated and primarily used for read scaling.
4. No. A 1P+2S database loses writes if its only primary is unavailable.
5. Because corruption/deletion/compromise can replicate to all copies; independent recovery artifacts and restore tests are still required.
Production judgment
| Decision surface | Questions that must be answered before production |
|---|---|
| graph/workload fit | Which writes need fault tolerance? How much read scaling is needed? Does graph fan-out make reader CPU/page-cache the bottleneck rather than topology? |
| correctness | Which operations need causal read-after-write? Which externally visible side effects require idempotency/reconciliation after an ambiguous client outcome? |
| topology/cardinality | How many primaries satisfy failure tolerance? Are secondaries justified by read demand? Are dense/hot entities producing lock contention that topology will not solve? |
| latency | What are write p50/p95/p99 with quorum across actual zones? What read tail latency is added by lag/bookmark waiting/routing refresh? |
| transactions/concurrency | How do leader changes, deadlocks, retries and long transactions interact? Is transaction work idempotent under managed retries? |
| memory/CPU/disk/network | Can remaining servers absorb a failed member? Is there space/bandwidth for store copy and catch-up while serving traffic? |
| indexes/constraints | Are access paths and uniqueness contracts identical/ONLINE across allocations after seeding or recovery? |
| driver | Is one maintained driver reused? Are routing URI, database selection, read/write modes, pool sizes, acquisition timeouts, max retry time and bookmarks intentional? |
| security | Are Bolt, HTTPS and intra-cluster links encrypted as required? Are server-management privileges and certificates isolated from application credentials? |
| backup/recovery | Replication is not backup. Where are off-host/immutable backups and when was an isolated restore last proven? |
| observability | Do alerts cover writer absence, unavailable primaries, replication lag, store copy, routing failures, retry spikes, queueing and capacity headroom? |
| testing/failure injection | Have host loss, reader loss, writer loss, majority loss, partition, slow network, maintenance and ambiguous commit outcomes been tested safely? |
| version/edition | Are server/Java/driver/APOC/GDS/Cypher/discovery settings compatible with the exact release and license? |
| Aura/self-managed | Which topology choices are operator-owned self-managed responsibilities versus managed-service controls/SLOs? |
| cost/migration | What is the cost of extra primaries/secondaries/zones and cross-zone traffic, and how will topology changes or rollback be staged? |
Summary and next step
You can now explain current Neo4j HA from database topology and quorum rather than server labels. Lesson 2 follows a real client request through discovery, the routing table, leadership changes, read scaling and bookmarks.
Authoritative references
- Current Neo4j versions — Current Neo4j 2026.07.1, 5.26.30 LTS, Cypher Shell and GDS release snapshot.
- Clustering architecture — Current server/database decoupling, primaries, secondaries, majority writes, and causal consistency.
- Deploy a basic cluster — Current discovery/bootstrap configuration and TOPOLOGY examples.
- Cluster server discovery — LIST/DNS/Kubernetes discovery and the current discovery-service model.
- Leadership, routing, and load balancing — Per-database Raft leadership, routing tables, readers/writers/routers, and routing policy.
- Managing databases in a cluster — Primary/secondary allocation and topology changes.
- Managing servers in a cluster — Server enablement, constraints, deallocation, removal, and operational lifecycle.
- Server management command syntax — SHOW/ENABLE/ALTER/DEALLOCATE/REALLOCATE/DROP server commands.
- Show databases — Role, writer, allocation counts, replication lag, and database-state evidence.
- Monitor databases in a cluster — SHOW DATABASES-based per-allocation monitoring and status interpretation.
- Resilient multi-data-center cluster — Failure-domain placement, latency/fault-tolerance tradeoffs, and recommended/anti-pattern layouts.
- Cluster disaster recovery — Recovery when allocations or whole failure domains are lost.
- Built-in procedures — Current routing, cordon, deallocation and related procedure names/deprecations.
- Python driver transactions/routing — Read/write routing and explicit database selection in the official driver.
- Python driver bookmarks — Bookmark propagation and causal-consistency coordination across sessions.