Chapter 17 · Clustering, High Availability, Routing, Consensus, and Failure Domains
Cluster Seeding, Membership Changes, Maintenance, Capacity Headroom, and Rebalancing Operational Risk
Operate a cluster without turning maintenance into an outage: add/enable servers, seed/recreate databases, cordon/deallocate/reallocate safely, preserve quorum, and reserve capacity for store copies and failure recovery.
AtlasMart needs to patch one cluster host and add capacity before a campaign. The unsafe plan is “stop server03, add server05, and let Neo4j figure it out.” That ignores quorum headroom, allocation movement, store-copy I/O, system-database availability, server state and capacity. Safe maintenance is a topology transition with explicit preconditions and rollback.
Learning outcomes
Explain cluster discovery, server enablement and database allocation as separate lifecycle steps.
Use current SHOW/ENABLE/cordon/deallocate/reallocate operations with dry-run evidence where available.
Distinguish seeding/recreation from backup retention and understand current seed URI boundaries.
Budget CPU, disk, network and quorum headroom for store copy, catch-up and maintenance.
Build a maintenance runbook that keeps one failure from combining with planned work into quorum loss.
Lab baseline and scope
The continuity environment remains
Neo4j Community 2026.07.1, database
neo4j, explicit CYPHER 25 where
language selection matters, container
atlasmart-neo4j, loopback Bolt
bolt://127.0.0.1:7687, synthetic local credential
neo4j / atlasmart-course-2026, Java 21 or 25, and
Python driver 6.3. Current 5.26 LTS is 5.26.30.
Community is intentionally retained for the mandatory path,
but it cannot run a Neo4j DBMS cluster; cluster output is
therefore never fabricated.
Neo4j DBMS clustering, multi-database topology management,
fine-grained server management, and the cluster administrative
evidence used in this chapter are Enterprise Edition
capabilities. The mandatory path is a deterministic local
simulation/trace plus application-level routing/bookmark
exercises that do not pretend Community is clustered. The
optional licensed lab assumes Neo4j Enterprise
2026.07.1 on isolated test servers and uses
commands documented for the current release. Aura manages
topology for you, so self-managed server allocation,
discovery, seeding, and failure-domain commands do not map
one-for-one to Aura operations.
AtlasMart remains the same customer/order/product graph used
throughout earlier chapters. This chapter changes deployment
topology, not business identity. The optional licensed topology
uses a separate database name atlasmart on
disposable test infrastructure so cluster administration cannot
damage the single-server continuity database.
| Term | Precise meaning in current Neo4j clustering |
|---|---|
| server | A Neo4j DBMS process/machine that can host allocations for multiple databases. A server is not permanently “a primary” or “a secondary”. |
| database topology | The set of primary and secondary allocations for one database. Each standard database has its own topology. |
| primary | A database copy that participates in fault-tolerant write processing and is eligible to become that database’s writer/leader. |
| writer / leader | Exactly one eligible primary at a time orders writes for a database. Different databases can have different leaders on the same physical cluster. |
| secondary | An asynchronously replicated database copy used primarily for read scaling; it does not vote in that database’s write quorum and can lag. |
| quorum / simple majority |
The minimum number of primaries needed to
acknowledge/advance fault-tolerant writes. For
N primaries, majority is
floor(N/2)+1.
|
| routing table | A driver-consumable list of routers/readers/writers for a database, cached for a TTL and refreshed when topology/leadership changes. |
| bookmark | A causal token used to ensure later work does not execute before the represented committed state is available. It is not global serializability. |
| failure domain | Infrastructure expected to fail together: host, rack, availability zone, data center, region, network path, or power domain. |
| Assumption | Pinned value / rule |
|---|---|
| current server | Neo4j 2026.07.1; 5.26.30 is the current LTS comparison line |
| Cypher | Cypher 25 for current administrative examples; Cypher 5 is the frozen compatibility line |
| mandatory path | Community/local deterministic simulation and trace; no successful Enterprise output is claimed |
| optional cluster | Enterprise 2026.07.1; at least three servers for primary quorum; fourth server used when demonstrating a secondary |
| driver |
Python driver 6.3 with neo4j:// routing URI
for a cluster; bolt:// is direct to one
address
|
| auth/TLS | synthetic credentials only; production/remote cluster transport requires properly verified TLS and intra-cluster encryption as appropriate |
| plugins | APOC/GDS not required for clustering lessons |
| measurement | all failover/lag/latency/alert values must be measured on the learner’s licensed environment; trace outputs are explicitly labeled deterministic simulations |
1. Joining a server does not make it a permanent role
A new server first needs discovery configuration so it can learn the cluster. Its initial constraints can restrict which database modes/names it may host, but those are placement constraints—not a global “primary server” identity. After discovery and enablement, database allocations can be moved or created on it according to topology and constraints.
# On a new Enterprise server, configure discovery to existing members.
server.default_listen_address=0.0.0.0
server.default_advertised_address=server05.example.test
dbms.cluster.discovery.resolver_type=LIST
dbms.cluster.endpoints=server01.example.test:6000,server02.example.test:6000,server03.example.test:6000,server04.example.test:6000
server.cluster.system_database_mode=SECONDARY
initial.server.mode_constraint=NONE
# Start server05. It is discovered, then inspect/enable it from system database.
| Layer | Question |
|---|---|
| discovery | Can server05 find the existing cluster and exchange topology state? |
| server state | Is it discovered/enabled/healthy and allowed to host the intended database? |
| allocation | Which database copies should move/create there, in which mode? |
| data transfer | How much store/catch-up data must cross the network and fit on disk? |
| routing | When can clients safely treat it as reader/writer/router? |
2. Cordon before evacuation
Cordoning marks a server unsuitable for new allocations; it does
not automatically force current allocations away. That
separation is useful: you first prevent new placement, then
preview/deallocate existing database copies while observing
quorum and capacity. Current Cypher 25 removes the old uncordon
procedure path; ENABLE SERVER is the current
replacement when a server is returned to service.
:use system
SHOW SERVERS YIELD serverId, name, address, state, health, hosting, requestedHosting, modeConstraint RETURN *;
# For a newly discovered server not already enabled, use its actual ID/name:
ENABLE SERVER 'server05';
# Preview rebalancing before moving allocations:
DRYRUN REALLOCATE DATABASES;
# For planned evacuation of a server:
CALL dbms.cluster.cordonServer('server04');
DRYRUN DEALLOCATE DATABASES FROM SERVER 'server04';
DEALLOCATE DATABASES FROM SERVER 'server04';
# After it hosts no databases, DROP SERVER is the removal step if decommissioning.
# DROP SERVER 'server04';
Never deallocate a primary just because another copy exists. Before execution, calculate the current primary majority for every affected database, ensure target servers have capacity, verify backups, and define what happens if another server fails mid-maintenance.
3. Reallocation moves load and consumes resources
REALLOCATE DATABASES can rebalance database
placement, but movement may trigger store copy and catch-up.
During that interval the cluster uses extra disk bandwidth,
network, CPU and page cache while still serving application
traffic. A maintenance window can therefore fail by overload
even when quorum arithmetic is correct.
| Resource | Why movement stresses it | Acceptance evidence |
|---|---|---|
| disk capacity | new allocation needs full store plus operational headroom | free bytes before/after; no low-disk alert |
| disk I/O | store copy competes with transaction/index/page-cache reads | I/O latency and query p99 remain acceptable |
| network | copy/catch-up can consume bandwidth used by Bolt and Raft | throughput, packet loss/retransmit, application latency |
| CPU | copy checksums, recovery, transactions and queries overlap | CPU saturation/queueing below agreed envelope |
| memory/page cache | newly started allocation has cold cache | page-cache hit behavior and tail latency stabilize |
4. Seeding is a creation/recreation mechanism
When a database must be created or recreated from known data,
current Neo4j supports seed URI workflows. The exact provider
can be file/cloud/server dependent. Starting in 2026.04,
server:// can reference a backup/dump staged in a
configured seeds directory. This is not an invitation to use
cluster servers as permanent backup storage: seeding artifacts
are operational inputs; durable backup retention remains the
Chapter 16 problem.
# Current 2026.07 concepts: seed from a validated backup/dump URI when creating/recreating.
# Exact provider/credentials depend on configured seed provider and deployment.
CREATE DATABASE atlasmart_restore
TOPOLOGY 3 PRIMARIES 1 SECONDARY
OPTIONS {seedURI: 'file:///validated-course-backup.dump'};
# Starting in Neo4j 2026.04, ServerSeedProvider also supports server:// URIs
# for a backup/dump placed in a configured seeds directory on a cluster server.
# Treat the seeds directory as transient staging, not long-term backup storage.
Recreation from surviving allocations or a backup chooses an authoritative source. If a lost allocation contained newer data than the chosen source, data loss is possible. Record recovery point and reconcile business invariants before promotion.
5. Maintenance headroom is a fault-tolerance budget
For a three-primary database, taking one primary down intentionally leaves two. You can still write, but another primary loss removes quorum. For five primaries, maintenance of one leaves four and still tolerates another failure. This is not a universal recommendation for five; it is a way to convert an operational requirement (“patch one host while tolerating another failure”) into topology math.
| Requirement during maintenance | Minimum topology reasoning |
|---|---|
| patch 1 primary; tolerate no concurrent primary failure | 3 primaries can remain write-available with 2 alive |
| patch 1 primary; still tolerate 1 additional arbitrary primary failure | requires enough primaries that two unavailable still leave majority; 5P satisfies 3-of-5 |
| move only secondaries | write quorum unchanged, but reader capacity/freshness may drop |
| patch system-primary hosts | must also preserve system-database operability, not only user DB quorum |
6. Controlled maintenance runbook
| Phase | Required evidence | Abort condition |
|---|---|---|
| preflight | backups/restores current; SHOW SERVERS/DATABASES healthy; free capacity; current writer/quorum recorded | any unexpected unavailable allocation or insufficient capacity |
| cordon | server stops receiving new allocations | server state/topology differs from plan |
| dry-run deallocate/reallocate | target placements and modes acceptable | move would consume last quorum/capacity headroom |
| execute move | store-copy/catch-up progresses; p99/error rate inside limit | sustained overload, quorum degradation, unexplained lag |
| maintenance | server intentionally unavailable only after evacuation/acceptance | another failure consumes remaining safety budget |
| return/enable | server healthy, eligible, allocations stable | version/config mismatch or repeated store-copy failure |
7. Wrong operations pattern and repair
Wrong: schedule simultaneous rolling restarts across three primaries because “the cluster has replicas.” Mechanism: overlapping restarts can remove the majority or force repeated elections/catch-up while capacity is already reduced. Repair: one controlled topology transition at a time, explicit current-primary count, health gate between steps, and a stop-the-line rule on unexpected failure.
A cluster that technically retains quorum but cannot absorb failed-member load may still violate availability SLOs. Keep compute, storage, connection-pool and network headroom for the degraded state—not only the healthy state.
8. Verification checklist
- New server discovery and enablement are distinct and observable.
- Server mode constraints are treated as placement constraints, not permanent server roles.
- Every deallocation/reallocation is previewed where supported and checked against per-database quorum.
- Store-copy bandwidth/disk/cache impact is included in maintenance acceptance criteria.
- Seed artifacts are validated and not confused with long-term backup storage.
- Maintenance includes a concurrent-failure stop rule.
Check your understanding
- What does cordoning do?
- What replaces the old uncordon procedure in current Cypher 25?
- Why use DRYRUN before deallocation/reallocation?
- Is a server:// seed directory a backup retention strategy?
- Why measure capacity during maintenance?
Review the answers
1. Prevents the server from being chosen for new allocations; it does not automatically move existing allocations away.
2. ENABLE SERVER.
3. To inspect planned movement and detect topology/capacity surprises before making changes.
4. No. It is staging for database creation/recreation; durable backup protection remains separate.
5. Degraded servers must absorb traffic plus store-copy/catch-up work; quorum alone does not guarantee SLOs.
Production judgment
| Decision surface | Questions that must be answered before production |
|---|---|
| graph/workload fit | Which writes need fault tolerance? How much read scaling is needed? Does graph fan-out make reader CPU/page-cache the bottleneck rather than topology? |
| correctness | Which operations need causal read-after-write? Which externally visible side effects require idempotency/reconciliation after an ambiguous client outcome? |
| topology/cardinality | How many primaries satisfy failure tolerance? Are secondaries justified by read demand? Are dense/hot entities producing lock contention that topology will not solve? |
| latency | What are write p50/p95/p99 with quorum across actual zones? What read tail latency is added by lag/bookmark waiting/routing refresh? |
| transactions/concurrency | How do leader changes, deadlocks, retries and long transactions interact? Is transaction work idempotent under managed retries? |
| memory/CPU/disk/network | Can remaining servers absorb a failed member? Is there space/bandwidth for store copy and catch-up while serving traffic? |
| indexes/constraints | Are access paths and uniqueness contracts identical/ONLINE across allocations after seeding or recovery? |
| driver | Is one maintained driver reused? Are routing URI, database selection, read/write modes, pool sizes, acquisition timeouts, max retry time and bookmarks intentional? |
| security | Are Bolt, HTTPS and intra-cluster links encrypted as required? Are server-management privileges and certificates isolated from application credentials? |
| backup/recovery | Replication is not backup. Where are off-host/immutable backups and when was an isolated restore last proven? |
| observability | Do alerts cover writer absence, unavailable primaries, replication lag, store copy, routing failures, retry spikes, queueing and capacity headroom? |
| testing/failure injection | Have host loss, reader loss, writer loss, majority loss, partition, slow network, maintenance and ambiguous commit outcomes been tested safely? |
| version/edition | Are server/Java/driver/APOC/GDS/Cypher/discovery settings compatible with the exact release and license? |
| Aura/self-managed | Which topology choices are operator-owned self-managed responsibilities versus managed-service controls/SLOs? |
| cost/migration | What is the cost of extra primaries/secondaries/zones and cross-zone traffic, and how will topology changes or rollback be staged? |
Summary and next step
You can now move topology deliberately rather than restarting machines blindly. Lesson 5 combines the chapter into an HA drill that injects failures and verifies driver recovery, transaction outcomes, lag, alerts and operator decisions.
Authoritative references
- Current Neo4j versions — Current Neo4j 2026.07.1, 5.26.30 LTS, Cypher Shell and GDS release snapshot.
- Clustering architecture — Current server/database decoupling, primaries, secondaries, majority writes, and causal consistency.
- Deploy a basic cluster — Current discovery/bootstrap configuration and TOPOLOGY examples.
- Cluster server discovery — LIST/DNS/Kubernetes discovery and the current discovery-service model.
- Leadership, routing, and load balancing — Per-database Raft leadership, routing tables, readers/writers/routers, and routing policy.
- Managing databases in a cluster — Primary/secondary allocation and topology changes.
- Managing servers in a cluster — Server enablement, constraints, deallocation, removal, and operational lifecycle.
- Server management command syntax — SHOW/ENABLE/ALTER/DEALLOCATE/REALLOCATE/DROP server commands.
- Show databases — Role, writer, allocation counts, replication lag, and database-state evidence.
- Monitor databases in a cluster — SHOW DATABASES-based per-allocation monitoring and status interpretation.
- Resilient multi-data-center cluster — Failure-domain placement, latency/fault-tolerance tradeoffs, and recommended/anti-pattern layouts.
- Cluster disaster recovery — Recovery when allocations or whole failure domains are lost.
- Built-in procedures — Current routing, cordon, deallocation and related procedure names/deprecations.
- Python driver transactions/routing — Read/write routing and explicit database selection in the official driver.
- Python driver bookmarks — Bookmark propagation and causal-consistency coordination across sessions.