Chapter 17 · Clustering, High Availability, Routing, Consensus, and Failure Domains

Cluster Seeding, Membership Changes, Maintenance, Capacity Headroom, and Rebalancing Operational Risk

Operate a cluster without turning maintenance into an outage: add/enable servers, seed/recreate databases, cordon/deallocate/reallocate safely, preserve quorum, and reserve capacity for store copies and failure recovery.

Advanced220–300 minutesMaintenance and allocation labNeo4j 2026.07.1 Community mandatory trace · Enterprise cluster optionalCypher 25 · current primary/secondary topology terminologyJava 21/25 · Python driver 6.3Last reviewed: September 2026

AtlasMart needs to patch one cluster host and add capacity before a campaign. The unsafe plan is “stop server03, add server05, and let Neo4j figure it out.” That ignores quorum headroom, allocation movement, store-copy I/O, system-database availability, server state and capacity. Safe maintenance is a topology transition with explicit preconditions and rollback.

Learning outcomes

01

Explain cluster discovery, server enablement and database allocation as separate lifecycle steps.

02

Use current SHOW/ENABLE/cordon/deallocate/reallocate operations with dry-run evidence where available.

03

Distinguish seeding/recreation from backup retention and understand current seed URI boundaries.

04

Budget CPU, disk, network and quorum headroom for store copy, catch-up and maintenance.

05

Build a maintenance runbook that keeps one failure from combining with planned work into quorum loss.

Lab baseline and scope

Chapter 17 baseline · reviewed 9 September 2026

The continuity environment remains Neo4j Community 2026.07.1, database neo4j, explicit CYPHER 25 where language selection matters, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic local credential neo4j / atlasmart-course-2026, Java 21 or 25, and Python driver 6.3. Current 5.26 LTS is 5.26.30. Community is intentionally retained for the mandatory path, but it cannot run a Neo4j DBMS cluster; cluster output is therefore never fabricated.

Edition and evidence boundary

Neo4j DBMS clustering, multi-database topology management, fine-grained server management, and the cluster administrative evidence used in this chapter are Enterprise Edition capabilities. The mandatory path is a deterministic local simulation/trace plus application-level routing/bookmark exercises that do not pretend Community is clustered. The optional licensed lab assumes Neo4j Enterprise 2026.07.1 on isolated test servers and uses commands documented for the current release. Aura manages topology for you, so self-managed server allocation, discovery, seeding, and failure-domain commands do not map one-for-one to Aura operations.

AtlasMart remains the same customer/order/product graph used throughout earlier chapters. This chapter changes deployment topology, not business identity. The optional licensed topology uses a separate database name atlasmart on disposable test infrastructure so cluster administration cannot damage the single-server continuity database.

Term Precise meaning in current Neo4j clustering
server A Neo4j DBMS process/machine that can host allocations for multiple databases. A server is not permanently “a primary” or “a secondary”.
database topology The set of primary and secondary allocations for one database. Each standard database has its own topology.
primary A database copy that participates in fault-tolerant write processing and is eligible to become that database’s writer/leader.
writer / leader Exactly one eligible primary at a time orders writes for a database. Different databases can have different leaders on the same physical cluster.
secondary An asynchronously replicated database copy used primarily for read scaling; it does not vote in that database’s write quorum and can lag.
quorum / simple majority The minimum number of primaries needed to acknowledge/advance fault-tolerant writes. For N primaries, majority is floor(N/2)+1.
routing table A driver-consumable list of routers/readers/writers for a database, cached for a TTL and refreshed when topology/leadership changes.
bookmark A causal token used to ensure later work does not execute before the represented committed state is available. It is not global serializability.
failure domain Infrastructure expected to fail together: host, rack, availability zone, data center, region, network path, or power domain.
Assumption Pinned value / rule
current server Neo4j 2026.07.1; 5.26.30 is the current LTS comparison line
Cypher Cypher 25 for current administrative examples; Cypher 5 is the frozen compatibility line
mandatory path Community/local deterministic simulation and trace; no successful Enterprise output is claimed
optional cluster Enterprise 2026.07.1; at least three servers for primary quorum; fourth server used when demonstrating a secondary
driver Python driver 6.3 with neo4j:// routing URI for a cluster; bolt:// is direct to one address
auth/TLS synthetic credentials only; production/remote cluster transport requires properly verified TLS and intra-cluster encryption as appropriate
plugins APOC/GDS not required for clustering lessons
measurement all failover/lag/latency/alert values must be measured on the learner’s licensed environment; trace outputs are explicitly labeled deterministic simulations

1. Joining a server does not make it a permanent role

A new server first needs discovery configuration so it can learn the cluster. Its initial constraints can restrict which database modes/names it may host, but those are placement constraints—not a global “primary server” identity. After discovery and enablement, database allocations can be moved or created on it according to topology and constraints.

neo4j.conf · optional new licensed server
# On a new Enterprise server, configure discovery to existing members.
server.default_listen_address=0.0.0.0
server.default_advertised_address=server05.example.test
dbms.cluster.discovery.resolver_type=LIST
dbms.cluster.endpoints=server01.example.test:6000,server02.example.test:6000,server03.example.test:6000,server04.example.test:6000
server.cluster.system_database_mode=SECONDARY
initial.server.mode_constraint=NONE

# Start server05. It is discovered, then inspect/enable it from system database.
Layer Question
discovery Can server05 find the existing cluster and exchange topology state?
server state Is it discovered/enabled/healthy and allowed to host the intended database?
allocation Which database copies should move/create there, in which mode?
data transfer How much store/catch-up data must cross the network and fit on disk?
routing When can clients safely treat it as reader/writer/router?

2. Cordon before evacuation

Cordoning marks a server unsuitable for new allocations; it does not automatically force current allocations away. That separation is useful: you first prevent new placement, then preview/deallocate existing database copies while observing quorum and capacity. Current Cypher 25 removes the old uncordon procedure path; ENABLE SERVER is the current replacement when a server is returned to service.

Cypher Shell · current server lifecycle operations
:use system
SHOW SERVERS YIELD serverId, name, address, state, health, hosting, requestedHosting, modeConstraint RETURN *;

# For a newly discovered server not already enabled, use its actual ID/name:
ENABLE SERVER 'server05';

# Preview rebalancing before moving allocations:
DRYRUN REALLOCATE DATABASES;

# For planned evacuation of a server:
CALL dbms.cluster.cordonServer('server04');
DRYRUN DEALLOCATE DATABASES FROM SERVER 'server04';
DEALLOCATE DATABASES FROM SERVER 'server04';

# After it hosts no databases, DROP SERVER is the removal step if decommissioning.
# DROP SERVER 'server04';
Safety gate

Never deallocate a primary just because another copy exists. Before execution, calculate the current primary majority for every affected database, ensure target servers have capacity, verify backups, and define what happens if another server fails mid-maintenance.

3. Reallocation moves load and consumes resources

REALLOCATE DATABASES can rebalance database placement, but movement may trigger store copy and catch-up. During that interval the cluster uses extra disk bandwidth, network, CPU and page cache while still serving application traffic. A maintenance window can therefore fail by overload even when quorum arithmetic is correct.

Resource Why movement stresses it Acceptance evidence
disk capacity new allocation needs full store plus operational headroom free bytes before/after; no low-disk alert
disk I/O store copy competes with transaction/index/page-cache reads I/O latency and query p99 remain acceptable
network copy/catch-up can consume bandwidth used by Bolt and Raft throughput, packet loss/retransmit, application latency
CPU copy checksums, recovery, transactions and queries overlap CPU saturation/queueing below agreed envelope
memory/page cache newly started allocation has cold cache page-cache hit behavior and tail latency stabilize

4. Seeding is a creation/recreation mechanism

When a database must be created or recreated from known data, current Neo4j supports seed URI workflows. The exact provider can be file/cloud/server dependent. Starting in 2026.04, server:// can reference a backup/dump staged in a configured seeds directory. This is not an invitation to use cluster servers as permanent backup storage: seeding artifacts are operational inputs; durable backup retention remains the Chapter 16 problem.

Cypher · seed concept for a disposable Enterprise recovery database
# Current 2026.07 concepts: seed from a validated backup/dump URI when creating/recreating.
# Exact provider/credentials depend on configured seed provider and deployment.
CREATE DATABASE atlasmart_restore
TOPOLOGY 3 PRIMARIES 1 SECONDARY
OPTIONS {seedURI: 'file:///validated-course-backup.dump'};

# Starting in Neo4j 2026.04, ServerSeedProvider also supports server:// URIs
# for a backup/dump placed in a configured seeds directory on a cluster server.
# Treat the seeds directory as transient staging, not long-term backup storage.
Data-loss boundary

Recreation from surviving allocations or a backup chooses an authoritative source. If a lost allocation contained newer data than the chosen source, data loss is possible. Record recovery point and reconcile business invariants before promotion.

5. Maintenance headroom is a fault-tolerance budget

For a three-primary database, taking one primary down intentionally leaves two. You can still write, but another primary loss removes quorum. For five primaries, maintenance of one leaves four and still tolerates another failure. This is not a universal recommendation for five; it is a way to convert an operational requirement (“patch one host while tolerating another failure”) into topology math.

Requirement during maintenance Minimum topology reasoning
patch 1 primary; tolerate no concurrent primary failure 3 primaries can remain write-available with 2 alive
patch 1 primary; still tolerate 1 additional arbitrary primary failure requires enough primaries that two unavailable still leave majority; 5P satisfies 3-of-5
move only secondaries write quorum unchanged, but reader capacity/freshness may drop
patch system-primary hosts must also preserve system-database operability, not only user DB quorum

6. Controlled maintenance runbook

Phase Required evidence Abort condition
preflight backups/restores current; SHOW SERVERS/DATABASES healthy; free capacity; current writer/quorum recorded any unexpected unavailable allocation or insufficient capacity
cordon server stops receiving new allocations server state/topology differs from plan
dry-run deallocate/reallocate target placements and modes acceptable move would consume last quorum/capacity headroom
execute move store-copy/catch-up progresses; p99/error rate inside limit sustained overload, quorum degradation, unexplained lag
maintenance server intentionally unavailable only after evacuation/acceptance another failure consumes remaining safety budget
return/enable server healthy, eligible, allocations stable version/config mismatch or repeated store-copy failure

7. Wrong operations pattern and repair

Wrong: schedule simultaneous rolling restarts across three primaries because “the cluster has replicas.” Mechanism: overlapping restarts can remove the majority or force repeated elections/catch-up while capacity is already reduced. Repair: one controlled topology transition at a time, explicit current-primary count, health gate between steps, and a stop-the-line rule on unexpected failure.

Capacity is part of HA

A cluster that technically retains quorum but cannot absorb failed-member load may still violate availability SLOs. Keep compute, storage, connection-pool and network headroom for the degraded state—not only the healthy state.

8. Verification checklist

  • New server discovery and enablement are distinct and observable.
  • Server mode constraints are treated as placement constraints, not permanent server roles.
  • Every deallocation/reallocation is previewed where supported and checked against per-database quorum.
  • Store-copy bandwidth/disk/cache impact is included in maintenance acceptance criteria.
  • Seed artifacts are validated and not confused with long-term backup storage.
  • Maintenance includes a concurrent-failure stop rule.

Check your understanding

  1. What does cordoning do?
  2. What replaces the old uncordon procedure in current Cypher 25?
  3. Why use DRYRUN before deallocation/reallocation?
  4. Is a server:// seed directory a backup retention strategy?
  5. Why measure capacity during maintenance?
Review the answers

1. Prevents the server from being chosen for new allocations; it does not automatically move existing allocations away.

2. ENABLE SERVER.

3. To inspect planned movement and detect topology/capacity surprises before making changes.

4. No. It is staging for database creation/recreation; durable backup protection remains separate.

5. Degraded servers must absorb traffic plus store-copy/catch-up work; quorum alone does not guarantee SLOs.

Production judgment

Decision surface Questions that must be answered before production
graph/workload fit Which writes need fault tolerance? How much read scaling is needed? Does graph fan-out make reader CPU/page-cache the bottleneck rather than topology?
correctness Which operations need causal read-after-write? Which externally visible side effects require idempotency/reconciliation after an ambiguous client outcome?
topology/cardinality How many primaries satisfy failure tolerance? Are secondaries justified by read demand? Are dense/hot entities producing lock contention that topology will not solve?
latency What are write p50/p95/p99 with quorum across actual zones? What read tail latency is added by lag/bookmark waiting/routing refresh?
transactions/concurrency How do leader changes, deadlocks, retries and long transactions interact? Is transaction work idempotent under managed retries?
memory/CPU/disk/network Can remaining servers absorb a failed member? Is there space/bandwidth for store copy and catch-up while serving traffic?
indexes/constraints Are access paths and uniqueness contracts identical/ONLINE across allocations after seeding or recovery?
driver Is one maintained driver reused? Are routing URI, database selection, read/write modes, pool sizes, acquisition timeouts, max retry time and bookmarks intentional?
security Are Bolt, HTTPS and intra-cluster links encrypted as required? Are server-management privileges and certificates isolated from application credentials?
backup/recovery Replication is not backup. Where are off-host/immutable backups and when was an isolated restore last proven?
observability Do alerts cover writer absence, unavailable primaries, replication lag, store copy, routing failures, retry spikes, queueing and capacity headroom?
testing/failure injection Have host loss, reader loss, writer loss, majority loss, partition, slow network, maintenance and ambiguous commit outcomes been tested safely?
version/edition Are server/Java/driver/APOC/GDS/Cypher/discovery settings compatible with the exact release and license?
Aura/self-managed Which topology choices are operator-owned self-managed responsibilities versus managed-service controls/SLOs?
cost/migration What is the cost of extra primaries/secondaries/zones and cross-zone traffic, and how will topology changes or rollback be staged?

Summary and next step

You can now move topology deliberately rather than restarting machines blindly. Lesson 5 combines the chapter into an HA drill that injects failures and verifies driver recovery, transaction outcomes, lag, alerts and operator decisions.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.