Chapter 17 · Clustering, High Availability, Routing, Consensus, and Failure Domains
Inject Node/Network Failures and Verify Driver Recovery, Transaction Outcomes, and Operational Alerts
Run a controlled high-availability failure exercise: change writer availability, remove a reader, simulate quorum loss, observe driver recovery and ambiguous outcomes, and prove that alerts and business invariants expose rather than hide failure.
The final exercise is not “kill a node and celebrate that the web page still loads.” AtlasMart must prove correct committed state, bounded client recovery, alerting, and operator understanding under several distinct failures. A driver timeout can occur even when the server committed; a secondary can lag after recovery; quorum can be lost while some processes remain alive. The drill must expose those nuances.
Learning outcomes
Design safe HA failure injections with explicit blast radius, preconditions, expected signals and reset paths.
Correlate writer/reader topology, routing refresh, retries and application request IDs during leader loss.
Treat client timeouts/disconnects as potentially ambiguous outcomes and reconcile by stable business identity.
Verify quorum-loss behavior refuses unsafe writes instead of forcing availability at the cost of divergent history.
Turn failure observations into SLOs, alerts, capacity actions and a repeatable AtlasMart HA runbook.
Lab baseline and scope
The continuity environment remains
Neo4j Community 2026.07.1, database
neo4j, explicit CYPHER 25 where
language selection matters, container
atlasmart-neo4j, loopback Bolt
bolt://127.0.0.1:7687, synthetic local credential
neo4j / atlasmart-course-2026, Java 21 or 25, and
Python driver 6.3. Current 5.26 LTS is 5.26.30.
Community is intentionally retained for the mandatory path,
but it cannot run a Neo4j DBMS cluster; cluster output is
therefore never fabricated.
Neo4j DBMS clustering, multi-database topology management,
fine-grained server management, and the cluster administrative
evidence used in this chapter are Enterprise Edition
capabilities. The mandatory path is a deterministic local
simulation/trace plus application-level routing/bookmark
exercises that do not pretend Community is clustered. The
optional licensed lab assumes Neo4j Enterprise
2026.07.1 on isolated test servers and uses
commands documented for the current release. Aura manages
topology for you, so self-managed server allocation,
discovery, seeding, and failure-domain commands do not map
one-for-one to Aura operations.
AtlasMart remains the same customer/order/product graph used
throughout earlier chapters. This chapter changes deployment
topology, not business identity. The optional licensed topology
uses a separate database name atlasmart on
disposable test infrastructure so cluster administration cannot
damage the single-server continuity database.
| Term | Precise meaning in current Neo4j clustering |
|---|---|
| server | A Neo4j DBMS process/machine that can host allocations for multiple databases. A server is not permanently “a primary” or “a secondary”. |
| database topology | The set of primary and secondary allocations for one database. Each standard database has its own topology. |
| primary | A database copy that participates in fault-tolerant write processing and is eligible to become that database’s writer/leader. |
| writer / leader | Exactly one eligible primary at a time orders writes for a database. Different databases can have different leaders on the same physical cluster. |
| secondary | An asynchronously replicated database copy used primarily for read scaling; it does not vote in that database’s write quorum and can lag. |
| quorum / simple majority |
The minimum number of primaries needed to
acknowledge/advance fault-tolerant writes. For
N primaries, majority is
floor(N/2)+1.
|
| routing table | A driver-consumable list of routers/readers/writers for a database, cached for a TTL and refreshed when topology/leadership changes. |
| bookmark | A causal token used to ensure later work does not execute before the represented committed state is available. It is not global serializability. |
| failure domain | Infrastructure expected to fail together: host, rack, availability zone, data center, region, network path, or power domain. |
| Assumption | Pinned value / rule |
|---|---|
| current server | Neo4j 2026.07.1; 5.26.30 is the current LTS comparison line |
| Cypher | Cypher 25 for current administrative examples; Cypher 5 is the frozen compatibility line |
| mandatory path | Community/local deterministic simulation and trace; no successful Enterprise output is claimed |
| optional cluster | Enterprise 2026.07.1; at least three servers for primary quorum; fourth server used when demonstrating a secondary |
| driver |
Python driver 6.3 with neo4j:// routing URI
for a cluster; bolt:// is direct to one
address
|
| auth/TLS | synthetic credentials only; production/remote cluster transport requires properly verified TLS and intra-cluster encryption as appropriate |
| plugins | APOC/GDS not required for clustering lessons |
| measurement | all failover/lag/latency/alert values must be measured on the learner’s licensed environment; trace outputs are explicitly labeled deterministic simulations |
1. Start with a failure hypothesis
Each injection needs one mechanism to test. “Node failure” is too broad. The secondary-loss experiment tests read-capacity reduction/catch-up without changing write quorum. Writer-loss tests election/routing/retry while majority survives. Two-primary loss tests safety when quorum disappears. Combining all faults at once makes diagnosis impossible.
| Experiment | Hypothesis | Primary acceptance criterion |
|---|---|---|
| secondary loss | write consensus unaffected | writes continue; reader capacity drops; secondary later catches up |
| writer loss with 2/3 primaries alive | new primary can be elected | writes recover through routing/retry; one committed business outcome |
| minority partition | minority cannot safely commit | only majority side retains write capability |
| 2 of 3 primaries unavailable | majority absent | new writes unavailable; alert fires; no forced split brain |
| member return | catch-up restores redundancy | allocation reaches healthy/online and lag returns inside stated objective |
2. Free deterministic trace first
Use a trace before a licensed chaos test. It forces you to state
the expected system and application outcomes in advance. The
none_temporarily writer state represents an
election interval in the model; its duration is deliberately
unspecified because the free trace cannot measure it.
# FREE deterministic trace — save as ha_trace.csv and reason about every row.
time,event,primaries_alive,writer,secondary_state,request,expected_business_outcome
00:00,healthy,3,P1,caught_up,write O-9100,commit once
00:01,stop secondary,3,P1,down,read catalog,served by eligible reader
00:02,restart secondary,3,P1,catching_up,read with bookmark,may wait until state available
00:03,stop writer P1,2,none_temporarily,cancelled client write,ambiguous until reconciled
00:04,elect P2,2,P2,caught_up,retry idempotent O-9101,commit once
00:05,stop P3,1,P2,caught_up,new write O-9102,write unavailable because no majority
00:06,start P1,2,P2,catching_up,new write O-9103,available after quorum restored
The timestamps are ordering labels, not measured election/failover durations. Real timings depend on server load, network, driver configuration, election state and failure mode.
3. Stable business keys make ambiguous outcomes reconcilable
Suppose the writer commits an order but the connection breaks
before the client receives success. The client sees a timeout,
yet “retry the same business operation” and “create a second
order” are not equivalent. Reuse the Chapter 6 idempotency
discipline: stable orderId, uniqueness constraint,
MERGE pattern scoped to identity, and post-error
reconciliation.
from neo4j import GraphDatabase
from neo4j.exceptions import Neo4jError
import uuid
URI = "neo4j://server01.example.test:7687"
AUTH = ("atlasmart_app", "<synthetic-test-secret>")
ORDER_ID = "O-HA-" + uuid.uuid4().hex[:10]
QUERY = """
MERGE (o:Order {orderId:$order_id})
ON CREATE SET o.status='PENDING', o.createdAt=datetime()
RETURN o.orderId AS orderId, o.status AS status
"""
with GraphDatabase.driver(URI, auth=AUTH, max_transaction_retry_time=15) as driver:
try:
rows, summary, keys = driver.execute_query(
QUERY,
order_id=ORDER_ID,
database_="atlasmart",
)
print("result", rows)
except Neo4jError as exc:
print("neo4j_code", exc.code)
print("retryable", exc.is_retryable())
# Do not blindly create a replacement business operation.
# Reconcile ORDER_ID against the graph before deciding what the client should do.
| Client observation | What it proves | Required next step |
|---|---|---|
| success response | driver received successful result | still record business invariant/telemetry |
| retryable Neo4j error | driver/server classified a retryable failure | managed retry may be safe if transaction function has no unsafe external side effect |
| timeout/disconnect | client did not receive a definitive result | query by stable orderId/operationId before issuing a new business action |
| constraint violation | identity invariant rejected duplicate/invalid state | diagnose source/retry logic; do not disable constraint |
4. Optional licensed failure sequence
This sequence assumes a disposable 3-primary + 1-secondary Enterprise test database. Do not use host firewall edits or real production networks. A process/container stop is sufficient for node-loss tests; a network-partition test should use an isolated test network with a documented reset command.
# OPTIONAL LICENSED LAB — execute only in an isolated test cluster.
# Before every step capture SHOW SERVERS, SHOW DATABASE atlasmart YIELD *, routing table,
# application success/error rate, retry count, and backup state.
1. Stop/restart only the secondary allocation host. Verify writes continue; measure catch-up/lag.
2. Stop the current writer primary. Verify another primary becomes writer while 2/3 majority remains.
3. During/around leader loss, issue an idempotent marker/order request with a stable business key.
Record whether the client saw success, a retryable error, timeout, or ambiguous outcome.
4. Stop a second primary. Verify new writes do NOT commit with only 1/3 primaries alive.
5. Restore one primary. Verify quorum and writer availability return; reconcile all marker/order IDs.
6. Restore all members. Wait for healthy/online allocations and acceptable replication lag.
7. Compare final business invariants and alerts with the pre-injection baseline.
Reset rule: if any unexpected database becomes degraded, stop the exercise and restore topology
before injecting another fault.
5. Capture four evidence planes
| Plane | Evidence | Why one plane is insufficient |
|---|---|---|
| database topology | SHOW SERVERS / SHOW DATABASE, writer, role, replicationLag | a healthy writer row does not prove the client refreshed routing |
| driver | routing refresh, retry classification/count, pool acquisition, bookmark state | driver recovery does not prove business state committed exactly once |
| application | request ID, order ID, response/error, reconciliation result | application success alone hides reduced redundancy/lag |
| infrastructure/ops | process health, CPU, disk, network, logs, alert timestamps | process survival does not prove quorum or database availability |
6. Alerts must reflect topology semantics
“Host down” is useful but not enough. AtlasMart needs alerts that encode the database risk: no writer, current primaries below requested, replication lag beyond the workload’s freshness objective, store-copy/deallocating stuck, routing/retry error spikes, and capacity saturation during degraded operation. The same host failure can be low severity for a secondary and urgent for the second lost primary in a 3P database.
| Alert | Severity context |
|---|---|
| secondary unavailable | read-capacity/freshness risk; write quorum unchanged |
| one of 3 primaries unavailable | writes still possible but fault headroom exhausted |
| writer absent during election | short transient may be expected; sustained absence is critical |
| two of 3 primaries unavailable | critical: write quorum unavailable |
| replication lag high after return | recovery not complete; causal reads may wait, redundancy/freshness objective not met |
| retry/timeout spike | client-visible instability; correlate with topology before increasing retry budgets |
7. Measure failover without hiding overload
Record request throughput and p50/p95/p99 latency before, during and after failure; transaction error/retry counts; time without an available writer; pool queueing; CPU/disk/network on survivors; replication lag/catch-up; and final business invariants. Do not report one average or only the requests that succeeded. A cluster can “fail over” while the remaining nodes are saturated and p99/error rate violate the SLO.
| Metric | Baseline | Failure window | Recovery acceptance |
|---|---|---|---|
| write success/error rate | record | record all errors/timeouts | returns inside defined objective |
| p50/p95/p99 | record warm workload | record distribution, not average | tail stabilizes within envelope |
| writer availability | one writer | measure no-writer interval | writer present with quorum |
| retry count | normal baseline | record spike and classifications | returns to baseline; no retry storm |
| replication lag | normal | may increase | inside stated freshness/redundancy objective |
| business invariant | known count/unique order IDs | track ambiguous requests | no duplicate/lost accepted operation after reconciliation |
8. Deliberately unsafe response: “increase retries”
When failover produces errors, a common reaction is to increase retry count/time globally. That can amplify overload, extend tail latency and repeat non-idempotent side effects. Repair the mechanism instead: current official driver, correct routing URI/access mode, bounded retry budget, small idempotent transaction functions, stable operation IDs, reconciliation, pool/queue limits and capacity headroom.
A timeout is not proof of rollback. Do not automatically create a second order/payment/shipment. Reconcile the stable operation key first.
9. Exit criteria and cleanup
-
All requested AtlasMart allocations are
onlineand healthy. - Exactly one writer exists for the database.
- Current primary/secondary counts match requested topology.
- Replication lag is inside the declared objective.
- Driver retry/error rates and pool queueing return to baseline.
- Every injected business operation is reconciled by stable ID with no unintended duplicates.
- Backup/restore status remains valid and independent of cluster copies.
- Every alert fired at the expected severity and cleared after recovery.
Check your understanding
- Why test secondary loss separately from writer loss?
- Does a client timeout prove the transaction rolled back?
- What should happen with only one of three primaries alive?
- What evidence proves failover is acceptable?
- Why not simply lengthen retry time?
Review the answers
1. They exercise different mechanisms: read capacity/catch-up versus election/routing/write continuity.
2. No. The commit outcome may be ambiguous to the client; reconcile by stable business identity.
3. New fault-tolerant writes should be unavailable because the simple majority is lost.
4. Topology + driver + application + infrastructure evidence, including tail latency/errors and final business invariants.
5. It can hide topology/capacity problems, amplify overload, and repeat unsafe side effects; retries must be bounded and semantically safe.
Production judgment
| Decision surface | Questions that must be answered before production |
|---|---|
| graph/workload fit | Which writes need fault tolerance? How much read scaling is needed? Does graph fan-out make reader CPU/page-cache the bottleneck rather than topology? |
| correctness | Which operations need causal read-after-write? Which externally visible side effects require idempotency/reconciliation after an ambiguous client outcome? |
| topology/cardinality | How many primaries satisfy failure tolerance? Are secondaries justified by read demand? Are dense/hot entities producing lock contention that topology will not solve? |
| latency | What are write p50/p95/p99 with quorum across actual zones? What read tail latency is added by lag/bookmark waiting/routing refresh? |
| transactions/concurrency | How do leader changes, deadlocks, retries and long transactions interact? Is transaction work idempotent under managed retries? |
| memory/CPU/disk/network | Can remaining servers absorb a failed member? Is there space/bandwidth for store copy and catch-up while serving traffic? |
| indexes/constraints | Are access paths and uniqueness contracts identical/ONLINE across allocations after seeding or recovery? |
| driver | Is one maintained driver reused? Are routing URI, database selection, read/write modes, pool sizes, acquisition timeouts, max retry time and bookmarks intentional? |
| security | Are Bolt, HTTPS and intra-cluster links encrypted as required? Are server-management privileges and certificates isolated from application credentials? |
| backup/recovery | Replication is not backup. Where are off-host/immutable backups and when was an isolated restore last proven? |
| observability | Do alerts cover writer absence, unavailable primaries, replication lag, store copy, routing failures, retry spikes, queueing and capacity headroom? |
| testing/failure injection | Have host loss, reader loss, writer loss, majority loss, partition, slow network, maintenance and ambiguous commit outcomes been tested safely? |
| version/edition | Are server/Java/driver/APOC/GDS/Cypher/discovery settings compatible with the exact release and license? |
| Aura/self-managed | Which topology choices are operator-owned self-managed responsibilities versus managed-service controls/SLOs? |
| cost/migration | What is the cost of extra primaries/secondaries/zones and cross-zone traffic, and how will topology changes or rollback be staged? |
Summary and next step
Chapter 17 has turned “clustering” into measurable per-database topology, quorum, routing, failure-domain and application-recovery behavior. Chapter 18 now focuses on the metrics, logs, query/transaction visibility, memory, store and capacity evidence needed to operate that architecture continuously.
Authoritative references
- Current Neo4j versions — Current Neo4j 2026.07.1, 5.26.30 LTS, Cypher Shell and GDS release snapshot.
- Clustering architecture — Current server/database decoupling, primaries, secondaries, majority writes, and causal consistency.
- Deploy a basic cluster — Current discovery/bootstrap configuration and TOPOLOGY examples.
- Cluster server discovery — LIST/DNS/Kubernetes discovery and the current discovery-service model.
- Leadership, routing, and load balancing — Per-database Raft leadership, routing tables, readers/writers/routers, and routing policy.
- Managing databases in a cluster — Primary/secondary allocation and topology changes.
- Managing servers in a cluster — Server enablement, constraints, deallocation, removal, and operational lifecycle.
- Server management command syntax — SHOW/ENABLE/ALTER/DEALLOCATE/REALLOCATE/DROP server commands.
- Show databases — Role, writer, allocation counts, replication lag, and database-state evidence.
- Monitor databases in a cluster — SHOW DATABASES-based per-allocation monitoring and status interpretation.
- Resilient multi-data-center cluster — Failure-domain placement, latency/fault-tolerance tradeoffs, and recommended/anti-pattern layouts.
- Cluster disaster recovery — Recovery when allocations or whole failure domains are lost.
- Built-in procedures — Current routing, cordon, deallocation and related procedure names/deprecations.
- Python driver transactions/routing — Read/write routing and explicit database selection in the official driver.
- Python driver bookmarks — Bookmark propagation and causal-consistency coordination across sessions.