Chapter 17 · Clustering, High Availability, Routing, Consensus, and Failure Domains

Inject Node/Network Failures and Verify Driver Recovery, Transaction Outcomes, and Operational Alerts

Run a controlled high-availability failure exercise: change writer availability, remove a reader, simulate quorum loss, observe driver recovery and ambiguous outcomes, and prove that alerts and business invariants expose rather than hide failure.

Advanced240–330 minutesHA failure-injection verification labNeo4j 2026.07.1 Community mandatory trace · Enterprise cluster optionalCypher 25 · current primary/secondary topology terminologyJava 21/25 · Python driver 6.3Last reviewed: September 2026

The final exercise is not “kill a node and celebrate that the web page still loads.” AtlasMart must prove correct committed state, bounded client recovery, alerting, and operator understanding under several distinct failures. A driver timeout can occur even when the server committed; a secondary can lag after recovery; quorum can be lost while some processes remain alive. The drill must expose those nuances.

Learning outcomes

01

Design safe HA failure injections with explicit blast radius, preconditions, expected signals and reset paths.

02

Correlate writer/reader topology, routing refresh, retries and application request IDs during leader loss.

03

Treat client timeouts/disconnects as potentially ambiguous outcomes and reconcile by stable business identity.

04

Verify quorum-loss behavior refuses unsafe writes instead of forcing availability at the cost of divergent history.

05

Turn failure observations into SLOs, alerts, capacity actions and a repeatable AtlasMart HA runbook.

Lab baseline and scope

Chapter 17 baseline · reviewed 9 September 2026

The continuity environment remains Neo4j Community 2026.07.1, database neo4j, explicit CYPHER 25 where language selection matters, container atlasmart-neo4j, loopback Bolt bolt://127.0.0.1:7687, synthetic local credential neo4j / atlasmart-course-2026, Java 21 or 25, and Python driver 6.3. Current 5.26 LTS is 5.26.30. Community is intentionally retained for the mandatory path, but it cannot run a Neo4j DBMS cluster; cluster output is therefore never fabricated.

Edition and evidence boundary

Neo4j DBMS clustering, multi-database topology management, fine-grained server management, and the cluster administrative evidence used in this chapter are Enterprise Edition capabilities. The mandatory path is a deterministic local simulation/trace plus application-level routing/bookmark exercises that do not pretend Community is clustered. The optional licensed lab assumes Neo4j Enterprise 2026.07.1 on isolated test servers and uses commands documented for the current release. Aura manages topology for you, so self-managed server allocation, discovery, seeding, and failure-domain commands do not map one-for-one to Aura operations.

AtlasMart remains the same customer/order/product graph used throughout earlier chapters. This chapter changes deployment topology, not business identity. The optional licensed topology uses a separate database name atlasmart on disposable test infrastructure so cluster administration cannot damage the single-server continuity database.

Term Precise meaning in current Neo4j clustering
server A Neo4j DBMS process/machine that can host allocations for multiple databases. A server is not permanently “a primary” or “a secondary”.
database topology The set of primary and secondary allocations for one database. Each standard database has its own topology.
primary A database copy that participates in fault-tolerant write processing and is eligible to become that database’s writer/leader.
writer / leader Exactly one eligible primary at a time orders writes for a database. Different databases can have different leaders on the same physical cluster.
secondary An asynchronously replicated database copy used primarily for read scaling; it does not vote in that database’s write quorum and can lag.
quorum / simple majority The minimum number of primaries needed to acknowledge/advance fault-tolerant writes. For N primaries, majority is floor(N/2)+1.
routing table A driver-consumable list of routers/readers/writers for a database, cached for a TTL and refreshed when topology/leadership changes.
bookmark A causal token used to ensure later work does not execute before the represented committed state is available. It is not global serializability.
failure domain Infrastructure expected to fail together: host, rack, availability zone, data center, region, network path, or power domain.
Assumption Pinned value / rule
current server Neo4j 2026.07.1; 5.26.30 is the current LTS comparison line
Cypher Cypher 25 for current administrative examples; Cypher 5 is the frozen compatibility line
mandatory path Community/local deterministic simulation and trace; no successful Enterprise output is claimed
optional cluster Enterprise 2026.07.1; at least three servers for primary quorum; fourth server used when demonstrating a secondary
driver Python driver 6.3 with neo4j:// routing URI for a cluster; bolt:// is direct to one address
auth/TLS synthetic credentials only; production/remote cluster transport requires properly verified TLS and intra-cluster encryption as appropriate
plugins APOC/GDS not required for clustering lessons
measurement all failover/lag/latency/alert values must be measured on the learner’s licensed environment; trace outputs are explicitly labeled deterministic simulations

1. Start with a failure hypothesis

Each injection needs one mechanism to test. “Node failure” is too broad. The secondary-loss experiment tests read-capacity reduction/catch-up without changing write quorum. Writer-loss tests election/routing/retry while majority survives. Two-primary loss tests safety when quorum disappears. Combining all faults at once makes diagnosis impossible.

Experiment Hypothesis Primary acceptance criterion
secondary loss write consensus unaffected writes continue; reader capacity drops; secondary later catches up
writer loss with 2/3 primaries alive new primary can be elected writes recover through routing/retry; one committed business outcome
minority partition minority cannot safely commit only majority side retains write capability
2 of 3 primaries unavailable majority absent new writes unavailable; alert fires; no forced split brain
member return catch-up restores redundancy allocation reaches healthy/online and lag returns inside stated objective

2. Free deterministic trace first

Use a trace before a licensed chaos test. It forces you to state the expected system and application outcomes in advance. The none_temporarily writer state represents an election interval in the model; its duration is deliberately unspecified because the free trace cannot measure it.

CSV · deterministic HA reasoning trace
# FREE deterministic trace — save as ha_trace.csv and reason about every row.
time,event,primaries_alive,writer,secondary_state,request,expected_business_outcome
00:00,healthy,3,P1,caught_up,write O-9100,commit once
00:01,stop secondary,3,P1,down,read catalog,served by eligible reader
00:02,restart secondary,3,P1,catching_up,read with bookmark,may wait until state available
00:03,stop writer P1,2,none_temporarily,cancelled client write,ambiguous until reconciled
00:04,elect P2,2,P2,caught_up,retry idempotent O-9101,commit once
00:05,stop P3,1,P2,caught_up,new write O-9102,write unavailable because no majority
00:06,start P1,2,P2,catching_up,new write O-9103,available after quorum restored
Do not turn trace times into an SLO

The timestamps are ordering labels, not measured election/failover durations. Real timings depend on server load, network, driver configuration, election state and failure mode.

3. Stable business keys make ambiguous outcomes reconcilable

Suppose the writer commits an order but the connection breaks before the client receives success. The client sees a timeout, yet “retry the same business operation” and “create a second order” are not equivalent. Reuse the Chapter 6 idempotency discipline: stable orderId, uniqueness constraint, MERGE pattern scoped to identity, and post-error reconciliation.

Python 6.3 · classify error and preserve reconciliation key
from neo4j import GraphDatabase
from neo4j.exceptions import Neo4jError
import uuid

URI = "neo4j://server01.example.test:7687"
AUTH = ("atlasmart_app", "<synthetic-test-secret>")
ORDER_ID = "O-HA-" + uuid.uuid4().hex[:10]

QUERY = """
MERGE (o:Order {orderId:$order_id})
ON CREATE SET o.status='PENDING', o.createdAt=datetime()
RETURN o.orderId AS orderId, o.status AS status
"""

with GraphDatabase.driver(URI, auth=AUTH, max_transaction_retry_time=15) as driver:
    try:
        rows, summary, keys = driver.execute_query(
            QUERY,
            order_id=ORDER_ID,
            database_="atlasmart",
        )
        print("result", rows)
    except Neo4jError as exc:
        print("neo4j_code", exc.code)
        print("retryable", exc.is_retryable())
        # Do not blindly create a replacement business operation.
        # Reconcile ORDER_ID against the graph before deciding what the client should do.
Client observation What it proves Required next step
success response driver received successful result still record business invariant/telemetry
retryable Neo4j error driver/server classified a retryable failure managed retry may be safe if transaction function has no unsafe external side effect
timeout/disconnect client did not receive a definitive result query by stable orderId/operationId before issuing a new business action
constraint violation identity invariant rejected duplicate/invalid state diagnose source/retry logic; do not disable constraint

4. Optional licensed failure sequence

This sequence assumes a disposable 3-primary + 1-secondary Enterprise test database. Do not use host firewall edits or real production networks. A process/container stop is sufficient for node-loss tests; a network-partition test should use an isolated test network with a documented reset command.

Runbook · optional licensed HA drill
# OPTIONAL LICENSED LAB — execute only in an isolated test cluster.
# Before every step capture SHOW SERVERS, SHOW DATABASE atlasmart YIELD *, routing table,
# application success/error rate, retry count, and backup state.

1. Stop/restart only the secondary allocation host. Verify writes continue; measure catch-up/lag.
2. Stop the current writer primary. Verify another primary becomes writer while 2/3 majority remains.
3. During/around leader loss, issue an idempotent marker/order request with a stable business key.
   Record whether the client saw success, a retryable error, timeout, or ambiguous outcome.
4. Stop a second primary. Verify new writes do NOT commit with only 1/3 primaries alive.
5. Restore one primary. Verify quorum and writer availability return; reconcile all marker/order IDs.
6. Restore all members. Wait for healthy/online allocations and acceptable replication lag.
7. Compare final business invariants and alerts with the pre-injection baseline.

Reset rule: if any unexpected database becomes degraded, stop the exercise and restore topology
before injecting another fault.

5. Capture four evidence planes

Plane Evidence Why one plane is insufficient
database topology SHOW SERVERS / SHOW DATABASE, writer, role, replicationLag a healthy writer row does not prove the client refreshed routing
driver routing refresh, retry classification/count, pool acquisition, bookmark state driver recovery does not prove business state committed exactly once
application request ID, order ID, response/error, reconciliation result application success alone hides reduced redundancy/lag
infrastructure/ops process health, CPU, disk, network, logs, alert timestamps process survival does not prove quorum or database availability

6. Alerts must reflect topology semantics

“Host down” is useful but not enough. AtlasMart needs alerts that encode the database risk: no writer, current primaries below requested, replication lag beyond the workload’s freshness objective, store-copy/deallocating stuck, routing/retry error spikes, and capacity saturation during degraded operation. The same host failure can be low severity for a secondary and urgent for the second lost primary in a 3P database.

Alert Severity context
secondary unavailable read-capacity/freshness risk; write quorum unchanged
one of 3 primaries unavailable writes still possible but fault headroom exhausted
writer absent during election short transient may be expected; sustained absence is critical
two of 3 primaries unavailable critical: write quorum unavailable
replication lag high after return recovery not complete; causal reads may wait, redundancy/freshness objective not met
retry/timeout spike client-visible instability; correlate with topology before increasing retry budgets

7. Measure failover without hiding overload

Record request throughput and p50/p95/p99 latency before, during and after failure; transaction error/retry counts; time without an available writer; pool queueing; CPU/disk/network on survivors; replication lag/catch-up; and final business invariants. Do not report one average or only the requests that succeeded. A cluster can “fail over” while the remaining nodes are saturated and p99/error rate violate the SLO.

Metric Baseline Failure window Recovery acceptance
write success/error rate record record all errors/timeouts returns inside defined objective
p50/p95/p99 record warm workload record distribution, not average tail stabilizes within envelope
writer availability one writer measure no-writer interval writer present with quorum
retry count normal baseline record spike and classifications returns to baseline; no retry storm
replication lag normal may increase inside stated freshness/redundancy objective
business invariant known count/unique order IDs track ambiguous requests no duplicate/lost accepted operation after reconciliation

8. Deliberately unsafe response: “increase retries”

When failover produces errors, a common reaction is to increase retry count/time globally. That can amplify overload, extend tail latency and repeat non-idempotent side effects. Repair the mechanism instead: current official driver, correct routing URI/access mode, bounded retry budget, small idempotent transaction functions, stable operation IDs, reconciliation, pool/queue limits and capacity headroom.

Ambiguity is an application state

A timeout is not proof of rollback. Do not automatically create a second order/payment/shipment. Reconcile the stable operation key first.

9. Exit criteria and cleanup

  • All requested AtlasMart allocations are online and healthy.
  • Exactly one writer exists for the database.
  • Current primary/secondary counts match requested topology.
  • Replication lag is inside the declared objective.
  • Driver retry/error rates and pool queueing return to baseline.
  • Every injected business operation is reconciled by stable ID with no unintended duplicates.
  • Backup/restore status remains valid and independent of cluster copies.
  • Every alert fired at the expected severity and cleared after recovery.

Check your understanding

  1. Why test secondary loss separately from writer loss?
  2. Does a client timeout prove the transaction rolled back?
  3. What should happen with only one of three primaries alive?
  4. What evidence proves failover is acceptable?
  5. Why not simply lengthen retry time?
Review the answers

1. They exercise different mechanisms: read capacity/catch-up versus election/routing/write continuity.

2. No. The commit outcome may be ambiguous to the client; reconcile by stable business identity.

3. New fault-tolerant writes should be unavailable because the simple majority is lost.

4. Topology + driver + application + infrastructure evidence, including tail latency/errors and final business invariants.

5. It can hide topology/capacity problems, amplify overload, and repeat unsafe side effects; retries must be bounded and semantically safe.

Production judgment

Decision surface Questions that must be answered before production
graph/workload fit Which writes need fault tolerance? How much read scaling is needed? Does graph fan-out make reader CPU/page-cache the bottleneck rather than topology?
correctness Which operations need causal read-after-write? Which externally visible side effects require idempotency/reconciliation after an ambiguous client outcome?
topology/cardinality How many primaries satisfy failure tolerance? Are secondaries justified by read demand? Are dense/hot entities producing lock contention that topology will not solve?
latency What are write p50/p95/p99 with quorum across actual zones? What read tail latency is added by lag/bookmark waiting/routing refresh?
transactions/concurrency How do leader changes, deadlocks, retries and long transactions interact? Is transaction work idempotent under managed retries?
memory/CPU/disk/network Can remaining servers absorb a failed member? Is there space/bandwidth for store copy and catch-up while serving traffic?
indexes/constraints Are access paths and uniqueness contracts identical/ONLINE across allocations after seeding or recovery?
driver Is one maintained driver reused? Are routing URI, database selection, read/write modes, pool sizes, acquisition timeouts, max retry time and bookmarks intentional?
security Are Bolt, HTTPS and intra-cluster links encrypted as required? Are server-management privileges and certificates isolated from application credentials?
backup/recovery Replication is not backup. Where are off-host/immutable backups and when was an isolated restore last proven?
observability Do alerts cover writer absence, unavailable primaries, replication lag, store copy, routing failures, retry spikes, queueing and capacity headroom?
testing/failure injection Have host loss, reader loss, writer loss, majority loss, partition, slow network, maintenance and ambiguous commit outcomes been tested safely?
version/edition Are server/Java/driver/APOC/GDS/Cypher/discovery settings compatible with the exact release and license?
Aura/self-managed Which topology choices are operator-owned self-managed responsibilities versus managed-service controls/SLOs?
cost/migration What is the cost of extra primaries/secondaries/zones and cross-zone traffic, and how will topology changes or rollback be staged?

Summary and next step

Chapter 17 has turned “clustering” into measurable per-database topology, quorum, routing, failure-domain and application-recovery behavior. Chapter 18 now focuses on the metrics, logs, query/transaction visibility, memory, store and capacity evidence needed to operate that architecture continuously.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.