Chapter 02 · Clusters, Datacenters, Racks, Snitches, Gossip, and Failure Detection

Failure Detection, Phi Accrual Concepts, Suspicions, DOWN State, and Client Impact

Treat failure detection as probabilistic evidence, not permanent truth: watch suspicion rise, see how consistency reacts, and separate server and driver perspectives.

Intermediate115–140 minutesFailure-detector + RF/CL labApache Cassandra 5.0.9 · 3-node dc1/rack1–3 labLast reviewed: September 2026

Learning outcomes

Gossip explains how state is exchanged; the failure detector explains how a node decides that another endpoint has become suspicious enough to mark DOWN. AtlasMart needs to understand that decision probabilistically so it can interpret failures, client retries and consistency errors without treating a detector threshold as proof of permanent death.

01

Explain the phi-accrual idea as a suspicion score derived from heartbeat arrival history rather than a fixed timeout.

02

Inspect nodetool failuredetector before, during and after a controlled peer pause.

03

Distinguish Cassandra peer suspicion, coordinator availability decisions and driver NodeState.

04

Relate one replica failure to RF/CL math without claiming all failures are equivalent.

05

Build an incident timeline that records signals instead of immediately removing the suspected node.

Chapter baseline reviewed 7 September 2026

Apache Cassandra 5.0.9 is the current GA 5.0 patch on the official download page. These labs pin cassandra:5.0.9. The chapter uses one logical datacenter, dc1, and three logical racks, rack1–rack3, on the isolated Docker network atlasmart-cassandra. Topology labels are learning metadata; they do not create real host/rack/availability-zone isolation by themselves.

Execution and safety note

The generation environment does not contain Docker or Cassandra, so commands were documentation- and syntax-checked rather than executed here. Record your actual addresses, host IDs, tokens, gossip generations, phi values, startup times, and driver events. Never treat the example output shapes as captured measurements. All failure injection is confined to disposable course containers; no host firewall, clock manipulation, or production endpoint is required.

1. Phi accrual asks “how surprising is this silence?”

A traditional fixed-timeout detector says “if no heartbeat arrives within X seconds, mark the peer down.” Cassandra's phi-accrual detector instead tracks heartbeat arrival behavior and produces a phi (φ) suspicion value: as the absence of an expected heartbeat becomes statistically less likely under the observed history, suspicion rises. When phi crosses phi_convict_threshold, the peer can be convicted as down by that observer.

In the current 5.0 configuration reference, the default threshold is 8 and the documentation says most users should not adjust it casually. The exact phi value is not a wall-clock guarantee and should not be copied as a universal tuning recipe. GC pauses, CPU starvation, network delay, packet loss and scheduler stalls can all change heartbeat timing.

Failure detectors are detectors, not oracles.

“DOWN” means this observer's evidence crossed its suspicion threshold. It does not prove the machine is destroyed, its disk is corrupt, or it should be removed from the ring.

2. Establish baseline failure-detector evidence

bash / PowerShell · reusable three-node topology lab
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait for all three nodes to become Up/Normal, then verify from two peers.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool status
baseline · compare detector views on two peers
docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool failuredetectordocker exec atlasmart-cass-2 nodetool failuredetector

Capture the output with a timestamp. Different observers can report different phi values because each has its own heartbeat observations. Healthy peers should remain Up/Normal in nodetool status even though numerical detector values fluctuate.

3. Cross the suspicion boundary safely

Pause node 3 and observe from node 1. Record three moments: immediately after the pause, after the detector marks the endpoint down, and after unpause/recovery. Do not hard-code how quickly the transition must occur.

failure timeline · phi/suspicion, DOWN state, recovery
docker pause atlasmart-cass-3# Wait and sample several times rather than assuming one fixed delay.docker exec atlasmart-cass-1 nodetool failuredetectordocker exec atlasmart-cass-1 nodetool status# Bash: sleep 10    PowerShell: Start-Sleep -Seconds 10docker exec atlasmart-cass-1 nodetool failuredetectordocker exec atlasmart-cass-1 nodetool statusdocker unpause atlasmart-cass-3# Wait for the peer to resume heartbeats and for state to converge.docker exec atlasmart-cass-1 nodetool failuredetectordocker exec atlasmart-cass-1 nodetool status

A useful report is a sequence such as “T0: three UN nodes; T1: node 3 paused; T2: node 1 marks node 3 DN after its local detector crosses threshold; T3: node 3 unpaused; T4: node 1 returns node 3 to UN.” Exact seconds and phi numbers are measurements, not lesson constants.

4. Client impact depends on RF, CL and coordinator path

Create the chapter's RF=3 keyspace and a deterministic row. At RF=3, a QUORUM operation generally needs two replicas. Losing one replica can still leave a quorum; losing two cannot. That arithmetic is necessary but not sufficient: the coordinator must also reach the required replicas before its timeout, and driver routing/retries affect which coordinator sees the request.

RF/CL boundary · one failure vs two failures
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_failure WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_failure.inventory_by_sku (sku text PRIMARY KEY, available int, updated_at timestamp);"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; INSERT INTO atlasmart_failure.inventory_by_sku (sku,available,updated_at) VALUES ('sku-9001',42,'2026-09-07T12:00:00Z');"docker pause atlasmart-cass-3# After failure detection converges, this RF=3 QUORUM read normally still has two reachable replicas.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_failure.inventory_by_sku WHERE sku='sku-9001';"docker pause atlasmart-cass-2# Now only one replica is responsive; QUORUM cannot normally be satisfied.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_failure.inventory_by_sku WHERE sku='sku-9001';"docker unpause atlasmart-cass-2docker unpause atlasmart-cass-3

Do not treat a timeout/unavailable error as evidence that no replica applied a write; ambiguous outcomes are a later driver/retry topic. This lab uses reads to make the replica-count boundary easier to reason about.

5. Driver state and incident response

The Java driver has its own node-state model. A node can be DOWN because the driver lost all connections, or UP while gossip on a Cassandra peer is suspicious if the driver still has active connections. This means an application incident can contain asymmetric evidence: “node A says B is down, but clients can still reach B.” That is a network-topology clue, not a contradiction to erase.

The deliberately wrong response is to run removenode, replace, or decommission immediately because one observer shows DOWN. First classify the failure: process stopped, host unavailable, peer-to-peer path impaired, client path impaired, overloaded JVM, or planned maintenance. Removal/lifecycle commands have ownership and streaming consequences and belong to later operational procedures.

Verification checklist

  • You sampled nodetool failuredetector from at least one observer.
  • You recorded a Down and recovery transition without changing the host clock/firewall.
  • You explained why the default phi threshold is not a universal latency timeout.
  • You demonstrated the RF=3/QUORUM one-failure versus two-failure boundary.
  • You did not confuse driver node state with Cassandra peer suspicion.

Check your understanding

  1. What does phi represent conceptually?
  2. Does phi=8 mean a peer has been unreachable for exactly eight seconds?
  3. Why can RF=3 QUORUM usually tolerate one unavailable replica?
  4. Why should a timeout not be blindly retried for every write?
  5. What is the first response to one peer showing another as DOWN?
Review the answers

1. A suspicion measure based on how unlikely the current heartbeat silence is relative to observed arrival history.

2. No. Phi is not seconds; it is a statistical suspicion score.

3. QUORUM requires two replicas for RF=3, so two reachable replicas can satisfy it.

4. A timed-out write can have been applied on some replicas, so retries require idempotency and error classification.

5. Collect correlated evidence and classify the failure; do not immediately remove the node from ownership.

Cleanup

remove failure-test keyspace
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_failure;"
cleanup · Bash form; remove only course-owned resources
docker unpause atlasmart-cass-1 2>/dev/null || truedocker unpause atlasmart-cass-2 2>/dev/null || truedocker unpause atlasmart-cass-3 2>/dev/null || truedocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
cleanup · PowerShell form
docker unpause atlasmart-cass-1 2>$nulldocker unpause atlasmart-cass-2 2>$nulldocker unpause atlasmart-cass-3 2>$nulldocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to Design Rack/Zone Placement That Survives Host, Rack, and Availability-Zone Failure.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.