Chapter 02 · Clusters, Datacenters, Racks, Snitches, Gossip, and Failure Detection
Failure Detection, Phi Accrual Concepts, Suspicions, DOWN State, and Client Impact
Treat failure detection as probabilistic evidence, not permanent truth: watch suspicion rise, see how consistency reacts, and separate server and driver perspectives.
Learning outcomes
Gossip explains how state is exchanged; the failure detector explains how a node decides that another endpoint has become suspicious enough to mark DOWN. AtlasMart needs to understand that decision probabilistically so it can interpret failures, client retries and consistency errors without treating a detector threshold as proof of permanent death.
Explain the phi-accrual idea as a suspicion score derived from heartbeat arrival history rather than a fixed timeout.
Inspect nodetool failuredetector before, during and after a controlled peer pause.
Distinguish Cassandra peer suspicion, coordinator availability decisions and driver NodeState.
Relate one replica failure to RF/CL math without claiming all failures are equivalent.
Build an incident timeline that records signals instead of immediately removing the suspected node.
Apache Cassandra 5.0.9 is the current GA 5.0 patch on the
official download page. These labs pin
cassandra:5.0.9. The chapter uses one logical
datacenter, dc1, and three logical racks,
rack1–rack3, on the isolated Docker
network atlasmart-cassandra. Topology labels are
learning metadata; they do not create real
host/rack/availability-zone isolation by themselves.
The generation environment does not contain Docker or Cassandra, so commands were documentation- and syntax-checked rather than executed here. Record your actual addresses, host IDs, tokens, gossip generations, phi values, startup times, and driver events. Never treat the example output shapes as captured measurements. All failure injection is confined to disposable course containers; no host firewall, clock manipulation, or production endpoint is required.
1. Phi accrual asks “how surprising is this silence?”
A traditional fixed-timeout detector says “if no heartbeat
arrives within X seconds, mark the peer down.” Cassandra's
phi-accrual detector instead tracks heartbeat arrival behavior
and produces a phi (φ) suspicion value: as the
absence of an expected heartbeat becomes statistically less
likely under the observed history, suspicion rises. When phi
crosses phi_convict_threshold, the peer can be
convicted as down by that observer.
In the current 5.0 configuration reference, the default threshold is 8 and the documentation says most users should not adjust it casually. The exact phi value is not a wall-clock guarantee and should not be copied as a universal tuning recipe. GC pauses, CPU starvation, network delay, packet loss and scheduler stalls can all change heartbeat timing.
“DOWN” means this observer's evidence crossed its suspicion threshold. It does not prove the machine is destroyed, its disk is corrupt, or it should be removed from the ring.
2. Establish baseline failure-detector evidence
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait for all three nodes to become Up/Normal, then verify from two peers.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-2 nodetool status
docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool failuredetectordocker exec atlasmart-cass-2 nodetool failuredetector
Capture the output with a timestamp. Different observers can
report different phi values because each has its own heartbeat
observations. Healthy peers should remain Up/Normal in
nodetool status even though numerical detector
values fluctuate.
3. Cross the suspicion boundary safely
Pause node 3 and observe from node 1. Record three moments: immediately after the pause, after the detector marks the endpoint down, and after unpause/recovery. Do not hard-code how quickly the transition must occur.
docker pause atlasmart-cass-3# Wait and sample several times rather than assuming one fixed delay.docker exec atlasmart-cass-1 nodetool failuredetectordocker exec atlasmart-cass-1 nodetool status# Bash: sleep 10 PowerShell: Start-Sleep -Seconds 10docker exec atlasmart-cass-1 nodetool failuredetectordocker exec atlasmart-cass-1 nodetool statusdocker unpause atlasmart-cass-3# Wait for the peer to resume heartbeats and for state to converge.docker exec atlasmart-cass-1 nodetool failuredetectordocker exec atlasmart-cass-1 nodetool status
A useful report is a sequence such as “T0: three UN nodes; T1: node 3 paused; T2: node 1 marks node 3 DN after its local detector crosses threshold; T3: node 3 unpaused; T4: node 1 returns node 3 to UN.” Exact seconds and phi numbers are measurements, not lesson constants.
4. Client impact depends on RF, CL and coordinator path
Create the chapter's RF=3 keyspace and a deterministic row. At RF=3, a QUORUM operation generally needs two replicas. Losing one replica can still leave a quorum; losing two cannot. That arithmetic is necessary but not sufficient: the coordinator must also reach the required replicas before its timeout, and driver routing/retries affect which coordinator sees the request.
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_failure WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_failure.inventory_by_sku (sku text PRIMARY KEY, available int, updated_at timestamp);"docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; INSERT INTO atlasmart_failure.inventory_by_sku (sku,available,updated_at) VALUES ('sku-9001',42,'2026-09-07T12:00:00Z');"docker pause atlasmart-cass-3# After failure detection converges, this RF=3 QUORUM read normally still has two reachable replicas.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_failure.inventory_by_sku WHERE sku='sku-9001';"docker pause atlasmart-cass-2# Now only one replica is responsive; QUORUM cannot normally be satisfied.docker exec atlasmart-cass-1 cqlsh -e "CONSISTENCY QUORUM; SELECT * FROM atlasmart_failure.inventory_by_sku WHERE sku='sku-9001';"docker unpause atlasmart-cass-2docker unpause atlasmart-cass-3
Do not treat a timeout/unavailable error as evidence that no replica applied a write; ambiguous outcomes are a later driver/retry topic. This lab uses reads to make the replica-count boundary easier to reason about.
5. Driver state and incident response
The Java driver has its own node-state model. A node can be
DOWN because the driver lost all connections, or
UP while gossip on a Cassandra peer is suspicious
if the driver still has active connections. This means an
application incident can contain asymmetric evidence: “node A
says B is down, but clients can still reach B.” That is a
network-topology clue, not a contradiction to erase.
The deliberately wrong response is to run
removenode, replace, or decommission immediately
because one observer shows DOWN. First classify the failure:
process stopped, host unavailable, peer-to-peer path impaired,
client path impaired, overloaded JVM, or planned maintenance.
Removal/lifecycle commands have ownership and streaming
consequences and belong to later operational procedures.
Verification checklist
-
You sampled
nodetool failuredetectorfrom at least one observer. - You recorded a Down and recovery transition without changing the host clock/firewall.
- You explained why the default phi threshold is not a universal latency timeout.
- You demonstrated the RF=3/QUORUM one-failure versus two-failure boundary.
- You did not confuse driver node state with Cassandra peer suspicion.
Check your understanding
- What does phi represent conceptually?
- Does phi=8 mean a peer has been unreachable for exactly eight seconds?
- Why can RF=3 QUORUM usually tolerate one unavailable replica?
- Why should a timeout not be blindly retried for every write?
- What is the first response to one peer showing another as DOWN?
Review the answers
1. A suspicion measure based on how unlikely the current heartbeat silence is relative to observed arrival history.
2. No. Phi is not seconds; it is a statistical suspicion score.
3. QUORUM requires two replicas for RF=3, so two reachable replicas can satisfy it.
4. A timed-out write can have been applied on some replicas, so retries require idempotency and error classification.
5. Collect correlated evidence and classify the failure; do not immediately remove the node from ownership.
Cleanup
docker exec atlasmart-cass-1 cqlsh -e "DROP KEYSPACE IF EXISTS atlasmart_failure;"
docker unpause atlasmart-cass-1 2>/dev/null || truedocker unpause atlasmart-cass-2 2>/dev/null || truedocker unpause atlasmart-cass-3 2>/dev/null || truedocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
docker unpause atlasmart-cass-1 2>$nulldocker unpause atlasmart-cass-2 2>$nulldocker unpause atlasmart-cass-3 2>$nulldocker rm -f atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3docker volume rm atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-datadocker network rm atlasmart-cassandra
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Design Rack/Zone Placement That Survives Host, Rack, and Availability-Zone Failure.
Authoritative references
- Apache Cassandra 5.0 documentation — Official documentation entry point for the current 5.0 line.
- Apache Cassandra downloads — Official release page used to verify Cassandra 5.0.9 as the current GA patch.
- Snitch — Official explanation of topology/proximity and rack-aware replica placement.
- cassandra-rackdc.properties — Official DC/rack configuration for GossipingPropertyFileSnitch.
- Dynamo architecture and replication — Official NetworkTopologyStrategy and replica-placement semantics.
- nodetool gossipinfo — Official gossip-state inspection command.
- nodetool failuredetector — Official failure-detector inspection command.
- Java Driver 4.19 node metadata — Apache Java Driver documentation for driver-side node state and topology metadata.
- cassandra.yaml phi_convict_threshold — Official current threshold description and tuning warning.
- Java Driver NodeState API — Exact driver-side UP/DOWN/UNKNOWN/FORCED_DOWN semantics.