Chapter 23 · Node Lifecycle: Bootstrap, Replace, Decommission, Remove, Rebuild, and Cleanup

Replace a Dead Node: Address / Token Identity, replace_address, and Failure Scenarios

Replace a dead Cassandra node using recorded endpoint identity and clean storage; observe replacement streaming and define when post-replacement repair is required.

Intermediate → Advanced135–180 minutesDead-node replacement labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · isolated RF=3 lifecycle topology · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart loses a node and its local disk. The replacement server must assume the dead member's ownership, but reusing an old volume or starting a generic new node can create duplicate/wrong membership. This lesson makes replacement identity explicit: identify the dead endpoint/host ID, prove it is actually down, start a clean replacement with the documented replacement property, observe streaming, then repair when hint/write-forwarding coverage is insufficient.

01

Distinguish endpoint/host/token identity from container/VM name and data-directory identity.

02

Use replace_address semantics only after proving the target member is dead and recording its identity/topology.

03

Observe replacement streaming, replacement logs, membership and replica coverage with a clean data volume.

04

Explain same-address versus different-address write-forwarding/hint-window implications.

05

Reject live-node replacement, stale data-volume reuse and insufficient-source replacement as unsafe failure handling.

Chapter 23 lifecycle lab baseline

Topology changes are intentionally isolated from the shared course cluster. The mandatory labs use a disposable Docker network atlasmart-cassandra-life, cluster atlasmart-lifecycle, nodes atlasmart-life-1..4 (plus explicitly named replacement/cross-DC nodes where a lesson needs them), pinned Docker Official Image cassandra:5.0.9, Java 17 inside the image, dc1, racks rack1..rack3, and 16 virtual nodes (vnodes) per node. The keyspace atlasmart_lifecycle uses NetworkTopologyStrategy and replication factor (RF) 3 in dc1; normal verification uses LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no default time-to-live (TTL), and Cassandra's normal gc_grace_seconds. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only on this isolated single-host learning network. The labs never expose JMX or native transport to an untrusted network. Exact IP addresses, host IDs, tokens, streaming sources, bytes, duration, ownership percentages, failure-detection timing, disk usage and p95/p99 application latency are learner-captured evidence. Docker examples are Bash-compatible; Windows Docker Desktop users can run the same docker commands individually in PowerShell if a shell loop is inconvenient.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and lifecycle mental model

A Cassandra node is one member of a peer-to-peer cluster. A datacenter (DC) and rack are logical topology labels used by replication placement to represent failure domains. A partition key is hashed by the partitioner to a token; with virtual nodes (vnodes), one physical node owns many token positions. A replica stores a copy of a token range according to the keyspace replication strategy. A coordinator is whichever node handles one request, not a permanent leader. A consistency level (CL) defines the replica acknowledgments/responses required for an operation.

A topology transition changes cluster membership or token ownership. Bootstrap is the join process in which a new node receives ranges by streaming from current replicas before becoming a normal owner. Pending ranges are ownership changes that are being prepared during a transition: Cassandra must keep writes safe while the future replica set is not yet fully ready. Streaming transfers SSTable data over the internode network. Decommission removes a healthy node and streams its ranges away. Replacement gives the logical ownership of a dead member to a fresh node identified by the dead endpoint. removenode removes an unreachable member from another live node and re-replicates from survivors. rebuild repopulates data on the current node by streaming from selected source nodes/DCs without changing that node's ownership. cleanup is a compaction-style rewrite that discards ranges a node no longer owns after a completed range movement. Repair is anti-entropy reconciliation between replicas; bootstrap/rebuild/cleanup do not make pre-existing replica inconsistencies disappear.

Node identity includes endpoint/broadcast address, host ID and token ownership recorded by cluster metadata. Headroom is spare disk, network, CPU, memory, compaction and replica availability capacity required to perform the transition without violating service objectives. A rollback/abort path is the documented action for a stalled or failed lifecycle sequence; it must be chosen from the current Cassandra version's supported topology state, not improvised by deleting system tables or reusing data directories.

1. Replacement is not “add a node with the old hostname”

The cluster already has an ownership record for the dead member. A replacement operation tells Cassandra that a fresh process is assuming that member's token ownership, so surviving replicas stream the ranges it needs. The current 5.0 documentation exposes the JVM system property cassandra.replace_address; topology guidance and Cassandra source also recognize the one-time replace_address_first_boot form. Deployment wrappers differ, so inspect the exact 5.0.x property accepted by your packaging before freezing automation. The Docker Official Image passes JVM_EXTRA_OPTS through Cassandra's JVM configuration, which gives this local lab an explicit, auditable property.

The replacement data directory must be clean. Reusing the dead node's partially failed/unknown disk state while also asking Cassandra to bootstrap a replacement mixes two recovery models. The replacement rack/DC must match intended placement, and the endpoint named for replacement must actually correspond to a dead cluster member—not a merely slow or partitioned live node.

2. Create four nodes, record identity, then simulate disk loss

Docker · create the isolated three-node RF=3 starting topology
# These names are dedicated to Chapter 23; the reset block never targets the shared course cluster.docker network inspect atlasmart-cassandra-life >/dev/null 2>&1 || docker network create atlasmart-cassandra-lifedocker volume create atlasmart-life-1-datadocker volume create atlasmart-life-2-datadocker volume create atlasmart-life-3-datadocker run -d --name atlasmart-life-1 --hostname atlasmart-life-1 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-life-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-life-1 nodetool statusdocker run -d --name atlasmart-life-2 --hostname atlasmart-life-2 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-life-3 --hostname atlasmart-life-3 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all three appear UN from at least two observers.docker exec atlasmart-life-1 nodetool statusdocker exec atlasmart-life-2 nodetool status
CQL · create a small deterministic AtlasMart ownership fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_lifecycleWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_lifecycle.order_probe (  order_id text PRIMARY KEY,  customer_id text,  status text,  total decimal,  updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1001','cust-42','PAID',129.90,'2026-09-08T06:00:00Z');INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1002','cust-77','PACKING',89.50,'2026-09-08T06:01:00Z');INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1003','cust-42','SHIPPED',42.00,'2026-09-08T06:02:00Z');SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';
Docker · bootstrap node 4 into dc1/rack1
docker volume create atlasmart-life-4-datadocker run -d --name atlasmart-life-4 --hostname atlasmart-life-4 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-4-data:/var/lib/cassandra cassandra:5.0.9# Poll from separate terminals while the node is joining; a tiny dataset may finish too quickly to catch UJ.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-4 nodetool netstats -Hdocker exec atlasmart-life-4 nodetool bootstrap# Final acceptance requires node 4 to be UN and schema versions to agree.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-1 nodetool describecluster
Docker/nodetool · record node 4 identity before failure
DEAD_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' atlasmart-life-4)echo "dead endpoint candidate: $DEAD_IP"docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-1 cqlsh -e "SELECT peer, host_id, data_center, rack, tokens FROM system.peers_v2;" docker exec atlasmart-life-1 nodetool getendpoints atlasmart_lifecycle order_probe ord-1001
Docker · crash the member and model total local-disk loss
docker stop atlasmart-life-4# Wait until multiple observers report the endpoint DN; failure detection is not instantaneous.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-2 nodetool status atlasmart_lifecycle# Remove the dead container and its lost data volume only after recording its endpoint/host identity.docker rm atlasmart-life-4docker volume rm atlasmart-life-4-data

3. Start a clean replacement and observe ownership recovery

Docker · replace the dead endpoint with a fresh data directory
DEAD_IP=<paste-the-recorded-dead-endpoint-here>docker volume create atlasmart-life-4r-datadocker run -d --name atlasmart-life-4r --hostname atlasmart-life-4r --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 \  -e JVM_EXTRA_OPTS="-Dcassandra.replace_address=$DEAD_IP" \  -v atlasmart-life-4r-data:/var/lib/cassandra cassandra:5.0.9# While replacement is active:docker exec atlasmart-life-4r nodetool netstats -Hdocker logs --tail 200 atlasmart-life-4r# After completion:docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-1 nodetool checktokenmetadatadocker exec atlasmart-life-4r cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';"

Exact host ID behavior and topology-sequence output are version-dependent; record what Cassandra 5.0.9 reports rather than assuming a UUID. The invariant is logical ownership continuity: the dead endpoint is no longer an active member, the replacement becomes normal only after required data arrives, and replica coverage for the keyspace remains valid.

4. Edge cases: live target, hint window, same address and missing sources

Case Why it changes the procedure Required judgment
target is still live/partitioned two processes could contend for one ownership identity prove death from several observers and infrastructure control plane before replace
different replacement address surviving nodes can forward writes to the joining replacement during replacement still verify missed-write window and repair requirements
same replacement address current docs warn writes are not forwarded in the same way during replacement repair if replacement interval exceeds hint coverage
dead longer than max_hint_window hints cannot cover the full outage window run repair after replacement
too few surviving replicas no trustworthy source may exist for required ranges restore source availability/backup; do not force a fake successful replace
Deliberately wrong: replace a node that is merely slow.

Cassandra's replacement checks are designed to reject replacing a live member. Do not work around that protection with IP tricks, stale gossip deletion or unsafe replacement flags. Diagnose the network/process first. A split-brain ownership mistake is far more dangerous than waiting for a controlled decision.

Evidence/repair · post-replacement convergence when the outage exceeded hint coverage
docker exec atlasmart-life-1 nodetool getmaxhintwindow# If the replacement timeframe or same-address behavior exceeded documented hint/write-forwarding coverage:docker exec atlasmart-life-4r nodetool repair --full atlasmart_lifecycle order_probe# Verify from multiple nodes after repair.docker exec atlasmart-life-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';"docker exec atlasmart-life-2 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';"

5. Replacement acceptance and reset

  • Old endpoint/identity was captured before destructive infrastructure cleanup.
  • Multiple observers confirmed the target was down.
  • The replacement used a clean volume and intended DC/rack settings.
  • netstats/logs showed replacement streaming or completion evidence.
  • Cluster metadata/token checks and representative RF3 reads pass.
  • Repair is run when downtime/write-forwarding/hint coverage cannot prove convergence.

Check your understanding

  1. Why must a replacement use a clean data directory?
  2. Why record the old endpoint and host ID before deleting infrastructure?
  3. When is post-replacement repair especially required?
  4. Can replacement safely target a node that is still live?
  5. What if surviving replicas cannot provide a required range?
Review the answers

1. The workflow reconstructs the dead member from surviving replicas; reusing uncertain partial data mixes identities/recovery models and can invalidate assumptions.

2. Replacement/removal decisions are made against Cassandra membership identity, not the cloud/VM/container name.

3. When the node was down longer than hint coverage, same-address replacement missed writes beyond the hint window, or any other evidence cannot prove full convergence.

4. No. Prove the old member is truly dead; current Cassandra explicitly protects against replacing a live member.

5. Stop forcing the workflow and restore a valid source or recovery path; replacement cannot manufacture data that no source has.

Docker · reset only the dedicated Chapter 23 lifecycle lab
docker rm -f atlasmart-life-1 atlasmart-life-2 atlasmart-life-3 atlasmart-life-4 atlasmart-life-4r atlasmart-life-dc2-1 2>/dev/null || truedocker volume rm atlasmart-life-1-data atlasmart-life-2-data atlasmart-life-3-data atlasmart-life-4-data atlasmart-life-4r-data atlasmart-life-dc2-1-data 2>/dev/null || truedocker network rm atlasmart-cassandra-life 2>/dev/null || true

Production judgment

Topology work consumes the same resources that serve customer traffic. Before a bootstrap, decommission, replacement, removal or rebuild, record Cassandra/JDK/driver versions, DC/rack layout, RF and query CLs, vnode count, per-node used/free disk, compaction backlog, streaming throughput, NIC saturation, JVM/GC pressure, repair age, hint window, backup status, SAI/vector indexes, tenant/security constraints, driver timeouts/retries/idempotency and representative p50/p95/p99 latency. Confirm that losing one additional host/rack during the operation still leaves the required replicas and operational headroom. Avoid concurrent range movements in the same failure domain unless the current version and runbook explicitly prove safety.

Streaming moves bytes; it does not cure bad partition keys, oversized partitions, old replica divergence, poor compaction headroom or missing repair. Cleanup reclaims no-longer-owned ranges; it is not anti-entropy. Replacement preserves ownership identity but still requires post-replacement convergence when downtime/write-forwarding exceeds the mechanisms that could cover missed writes. Managed Cassandra services may hide or forbid these commands; use the provider's documented replacement/scale workflow but keep the same evidence and acceptance model. Lesson 4 separates two more operations that are often conflated with replacement: removing a permanently unreachable member from surviving nodes, and rebuilding the current node’s replicas from a chosen source datacenter.

Summary and next bridge

Replacement preserves logical ownership for a dead member while rebuilding a fresh local data set from surviving replicas. Correct endpoint identity, clean storage, source availability and post-replacement convergence are part of the operation. Next, remove a dead member when it will not be replaced and contrast that membership change with a rebuild that does not change membership.

Authoritative references

Lifecycle commands are version-sensitive and state-sensitive. Re-check these sources, the release notes and your deployment's orchestration/managed-service rules before applying a topology procedure.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.