Chapter 23 · Node Lifecycle: Bootstrap, Replace, Decommission, Remove, Rebuild, and Cleanup

Remove Unreachable Nodes, Rebuild from Another Datacenter, and Understand Streaming Sources

Contrast removenode of an unreachable member with rebuild of a live member from a selected source datacenter, including source and WAN-capacity tradeoffs.

Intermediate → Advanced140–190 minutesremovenode + cross-DC rebuild labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · isolated RF=3 lifecycle topology · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart has two different incidents: one dead node will never return and should be removed; separately, a surviving node needs to repopulate local replica data from a remote datacenter. Both involve streaming, but their ownership semantics are opposite. removenode changes membership; rebuild keeps membership and repopulates the current node from selected sources.

01

Use host ID—not a guessed IP/container name—to remove an unreachable Cassandra member.

02

Explain how removenode re-replicates from surviving replicas and why force/assassinate are last-resort paths.

03

Explain rebuild as local replica repopulation from a source DC without a token-ownership change.

04

Create a free local second-DC extension and observe a bounded cross-DC rebuild.

05

Compare source-selection, bandwidth, RF/CL and failure-domain risks for removal versus rebuild.

Chapter 23 lifecycle lab baseline

Topology changes are intentionally isolated from the shared course cluster. The mandatory labs use a disposable Docker network atlasmart-cassandra-life, cluster atlasmart-lifecycle, nodes atlasmart-life-1..4 (plus explicitly named replacement/cross-DC nodes where a lesson needs them), pinned Docker Official Image cassandra:5.0.9, Java 17 inside the image, dc1, racks rack1..rack3, and 16 virtual nodes (vnodes) per node. The keyspace atlasmart_lifecycle uses NetworkTopologyStrategy and replication factor (RF) 3 in dc1; normal verification uses LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no default time-to-live (TTL), and Cassandra's normal gc_grace_seconds. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only on this isolated single-host learning network. The labs never expose JMX or native transport to an untrusted network. Exact IP addresses, host IDs, tokens, streaming sources, bytes, duration, ownership percentages, failure-detection timing, disk usage and p95/p99 application latency are learner-captured evidence. Docker examples are Bash-compatible; Windows Docker Desktop users can run the same docker commands individually in PowerShell if a shell loop is inconvenient.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and lifecycle mental model

A Cassandra node is one member of a peer-to-peer cluster. A datacenter (DC) and rack are logical topology labels used by replication placement to represent failure domains. A partition key is hashed by the partitioner to a token; with virtual nodes (vnodes), one physical node owns many token positions. A replica stores a copy of a token range according to the keyspace replication strategy. A coordinator is whichever node handles one request, not a permanent leader. A consistency level (CL) defines the replica acknowledgments/responses required for an operation.

A topology transition changes cluster membership or token ownership. Bootstrap is the join process in which a new node receives ranges by streaming from current replicas before becoming a normal owner. Pending ranges are ownership changes that are being prepared during a transition: Cassandra must keep writes safe while the future replica set is not yet fully ready. Streaming transfers SSTable data over the internode network. Decommission removes a healthy node and streams its ranges away. Replacement gives the logical ownership of a dead member to a fresh node identified by the dead endpoint. removenode removes an unreachable member from another live node and re-replicates from survivors. rebuild repopulates data on the current node by streaming from selected source nodes/DCs without changing that node's ownership. cleanup is a compaction-style rewrite that discards ranges a node no longer owns after a completed range movement. Repair is anti-entropy reconciliation between replicas; bootstrap/rebuild/cleanup do not make pre-existing replica inconsistencies disappear.

Node identity includes endpoint/broadcast address, host ID and token ownership recorded by cluster metadata. Headroom is spare disk, network, CPU, memory, compaction and replica availability capacity required to perform the transition without violating service objectives. A rollback/abort path is the documented action for a stalled or failed lifecycle sequence; it must be chosen from the current Cassandra version's supported topology state, not improvised by deleting system tables or reusing data directories.

1. removenode: the member is dead and will not stream

For an unreachable node, execute nodetool removenode <host-id> from another live member. Surviving replicas stream data needed to restore the configured replica placement. The host ID is the membership identity shown by nodetool status/system metadata; it is not interchangeable with an IP string. removenode status reports an in-progress removal and removenode force exists to force completion of a stuck removal, not as the normal first command. Cassandra lists assassinate as a last resort when removenode cannot be used because it removes membership without normal re-replication safety.

2. Controlled dead-node removal

Docker · create the isolated three-node RF=3 starting topology
# These names are dedicated to Chapter 23; the reset block never targets the shared course cluster.docker network inspect atlasmart-cassandra-life >/dev/null 2>&1 || docker network create atlasmart-cassandra-lifedocker volume create atlasmart-life-1-datadocker volume create atlasmart-life-2-datadocker volume create atlasmart-life-3-datadocker run -d --name atlasmart-life-1 --hostname atlasmart-life-1 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-life-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-life-1 nodetool statusdocker run -d --name atlasmart-life-2 --hostname atlasmart-life-2 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-life-3 --hostname atlasmart-life-3 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all three appear UN from at least two observers.docker exec atlasmart-life-1 nodetool statusdocker exec atlasmart-life-2 nodetool status
CQL · create a small deterministic AtlasMart ownership fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_lifecycleWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_lifecycle.order_probe (  order_id text PRIMARY KEY,  customer_id text,  status text,  total decimal,  updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1001','cust-42','PAID',129.90,'2026-09-08T06:00:00Z');INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1002','cust-77','PACKING',89.50,'2026-09-08T06:01:00Z');INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1003','cust-42','SHIPPED',42.00,'2026-09-08T06:02:00Z');SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';
Docker · bootstrap node 4 into dc1/rack1
docker volume create atlasmart-life-4-datadocker run -d --name atlasmart-life-4 --hostname atlasmart-life-4 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-4-data:/var/lib/cassandra cassandra:5.0.9# Poll from separate terminals while the node is joining; a tiny dataset may finish too quickly to catch UJ.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-4 nodetool netstats -Hdocker exec atlasmart-life-4 nodetool bootstrap# Final acceptance requires node 4 to be UN and schema versions to agree.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-1 nodetool describecluster
Docker/nodetool · record host ID, kill node 4 and wait for DN
docker exec atlasmart-life-1 nodetool status atlasmart_lifecycle# Copy the Host ID shown for atlasmart-life-4 / its endpoint before continuing.DEAD_HOST_ID=<paste-node-4-host-id-here>docker stop atlasmart-life-4docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-2 nodetool status atlasmart_lifecycle
Docker/nodetool · remove the unreachable membership from a live node
DEAD_HOST_ID=<paste-node-4-host-id-here># Terminal A:docker exec atlasmart-life-1 nodetool removenode "$DEAD_HOST_ID"# Terminal B, while removal is active:docker exec atlasmart-life-1 nodetool removenode statusdocker exec atlasmart-life-1 nodetool netstats -Hdocker exec atlasmart-life-2 nodetool netstats -H# Acceptance:docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-1 nodetool checktokenmetadatadocker exec atlasmart-life-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';"
Do not switch immediately to removenode force.

If removal stalls, identify the missing streaming source or failed participant. Forcing metadata completion can leave replica coverage incomplete. Restore source availability or use the current-version documented recovery path before bypassing the transition.

3. rebuild: same node, same ownership, chosen streaming source

nodetool rebuild repopulates data on the node on which it runs by streaming from other nodes. It is commonly used when populating/restoring a datacenter or when a node needs data rebuilt from another DC. The command accepts a source DC and current 5.0 options for a keyspace, explicit sources/tokens, or excluding the local DC. Unlike bootstrap/decommission/removenode, rebuild does not itself transfer token ownership away from or to the current member.

The local extension below adds one dc2 source and changes only the disposable keyspace to RF dc1:3, dc2:1. A repair populates the new DC before the rebuild test; production multi-DC expansion requires a deliberate DC bootstrap/repair plan rather than this compressed lab sequence.

Docker/CQL · add a single local dc2 source for the rebuild experiment
docker volume create atlasmart-life-dc2-1-datadocker run -d --name atlasmart-life-dc2-1 --hostname atlasmart-life-dc2-1 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc2 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-dc2-1-data:/var/lib/cassandra cassandra:5.0.9# Wait for the dc2 node to be UN, then change this disposable keyspace only.docker exec atlasmart-life-1 cqlsh -e "ALTER KEYSPACE atlasmart_lifecycle WITH replication = {'class':'NetworkTopologyStrategy','dc1':3,'dc2':1};"# Populate new replicas; exact repair flags/cadence are Chapter 24 material.docker exec atlasmart-life-1 nodetool repair --full atlasmart_lifecycle order_probedocker exec atlasmart-life-dc2-1 cqlsh -e "CONSISTENCY LOCAL_ONE; SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';"
Docker/nodetool · rebuild node 1 using dc2 as the streaming source
# Terminal A:docker exec atlasmart-life-1 nodetool rebuild -ks atlasmart_lifecycle dc2# Terminal B:docker exec atlasmart-life-1 nodetool netstats -Hdocker stats --no-stream atlasmart-life-1 atlasmart-life-dc2-1# Rebuild should not change node 1 token ownership/membership.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycle

4. Source-selection and network judgment

Workflow Membership changes? Primary source Network implication
removenode dead node Yes surviving replicas remaining nodes both serve traffic and re-replicate
rebuild from dc2 No replicas in selected source DC WAN/cross-DC bandwidth and egress can dominate
bootstrap new node Yes current replicas selected for pending ranges incoming node + sources share load
replace dead node Yes, preserving old ownership role surviving replicas source availability determines whether replacement can finish

A rebuild from another DC is not automatically a good idea because it is syntactically available. Check source DC health, WAN capacity/cost, encryption, throttles, locality and whether the chosen source has been repaired. On a managed service, direct rebuild/removenode may be unavailable; use provider workflows and retain the same evidence model.

5. Verification and reset

Check your understanding

  1. Why does removenode take a host ID?
  2. Where does data come from when the dead node cannot stream?
  3. What is the key semantic difference between rebuild and bootstrap?
  4. Why can cross-DC rebuild be operationally expensive?
  5. When is assassinate appropriate?
Review the answers

1. It removes a Cassandra membership identity; IP/container names are not sufficient stable identity for the operation.

2. The surviving replicas stream/re-replicate the ranges needed by the future replica placement.

3. Rebuild repopulates data for the current node without assigning it new ownership; bootstrap is a membership/ownership transition for a joining node.

4. It can consume WAN bandwidth/egress, source-node disk/network capacity and foreground latency headroom.

5. Only as a documented last resort when normal removenode cannot be completed; it bypasses normal re-replication safety and therefore demands explicit recovery verification.

Docker · reset only the dedicated Chapter 23 lifecycle lab
docker rm -f atlasmart-life-1 atlasmart-life-2 atlasmart-life-3 atlasmart-life-4 atlasmart-life-4r atlasmart-life-dc2-1 2>/dev/null || truedocker volume rm atlasmart-life-1-data atlasmart-life-2-data atlasmart-life-3-data atlasmart-life-4-data atlasmart-life-4r-data atlasmart-life-dc2-1-data 2>/dev/null || truedocker network rm atlasmart-cassandra-life 2>/dev/null || true

Production judgment

Topology work consumes the same resources that serve customer traffic. Before a bootstrap, decommission, replacement, removal or rebuild, record Cassandra/JDK/driver versions, DC/rack layout, RF and query CLs, vnode count, per-node used/free disk, compaction backlog, streaming throughput, NIC saturation, JVM/GC pressure, repair age, hint window, backup status, SAI/vector indexes, tenant/security constraints, driver timeouts/retries/idempotency and representative p50/p95/p99 latency. Confirm that losing one additional host/rack during the operation still leaves the required replicas and operational headroom. Avoid concurrent range movements in the same failure domain unless the current version and runbook explicitly prove safety.

Streaming moves bytes; it does not cure bad partition keys, oversized partitions, old replica divergence, poor compaction headroom or missing repair. Cleanup reclaims no-longer-owned ranges; it is not anti-entropy. Replacement preserves ownership identity but still requires post-replacement convergence when downtime/write-forwarding exceeds the mechanisms that could cover missed writes. Managed Cassandra services may hide or forbid these commands; use the provider's documented replacement/scale workflow but keep the same evidence and acceptance model. Lesson 5 closes the lifecycle by reclaiming ranges that old owners intentionally retained after movement and proving that cleanup changes disk state without being mistaken for repair.

Summary and next bridge

Removing a dead member and rebuilding a live member can both stream bytes but have different ownership semantics. removenode changes membership and relies on survivors; rebuild keeps membership while choosing source replicas. The final lifecycle lesson handles the data deliberately left behind after range movement.

Authoritative references

Lifecycle commands are version-sensitive and state-sensitive. Re-check these sources, the release notes and your deployment's orchestration/managed-service rules before applying a topology procedure.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.