Chapter 23 · Node Lifecycle: Bootstrap, Replace, Decommission, Remove, Rebuild, and Cleanup
Decommission a Healthy Node and Stream Its Ranges Safely
Remove a healthy Cassandra member with decommission, monitor outgoing range streaming, preserve RF/failure-domain safety and destroy the old instance only after acceptance.
Learning outcomes
AtlasMart must retire a healthy host for hardware replacement.
Killing it and running removenode would throw away
the strongest source of its data. Because the node is alive, the
correct workflow is decommission: change ownership while
allowing that node to stream its ranges directly to future
replicas.
Distinguish decommission of a healthy member from removenode of an unreachable member.
Explain outgoing range streaming, failure-domain sequencing and why the force flag can violate RF expectations.
Capture before/during/after ownership, netstats, application latency, schema and replica coverage evidence.
Explain why decommission does not erase the old node data directory automatically.
Define a safe removal acceptance gate before deleting the container/volume.
Topology changes are intentionally isolated from the shared
course cluster. The mandatory labs use a disposable Docker
network atlasmart-cassandra-life, cluster
atlasmart-lifecycle, nodes
atlasmart-life-1..4 (plus explicitly named
replacement/cross-DC nodes where a lesson needs them), pinned
Docker Official Image cassandra:5.0.9, Java 17
inside the image, dc1, racks
rack1..rack3, and 16 virtual nodes (vnodes) per
node. The keyspace atlasmart_lifecycle uses
NetworkTopologyStrategy and replication factor
(RF) 3 in dc1; normal verification uses
LOCAL_QUORUM. New tables explicitly use
UnifiedCompactionStrategy (UCS), no default time-to-live
(TTL), and Cassandra's normal gc_grace_seconds.
Authentication, client/internode Transport Layer Security
(TLS), and remote Java Management Extensions (JMX) are
disabled only on this isolated single-host learning network.
The labs never expose JMX or native transport to an untrusted
network. Exact IP addresses, host IDs, tokens, streaming
sources, bytes, duration, ownership percentages,
failure-detection timing, disk usage and p95/p99 application
latency are learner-captured evidence. Docker examples are
Bash-compatible; Windows Docker Desktop users can run the same
docker commands individually in PowerShell if a
shell loop is inconvenient.
Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.
Terms and lifecycle mental model
A Cassandra node is one member of a peer-to-peer cluster. A datacenter (DC) and rack are logical topology labels used by replication placement to represent failure domains. A partition key is hashed by the partitioner to a token; with virtual nodes (vnodes), one physical node owns many token positions. A replica stores a copy of a token range according to the keyspace replication strategy. A coordinator is whichever node handles one request, not a permanent leader. A consistency level (CL) defines the replica acknowledgments/responses required for an operation.
A topology transition changes cluster membership or token ownership. Bootstrap is the join process in which a new node receives ranges by streaming from current replicas before becoming a normal owner. Pending ranges are ownership changes that are being prepared during a transition: Cassandra must keep writes safe while the future replica set is not yet fully ready. Streaming transfers SSTable data over the internode network. Decommission removes a healthy node and streams its ranges away. Replacement gives the logical ownership of a dead member to a fresh node identified by the dead endpoint. removenode removes an unreachable member from another live node and re-replicates from survivors. rebuild repopulates data on the current node by streaming from selected source nodes/DCs without changing that node's ownership. cleanup is a compaction-style rewrite that discards ranges a node no longer owns after a completed range movement. Repair is anti-entropy reconciliation between replicas; bootstrap/rebuild/cleanup do not make pre-existing replica inconsistencies disappear.
Node identity includes endpoint/broadcast address, host ID and token ownership recorded by cluster metadata. Headroom is spare disk, network, CPU, memory, compaction and replica availability capacity required to perform the transition without violating service objectives. A rollback/abort path is the documented action for a stalled or failed lifecycle sequence; it must be chosen from the current Cassandra version's supported topology state, not improvised by deleting system tables or reusing data directories.
1. Healthy departure uses the departing node as a source
nodetool decommission is executed against the node
being removed. Cassandra prepares a topology transition,
identifies the replicas that must inherit its ranges, and
streams from the departing node while cluster metadata moves
ownership. Current Cassandra protects RF by default;
--force exists to decommission even when doing so
would reduce replicas below configured RF and therefore belongs
only in an explicitly justified emergency procedure.
| Action | Node condition | Where data comes from | Typical operator risk |
|---|---|---|---|
| decommission | alive and manageable | departing node streams its ranges | removing too many failure-domain peers / insufficient destination headroom |
| removenode | dead/unreachable | remaining replicas re-replicate | insufficient surviving sources / wrong host ID |
| docker rm/VM delete only | any | nowhere | membership remains and replica coverage is wrong |
| decommission --force | alive | departing node, but RF safety can be bypassed | availability/durability below design |
2. Create a healthy four-node starting point
# These names are dedicated to Chapter 23; the reset block never targets the shared course cluster.docker network inspect atlasmart-cassandra-life >/dev/null 2>&1 || docker network create atlasmart-cassandra-lifedocker volume create atlasmart-life-1-datadocker volume create atlasmart-life-2-datadocker volume create atlasmart-life-3-datadocker run -d --name atlasmart-life-1 --hostname atlasmart-life-1 --network atlasmart-cassandra-life \ -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -v atlasmart-life-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-life-1 nodetool statusdocker run -d --name atlasmart-life-2 --hostname atlasmart-life-2 --network atlasmart-cassandra-life \ -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-life-3 --hostname atlasmart-life-3 --network atlasmart-cassandra-life \ -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all three appear UN from at least two observers.docker exec atlasmart-life-1 nodetool statusdocker exec atlasmart-life-2 nodetool status
CREATE KEYSPACE IF NOT EXISTS atlasmart_lifecycleWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_lifecycle.order_probe ( order_id text PRIMARY KEY, customer_id text, status text, total decimal, updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1001','cust-42','PAID',129.90,'2026-09-08T06:00:00Z');INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1002','cust-77','PACKING',89.50,'2026-09-08T06:01:00Z');INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1003','cust-42','SHIPPED',42.00,'2026-09-08T06:02:00Z');SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';
docker volume create atlasmart-life-4-datadocker run -d --name atlasmart-life-4 --hostname atlasmart-life-4 --network atlasmart-cassandra-life \ -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \ -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \ -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-4-data:/var/lib/cassandra cassandra:5.0.9# Poll from separate terminals while the node is joining; a tiny dataset may finish too quickly to catch UJ.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-4 nodetool netstats -Hdocker exec atlasmart-life-4 nodetool bootstrap# Final acceptance requires node 4 to be UN and schema versions to agree.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-1 nodetool describecluster
docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-4 nodetool statusbinarydocker exec atlasmart-life-4 df -h /var/lib/cassandradocker exec atlasmart-life-1 nodetool getendpoints atlasmart_lifecycle order_probe ord-1001docker exec atlasmart-life-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; TRACING ON; SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001'; TRACING OFF;"docker stats --no-stream atlasmart-life-1 atlasmart-life-2 atlasmart-life-3 atlasmart-life-4
Before removing a production node, also verify repair age, pending compactions, backups and that no simultaneous rack/host maintenance has consumed the redundancy budget. With RF=3, four nodes provide enough membership to decommission one while retaining three replicas, but the exact ranges and destinations still depend on tokens and racks.
3. Decommission and watch the handoff
docker exec atlasmart-life-4 nodetool decommission
docker exec atlasmart-life-4 nodetool netstats -Hdocker exec atlasmart-life-1 nodetool netstats -Hdocker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker stats --no-stream atlasmart-life-1 atlasmart-life-2 atlasmart-life-3 atlasmart-life-4
Depending on the dataset, the transition may be too fast to see intermediate states. The durable evidence is the before/after token membership plus logs/streaming records. After completion, node 4 should no longer be a normal ring member. Cassandra's topology documentation notes that data is not automatically removed from the decommissioned node; that is a deliberate safety property.
docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-1 nodetool describeclusterdocker exec atlasmart-life-1 nodetool checktokenmetadatadocker exec atlasmart-life-1 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';"# Inspect the decommissioned volume before deciding to destroy it.docker exec atlasmart-life-4 sh -lc 'du -sh /var/lib/cassandra/data 2>/dev/null || true'
4. Broken approaches and rollback thinking
Topology transitions consume redundancy and streaming
capacity. A second failure/decommission can leave insufficient
sources or violate the availability assumed by
LOCAL_QUORUM. Sequence one failure domain at a
time, wait for stream/compaction/application stabilization,
and re-run acceptance gates before the next change.
Once a decommission transition has committed and completed,
“rollback” is not “restart the old process with its old data.”
That data directory contains a former membership identity.
Returning capacity is a new bootstrap with a clean identity/data
path unless a version-specific documented procedure says
otherwise. If a decommission is still in progress or fails,
inspect current topology sequence/logs and current-version
tooling; do not switch to removenode force just to
make metadata disappear.
5. Safe destruction and reset
# Only after node 4 is absent from ring membership and reads/coverage are accepted:docker rm -f atlasmart-life-4docker volume rm atlasmart-life-4-data# Verify the surviving topology again.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-2 nodetool status atlasmart_lifecycle
Check your understanding
- Why prefer decommission over removenode for a healthy member?
- Why is --force dangerous?
- Does decommission wipe the retired node data directory?
- What should gate a second node removal?
- Can a decommissioned data volume simply be restarted as the old member?
Review the answers
1. The departing node can stream the ranges it owns directly, preserving the strongest available source and a coordinated ownership transition.
2. It can allow a decommission even when the resulting topology would have fewer replicas than configured RF.
3. No. Cassandra intentionally leaves it; destroy/reimage it only after the topology change is accepted.
4. Stabilized topology, replica coverage, streaming/compaction/resource metrics, application SLOs and restored failure-domain headroom.
5. No. Treat re-entry as a documented new lifecycle operation with clean identity unless the exact Cassandra version provides a specific supported recovery path.
docker rm -f atlasmart-life-1 atlasmart-life-2 atlasmart-life-3 atlasmart-life-4 atlasmart-life-4r atlasmart-life-dc2-1 2>/dev/null || truedocker volume rm atlasmart-life-1-data atlasmart-life-2-data atlasmart-life-3-data atlasmart-life-4-data atlasmart-life-4r-data atlasmart-life-dc2-1-data 2>/dev/null || truedocker network rm atlasmart-cassandra-life 2>/dev/null || true
Production judgment
Topology work consumes the same resources that serve customer traffic. Before a bootstrap, decommission, replacement, removal or rebuild, record Cassandra/JDK/driver versions, DC/rack layout, RF and query CLs, vnode count, per-node used/free disk, compaction backlog, streaming throughput, NIC saturation, JVM/GC pressure, repair age, hint window, backup status, SAI/vector indexes, tenant/security constraints, driver timeouts/retries/idempotency and representative p50/p95/p99 latency. Confirm that losing one additional host/rack during the operation still leaves the required replicas and operational headroom. Avoid concurrent range movements in the same failure domain unless the current version and runbook explicitly prove safety.
Streaming moves bytes; it does not cure bad partition keys, oversized partitions, old replica divergence, poor compaction headroom or missing repair. Cleanup reclaims no-longer-owned ranges; it is not anti-entropy. Replacement preserves ownership identity but still requires post-replacement convergence when downtime/write-forwarding exceeds the mechanisms that could cover missed writes. Managed Cassandra services may hide or forbid these commands; use the provider's documented replacement/scale workflow but keep the same evidence and acceptance model. Lesson 3 changes the failure premise: the old member is dead. You will preserve its logical ownership with a clean replacement node and verify endpoint/host identity before and after streaming.
Summary and next bridge
Decommission is an orderly departure: the live node streams its ranges, membership changes, and the physical instance is destroyed only after acceptance. A dead host cannot do that. The next lesson uses replacement identity to recover a failed member without treating a fresh empty process as an unrelated bootstrap.
Authoritative references
Lifecycle commands are version-sensitive and state-sensitive. Re-check these sources, the release notes and your deployment's orchestration/managed-service rules before applying a topology procedure.
- Apache Cassandra downloads / current 5.0 patch
- Adding, replacing, moving and removing nodes
- nodetool command reference
- nodetool bootstrap
- nodetool decommission
- nodetool removenode
- nodetool rebuild
- nodetool cleanup
- nodetool netstats
- cassandra-env JVM/system properties including replacement
- Compaction overview / cleanup semantics
- Docker Official Cassandra 5.0 packaging