Chapter 23 · Node Lifecycle: Bootstrap, Replace, Decommission, Remove, Rebuild, and Cleanup

Add Capacity: Bootstrap Streaming, Token Ownership, Pending Ranges, and Headroom

Add a fourth Cassandra node as an observable range movement: prove headroom, pending/join state, streaming, ownership changes and post-bootstrap obligations.

Intermediate → Advanced125–170 minutesBootstrap + headroom labApache Cassandra 5.0.9 · Java 17 · cqlsh/nodetool · isolated RF=3 lifecycle topology · UCSLast reviewed: September 2026

Learning outcomes

AtlasMart's order workload has outgrown three nodes. Adding a fourth server looks easy—start Cassandra and wait for UN—but that shorthand hides the actual risk: ownership changes while real writes continue, existing replicas must stream ranges, compaction/network load rises, and old owners retain files until cleanup. This lesson treats capacity expansion as an observable topology transition.

01

Explain bootstrap, future/pending ownership, streaming sources, vnodes and why a joining node is not immediately a normal replica.

02

Estimate disk/network/failure headroom before adding capacity and identify why “more disk on the new node” is insufficient.

03

Observe bootstrap using status, netstats, bootstrap state, schema agreement, endpoint placement and per-node resource evidence.

04

Distinguish bootstrap completion from repair/convergence and cleanup obligations.

05

Use supported resume/reset guidance rather than deleting system metadata when a bootstrap stalls.

Chapter 23 lifecycle lab baseline

Topology changes are intentionally isolated from the shared course cluster. The mandatory labs use a disposable Docker network atlasmart-cassandra-life, cluster atlasmart-lifecycle, nodes atlasmart-life-1..4 (plus explicitly named replacement/cross-DC nodes where a lesson needs them), pinned Docker Official Image cassandra:5.0.9, Java 17 inside the image, dc1, racks rack1..rack3, and 16 virtual nodes (vnodes) per node. The keyspace atlasmart_lifecycle uses NetworkTopologyStrategy and replication factor (RF) 3 in dc1; normal verification uses LOCAL_QUORUM. New tables explicitly use UnifiedCompactionStrategy (UCS), no default time-to-live (TTL), and Cassandra's normal gc_grace_seconds. Authentication, client/internode Transport Layer Security (TLS), and remote Java Management Extensions (JMX) are disabled only on this isolated single-host learning network. The labs never expose JMX or native transport to an untrusted network. Exact IP addresses, host IDs, tokens, streaming sources, bytes, duration, ownership percentages, failure-detection timing, disk usage and p95/p99 application latency are learner-captured evidence. Docker examples are Bash-compatible; Windows Docker Desktop users can run the same docker commands individually in PowerShell if a shell loop is inconvenient.

Execution and safety note

Run commands only against the disposable Apache Cassandra course lab or another explicitly approved non-production environment. Confirm node, keyspace, table, container, volume, path, and datacenter targets before destructive, failure-injection, cleanup, repair, restore, security, or topology operations. Capture current state and expected rollback/recovery evidence first; output and timings can differ by host, operating system, Java runtime, Docker/runtime, driver, and Cassandra configuration.

Terms and lifecycle mental model

A Cassandra node is one member of a peer-to-peer cluster. A datacenter (DC) and rack are logical topology labels used by replication placement to represent failure domains. A partition key is hashed by the partitioner to a token; with virtual nodes (vnodes), one physical node owns many token positions. A replica stores a copy of a token range according to the keyspace replication strategy. A coordinator is whichever node handles one request, not a permanent leader. A consistency level (CL) defines the replica acknowledgments/responses required for an operation.

A topology transition changes cluster membership or token ownership. Bootstrap is the join process in which a new node receives ranges by streaming from current replicas before becoming a normal owner. Pending ranges are ownership changes that are being prepared during a transition: Cassandra must keep writes safe while the future replica set is not yet fully ready. Streaming transfers SSTable data over the internode network. Decommission removes a healthy node and streams its ranges away. Replacement gives the logical ownership of a dead member to a fresh node identified by the dead endpoint. removenode removes an unreachable member from another live node and re-replicates from survivors. rebuild repopulates data on the current node by streaming from selected source nodes/DCs without changing that node's ownership. cleanup is a compaction-style rewrite that discards ranges a node no longer owns after a completed range movement. Repair is anti-entropy reconciliation between replicas; bootstrap/rebuild/cleanup do not make pre-existing replica inconsistencies disappear.

Node identity includes endpoint/broadcast address, host ID and token ownership recorded by cluster metadata. Headroom is spare disk, network, CPU, memory, compaction and replica availability capacity required to perform the transition without violating service objectives. A rollback/abort path is the documented action for a stalled or failed lifecycle sequence; it must be chosen from the current Cassandra version's supported topology state, not improvised by deleting system tables or reusing data directories.

1. What changes when node 4 joins?

With Murmur3Partitioner and 16 vnodes, the joining node obtains multiple token positions. For each keyspace, NetworkTopologyStrategy derives a replica set from token ownership and rack/DC topology. During bootstrap Cassandra calculates ranges that the joining node will own, selects existing replicas as streaming sources, and keeps the transition pending until required data is available. Reads avoid treating the joining node as a fully ready source for data it has not received, while writes must account for the pending replica state so post-transition reads are safe.

State/evidence Interpretation Do not infer
UJ / joining in status membership transition is in progress all assigned ranges are fully present
netstats streams files/ranges are moving between endpoints all pre-existing replica divergence is repaired
UN after bootstrap node completed join and is normal old owners already removed no-longer-owned files
schema agreement nodes agree on schema version data replicas are byte-for-byte converged
lower ownership/load per old node distribution changed hot-key or partition-model problems are fixed

2. Preflight: prove headroom before movement

Docker · create the isolated three-node RF=3 starting topology
# These names are dedicated to Chapter 23; the reset block never targets the shared course cluster.docker network inspect atlasmart-cassandra-life >/dev/null 2>&1 || docker network create atlasmart-cassandra-lifedocker volume create atlasmart-life-1-datadocker volume create atlasmart-life-2-datadocker volume create atlasmart-life-3-datadocker run -d --name atlasmart-life-1 --hostname atlasmart-life-1 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -v atlasmart-life-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 is UN before starting peers.docker exec atlasmart-life-1 nodetool statusdocker run -d --name atlasmart-life-2 --hostname atlasmart-life-2 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-life-3 --hostname atlasmart-life-3 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-3-data:/var/lib/cassandra cassandra:5.0.9# Continue only after all three appear UN from at least two observers.docker exec atlasmart-life-1 nodetool statusdocker exec atlasmart-life-2 nodetool status
CQL · create a small deterministic AtlasMart ownership fixture
CREATE KEYSPACE IF NOT EXISTS atlasmart_lifecycleWITH replication = {'class':'NetworkTopologyStrategy','dc1':3};CREATE TABLE IF NOT EXISTS atlasmart_lifecycle.order_probe (  order_id text PRIMARY KEY,  customer_id text,  status text,  total decimal,  updated_at timestamp) WITH compaction = {'class':'UnifiedCompactionStrategy'};CONSISTENCY LOCAL_QUORUM;INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1001','cust-42','PAID',129.90,'2026-09-08T06:00:00Z');INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1002','cust-77','PACKING',89.50,'2026-09-08T06:01:00Z');INSERT INTO atlasmart_lifecycle.order_probe (order_id,customer_id,status,total,updated_at)VALUES ('ord-1003','cust-42','SHIPPED',42.00,'2026-09-08T06:02:00Z');SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';
Docker/nodetool · capture the before-state and headroom
docker exec atlasmart-life-1 nodetool version -vdocker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-1 nodetool describering atlasmart_lifecycledocker exec atlasmart-life-1 nodetool getendpoints atlasmart_lifecycle order_probe ord-1001docker exec atlasmart-life-1 nodetool compactionstatsdocker exec atlasmart-life-2 nodetool compactionstatsdocker exec atlasmart-life-3 nodetool compactionstatsdocker exec atlasmart-life-1 df -h /var/lib/cassandradocker exec atlasmart-life-2 df -h /var/lib/cassandradocker exec atlasmart-life-3 df -h /var/lib/cassandradocker stats --no-stream atlasmart-life-1 atlasmart-life-2 atlasmart-life-3

For a longer observable stream, an optional synthetic dataset may be created with the bundled Cassandra stress tool; it is deliberately separate from the AtlasMart schema and may need substantial disk/time. The mandatory correctness fixture remains the small AtlasMart table. In production, use real table size distributions and expected outgoing/incoming stream bytes rather than a synthetic row count.

Optional local-only load · make bootstrap long enough to observe
# Optional: synthetic data only; skip on resource-constrained machines.docker exec atlasmart-life-1 /opt/cassandra/tools/bin/cassandra-stress write n=50000 -rate threads=8 -node atlasmart-life-1# Recheck free disk and compaction backlog before adding the node.docker exec atlasmart-life-1 df -h /var/lib/cassandradocker exec atlasmart-life-1 nodetool compactionstats

3. Bootstrap node 4 and capture pending/streaming evidence

Docker · bootstrap node 4 into dc1/rack1
docker volume create atlasmart-life-4-datadocker run -d --name atlasmart-life-4 --hostname atlasmart-life-4 --network atlasmart-cassandra-life \  -e CASSANDRA_CLUSTER_NAME=atlasmart-lifecycle -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 \  -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 \  -e CASSANDRA_SEEDS=atlasmart-life-1 -v atlasmart-life-4-data:/var/lib/cassandra cassandra:5.0.9# Poll from separate terminals while the node is joining; a tiny dataset may finish too quickly to catch UJ.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-4 nodetool netstats -Hdocker exec atlasmart-life-4 nodetool bootstrap# Final acceptance requires node 4 to be UN and schema versions to agree.docker exec atlasmart-life-1 nodetool status atlasmart_lifecycledocker exec atlasmart-life-1 nodetool describecluster

On a tiny lab, the transition may finish before a polling command catches UJ. That is not a failure of the mechanism; it means the dataset is too small for a long observation window. Capture container logs and netstats while the node starts. Compare endpoint placement before and after: a given partition may or may not move to node 4 because token ownership is hash-dependent.

Docker/nodetool · post-bootstrap ownership and service evidence
docker exec atlasmart-life-1 nodetool checktokenmetadatadocker exec atlasmart-life-1 nodetool getendpoints atlasmart_lifecycle order_probe ord-1001docker exec atlasmart-life-4 cqlsh -e "CONSISTENCY LOCAL_QUORUM; SELECT * FROM atlasmart_lifecycle.order_probe WHERE order_id='ord-1001';"docker exec atlasmart-life-4 df -h /var/lib/cassandradocker stats --no-stream atlasmart-life-1 atlasmart-life-2 atlasmart-life-3 atlasmart-life-4

4. Broken approach: add capacity during a headroom emergency

Failure pattern: “The old disks are nearly full, so bootstrap immediately.”

The new node must receive data, but the old owners simultaneously read/stream SSTables, continue foreground reads/writes, compact, and keep obsolete range data until cleanup. An almost-full source can run out of disk or I/O headroom during the very operation intended to save it. First reduce immediate risk, quantify per-node free space and compaction debt, lower competing work if supported by the runbook, then expand with an explicit abort/rollback threshold.

A stalled bootstrap should be diagnosed from logs and streaming state. Cassandra 5.0 exposes nodetool bootstrap resume; documentation also describes resetting bootstrap progress with -Dcassandra.reset_bootstrap_progress=true when intentionally starting fresh. The --force resume option is explicitly dangerous. Never “fix” a joining node by copying another node's system tables or reusing an unrelated data volume.

Evidence only · inspect bootstrap recovery surfaces before acting
docker exec atlasmart-life-4 nodetool bootstrapdocker exec atlasmart-life-4 nodetool netstats -Hdocker logs --tail 200 atlasmart-life-4# If the documented state calls for resume:# docker exec atlasmart-life-4 nodetool bootstrap resume# Do NOT add --force merely to silence a safety check.

5. Acceptance, verification and reset

  • All four intended nodes are UN from more than one observer.
  • Schema versions agree and checktokenmetadata reports no unexpected mismatch.
  • Representative RF=3/LOCAL_QUORUM reads succeed.
  • No stream remains unexpectedly active; compaction backlog and disk/network utilization return toward baseline.
  • Old owners are marked for cleanup only after the new node is stable; cleanup itself is Lesson 5.
  • A repair plan remains intact because bootstrap is not a substitute for anti-entropy.

Check your understanding

  1. Why is a joining node not immediately equivalent to an existing replica?
  2. What does UN prove after bootstrap?
  3. Why can adding a node increase short-term disk and network pressure on old nodes?
  4. Does bootstrap repair old inconsistencies between the source replicas?
  5. What is the safer response to a stalled bootstrap?
Review the answers

1. Its future token ranges must be streamed and made ready before the topology transition can safely expose it as a normal owner.

2. Membership/join completion, not universal replica convergence, cleanup completion or good application latency.

3. They must read and stream SSTables while continuing foreground traffic and retaining obsolete ranges until cleanup.

4. No. Streaming obtains data from selected sources; scheduled/targeted repair remains the anti-entropy mechanism.

5. Inspect state/logs/streams and use version-documented resume or reset procedures; do not delete metadata or force past safety checks without understanding the state.

Docker · reset only the dedicated Chapter 23 lifecycle lab
docker rm -f atlasmart-life-1 atlasmart-life-2 atlasmart-life-3 atlasmart-life-4 atlasmart-life-4r atlasmart-life-dc2-1 2>/dev/null || truedocker volume rm atlasmart-life-1-data atlasmart-life-2-data atlasmart-life-3-data atlasmart-life-4-data atlasmart-life-4r-data atlasmart-life-dc2-1-data 2>/dev/null || truedocker network rm atlasmart-cassandra-life 2>/dev/null || true

Production judgment

Topology work consumes the same resources that serve customer traffic. Before a bootstrap, decommission, replacement, removal or rebuild, record Cassandra/JDK/driver versions, DC/rack layout, RF and query CLs, vnode count, per-node used/free disk, compaction backlog, streaming throughput, NIC saturation, JVM/GC pressure, repair age, hint window, backup status, SAI/vector indexes, tenant/security constraints, driver timeouts/retries/idempotency and representative p50/p95/p99 latency. Confirm that losing one additional host/rack during the operation still leaves the required replicas and operational headroom. Avoid concurrent range movements in the same failure domain unless the current version and runbook explicitly prove safety.

Streaming moves bytes; it does not cure bad partition keys, oversized partitions, old replica divergence, poor compaction headroom or missing repair. Cleanup reclaims no-longer-owned ranges; it is not anti-entropy. Replacement preserves ownership identity but still requires post-replacement convergence when downtime/write-forwarding exceeds the mechanisms that could cover missed writes. Managed Cassandra services may hide or forbid these commands; use the provider's documented replacement/scale workflow but keep the same evidence and acceptance model. Lesson 2 starts from a healthy four-node topology and removes one member deliberately, so you can contrast bootstrap’s incoming streams with decommission’s outgoing ownership handoff.

Summary and next bridge

Adding capacity is a range movement with pending ownership, source selection and temporary resource amplification. A node reaching UN is necessary but not sufficient evidence of a completed operational change. Next, remove a healthy node with decommission and prove that its ranges are safely transferred before destroying the old instance.

Authoritative references

Lifecycle commands are version-sensitive and state-sensitive. Re-check these sources, the release notes and your deployment's orchestration/managed-service rules before applying a topology procedure.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.