Chapter 03 · Partitioners, Tokens, Vnodes, Token Rings, and Data Distribution

Inspect Token Ownership and Distribution with nodetool and System Tables

Close the chapter with an evidence-backed before/after distribution audit instead of trusting a ring picture or one ownership percentage.

Intermediate125–150 minutesToken-distribution and node-bootstrap auditApache Cassandra 5.0.9 · Murmur3 · 16 vnodes/nodeLast reviewed: September 2026

Learning outcomes

The chapter ends by combining token metadata, keyspace replication, vnode ownership, request-key evidence, and streaming into one operator workflow. The goal is not to memorize nodetool columns; it is to know which command or system table answers each placement question and how to avoid false conclusions.

01

Inspect local/peer token collections, nodetool status/ring output, and concrete token() values as complementary evidence.

02

Use keyspace-aware ownership and SHOW REPLICAS to explain where AtlasMart partitions are stored.

03

Capture before/during/after evidence for adding a disposable fourth node and its streaming activity.

04

Explain why old data can remain on former owners until cleanup after range movement.

05

Produce a concise token-distribution acceptance report that separates measured facts from assumptions.

Chapter baseline reviewed 7 September 2026

Apache Cassandra 5.0.9 is the current GA 5.0 patch on the official download page. The labs pin cassandra:5.0.9 and explicitly set CASSANDRA_NUM_TOKENS=16 so vnode behavior is reproducible instead of inheriting an unnoticed image/configuration default. Current Cassandra 5.0 documentation uses num_tokens: 16 as the modern baseline; older Cassandra material often mentions 256 random vnodes, so this chapter treats token count as a version- and deployment-sensitive design choice rather than folklore.

Execution and evidence note

The generation environment does not contain Docker or Cassandra, so Cassandra commands were checked against current official documentation but were not executed here. Exact token values, IP addresses, host IDs, ownership percentages, load, stream sizes, latency, and hot-partition samples must be captured on the learner's machine. Expected output is described as invariants or shapes, never presented as measured output.

1. Build an evidence matrix before touching topology

bash / PowerShell · reproducible 3-node, 16-vnode lab
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait for all nodes to become Up/Normal.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool version
AtlasMart fixture · RF=3 products_by_id
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_tokens WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_tokens.products_by_id (product_id text PRIMARY KEY, category text, name text, price_cents int);"docker exec atlasmart-cass-1 cqlsh -e "INSERT INTO atlasmart_tokens.products_by_id (product_id,category,name,price_cents) VALUES ('p-1001','laptop','AtlasBook 14',129900); INSERT INTO atlasmart_tokens.products_by_id (product_id,category,name,price_cents) VALUES ('p-1002','camera','AtlasCam X',89900); INSERT INTO atlasmart_tokens.products_by_id (product_id,category,name,price_cents) VALUES ('p-1003','audio','AtlasPods',14900); INSERT INTO atlasmart_tokens.products_by_id (product_id,category,name,price_cents) VALUES ('p-2001','home','AtlasLamp',7900);"

Each source answers a different question. Capture all of them before adding a node so the post-bootstrap comparison has a baseline.

baseline report · token map, peers, key tokens, network state
docker exec atlasmart-cass-1 nodetool status atlasmart_tokensdocker exec atlasmart-cass-1 sh -lc "nodetool ring atlasmart_tokens > /tmp/ring-before.txt"docker exec atlasmart-cass-1 cqlsh -e "SELECT host_id, partitioner, tokens FROM system.local;"docker exec atlasmart-cass-1 cqlsh -e "SELECT peer,peer_port,host_id,data_center,rack,tokens FROM system.peers_v2;"docker exec atlasmart-cass-1 cqlsh -e "SELECT product_id,token(product_id) AS token_value FROM atlasmart_tokens.products_by_id;"docker exec atlasmart-cass-1 nodetool netstats -H
Evidence source Best use Important caveat
system.local.tokens Local vnode positions Not keyspace replica ownership
system.peers_v2.tokens Peer token metadata visible locally Observer's metadata snapshot; not a traffic metric
nodetool status <ks> Node state, load, token count, effective ownership Load is not request rate
nodetool ring <ks> Verbose token-range map Very noisy with vnodes
token(pk) Map a business partition key to a token Still needs replication metadata for replicas
SHOW REPLICAS Resolve a token to keyspace replicas Keyspace context matters
nodetool netstats Streaming/network activity Tiny fixtures may show little data

2. Build a deterministic partition-placement report

For each sample product, capture its token and replica set. This is stronger evidence than saying “RF=3 so it must be everywhere,” because it proves the actual metadata path on the running cluster.

acceptance evidence · product → token → replica set
# 1) Capture the business-key/token pairs.docker exec atlasmart-cass-1 cqlsh -e "SELECT product_id,token(product_id) AS token_value FROM atlasmart_tokens.products_by_id;"# 2) For each captured token, query replicas in an interactive cqlsh session.docker exec -it atlasmart-cass-1 cqlsh# SHOW REPLICAS <captured_token> atlasmart_tokens;# 3) Cross-check effective ownership and racks.docker exec atlasmart-cass-1 nodetool status atlasmart_tokens

Record host IDs/racks rather than relying only on ephemeral container IP addresses. In production, addresses can change while host identity and topology metadata are the more durable correlation keys.

Interpret ownership with RF.

At RF=3 on three nodes, every sample should normally resolve to all three nodes, and effective ownership can look like 100% on each. That is expected replication, not a failure of hashing.

3. Add node 4 and observe the ownership transition

Adding capacity is called bootstrap. The joining node receives token positions and streams the data needed for the ranges it will replicate. Define the blast radius first: this lab changes only disposable containers/volumes on atlasmart-cassandra; it does not remove production nodes, alter host networking, or modify an unrelated keyspace.

topology change · bootstrap node 4 and capture streaming
docker volume create atlasmart-cass-4-datadocker run -d --name atlasmart-cass-4 --hostname atlasmart-cass-4 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-4-data:/var/lib/cassandra cassandra:5.0.9# Sample repeatedly while bootstrap is active.docker exec atlasmart-cass-4 nodetool netstats -Hdocker exec atlasmart-cass-1 nodetool status atlasmart_tokensdocker exec atlasmart-cass-4 nodetool status atlasmart_tokens# After node 4 is Up/Normal, capture post-state.docker exec atlasmart-cass-1 cqlsh -e "SELECT peer,peer_port,host_id,data_center,rack,tokens FROM system.peers_v2;"docker exec atlasmart-cass-1 sh -lc "nodetool ring atlasmart_tokens > /tmp/ring-after.txt"

Do not assume the tiny dataset keeps the bootstrap state visible long enough for a single sample. On a fast laptop it may finish before netstats is run. Record that honestly and rely on before/after token metadata if the stream window is too short; for performance work, load a controlled synthetic dataset and measure the stream.

4. Compare before/after without confusing disk leftovers with ownership

After node 4 joins, some existing nodes lose token ranges. Cassandra's topology-change guidance notes that old range data is not automatically removed from nodes that lost ownership; operators run nodetool cleanup after they are satisfied the new node is healthy. This safety behavior means raw disk load can remain temporarily higher than the new ownership map implies.

post-bootstrap · ownership versus obsolete local range data
# Compare current ownership to the saved ring files.docker exec atlasmart-cass-1 nodetool status atlasmart_tokensdocker exec atlasmart-cass-1 sh -lc "diff -u /tmp/ring-before.txt /tmp/ring-after.txt | head -120 || true"# Learn the cleanup scope first. Do not run blindly on unrelated clusters.docker exec atlasmart-cass-1 nodetool help cleanup# In this disposable lab, cleanup can be applied to the course keyspace after verification:docker exec atlasmart-cass-1 nodetool cleanup atlasmart_tokens

cleanup is not a rebalance command and does not “fix the ring.” It removes data that the node no longer owns after a range movement. Run it only after validating the topology change and with disk/compaction capacity awareness.

5. Chapter acceptance report and production judgment

A useful operator report names the exact Cassandra patch, partitioner, vnode policy, node/DC/rack map, keyspace RF, evidence time, sample partition-key tokens, replica sets, ownership/load before and after topology change, streaming evidence, and whether cleanup has run. It also states what was not measured: production latency, large-data stream throughput, repair duration, and real AZ failure tolerance.

Do not use manual token assignment or repeated scale-out as a substitute for schema review. If token ownership is balanced but toppartitions shows one key dominating traffic, Chapter 03's conclusion is a data-model action. If request distribution is healthy but ownership/load are uneven, investigate token allocation, hardware asymmetry, data-size distribution, compaction, and topology state separately.

Verification checklist

  • Exactly 16 vnode positions per course node were explicitly configured and observed.
  • You captured system/local peer token metadata and keyspace-aware nodetool output.
  • You mapped sample product keys to tokens and replicas.
  • You captured before/during/after evidence for the fourth-node bootstrap.
  • You can explain why disk bytes can remain on former owners until cleanup.
  • You did not present ownership percentages as request-distribution measurements.

Check your understanding

  1. Why inspect both system token tables and nodetool?
  2. Why can disk usage lag behind token ownership after bootstrap?
  3. What should you record if bootstrap completes before netstats catches it?
  4. Does balanced ownership prove balanced requests?
  5. What is the next design dependency after token placement?
Review the answers

1. They provide complementary queryable metadata and operator views; cross-checking reduces the chance of misreading one presentation.

2. Cassandra keeps data on nodes that lost ranges as a safety measure until cleanup removes no-longer-owned data.

3. State that no live streaming sample was captured and use before/after token/ownership evidence; never invent stream numbers.

4. No. Request distribution depends on partition-key frequency and application traffic.

5. Keyspace replication strategy and replication factor, which Chapter 04 teaches explicitly.

Cleanup

cleanup · Bash; remove only course-owned resources
docker rm -f atlasmart-cass-4 atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>/dev/null || truedocker volume rm atlasmart-cass-4-data atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra 2>/dev/null || true
cleanup · PowerShell; remove only course-owned resources
docker rm -f atlasmart-cass-4 atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>$nulldocker volume rm atlasmart-cass-4-data atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>$nulldocker network rm atlasmart-cassandra 2>$null

Summary and next step

This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.

Next, continue to CREATE KEYSPACE and Replication Configuration: Durable Writes and Namespaces.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.