Chapter 03 · Partitioners, Tokens, Vnodes, Token Rings, and Data Distribution
Token Ranges, Ownership, Replicas, and Why the Ring Is a Useful Mental Model
Read the token ring as a logical ownership map, distinguish primary and effective ownership, and prove replica placement from concrete AtlasMart tokens.
Learning outcomes
Tokens become operationally useful only when they are treated as boundaries. Each owned token closes a token range; each physical node owns many such ranges when vnodes are enabled; each keyspace then replicates those ranges according to its own placement policy. This lesson turns the circular “ring” picture into a precise ownership model and explains when ownership percentages are meaningful.
Interpret a token range using the conventional (previous_token, token] boundary model and the wrap-around edge.
Distinguish primary token ownership from keyspace-specific effective replica ownership.
Use nodetool ring/status with and without a keyspace and explain why ownership context matters.
Use SHOW REPLICAS to connect a sampled token to a rack-aware replica set.
Explain why the ring is a logical hash-space model rather than the cluster network topology.
Apache Cassandra 5.0.9 is the current GA 5.0 patch on the
official download page. The labs pin
cassandra:5.0.9 and explicitly set
CASSANDRA_NUM_TOKENS=16 so vnode behavior is
reproducible instead of inheriting an unnoticed
image/configuration default. Current Cassandra 5.0
documentation uses num_tokens: 16 as the modern
baseline; older Cassandra material often mentions 256 random
vnodes, so this chapter treats token count as a version- and
deployment-sensitive design choice rather than folklore.
The generation environment does not contain Docker or Cassandra, so Cassandra commands were checked against current official documentation but were not executed here. Exact token values, IP addresses, host IDs, ownership percentages, load, stream sizes, latency, and hot-partition samples must be captured on the learner's machine. Expected output is described as invariants or shapes, never presented as measured output.
1. Tokens are points; ranges are the owned intervals between them
Sort all tokens in the cluster. A token at position
t2 conventionally closes the range
(t1, t2], where t1 is the previous
token around the ring. A partition whose hash falls inside that
interval has its primary token range associated
with the endpoint that owns t2. Because the hash
space wraps, the largest token connects back to the smallest;
the word “ring” is convenient for that wrap-around arithmetic.
With vnodes, a single physical node appears at many positions. Therefore a three-node cluster does not mean three large contiguous thirds. At 16 tokens per node the lab has roughly 48 vnode positions distributed around the logical hash space. The exact boundaries are discovered from token metadata.
Requests do not travel clockwise through every node. A coordinator can send a replica request directly to an endpoint. “Walking the ring” is a useful way to explain ownership/replica selection over sorted token positions, not packet routing.
2. Inspect primary token positions and keyspace-aware ownership
docker network create atlasmart-cassandradocker volume create atlasmart-cass-1-datadocker volume create atlasmart-cass-2-datadocker volume create atlasmart-cass-3-datadocker run -d --name atlasmart-cass-1 --hostname atlasmart-cass-1 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack1 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-1-data:/var/lib/cassandra cassandra:5.0.9# Wait until node 1 accepts CQL before starting peers.docker run -d --name atlasmart-cass-2 --hostname atlasmart-cass-2 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack2 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-2-data:/var/lib/cassandra cassandra:5.0.9docker run -d --name atlasmart-cass-3 --hostname atlasmart-cass-3 --network atlasmart-cassandra -e CASSANDRA_CLUSTER_NAME=atlasmart-course -e CASSANDRA_SEEDS=atlasmart-cass-1 -e CASSANDRA_DC=dc1 -e CASSANDRA_RACK=rack3 -e CASSANDRA_ENDPOINT_SNITCH=GossipingPropertyFileSnitch -e CASSANDRA_NUM_TOKENS=16 -v atlasmart-cass-3-data:/var/lib/cassandra cassandra:5.0.9# Wait for all nodes to become Up/Normal.docker exec atlasmart-cass-1 nodetool statusdocker exec atlasmart-cass-1 nodetool version
docker exec atlasmart-cass-1 cqlsh -e "CREATE KEYSPACE IF NOT EXISTS atlasmart_tokens WITH replication = {'class':'NetworkTopologyStrategy','dc1':3};"docker exec atlasmart-cass-1 cqlsh -e "CREATE TABLE IF NOT EXISTS atlasmart_tokens.products_by_id (product_id text PRIMARY KEY, category text, name text, price_cents int);"docker exec atlasmart-cass-1 cqlsh -e "INSERT INTO atlasmart_tokens.products_by_id (product_id,category,name,price_cents) VALUES ('p-1001','laptop','AtlasBook 14',129900); INSERT INTO atlasmart_tokens.products_by_id (product_id,category,name,price_cents) VALUES ('p-1002','camera','AtlasCam X',89900); INSERT INTO atlasmart_tokens.products_by_id (product_id,category,name,price_cents) VALUES ('p-1003','audio','AtlasPods',14900); INSERT INTO atlasmart_tokens.products_by_id (product_id,category,name,price_cents) VALUES ('p-2001','home','AtlasLamp',7900);"
docker exec atlasmart-cass-1 nodetool status# Add a keyspace argument when you need topology-aware effective ownership.docker exec atlasmart-cass-1 nodetool status atlasmart_tokensdocker exec atlasmart-cass-1 nodetool ring atlasmart_tokens# This is verbose with vnodes; save it for analysis rather than treating it as the default dashboard.docker exec atlasmart-cass-1 sh -lc "nodetool ring atlasmart_tokens > /tmp/ring-atlasmart.txt"docker exec atlasmart-cass-1 sh -lc "wc -l /tmp/ring-atlasmart.txt && head -40 /tmp/ring-atlasmart.txt"
Official nodetool ring documentation explicitly
accepts a keyspace “for accurate ownership information (topology
awareness).” That matters because physical token ownership and
effective replicated ownership are not the same percentage. In
this three-node RF=3 lab, every partition is replicated to all
three physical nodes, so effective keyspace ownership can be
100% even though each node is primary owner for only a fraction
of the token ranges.
| Metric/view | Question it answers | Common misread |
|---|---|---|
| Tokens per node | How many vnode positions map to this endpoint? | “More tokens means more replicas.” |
| Primary token ranges | Which ranges start their natural placement at this endpoint? | “This is the only copy.” |
| Effective ownership for keyspace | How much replicated keyspace data is attributed to this node? | “Percent must sum to 100.” With replication it may not. |
| Load | Current disk data reported by the node | “Load equals request traffic.” It does not. |
3. Resolve concrete products to token ranges and replicas
Sample tokens from AtlasMart, then ask cqlsh for replica placement. The goal is to connect a business key to a replica set without manually eyeballing dozens of vnode lines.
docker exec atlasmart-cass-1 cqlsh -e "SELECT product_id, token(product_id) AS token_value FROM atlasmart_tokens.products_by_id;"# Copy one token value, then open cqlsh:docker exec -it atlasmart-cass-1 cqlsh# Example form; substitute the actual bigint value:SHOW REPLICAS TOKEN_VALUE atlasmart_tokens;
Repeat for several product IDs. With RF=3 and one node per rack, all three lab endpoints should normally appear in each natural replica set. That does not make the token map irrelevant: token ownership still determines the primary range and would matter immediately at RF=1 or RF=2, during topology changes, for token-aware coordination, and for ownership/streaming decisions.
4. Why replication changes the meaning of “owns”
Suppose a node is the primary owner of approximately one third of the token space. At RF=3 in a three-node datacenter, it can still store replicas for essentially the entire keyspace. The two statements are not contradictory: one describes primary range ownership; the other describes replicated data placement. This is why reading a bare ownership number without naming the keyspace and RF is dangerous.
NetworkTopologyStrategy also applies rack
awareness. If there are at least as many racks as replicas,
Cassandra attempts to put replicas in distinct racks. Therefore
“next tokens clockwise” is only an introductory mental model;
the real strategy skips vnode positions that would duplicate a
physical node and considers topology when choosing distinct
replicas.
Disk usage also depends on replication factor, table sizes, compaction state, tombstones, indexes, compression, unreclaimed streamed data, and uneven application key distributions. Use token ownership as one input, not a capacity oracle.
5. Wrong model: “the ring is three nodes in a circle”
Single-token diagrams are valuable to explain wrap-around, but
operational Cassandra with vnodes maps each host to many ring
positions. A learner who memorizes one-node-one-token will
misunderstand bootstrap streaming, repair neighbors,
availability coupling, and why
nodetool ring becomes long.
A better production mental model has three layers: (1) a sorted token space divided into ranges, (2) a token map from many positions to physical endpoints, and (3) a keyspace-specific replica map derived from token ownership plus topology-aware replication. The network topology is a separate graph of endpoints and links.
Verification checklist
- You can explain the open/closed boundaries of a token range and wrap-around.
-
You compared
nodetool statuswith keyspace-awarestatus/ring. - You can distinguish primary ownership from effective replicated ownership.
- You resolved multiple AtlasMart tokens to replicas.
- You can explain why requests do not physically hop clockwise around the ring.
Check your understanding
- What does the interval (t1, t2] mean?
- Why can three RF=3 nodes each show very high effective ownership?
- Why pass a keyspace to nodetool ring?
- Does walking clockwise describe actual network hops?
- Why do vnodes make the ring output long?
Review the answers
1. Keys whose partitioner tokens are greater than t1 and less than or equal to t2 belong to that token range, with wrap-around at the hash-space boundary.
2. Because each partition has three replicas; effective replicated ownership is not a partition of 100% across nodes.
3. Replication strategy/topology are keyspace-specific, and the official tool requires that context for accurate effective ownership.
4. No. It describes sorted token/replica selection logic, not packet routing.
5. Each physical node owns many token positions/ranges rather than a single position.
Cleanup
docker rm -f atlasmart-cass-4 atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>/dev/null || truedocker volume rm atlasmart-cass-4-data atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>/dev/null || truedocker network rm atlasmart-cassandra 2>/dev/null || true
docker rm -f atlasmart-cass-4 atlasmart-cass-1 atlasmart-cass-2 atlasmart-cass-3 2>$nulldocker volume rm atlasmart-cass-4-data atlasmart-cass-1-data atlasmart-cass-2-data atlasmart-cass-3-data 2>$nulldocker network rm atlasmart-cassandra 2>$null
Summary and next step
This lesson’s concepts, evidence path, failure boundaries, and production judgment should now be explicit enough to verify rather than assume. Re-run the check-your-understanding prompts and preserve any lab evidence you need before changing or cleaning up the environment.
Next, continue to Virtual Nodes: Many Tokens per Node, Distribution, Streaming, and Operational Tradeoffs.
Authoritative references
- Apache Cassandra 5.0 documentation — Official documentation entry point for the current 5.0 line.
- Apache Cassandra downloads — Official release page used to verify Cassandra 5.0.9 as the current GA patch.
- Dynamo architecture: token ring, vnodes, replication — Official explanation of consistent hashing, token ranges, vnodes, natural replicas, and ring membership.
- cassandra.yaml configuration — Official partitioner, num_tokens, token-allocation, and related configuration reference.
- Production token recommendations — Current guidance for vnode counts and token-allocation tradeoffs.
- CQL token() function — Official semantics for mapping partition-key values to partitioner tokens.
- cqlsh SHOW REPLICAS — Official cqlsh command for resolving a token to replicas for a keyspace.
- nodetool ring — Official ring command and keyspace requirement for topology-aware ownership.
- Topology changes and bootstrap streaming — Official bootstrap/token-allocation/streaming and cleanup guidance.
- nodetool netstats — Official command for observing streaming/network activity.
- Java Driver 4.19 TokenMap — Driver-side token ranges, node tokens, and replica lookup API.